A method and device for multi-aggregator day-ahead bidding game collaborative optimization decision of electric vehicles considering carbon emission reduction benefits
By using a multi-agent deep deterministic policy gradient algorithm to model Markov games, we constructed the carbon emission reduction revenue objective function and constraints for multiple aggregators of electric vehicles. This solved the strategic game problem in a competitive-cooperative environment for multiple aggregators, and enabled efficient and robust decision-making in the electricity market.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- STATE GRID FUJIAN POWER ELECTRIC CO ECONOMIC RESEARCH INSTITUTE
- Filing Date
- 2026-05-19
- Publication Date
- 2026-07-31
AI Technical Summary
Existing technologies struggle to effectively characterize the strategic game scenarios of electric vehicle aggregators in a multi-aggregator competitive-cooperative environment. Furthermore, existing methods are unable to accurately model the randomness of electric vehicle travel and market price fluctuations, and lack adaptive capabilities.
A Markov game model is constructed using the Multi-Agent Deep Deterministic Policy Gradient Algorithm (MADDPG), which includes an objective function and constraints that incorporate carbon emission reduction benefits. A unified clearing model for electric vehicle aggregators participating in the distribution network electricity market is used to conduct day-ahead power purchase and sale bidding.
It improves the robustness and accuracy of bidding decisions, coordinates and optimizes the economic benefits of aggregators with the low-carbon operation goals of the distribution network, enhances the adaptability to changes in market signals, and strengthens the foresight and engineering adaptability of decisions.
Smart Images

Figure CN122491603A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of electricity market optimization, and more particularly to a collaborative optimization decision-making method and apparatus for multi-aggregator day-ahead bidding game of electric vehicles that takes into account carbon emission reduction benefits. Background Technology
[0002] Electric Vehicle Aggregators (EVAs) are playing a crucial role in the low-carbon transformation of the power system as a scalable and flexible resource. Multiple EVAs aggregate a large number of distributed electric vehicles to form dispatchable capacity with dual "source-load" attributes. They can participate in both the electricity and carbon markets simultaneously, achieving deep synergy in charge / discharge scheduling, energy trading, carbon quota optimization, and demand response. This provides strong support for renewable energy consumption and the low-carbon operation of the system. Electric vehicles, relying on aggregators, implement V2G (Vehicle-to-Grid) technology, discharging during periods of high electricity prices and charging during periods of low carbon prices, or participating in carbon emission reduction trading, resulting in significant peak-shaving reserves and carbon emission reduction benefits.
[0003] Currently, a relatively mature methodology has been developed for optimizing the scheduling of electricity and carbon markets under uncertainty. Mainstream model-based methods include stochastic optimization, robust optimization, and partial Bruker optimization. Stochastic optimization explicitly characterizes price and EV uncertainty through scenario sampling, but the computational burden increases exponentially with the number of scenarios. Robust optimization seeks conservative solutions under the worst-case scenario, which can easily lead to over-conservatism and economic losses. Partial Bruker optimization seeks optimal decisions within a pre-defined fuzzy distribution set, but still relies on analytical modeling of the uncertain distribution. Furthermore, existing methods mostly focus on independent optimization problems of single EVAs, such as bidding strategies in the electricity spot market or quota management in the carbon market, making it difficult to effectively characterize the systemic challenges of complex dynamic interactions among multiple stakeholders. When facing collaborative optimization scenarios with multiple aggregators in a competitive-cooperative mixed environment, game theory-based methods, while capable of describing the strategic behavior of each stakeholder, are highly sensitive to model parameters and often involve high solution complexity.
[0004] In summary, existing model-based methods have significant advantages in theoretical rigor and interpretability, but their common limitation lies in their heavy reliance on accurate mathematical models. Faced with highly nonlinear and non-stationary decision-making environments characterized by the stochasticity of electric vehicle travel, market price fluctuations, and multi-agent strategic interactions, these methods often struggle to achieve accurate modeling and lack adaptability to multi-agent strategic game scenarios. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a collaborative optimization decision-making method and device for multi-aggregator day-ahead bidding game of electric vehicles that takes into account carbon emission reduction benefits, which can provide an efficient, robust and scalable intelligent decision-making paradigm for the electricity market.
[0006] To solve the above-mentioned technical problems, the present invention adopts the following technical solution: A collaborative optimization decision-making method based on a competitive game among multiple aggregators of electric vehicles, taking into account carbon emission reduction benefits, includes the following steps: S1. Introduce multiple electric vehicle aggregators into the bidding system and construct an objective function that includes carbon emission reduction benefits and corresponding constraints. S2. Construct a unified clearing model for electric vehicle multi-aggregator participation in the power distribution network market within the bidding system; S3. Based on the unified clearing model, perform Markov game modeling for day-ahead power purchase and sale transactions; S4. Under the constraints, with the maximization of the objective function as the reward signal, the strategy of the Markov game is updated using a multi-agent deep deterministic policy gradient algorithm.
[0007] To solve the above-mentioned technical problems, another technical solution adopted by the present invention is as follows: A collaborative optimization decision-making device for multi-aggregator day-ahead bidding game of electric vehicles that takes into account carbon emission reduction benefits includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the aforementioned collaborative optimization method for multi-aggregator day-ahead bidding game of electric vehicles that takes into account carbon emission reduction benefits.
[0008] The beneficial effects of this invention are as follows: Through interactive learning between the agent and the environment, it can better adapt to the multi-source uncertainties brought about by fluctuations in renewable energy output and the randomness of user charging behavior, thereby improving the robustness and accuracy of bidding decisions; by constructing a bilateral power trading model that considers the revenue from carbon emission reduction (CCER), it can synergistically optimize the economic benefits of aggregators and the low-carbon operation goals of the distribution network, and improve the adaptability to changes in market signals while ensuring user charging needs; by introducing Markov game modeling and a unified clearing mechanism, it helps to improve the stability and convergence performance of the training process of the multi-agent deep deterministic policy gradient (MADDPG) algorithm, enhance the foresight of the day-ahead scheduling strategy, and better characterize the competition and cooperation relationships among multiple agents; at the same time, the method has a clear structure and good engineering adaptability. Combined with supply and demand balance constraints and carbon emission reduction incentive mechanisms, it can provide quantitative basis for aggregator market participation and user response behavior, enhancing the practical application value of the method. Therefore, it has the advantages of strong adaptability to multi-source uncertainties, good scheduling coordination, high decision-making efficiency, balance between economy and low carbon, and high engineering application feasibility. Attached Figure Description
[0009] Figure 1 This is a flowchart illustrating the steps of a collaborative optimization decision-making method for multi-aggregator day-ahead bidding game in electric vehicles that takes into account carbon emission reduction benefits, according to an embodiment of the present invention. Figure 2 This is a schematic diagram of a collaborative optimization decision-making device for multi-aggregator day-ahead bidding game of electric vehicles that takes into account carbon emission reduction benefits, according to an embodiment of the present invention. Figure 3 This is a convergence analysis diagram of the MADDPG algorithm according to an embodiment of the present invention; Figure 4 This is a convergence analysis diagram of the IDPG algorithm in the prior art; Figure 5 This invention provides a comparison of surplus renewable energy power in the distribution network under the MADDPG model. Figure 6 A comparison of surplus renewable energy power in distribution networks under existing IDPG technology; Figure 7 This is a typical daily distribution of new energy power generation, load demand, and EVA demand according to an embodiment of the present invention; Figure 8 The MADDPG-EVAs energy state and market trading energy distribution in this embodiment of the invention; Figure 9 The energy state and market trading energy distribution of IDPPG-EVAs are shown in this embodiment of the invention. Detailed Implementation
[0010] To explain in detail the technical content, objectives, and effects of the present invention, the following description is provided in conjunction with the embodiments and accompanying drawings.
[0011] The above-described method and apparatus for collaborative optimization decision-making in a day-ahead bidding game among multiple electric vehicle aggregators, taking into account carbon emission reduction benefits, is applicable to application scenarios where multiple electric vehicle aggregators participate in the collaborative optimization of the power distribution network market. The following detailed implementation method illustrates this: In one alternative implementation, such as Figure 1 As shown, a collaborative optimization decision-making method for multi-aggregator bidding game of electric vehicles that takes into account carbon emission reduction benefits includes the following steps: S1. Introduce multiple electric vehicle aggregators into the bidding system and construct an objective function that includes carbon emission reduction benefits and corresponding constraints. First, the application scenarios and core objectives of the technology are clearly defined, and the participating entities and core technologies of multiple aggregators are identified: For the day-ahead bidding scenario of distribution networks with multiple aggregators and fluctuating renewable energy output, the core objectives are clearly defined as improving the overall revenue of aggregators and promoting renewable energy consumption, while simultaneously meeting market clearing and power balance constraints. Multiple electric vehicle aggregators (EVAs) are introduced into the bidding system as market participants, focusing on exploring the dual "source-load" attributes of electric vehicles. Combined with carbon emission reduction trading mechanisms, a game-theoretic bidding and collaborative optimization framework involving multiple EVAs is constructed to effectively respond to the needs of the aforementioned scenario. The objective function in S1 also includes revenue from electric vehicle charging services, electricity purchase and sale costs in the electricity market, and compensation costs for vehicle owners. The constraints include constraints on the amount of electricity bid and won by aggregators, constraints on the energy of aggregators, constraints on the amount of electricity purchased and sold by aggregators to the distribution network, and constraints on carbon emission reduction revenue. S2. Construct a unified clearing model for electric vehicle multi-aggregator participation in the power distribution network market within the bidding system; S3. Based on the unified clearing model, perform Markov game modeling for day-ahead power purchase and sale transactions; S4. Under the constraints, with the maximization of the objective function as the reward signal, the strategy of the Markov game is updated using a multi-agent deep deterministic policy gradient algorithm.
[0012] In another alternative implementation, the objective function is:
[0013] In the formula, This represents the total revenue of the nth electric vehicle aggregator within the scheduling period T. The total revenue of the nth electric vehicle aggregator in time period t can be expressed as:
[0014] In the formula, This indicates revenue from electric vehicle charging services. Indicates the benefits of carbon emission reduction. This indicates the cost of buying and selling electricity in the electricity market. This indicates the cost of temporary electricity purchases. This indicates the cost of compensation for the car owner; In this embodiment, the following settings are provided: , The time interval between adjacent SOC moments is calculated, where h is in hours; in,
[0015] In the formula, The EV charging price coefficient set by the nth electric vehicle aggregator in the tth time period represents the revenue that the aggregator obtains for each unit of electricity charged to the EV. Let be the total energy replenishment demand of the nth electric vehicle aggregator in time period t, which represents the total electricity obtained by the EVs from the aggregator. Clearly, we have:
[0016] Typically determined by electricity tariff coefficients Service fee coefficient It consists of two parts, which vary according to different time periods, and generally exhibit the pattern of "high electricity rate, low service fee rate, and low electricity rate, high service fee rate":
[0017] carbon emission benefits The formula is as follows:
[0018]
[0019] In the formula, The service fee charged to the nth electric vehicle aggregator for selling a unit of electricity generated by the CCER (Carbon Credit Emission Revenue), is typically based on carbon pricing. The percentage-based commission method, with a commission rate of [percentage missing]. CCER stands for Chinese Certified Emission Reduction, which is the amount of carbon emissions reduced compared to traditional technologies. The CCER (Carbon Credit Equivalent) is a credit that can be transferred to an electric vehicle aggregator after the electric vehicle has been charged with a unit of electricity. The electric vehicle aggregator then trades it on the carbon market on behalf of the EV. This represents the winning bid power obtained by the nth aggregator in the day-ahead market bidding during the t-th time period. When it is positive, it means that the aggregator purchases electricity from the distribution network, and when it is negative, it means that the aggregator sells electricity to the distribution network. The efficiency of the nth aggregator when charging the EV; Let be the efficiency of the EV when discharging to the nth polymerizer; Electricity purchase and sale costs in the electricity market and It is expressed as follows:
[0020]
[0021] In the formula, Let be the unified electricity price after market clearing in the t-th time period. The cost of purchasing and selling electricity generated from this price and the corresponding electrical energy is: , To punish electricity prices, Let be the net power surplus or deficit of the nth aggregator in the t-th time period, used to characterize the power balance state of that aggregator in that time period; when When t is a power shortage, it indicates that aggregator n has a power deficit during time period t and needs to purchase electricity from external sources to make up the difference in order to meet the charging needs of its EVs; the temporary electricity purchase cost incurred at this time is ; Car owner compensation costs It is expressed as follows:
[0022] In the formula, The compensation costs that need to be paid to the owners of its EVs for the nth time are used to offset depreciation losses such as those caused by battery charging and discharging. The constraints on the electricity volume awarded in the aggregated commercial bidding include:
[0023] This represents the bid power submitted by electric vehicle aggregator n to the distribution network at time t. A positive value indicates that the aggregator purchases electricity from the distribution network, while a negative value indicates that the aggregator sells electricity to the distribution network. This represents the winning power obtained by the nth electric vehicle aggregator in the day-ahead market bidding during the t-th time period; The aggregator energy constraint includes:
[0024] This represents the maximum amount of electricity that electric vehicle aggregator n can aggregate. Let n be the amount of electricity possessed by electric vehicle aggregator n during time period t; The constraints on the amount of electricity purchased and sold by the aggregator from the distribution network include:
[0025] electric vehicle aggregator During the period The state variable for electricity purchase and sale is 1 when the aggregator purchases electricity and 0 when the aggregator sells electricity. , Electric vehicle aggregators During the period The power volume submitted to the power distribution network for power purchase and sale decisions must satisfy the following:
[0026] and Electric vehicle aggregators During the period Maximum power purchase capacity and maximum power sales capacity, these two parameters vary with electric vehicle aggregators During the period Energy Total energy demand Maximum charging power of electric vehicle aggregators Maximum discharge power The upper limit of energy that electric vehicle aggregators can aggregate. Related, expressed as:
[0027] That is, the maximum power purchased by aggregator n in time period t. Not only affected by the maximum charging power of charging stations and charging piles under the aggregator Restrictions, and also subject to its time period Remaining battery capacity limitations; and the maximum power sold by aggregator n during time period t. It is limited not only by its maximum discharge power, but also by the energy remaining after the inherent losses of its EVs are met in time period t.
[0028] The carbon emission reduction benefit constraints include: In the carbon market, in addition to trading carbon emission allowances, the carbon market will issue CCERs to entities that achieve emission reductions through the use of new technologies, such as the EV in this application and the EVA that aggregates these resources. The CCERs generated by the EV represent the carbon emission reductions achieved by replacing traditional gasoline-powered vehicles, and the calculation formula is as follows:
[0029] The carbon emission benefits generated by electric vehicles (EVs) The number of kilometers an EV travels per unit of electricity consumed. This refers to the average fuel consumption per unit kilometer of a traditional gasoline-powered vehicle. This refers to the carbon emission factor of traditional gasoline-powered vehicles, specifically the carbon emissions produced per unit of fuel consumed by a traditional gasoline-powered vehicle. The average power consumption per kilometer driven by the EV. The carbon emission factor of an EV is the amount of carbon emissions generated per unit of electricity consumed by the EV. Carbon emission reduction benefits per unit of electricity consumed by EVs The calculation formula is as follows:
[0030] In the formula: This refers to the CCER value that an EV can transfer to an EVA after charging a unit of electricity. It should be noted that when electric vehicle aggregators participate in the carbon market, they should only calculate the service fee for clearing emission reductions for the EVs they represent in the carbon market, as they do not have their own carbon allowances and do not directly generate CCERs.
[0031] This implementation method establishes an optimized scheduling model by determining the objective function and defining the corresponding constraints.
[0032] In another alternative implementation, S2 includes: A centralized trading market is established with electric vehicle aggregators as the main bidding entities. The buyers are the electric vehicle aggregators that submit electricity purchase bids, and the sellers are the distribution network in a state of renewable energy surplus and the electric vehicle aggregators that submit electricity sales bids. When the renewable energy output of the system exceeds the load demand, the distribution network, as the highest priority seller, enters the market with a zero bid. A unified clearing mechanism is executed at preset time intervals. Taking into account the power and price of each electric vehicle aggregator's bids, combined with the new energy surplus, the market clearing is completed based on the principles of price priority and proportional allocation, and the actual transaction power and market settlement price of each electric vehicle aggregator are determined. The market clearing process based on the principles of price priority and proportional allocation includes: The clearing priority is sorted for both buyers and sellers based on price. For buyers, the order is sorted in descending order of bid price. When sorting buyers, if the energy storage of an electric vehicle aggregator is lower than a preset threshold, the emergency power purchase process is initiated and its bid priority is increased. For sellers, the order is sorted in ascending order of bid price. Market participants with the same price are grouped into the same price group and matched on a group basis. By extracting the same price groups of buyers and sellers in rounds, the transaction conditions are determined, and the transaction volume and unified settlement price are determined based on the principle of supply and demand balance. The transaction volume is allocated according to the proportion of the remaining tradable power of each entity in the group, and the unified settlement price is determined based on the marginal buy and sell quotes corresponding to the last successful match in the corresponding time period.
[0033] In practice: Multiple electric vehicle aggregators (EVA) operate as independent market players, jointly participating in short-term (e.g., day-ahead and real-time) wholesale electricity market bidding, and coordinating resources under the organization and coordination of distribution network operators. To simulate and analyze the complex strategic interactions among these aggregators at the distribution network level as they compete for limited network resources, this implementation constructs a local market bidding model suitable for the distribution network environment. This model achieves electricity transaction matching between the distribution network and EVAs by executing a uniform clearing mechanism every hour. This mechanism comprehensively considers the power and bid price of each EVA, combined with the surplus of new energy sources, and completes market clearing based on price priority and proportional allocation principles, determining the actual transaction power and market clearing price for each aggregator, and assisting the distribution network in absorbing surplus new energy sources.
[0034] (1) Market aggregation construction To depict the clearing process of multiple electric vehicle aggregators (EVA) participating in power purchase and sale transactions on the distribution network, this implementation method constructs a centralized trading market for multiple aggregators. The bidding entities in the market are the individual EVAs, who can submit power purchase or sale applications based on their own energy status and trading needs. When the renewable energy output within the distribution network exceeds the local load, the remaining renewable energy is included in the market clearing as additional supply to reflect the priority consumption of renewable energy. It should be noted that the distribution network side only provides the remaining renewable energy and does not participate in the bidding process among aggregators as a market operator. Buyers submit power purchase bids. The various EVAs, among which The seller is a distribution network in a state of renewable energy surplus and submits a bid for electricity sales. The various EVAs, among which ; This represents the electricity purchase and sale price declared by the EV aggregator at time t, and is a non-negative number. When the system's renewable energy output exceeds load demand, i.e. At that time, the distribution network, as the highest priority seller, enters the market with a zero bid to ensure priority consumption of new energy.
[0035]
[0036]
[0037]
[0038] in: Let be the set of buyers in time period t; Let be the set of sellers for time period t; Index 0 in the formula represents the bid for the distribution network, and its price. The value is 0, and the renewable energy power margin (power for distribution network bidding) is... ; The power output of new energy sources during time period t. Let t represent the load demand of the distribution network during time period t.
[0039] (2) Priority sorting rules To ensure fairness in market clearing, buyers and sellers must be prioritized according to price. Buyer ranking: Buyers are ranked in descending order of bid price, reflecting the "highest bidder wins" principle, prioritizing the electricity purchase needs of aggregators with stronger willingness to pay. Specifically, when ranking buyers, EVA needs to consider its own energy reserves; if its energy storage falls below a pre-set threshold... At that time, the "emergency power purchase process" will be initiated, and the priority of its bid will be increased:
[0040] in: This represents the lower limit of energy storage safety for the nth EVA. Its physical meaning can be understood as "the minimum amount of electricity required to ensure that the EVA can provide normal service to the EV". Let M represent the energy of the nth EVA at time t; M is a large parameter used to ensure that EVAs with low energy are prioritized in the buyer bidding order. Therefore, the buyer priorities can be arranged in descending order as follows:
[0041] Seller sorting: Sellers are ranked in ascending order of bid price, prioritizing those with lower-cost discharge resources. Therefore, seller priority can be arranged in ascending order as follows:
[0042] In the formula: The sequence of buyers in time period t, sorted in descending order of bid prices; The sequence of sellers in time period t, sorted in ascending order of bid prices; and These represent the electricity purchase price offered by the i-th and j-th entities in the buyer sequence, respectively. and These represent the electricity purchase price offered by the i-th and j-th entities in the seller sequence, respectively.
[0043] (3) Calculation of transaction volume To improve settlement efficiency and ensure fairness, the algorithm groups market participants with the same price quote into the same "price group" and matches them on a group-by-group basis. By extracting groups of buyers and sellers with the same price quote round by round, the algorithm determines the transaction conditions and establishes a unified settlement price based on the principle of supply and demand balance, achieving efficient and transparent bilateral transaction clearing. In the k-th round of matching, the highest priority buyer quote is defined as... The corresponding buyer-same-price group is:
[0044] The lowest cost seller's offer is The corresponding seller's same-price group is:
[0045] In the formula: This represents the buyer's same-price group corresponding to the highest priority bid in the k-th round of matching; This represents the seller's lowest bid matching group in the k-th round of matching.
[0046] For buyer n and seller m, record their respective first and second... The remaining tradable power before the start of the round is and The initial value is equal to the absolute value of its original bid power. Therefore, the total residual demand and supply for this round of matching are respectively:
[0047] A transaction can only be completed if the buyer's offer is not lower than the seller's offer, that is:
[0048] If the conditions are not met, the matching process terminates. If the conditions are met, the maximum feasible transaction volume in this round is determined by the shortcomings of both supply and demand:
[0049] In the formula: This represents the maximum energy transaction volume that can be completed between groups of the same price.
[0050] (4) Allocation of liquidation price and winning bid ratio In the aforementioned round-by-round matching process, each round of matching is primarily used to determine the transaction volume between buyers and sellers. To ensure consistency and transparency in market settlement, this implementation method does not use different prices for each round of matching. Instead, it uses the marginal buy / sell quote corresponding to the successful match in the last round as the basis to determine a unique, unified clearing price for that period, ensuring a balance of interests between buyers and sellers. The matching round in time period t that results in a successful transaction in the last round is defined as... Then we have:
[0051] If there exists a satisfying The matching round, then the time period The uniform clearing price can be expressed as:
[0052] In the formula: Let $T$ be the unified electricity price after market clearing by the day-end of time period $t$. If no such price exists... If the matching round is zero, it means there were no valid transactions during that period. At this time, the clearing price is set to 0, and the transaction volume for each entity is 0. To ensure fairness within the group, the transaction volume in the k-th round is allocated according to the proportion of remaining tradable power for each entity within the group: For each buyer:
[0053] For each seller:
[0054] In the formula: The transaction power allocated to the buyer entity in time period t at the current round k; Assign the allocated trading power to the seller in the current round at time period t, and update the remaining power of each trading entity:
[0055] In the formula: Let n be the remaining power of buyer entity n in the (k+1)th round of time period t; Let n be the remaining power of buyer entity n in the k-th round of time period t; The remaining power of seller m in the (k+1)th round of time period t; Let m be the remaining power of seller entity m in round k of time period t. If a group fails to complete a transaction, its unfinished entities must be kept in the queue for matching in the next round. This mechanism allows for multiple rounds of cross-matching until no matching can be found. The price pair is used to maximize market turnover.
[0056] This implementation method enables the construction of a unified clearing model for electric vehicle multi-aggregator participation in the power distribution network market.
[0057] In another alternative implementation, S3 includes: Based on the scenario of the day-ahead purchase and sale price of e-sports in the distribution network, a corresponding Markov game decision chain is set; Based on the Markov game decision chain, at time t, the nth electric vehicle aggregator observes the local state information and takes action according to its own strategy. Apply the joint bidding actions of all electric vehicle aggregators to the scenario of the day-ahead purchase and sale e-sports price of the distribution network; The unified clearing model completes the clearing calculation for the bidding of multiple electric vehicle aggregators, provides real-time rewards to each electric vehicle aggregator, and transitions to the next state, providing observational basis for the decision-making of each electric vehicle aggregator at time t+1.
[0058] The Markov game decision chain set according to the scenario of the day-ahead purchase and sale price of the distribution network includes: A preset period is defined as a complete decision chain, and the period is divided into preset time periods; Each electric vehicle aggregator is set to complete a bidding decision once in each time period. After the bidding decision is completed, the immediate reward and the clearing result for the next time period are calculated through the unified clearing model.
[0059] In practice, the elements in a Markov game process include a set of agents. Observation state space Action space Reward function and discount factor Based on the actual scenario of day-ahead bidding in the distribution network, this implementation method sets the corresponding Markov game decision chain as follows: One day is considered a complete decision chain, divided into 24 time periods (the time period division standard is 0:00-1:00 as the first time period, and so on up to 23:00-24:00 as the 24th time period); each electric vehicle aggregator (agent) completes a bidding decision every hour, that is, at time t... Bidding strategies for specific time periods. After a decision is made, an immediate reward is obtained through market clearing model calculations. Results will be released immediately.
[0060] Based on the above Markov game decision chain, the specific interaction process between each agent and the market at time t is as follows: First, the nth agent observes its local state information. And based on their own strategies Generate action Secondly, the joint bidding actions of all intelligent agents. The market environment is then applied; finally, the market environment completes the clearing calculation of multi-entity bidding through the market clearing model preset in the above implementation method, and provides real-time rewards to each agent. and transition to the next state. This provides observational basis for the agent's decision-making at time t+1.
[0061] The key to solving the day-ahead bidding problem for electric vehicle aggregators in the distribution network using deep reinforcement learning lies in the rationality and adaptability of the designed Markov game process. Therefore, it is necessary to systematically design the observation state space, action space, and reward function for the electric vehicle aggregator (agent) to ensure that the MADDPG algorithm can effectively train and optimize the agent's bidding decision strategy.
[0062] The step of updating the strategy of the Markov game using a multi-agent deep deterministic policy gradient algorithm includes: The multi-agent deep deterministic policy gradient algorithm updates multi-agent game strategies based on a mechanism of global state-local observation-joint action-centralized value evaluation. In the multi-agent deep deterministic policy gradient algorithm, the state space Observations from each agent Composition, corresponding to the first The local information received by each electric vehicle aggregator is represented by the following state set:
[0063] In the formula: For each electric vehicle aggregator, a local information vector containing 12 dimensions is provided, including the current battery level of the electric vehicle. System power difference Total energy replenishment demand for electric vehicles Current electricity price Maximum power purchase capacity Maximum power sales Current period Number of remaining time periods Forecast values of new energy power output Distribution network load demand Carbon market prices Electricity price forecast for the next period The observation set for the single aggregator quotient is as follows:
[0064] In the formula, the number of remaining time periods ; Action space The actions of each agent Composition, specifically representing the first The electricity purchase and sale decision of an electric vehicle aggregator includes two types of decision variables: the electricity purchased and sold and the bid electricity price. The action set is defined as follows:
[0065]
[0066] Among them, a dynamic action boundary calculation method is adopted to calculate the feasible region of the action space for each electric vehicle aggregator based on real-time state information and the constraints. For the bidding power... The action boundary is calculated using the constraints defined in the aforementioned implementation method, for the electricity purchase and sale price declared by the aggregator. It should satisfy:
[0067] The agent's reward set Rewards for each agent The composition is specifically defined as the reward value obtained by electric vehicle aggregator n within the preset period (1 hour in this embodiment) through the unified clearing model:
[0068]
[0069] In the formula: Indicates in During the time period, the nth electric vehicle aggregator participating in the game participates in market bidding and the profit value is calculated after market clearing.
[0070] Based on the above Markov game process, in the multi-agent deep deterministic policy gradient algorithm, each agent independently maintains a set of Actor-Critic networks. The Actor network outputs policy actions using only local observation information and fits the policy function, where For the first Policy network parameters for each agent; The Critic network collects global state information and evaluates the value of joint actions, fitting a value function with the following parameters: To guide network parameter updates; In the multi-agent deep deterministic policy gradient algorithm, using Indicates the first The strategy of an agent Then the first The policy gradient of an agent can be expressed as:
[0071] In the formula, For information about the global environment, It is the first A centralized Critic network is used to fit the value function, and its input is the th... Individual agents based on deep deterministic policies And the actions taken and environmental information The output is the first... A single agent value, This represents the experience replay pool, which contains tuples. Record all training samples of the agents; use Indicates the first The objective policy function of each agent, using Let the parameters of the target Critic network be represented. Then, the loss function of the Critic is defined as follows:
[0072]
[0073] In the formula, This represents the parameters of the current Critic network. This represents the reward value obtained by the agent. This represents the state vector at the next moment. This represents the action vector at the next moment. Indicates the discount factor. This represents the m-th centralized target Critic network; The parameters of the Actor network and the Critic network can be updated using the following two formulas:
[0074]
[0075] In the formula, , These represent the learning rates of the Actor network and the Critic network, respectively. Represents the parameters of the current Actor network; The parameters of both the target Actor network and the target Critic network are updated using a soft update method:
[0076]
[0077] In the formula, Indicates the soft update rate. , These represent the parameters of the target Actor and the target Critic network, respectively.
[0078] In another alternative implementation, such as Figure 2As shown, a collaborative optimization decision-making device for a multi-aggregator day-ahead bidding game of electric vehicles that takes into account carbon emission reduction benefits includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the collaborative optimization method for a multi-aggregator day-ahead bidding game of electric vehicles that takes into account carbon emission reduction benefits as described in any of the above embodiments.
[0079] To verify the performance of this application, a case study analysis was conducted in another optional implementation. In this implementation, all case studies were solved and analyzed on a computer with an Intel Ultra 7 265K CPU, an NVIDIA GeForce RTX 4060ti 16GB GPU, and 48GB DDR5 8200MHz memory. The program code was written using PyCharm and Python 3.9. To further verify the effectiveness of the proposed multi-agent cooperative bidding method, this implementation introduced a single-agent deep deterministic policy gradient algorithm (IDDPG) as a baseline for comparison in the case studies. In this method, each aggregator independently learns policies and updates actions based only on its own local observations, and its critic network only receives the state-action information of its own agent. Both algorithms were trained in the same environment to compare the performance improvement of multi-agent cooperative decision-making compared to independent decision-making. (1) Analysis of algorithm convergence and scalability During the training process of multi-agent deep reinforcement learning, the rewards for the original rounds may fluctuate significantly due to factors such as randomly sampling different daily market data in each round, adding Gaussian noise to the actor network output to ensure policy exploration, and environmental non-stationarity caused by policy updates by different agents. Therefore, to accurately assess the convergence trend and eliminate short-term fluctuation interference, this implementation method uses the Moving Average (MA) method to smooth the original rewards:
[0080] In the formula: Let be the original reward for agent n in round j. Let W be the smoothed reward for agent n in the k-th round. W is the smoothing window size, which is set to 20 in this implementation. To analyze the convergence characteristics of the algorithm under different numbers of aggregators, this implementation was trained using the MADDPG and IDDPG algorithms in scenarios with 3, 6, and 9 aggregators, respectively. The curves showing the change in reward value with training rounds are shown below. Figure 3 and Figure 4 As shown.
[0081] Through analysis Figure 3 and Figure 4 As can be seen, with the increase in training rounds, both the MADDPG and IDDPG algorithms used in this implementation can gradually converge within 2000 rounds in scenarios with 3, 6, and 9 aggregators, indicating that both algorithms have good trainability and convergence. Further observation shows that as the number of aggregators increases from 3 to 9, the fluctuation range of the reward curve in the early stages of training generally increases, and the number of rounds required to reach a stable state also increases. This indicates that as the number of participating aggregators increases, the game coupling relationship between multiple agents further strengthens, and the system state transition and policy search space become more complex, thus increasing the training difficulty of the algorithm. Further comparison of the reward curve characteristics of the two algorithms shows that IDDPG enters a relatively stable range earlier in some scenarios, exhibiting a faster early convergence speed; in contrast, MADDPG needs to undergo a more thorough exploration and multi-agent policy coordination process in the early stages of training, therefore its convergence speed is slightly slower.
[0082] The average rewards of the two algorithms during the last 20 training rounds are recorded in Table 1: Table 1 Comparison of average rewards in the last 20 rounds
[0083] As shown in Table 1, MADDPG's average reward in the last 20 rounds is higher than IDDPG in scenarios with 3, 6, and 9 aggregators, improving by 25.44%, 32.41%, and 3.17% respectively, indicating that MADDPG is superior in terms of stable returns after convergence. The improvements are more significant in scenarios with 3 and 6 aggregators, suggesting that in small-to-medium-scale multi-agent game environments, MADDPG can more effectively uncover interactions between agents and achieve collaborative optimization. However, in the scenario with 9 aggregators, the marginal impact of a single agent on the overall clearing result is relatively weakened, and the gains from multi-agent collaboration exhibit diminishing marginal returns, thus narrowing MADDPG's advantage over IDDPG. Overall, although MADDPG requires more rounds to coordinate policies in the early stages of training, it ultimately achieves better collaborative decision-making results through more comprehensive multi-agent information interaction, resulting in higher economic benefits.
[0084] (2) Comparison of the renewable energy consumption effect of typical daily power distribution networks To further examine the performance differences between MADDPG and IDPG in promoting the absorption of renewable energy on the distribution network side, this implementation method compares the adjustment effects of the two algorithms on the surplus renewable energy power of the distribution network under different numbers of aggregators based on a typical daily scenario. By analyzing the changes in the surplus power of the distribution network before and after the transaction, the role of the multi-aggregator collaborative bidding strategy in absorbing surplus renewable energy and alleviating the pressure of power curtailment is revealed. The surplus power curves of the MADDPG and IDPG algorithms after transactions with 3, 6, and 9 aggregators are plotted as follows. Figure 5 and Figure 6 As shown.
[0085] analyze Figure 5 and Figure 6 As can be seen, compared with the pre-trade scenario, the remaining surplus renewable energy power on the distribution network side decreased significantly after the transaction, indicating that the participation of aggregators in the market can effectively promote renewable energy consumption. With the increase in the number of aggregators, the remaining surplus power after the transaction generally shows a downward trend, indicating that the expansion of adjustable load resources helps enhance the system's ability to absorb surplus renewable energy. To quantitatively compare the renewable energy consumption effects under different algorithms and aggregator scales, Table 2 presents typical daily distribution network surplus renewable energy statistical indicators for different numbers of aggregators, including average surplus power, fluctuation degree, and curtailment rate. The fluctuation degree is measured using standard deviation.
[0086] Table 2 Comparison of Surplus New Energy Indicators in Typical Daily Power Distribution Networks
[0087] Analysis of Table 2 shows that both algorithms effectively reduce the curtailment rate of renewable energy in the distribution network and suppress its fluctuations. Further comparison reveals that, under the same aggregator scale, MADDPG generally has lower surplus power than IDDPG, and its reduction effect is more significant during periods of high renewable energy output, indicating that it can more effectively coordinate the collaborative response of multiple aggregators. Specifically, MADDPG can effectively reduce response dispersion and timing mismatch caused by independent decision-making, thereby improving the absorption capacity of surplus renewable energy and alleviating curtailment pressure. This result corresponds to the conclusion that MADDPG has a higher average reward after convergence, indicating that it can more accurately characterize the strategic coupling relationship between multiple aggregators and form a more effective collaborative adjustment effect in renewable energy consumption scenarios involving multiple stakeholders.
[0088] (3) Analysis of typical daily market behavior and benefits of aggregators To further analyze the impact of aggregator market behavior on the consumption of renewable energy in the distribution network and overall efficiency, this implementation method uses three aggregator scenarios as examples to compare the market behavior of the two algorithms under typical days. Typical daily renewable energy output, load demand, and EVA energy demand in the distribution network are as follows: Figure 7 As shown.
[0089] analyze Figure 7 It is evident that there is a significant time-varying mismatch between renewable energy generation and distribution network load demand during a typical intraday period. Overall, renewable energy generation fluctuates considerably, while load demand remains relatively stable. Starting from period 10, renewable energy output rises rapidly, then falls back at period 17:00, creating a surplus of electricity available for aggregators to absorb during this period.
[0090] From the perspective of aggregator load demand, all three aggregators exhibited certain fluctuations throughout the day, with varying scales among them, showing a clear "commuter-responsive" characteristic. For example, aggregator load demand was low during periods 6-8 because a large number of EVs were disconnected from the grid and put into operation during commuting hours, leading to a temporary decrease in charging demand. However, it gradually increased from period 9 onwards until the lunch break, indicating that users were reconnecting their EVs to the grid, thus restoring load demand. The following graphs are used to plot the EVA energy state and market trading energy distribution on a typical day under the models trained by the two algorithms. Figure 8 and Figure 9 .
[0091] analyze Figure 8 and Figure 9 It is evident that both algorithms exhibit certain responsive characteristics in terms of EVA energy state and market trading behavior. Specifically, during periods of high renewable energy output, EVA energy state is increased through electricity purchases, and then gradually released in subsequent periods, thereby absorbing surplus renewable energy while meeting subsequent load demands. However, compared to IDPG, MADDPG shows a closer match between the trading behavior of each EVA and the periods of renewable energy surplus. Trading volume is more concentrated in the high-output renewable energy range, and the energy state rises faster and reaches a higher peak, indicating that it can more effectively guide multiple aggregators to form a coordinated response during periods of surplus renewable energy.
[0092] It is worth noting that under the IDDPG algorithm, some aggregators exhibit significantly insufficient trading activity during certain time periods, even displaying a strategy characteristic of near-non-participation in trading. This is because IDDPG employs an independent learning mechanism, which struggles to accurately characterize the strategic coupling relationships among multiple aggregators. This can easily lead individual entities to converge to conservative, locally suboptimal strategies during competition, thereby weakening the overall collaborative adjustment capability. In contrast, MADDPG, through centralized training, can more effectively coordinate the behavior of multiple aggregators, thus preventing some entities from being marginalized. EVA income and expenditure are recorded in Tables 3 and 4, with all income items in the tables expressed in yuan.
[0093] Table 3 MADDPG - EVAs Profitability by Item
[0094] Table 4 IDDPG - EVAs Profitability
[0095] Analysis of Tables 3 and 4 shows that under both algorithms, the main revenue of EVA comes from electric vehicle charging services. Although CCER revenue is positive, its overall proportion is relatively small, at 2.24% and 2.11% under the MADDPG and IDDPG algorithms, respectively. This indicates that under the current parameter settings, the direct contribution of carbon revenue to total revenue is still relatively limited. It is worth noting that in the scenario with three aggregators, the two algorithms correspond to... All values are 0, indicating that the surplus renewable energy on the distribution network side during a typical day can meet the power purchase demand of aggregators, and the market clearing price drops to the distribution network bidding price, hence this item is 0. In the IDPG algorithm, since EVA 2 has no valid bidding activity, its CCER revenue is 0. Table 5 records the total system revenue for different aggregators, with all revenue items in the table in yuan.
[0096] Table 5 Comparison of System Revenue Indicators under Different Aggregators
[0097] As shown in Table 5, the revenue from system charging services increases significantly with the increase in the number of aggregators, indicating that the expansion of aggregator scale can effectively drive the growth of market transaction volume and improve the overall revenue level. At the same time, surplus renewable energy on the distribution network side is further absorbed, and the surplus electricity available for aggregators to obtain at low prices gradually decreases. Under these circumstances, some aggregators can no longer meet their own needs by purchasing low-priced surplus renewable energy, thus needing to rely on more expensive alternatives, leading to an increase in system penalty costs.
[0098] In conclusion, compared to IDDPG, MADDPG can more effectively guide multiple aggregators to respond to periods of renewable energy surplus, thereby improving the overall system revenue and enhancing the renewable energy absorption capacity.
[0099] Deep reinforcement learning, as a data-driven, model-free method, learns the optimal policy simply by interacting with the market environment. It abstracts the multi-agent collaborative optimization problem into a partially observable Markov game, gradually approximating the globally optimal joint policy through the interaction of multiple agents with the market environment, without explicitly modeling the uncertainty distribution or game equilibrium. While Deep Deterministic Policy Gradient (DDPG) is suitable for continuous action spaces, it cannot effectively handle the environmental non-stationarity caused by other agent behaviors within a single-agent framework. Multi-Agent Deep Deterministic Policy Gradient (MADDPG) is a multi-agent extension of this approach. Its core idea is as follows: each aggregator's actor network outputs a deterministic policy to generate actions; a centralized evaluator network simultaneously observes the state-action information of all agents during the training phase to calculate Q-values and estimate the advantage function; and during the execution phase, each agent relies only on local observations to achieve decentralized decision-making. This centralized training-decentralized execution (CTDE) paradigm effectively mitigates multi-agent non-stationarity while retaining the high sample efficiency of deterministic policy gradients. Compared to Proximal Policy Optimization (PPO, with low sample utilization for updating with the same policy), Trust Region Policy Optimization (TRPO, requiring a large computational cost of a second-order Hessian matrix), and Single-Agent DDPG (ignoring interactions leads to policy collapse), MADDPG offers advantages such as stable training, high sample efficiency, and natural adaptation to continuous high-dimensional action spaces: the centralized evaluator solves the credit allocation problem, the deterministic policy avoids random sampling noise, and only a first-order gradient is needed for efficient optimization. In the collaborative optimization of the power distribution network market involving multiple aggregators of electric vehicles, MADDPG can use the available capacity of EVs and the net load of the system as state inputs, the bidding volume and bidding price of each aggregator as continuous actions, and the maximization of aggregator revenue as a reward signal to achieve multi-agent adaptive collaborative optimization, providing an efficient, robust and scalable intelligent decision-making paradigm for the power market.
[0100] In summary, this invention provides a collaborative optimization decision-making method and apparatus for day-ahead bidding game of electric vehicles that takes into account carbon emission reduction benefits. Addressing the complexities of day-ahead bidding decisions among multiple aggregators and the difficulty of traditional optimization methods effectively handling multi-agent games in the market, this invention proposes a collaborative optimization method for day-ahead bidding among multiple aggregators based on the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm. This method, through interactive learning between the agent and the environment, can better adapt to the multi-source uncertainties brought about by fluctuations in renewable energy output and the randomness of user charging behavior, thereby improving the robustness and accuracy of bidding decisions. By constructing a bilateral power trading model that considers CCER revenue, it can synergistically optimize the economic benefits of aggregators and the low-carbon operation goals of the distribution network, and improve the adaptability to market signal changes while ensuring user charging needs. By introducing Markov game modeling and a unified clearing mechanism, it helps to improve the stability and convergence performance of the MADDPG training process, enhance the foresight of the day-ahead scheduling strategy, and better characterize the competition and cooperation relationships among multiple agents. At the same time, the method has a clear structure and good engineering adaptability. Combined with supply and demand balance constraints and carbon emission reduction incentive mechanisms, it can provide quantitative basis for aggregator market participation and user response behavior, enhancing the practical application value of the method. Therefore, this invention has the advantages of strong adaptability to multi-source uncertainties, good scheduling coordination, high decision-making efficiency, balance between economy and low carbon, and high engineering application feasibility.
[0101] The above description is merely an embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent modifications made based on the content of the present invention specification and drawings, or direct or indirect applications in related technical fields, are similarly included within the patent protection scope of the present invention.
Claims
1. A method for coordinated optimization decision of day-ahead bidding game of electric vehicle multi-aggregator considering carbon emission reduction benefits, characterized in that, Including the following steps: S1. Introduce multiple electric vehicle aggregators into the bidding system and construct an objective function that includes carbon emission reduction benefits and corresponding constraints. S2. Construct a unified clearing model for electric vehicle multi-aggregator participation in the power distribution network market within the bidding system; S3. Based on the unified clearing model, perform Markov game modeling for day-ahead power purchase and sale transactions; S4. Under the constraints, with the maximization of the objective function as the reward signal, the strategy of the Markov game is updated using a multi-agent deep deterministic policy gradient algorithm. 2.The method of claim 1, wherein, The objective function in S1 also includes electric vehicle charging service revenue, electricity purchase and sale costs in the electricity market, and vehicle owner compensation costs. The constraints include constraints on the amount of electricity won by aggregators, constraints on the energy of aggregators, constraints on the amount of electricity purchased and sold by aggregators to the distribution network, and constraints on carbon emission reduction revenue. 3.The method of claim 2, wherein, The objective function is: wherein, denotes the total revenue of the nth electric vehicle aggregator over the dispatching period T, denotes the total revenue of the nth electric vehicle aggregator at time period t, which can be expressed as: In the formula, This indicates revenue from electric vehicle charging services. Indicates the benefits of carbon emission reduction. This indicates the cost of buying and selling electricity in the electricity market. This indicates the cost of temporary electricity purchases. This indicates the cost of compensation for the car owner; The constraints on the electricity volume awarded in the aggregated commercial bidding include: This represents the bid power submitted by electric vehicle aggregator n to the power distribution network at time t. This represents the winning power obtained by the nth electric vehicle aggregator in the day-ahead market bidding during the t-th time period; The aggregator energy constraint includes: This represents the maximum amount of electricity that electric vehicle aggregator n can aggregate. Let n be the amount of electricity possessed by electric vehicle aggregator n during time period t; The constraints on the amount of electricity purchased and sold by the aggregator from the distribution network include: electric vehicle aggregator During the period The state variables of electricity purchase and sale, , Electric vehicle aggregators During the period The power volume submitted to the power distribution network for power purchase and sale decisions must satisfy the following: and Electric vehicle aggregators During the period Maximum power purchase capacity and maximum power sales capacity, these two parameters vary with electric vehicle aggregators During the period Energy Total energy demand Maximum charging power of electric vehicle aggregators Maximum discharge power The upper limit of energy that electric vehicle aggregators can aggregate. Related, expressed as: To calculate the time interval between adjacent SOC moments; The carbon emission reduction benefit constraints include: The carbon emission benefits generated by electric vehicles (EVs) The number of kilometers an EV travels per unit of electricity consumed. This refers to the average fuel consumption per unit kilometer of a traditional gasoline-powered vehicle. For carbon emission factors of traditional gasoline vehicles, The average power consumption per kilometer driven by the EV. Carbon emission factor for EVs; Carbon emission reduction benefits per unit of electricity consumed by EVs The calculation formula is as follows: 。 4. The collaborative optimization decision-making method for multi-aggregator day-ahead bidding game of electric vehicles taking into account carbon emission reduction benefits, as described in claim 1, is characterized in that... S2 includes: A centralized trading market is established with electric vehicle aggregators as the main bidding entities. The buyers are the electric vehicle aggregators that submit electricity purchase bids, and the sellers are the distribution network in a state of renewable energy surplus and the electric vehicle aggregators that submit electricity sales bids. When the renewable energy output of the system exceeds the load demand, the distribution network, as the highest priority seller, enters the market with a zero bid. A unified clearing mechanism is executed at preset time intervals. Taking into account the power and price of each electric vehicle aggregator's bids, and combined with the new energy surplus, the market clearing is completed based on the principles of price priority and proportional allocation, and the actual transaction power and market settlement price of each electric vehicle aggregator are determined.
5. The collaborative optimization decision-making method for multi-aggregator day-ahead bidding game of electric vehicles taking into account carbon emission reduction benefits, as described in claim 4, is characterized in that... The market clearing process based on the principles of price priority and proportional allocation includes: The clearing priority is sorted for both buyers and sellers based on price. For buyers, the order is sorted in descending order of bid price. When sorting buyers, if the energy storage of an electric vehicle aggregator is lower than a preset threshold, the emergency power purchase process is initiated and its bid priority is increased. For sellers, the order is sorted in ascending order of bid price. Market participants with the same price are grouped into the same price group and matched on a group basis. By extracting the same price groups of buyers and sellers in rounds, the transaction conditions are determined, and the transaction volume and unified settlement price are determined based on the principle of supply and demand balance. The transaction volume is allocated according to the proportion of the remaining tradable power of each entity in the group, and the unified settlement price is determined based on the marginal buy and sell quotes corresponding to the last successful match in the corresponding time period.
6. A collaborative optimization decision-making method for multi-aggregator day-ahead bidding game of electric vehicles taking into account carbon emission reduction benefits, as described in any one of claims 1 to 5, characterized in that, S3 includes: Based on the scenario of the day-ahead purchase and sale price of e-sports in the distribution network, a corresponding Markov game decision chain is set; Based on the Markov game decision chain, at time t, the nth electric vehicle aggregator observes the local state information and takes action according to its own strategy. Apply the joint bidding actions of all electric vehicle aggregators to the scenario of the day-ahead purchase and sale e-sports price of the distribution network; The unified clearing model completes the clearing calculation for the bidding of multiple electric vehicle aggregators, provides real-time rewards to each electric vehicle aggregator, and transitions to the next state, providing observational basis for the decision-making of each electric vehicle aggregator at time t+1.
7. The collaborative optimization decision-making method for multi-aggregator day-ahead bidding game of electric vehicles taking into account carbon emission reduction benefits, as described in claim 6, is characterized in that... The Markov game decision chain corresponding to the scenario of purchasing and selling e-sports prices based on the distribution network day-ahead includes: A preset period is defined as a complete decision chain, and the period is divided into preset time periods; Each electric vehicle aggregator is set to complete a bidding decision once in each time period. After the bidding decision is completed, the immediate reward and the clearing result for the next time period are calculated through the unified clearing model.
8. A collaborative optimization decision-making method for multi-aggregator day-ahead bidding game of electric vehicles taking into account carbon emission reduction benefits, as described in any one of claims 1 to 5, characterized in that, The process of updating the strategy of the Markov game using a multi-agent deep deterministic policy gradient algorithm includes: The multi-agent deep deterministic policy gradient algorithm updates multi-agent game strategies based on a mechanism of global state-local observation-joint action-centralized value evaluation. In the multi-agent deep deterministic policy gradient algorithm, the state space Observations from each agent Composition, corresponding to the first The local information received by each electric vehicle aggregator is represented by the following state set: In the formula: For each electric vehicle aggregator, a local information vector containing 12 dimensions is provided, including the current battery level of the electric vehicle. System power difference Total energy replenishment demand for electric vehicles Current electricity price Maximum power purchase capacity Maximum power sales Current period Number of remaining time periods Forecast values of new energy power output Distribution network load demand Carbon market prices Electricity price forecast for the next period The observation set for the single aggregator quotient is as follows: In the formula, the number of remaining time periods ; Action space The actions of each agent The composition specifically represents the electricity purchase and sale decision of the nth electric vehicle aggregator, which includes two types of decision variables: the electricity purchased and sold and the declared electricity price. The action set is defined as follows: Among them, a dynamic action boundary calculation method is adopted to calculate the feasible domain of the action space for each electric vehicle aggregator based on real-time state information and the constraints. The agent's reward set Rewards for each agent The composition, specifically defined as the reward value obtained by electric vehicle aggregator n within the preset period after calculation using the unified clearing model: In the formula: Indicates in During the time period, the nth electric vehicle aggregator participating in the game participates in market bidding and the profit value is calculated after market clearing.
9. A collaborative optimization decision-making method for multi-aggregator day-ahead bidding game of electric vehicles taking into account carbon emission reduction benefits, as described in claim 8, is characterized in that... In the multi-agent deep deterministic policy gradient algorithm, each agent independently maintains a set of Actor-Critic networks; The Actor network outputs policy actions using only local observation information and fits the policy function, where For the first Policy network parameters for each agent; The Critic network collects global state information and evaluates the value of joint actions, fitting a value function with the following parameters: To guide network parameter updates; In the multi-agent deep deterministic policy gradient algorithm, using Indicates the first The strategy of an agent Then the first The policy gradient of an agent can be expressed as: In the formula, For information about the global environment, It is the first A centralized Critic network, whose input is the first... Individual agents based on deep deterministic policies And the actions taken and environmental information The output is the first... A single agent value, This represents the experience replay pool, which contains tuples. Record all training samples of the agents; use Indicates the first The objective policy function of each agent, using Let the parameters of the target Critic network be represented. Then, the loss function of the Critic is defined as follows: In the formula, This represents the parameters of the current Critic network. This represents the reward value obtained by the agent. This represents the state vector at the next moment. This represents the action vector at the next moment. Indicates the discount factor. This represents the m-th centralized target Critic network; The parameters of the Actor network and the Critic network can be updated using the following two formulas: In the formula, , These represent the learning rates of the Actor network and the Critic network, respectively. This represents the parameters of the current Actor network; The parameters of both the target Actor network and the target Critic network are updated using a soft update method: In the formula, Indicates the soft update rate. , These represent the parameters of the target Actor and the target Critic network, respectively.
10. A collaborative optimization decision-making device for multi-aggregator day-ahead bidding game of electric vehicles that takes into account carbon emission reduction benefits, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the day-ahead competitive game collaborative optimization method for electric vehicles that takes into account carbon emission reduction benefits, as described in any one of claims 1 to 9.