A multi-agent collaborative game decision method and system

By constructing a multi-agent collaborative game decision-making system, utilizing deep reinforcement learning to optimize bidding strategies and combining cooperative game algorithms to allocate benefits, the problems of isolated decision-making and unreasonable benefit distribution in the frequency modulation market are solved, thereby improving market efficiency and alliance stability.

CN121602392BActive Publication Date: 2026-05-15RES INST OF ECONOMICS & TECH STATE GRID SHANDONG ELECTRIC POWER
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
RES INST OF ECONOMICS & TECH STATE GRID SHANDONG ELECTRIC POWER
Filing Date
2026-01-28
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

The existing frequency regulation market suffers from isolated decision-making among multiple stakeholders, distorted modeling, unreasonable revenue distribution, and poor policy adaptability, resulting in low market efficiency and difficulty in coping with the systemic risks brought by new energy sources. Furthermore, existing models are insufficient to meet real-time decision-making needs.

Method used

A multi-agent collaborative game decision-making system is constructed. Deep reinforcement learning (DRL) agents optimize bidding strategies within a price limit range, and cooperative game algorithms are used to distribute profits, thereby achieving the aggregation and fair allocation of flexible resources.

Benefits of technology

It has achieved economies of scale and complementary effects in market competition, increased the probability of winning bids and expected total benefits, reduced the update costs caused by policy changes, and ensured the stability and fairness of the alliance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121602392B_ABST
    Figure CN121602392B_ABST
Patent Text Reader

Abstract

The application provides a multi-agent cooperative game decision method and system, relates to the technical field of power market, and comprises the following steps: acquiring sensing data required for decision; constructing a cooperation alliance by a thermal power-energy storage-VPP multi-agent, participating in frequency modulation market bidding as a whole, under the bidding constraint, taking the maximization of expected total income of the alliance as a target, and searching for an optimal bidding strategy through a deep reinforcement learning intelligent agent to assist the alliance bidding; and after the alliance wins the income in the frequency modulation service, an optimal income distribution scheme is determined by combining a cooperative game algorithm to assist the distribution of the alliance income among the alliance members. The application aggregates the dispersed flexible resources (i.e. the thermal power-energy storage-VPP) into a whole (i.e. the cooperative alliance) to participate in the market competition of the frequency modulation auxiliary service, and realizes the scale effect and complementary effect of "1+1>2".
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of electricity market technology, specifically to a multi-stakeholder collaborative game decision-making method and system. Background Technology

[0002] With the deepening of the construction of new power systems, the integration of a high proportion of renewable energy into the grid has become an inevitable trend. The randomness, volatility, and intermittency of power output from new energy sources such as wind and solar power greatly increase the difficulty of real-time grid balancing, placing higher demands on the quantity and quality of ancillary services such as frequency regulation. Against this backdrop, the frequency regulation market, as a key mechanism for discovering the value of frequency regulation services and incentivizing flexible resources to participate in grid regulation, is becoming increasingly important.

[0003] Traditionally, the frequency regulation market has primarily consisted of thermal power units, which offer stable regulation performance but relatively slow response times. However, the market is rapidly diversifying, with energy storage (ES), virtual power plants (VPPs), and even renewable energy power plants with regulation capabilities actively participating, in addition to thermal power. These emerging players offer advantages such as fast response times and high regulation accuracy, but their cost structures (e.g., cycle life degradation of energy storage) and output characteristics (e.g., load aggregation uncertainty of VPPs) differ significantly from traditional thermal power, presenting new challenges for market decision-making.

[0004] Currently, both domestic and international frequency regulation markets generally adopt a price cap / floor policy, where regulatory agencies exogenously set upper and lower limits for capacity and performance pricing to prevent market manipulation and drastic price fluctuations, thus ensuring system security. In existing technologies, decision-making by multiple stakeholders under a given price cap mainly faces the following dilemmas:

[0005] (1) Decision-making isolation and inefficiency: Market players usually make decisions independently and bid separately, and their strategies are based solely on their own costs and incomplete market information. This "fighting alone" model cannot form a synergy and is difficult to cope with the systemic risks brought by new energy, resulting in overall market inefficiency and the failure to fully tap the value of a large number of flexible resources.

[0006] (2) Modeling distortion and single strategy: Existing bidding models are mostly based on cost-plus pricing or simple game theory, which makes it difficult to accurately depict the real costs and constraints of complex entities such as energy storage and VPP (such as nonlinear life loss and aggregation effect); their strategies are often limited to passively accepting price limits and lack the ability to actively and meticulously bid within the price limit range, thus failing to maximize profits.

[0007] (3) Unreasonable profit distribution mechanism: In existing cooperative bidding studies, traditional cooperative game methods such as the Shapley value method are commonly used for profit distribution. However, the Shapley value method requires calculating all possible permutations and combinations of members, and the computational workload increases exponentially with the number of members (combinatorial explosion problem), which is difficult to meet the needs of real-time decision-making in the electricity market. At the same time, its distribution results may not effectively incentivize key contributors and damage the long-term stability of the alliance.

[0008] (4) Poor policy adaptability: Existing decision-making systems usually treat price limits as fixed boundary conditions, resulting in rigid strategies. When regulatory agencies adjust the price limit range according to market development dynamics, existing models lack the ability to adapt quickly and often require re-adjusting parameters or even reconstructing the model, which is costly and time-consuming.

[0009] Therefore, the aforementioned shortcomings of existing technologies hinder the improvement of market efficiency and make it difficult to ensure the safe and stable operation of the power grid. Summary of the Invention

[0010] To address the aforementioned issues, this invention proposes a multi-agent collaborative game decision-making method and system. By constructing a collaborative alliance, dispersed flexible resources are aggregated into a whole to participate in the market competition for frequency modulation auxiliary services, achieving a scale effect and complementary effect of "1+1>2". The intelligent bidding strategy based on DRL (deep reinforcement learning) can find the optimal balance between returns and risks within a given price limit range, maximizing the alliance's winning probability and expected total returns.

[0011] According to some embodiments, the present invention adopts the following technical solution:

[0012] A multi-agent collaborative game decision-making method includes:

[0013] Acquire the perceived data needed for decision-making; perceived data includes market data, entity data, and external environment data.

[0014] Construct a multi-entity cooperative alliance. Based on perception data, under given bidding constraints, with the goal of maximizing the expected total revenue of the multi-entity cooperative alliance, use a deep reinforcement learning agent to find the best bidding strategy to participate in the FM market bidding.

[0015] After a multi-entity cooperative alliance wins a bid for FM service and obtains revenue, the optimal revenue distribution scheme is determined based on the value of each member in the alliance and the revenue obtained by the alliance from winning the bid for FM service, combined with a cooperative game theory algorithm. The benefits are then distributed to each member according to the optimal revenue distribution scheme.

[0016] According to some embodiments, the present invention adopts the following technical solution:

[0017] A multi-agent collaborative game decision-making system, comprising:

[0018] The perception data acquisition module is configured to acquire the perception data required for decision-making, including market data, entity data, and external environment data.

[0019] The bidding strategy optimization module is configured to: construct a multi-entity cooperative alliance; based on perception data, under given bidding constraints, with the goal of maximizing the expected total revenue of the multi-entity cooperative alliance, use a deep reinforcement learning agent to find the best bidding strategy to participate in the FM market bidding.

[0020] The revenue distribution optimization module is configured as follows: After a multi-entity cooperative alliance wins a bid for frequency modulation services and obtains revenue, it determines the optimal revenue distribution scheme based on the value of each member in the multi-entity cooperative alliance and the revenue obtained by the alliance from winning the bid for frequency modulation services, combined with a cooperative game algorithm, and distributes the benefits to each member according to the optimal revenue distribution scheme.

[0021] According to some embodiments, the present invention adopts the following technical solution:

[0022] A computer program product includes a computer program that, when executed by a processor, implements the aforementioned multi-agent collaborative game decision-making method.

[0023] According to some embodiments, the present invention adopts the following technical solution:

[0024] A non-transitory computer-readable storage medium is provided for storing computer instructions, which, when executed by a processor, implement the aforementioned multi-agent collaborative game decision-making method.

[0025] According to some embodiments, the present invention adopts the following technical solution:

[0026] An electronic device includes a processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to enable the electronic device to implement the multi-agent collaborative game decision-making method.

[0027] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0028] 1. This invention, by constructing a collaborative alliance, aggregates dispersed flexible resources into a whole to participate in market competition, achieving a scale effect and complementary effect of "1+1>2". The intelligent bidding strategy based on DRL can find the optimal balance between returns and risks within a given price limit range, maximizing the alliance's winning probability and expected total returns. The fair profit distribution mechanism ensures that each member can obtain higher returns than when bidding independently, thereby greatly enhancing the attractiveness and stability of the alliance and effectively smoothing out abnormal fluctuations in market prices.

[0029] 2. This invention directly inputs the price limit range as an exogenous variable into the state space of the DRL model; when market regulators adjust price limit policies, only the input parameters need to be updated, and the DRL agent can quickly adapt to the new price environment and automatically adjust its bidding strategy through a small amount of online learning or fine-tuning; this flexibility means that the system does not need to redesign the core algorithm, which greatly reduces the update cost and delay caused by policy changes;

[0030] 3. This invention completely abandons the computationally complex Shapley value method and innovatively adopts the nucleolus method for profit distribution; the nucleolus pursues the greatest possible fairness (minimizing the maximum complaint), ensuring the stability of the alliance; it accurately calculates the value complaint value of each sub-alliance, changing the distribution basis from "theoretical capacity" to "actual contribution", making the distribution result more scientific, reasonable and easily accepted by all parties. Attached Figure Description

[0031] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0032] Figure 1 This is a flowchart of a multi-agent collaborative game decision-making method as shown in Example 1. Detailed Implementation

[0033] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0034] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0035] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0036] Example 1

[0037] One embodiment of the present invention provides a multi-agent collaborative game decision-making method, comprising:

[0038] Step S1: Obtain the perception data required for decision-making. The perception data includes market data, entity data, and external environment data.

[0039] Step S2: Construct a multi-agent cooperative alliance. Based on the perception data, under the given bidding constraints, with the goal of maximizing the expected total revenue of the multi-agent cooperative alliance, the best bidding strategy is found through a deep reinforcement learning agent, so as to participate in the bidding in the frequency modulation market using the best bidding strategy.

[0040] Furthermore, the process of finding the optimal bidding strategy using a deep reinforcement learning agent specifically involves:

[0041] A coalition bidding model is defined using Markov decision processes;

[0042] Based on the alliance bidding model, the agent is trained through deep reinforcement learning. Under bidding constraints including FM capacity price limit, FM performance price limit and FM capacity limit, the optimal bidding strategy is sought with the goal of maximizing the expected total revenue constructed based on the winning bid revenue and deviation penalty.

[0043] The bidding strategy mentioned above involves the alliance offering FM capacity, FM performance, and FM capacity quotes for specific time periods.

[0044] Furthermore, the alliance bidding model is defined, including defining the state space, action space, and reward function;

[0045] The states in the state space This can be expressed as a formula:

[0046]

[0047] in, This represents the total frequency modulation capacity that the alliance can provide during time period t; This represents the overall actual performance index of the alliance in time period t. This represents the historical frequency-adjusted clearing price series over the past N time periods. This indicates the exogenous policy constraints corresponding to time period t. This is timestamp information;

[0048] Actions in the action space This can be expressed as a formula:

[0049]

[0050] in, This indicates the frequency modulation capacity quotation for the alliance during time period t; This indicates the alliance's frequency modulation performance quote for time period t; This represents the frequency modulation capacity of the alliance during time period t;

[0051] The reward function This can be expressed as a formula:

[0052]

[0053] in, This indicates the portion of the revenue the alliance receives from participating in frequency modulation. This indicates the penalty for deviation.

[0054] Furthermore, the agent is trained using deep reinforcement learning, employing a soft actor-commentator algorithm for training, with the training objective being to maximize the expected total revenue of the multi-agent cooperative alliance, expressed by the formula:

[0055]

[0056] in, The objective function is represented by t; the iteration step is represented by t, and the maximum iteration step is represented by T. This represents the curve trajectory formed by all strategies; Indicates the discount factor; Represents the reward function, This is the current state. This is the current action.

[0057] Step S3: After the multi-entity cooperative alliance wins the bid for the frequency modulation service and obtains the revenue, the optimal revenue distribution scheme is determined based on the value of each member in the multi-entity cooperative alliance and the revenue obtained by the alliance from winning the bid for the frequency modulation service, combined with the cooperative game algorithm, and the benefits are distributed to each member according to the optimal revenue distribution scheme.

[0058] Furthermore, the optimal profit distribution scheme determined by combining cooperative game theory algorithms is used to assist in distributing alliance profits among alliance members, specifically as follows:

[0059] Initialize the current profit distribution plan;

[0060] Based on the trained agent, the individual earnings of each member are predicted to quantify the member's value. Then, based on this quantified value, the complaint value of each member under the current earnings distribution scheme is calculated. This constitutes the current profit distribution plan. Complaint value vector ;

[0061] To minimize the maximum complaint value in the complaint value vector With the goal of improving the current profit distribution plan... Perform iterative optimization until the condition for stopping the iteration is met.

[0062] Furthermore, the calculation of each member's complaint value under the current profit distribution scheme, whereby the complaint value represents how much more profit the sub-alliance could gain by acting alone compared to the current distribution scheme, is used to calculate the profit of the sub-alliance acting alone. Alliance with children Under the current allocation scheme Total revenue The difference is used to calculate, and the formula is as follows:

[0063]

[0064] in, For the Alliance Under the current allocation scheme The complaint value below, For the current profit distribution plan, Sub-alliances can be composed of several members, including special sub-alliances composed of a single member. To predict the profits of a sub-alliance acting alone, Indicates the current profit distribution plan The revenue to be allocated to the i-th member of the neutron alliance.

[0065] As one embodiment, the multi-agent collaborative game decision-making method of the present invention provides a two-stage decision-making framework:

[0066] Phase 1 (Bidding Strategy Phase): Construct a cooperative alliance (e.g., led by a VPP aggregator, comprising energy storage, controllable loads, and even thermal power plants with good regulation performance). The alliance participates in frequency regulation market bidding as a whole, and its bidding strategy is optimized by a deep reinforcement learning agent, with the goal of maximizing the alliance's expected total revenue under given exogenous (i.e., determined by the external environment) bidding constraints.

[0067] The second stage (profit distribution stage): Based on the bidding results and actual performance data of the first stage, a distribution mechanism based on marginal contribution and incentive compatibility is adopted instead of Shapley value to fairly and efficiently distribute the total profit of the alliance and ensure the stability of the alliance. Marginal contribution is often used to measure whether an interest group should accept new members, that is, to increase the contribution of unit members to the group, while incentive compatibility refers to aligning the interests of agents and owners through certain systems and technologies.

[0068] Based on a two-stage decision-making framework, the specific implementation scheme of this embodiment will be described in detail:

[0069] I. System Architecture Implementation

[0070] The implementation system corresponding to the method described in this embodiment adopts a layered design, follows the data flow orientation principle, and realizes full-process automation from data acquisition, processing, intelligent decision-making to application output. The system architecture is mainly divided into three core layers: the data perception layer, the model algorithm layer, and the decision display layer. Its overall architecture and data flow are as follows: Figure 1 As shown:

[0071] 1. Data perception layer

[0072] The data perception layer is the foundation of the entire system, responsible for collecting, cleaning, and standardizing data from diverse and heterogeneous data sources in real time or near real time. This layer mainly includes the following modules:

[0073] (1) Market Data Module:

[0074] The data source is the publicly available data interface released by the power trading center. The collected data includes historical and real-time frequency regulation market clearing prices (capacity price, performance price), cleared volume, frequency regulation mileage, total market demand, and most importantly, the exogenous price limit range. (Frequency modulation capacity price range) and (Price range for frequency modulation performance).

[0075] (2) Main data module:

[0076] The data source is a data acquisition agent deployed locally on the premises of each alliance member (thermal power, energy storage, VPP, and new energy). The collected content includes:

[0077] Energy storage: Real-time state of charge, rated power, charge / discharge efficiency, maximum adjustable frequency capacity, and health status.

[0078] Virtual power plant: aggregated total interruptible capacity of flexible loads, current controllable distributed energy output, and predicted response potential.

[0079] Thermal power: Current output, ramp-up rate, minimum technical output, and historical data on frequency regulation performance.

[0080] New energy: Ultra-short-term power prediction curves and prediction error distribution.

[0081] (3) External environment data module:

[0082] The data sources are meteorological departments and time servers. The collected content includes future weather conditions (which affect the output and load of new energy sources) and current timestamps (date, hour, and whether it is a holiday, used to capture the periodic patterns of load and frequency regulation demand).

[0083] 2. Model Algorithm Layer

[0084] The model algorithm layer is the "brain" of the system, carrying the core computational intelligence. It receives data processed by the data perception layer, runs advanced algorithms, and outputs decision results, specifically:

[0085] (1) Data preprocessing and feature engineering module:

[0086] Receive sensing data, perform data cleaning (handling missing values ​​and outliers), and normalize / standardize; fuse and reconstruct multi-source data into the state vector required by the DRL agent, for example, convert historical price series into time series features, and aggregate data from each member into the overall capability index of the alliance.

[0087] (2) DRL Collaborative Bidding Agent Module:

[0088] The system employs a soft actor-commentator algorithm to perform model training, policy inference, and experience replay. Model training involves continuously interacting with the environment (exploration-exploitation) in a simulated environment or historical data to update and optimize the policy network and Q-network parameters, maximizing long-term cumulative rewards. Policy inference, at the bidding moment, calculates the optimal bidding action (i.e., capacity bid, performance bid, and frequency modulation capacity) based on the current state, ensuring that its output strictly falls within the price limit range. Experience replay involves storing and transferring samples for offline learning, improving data utilization and training stability.

[0089] (3) Profit distribution calculation engine module:

[0090] After the alliance wins the bid, a cooperative game algorithm is run based on the value of each sub-alliance. The algorithm library uses two optimization solvers: a built-in kernel and a least squares value algorithm. The kernel is implemented by solving a linear programming sequence, while the least squares value is quickly calculated through analytical solutions or the least squares method, outputting a fair and reasonable payout vector.

[0091] 3. Decision-making display layer

[0092] The decision presentation layer provides users (alliance operators) with a way to view decision-making solutions. It transforms the output of the algorithm layer into an understandable presentation to assist users in bidding and profit distribution, including:

[0093] (1) The joint bidding strategy generated by the DRL agent The tender documents are automatically packaged into a format that meets the data requirements of the trading center for users to view.

[0094] (2) Display the output results of the revenue distribution calculation engine and generate a revenue list for each member; provide support functions such as data sharing, contract management, and dispute arbitration (if there is any objection to the distribution results) to maintain the daily operation of the alliance.

[0095] After completing the above system architecture, the system's workflow can be summarized as follows: the data perception layer continuously collects data and sends it to the model algorithm layer; the model algorithm layer constructs it into a state vector for the DRL agent to make real-time decisions; the decision display layer presents the decision results to the user or directly executes the bidding; after winning the bid, the data perception layer continues to collect actual performance data and feeds it back to the algorithm layer, triggering the profit distribution calculation engine to run, and finally publishing the distribution plan in the decision display layer; the entire system forms a closed-loop intelligent decision-making ecosystem.

[0096] The implementation methods of the DRL collaborative bidding agent module and the profit distribution calculation engine module in the model algorithm layer are explained below:

[0097] II. DRL Collaborative Bidding Agent Module: This module uses a deep reinforcement learning agent to find the optimal bidding strategy to assist in alliance bidding.

[0098] A coalition bidding model is defined using Markov decision processes. Based on this model, an agent is trained using deep reinforcement learning to automatically generate the optimal bidding strategy for the coalition bidding, including:

[0099] 1. MDP-based consortium bidding model:

[0100] The core of this phase is to train a DRL agent to act as the "super brain" of the alliance, making optimal joint bidding decisions under given price limits.

[0101] Therefore, the problem is first modeled as a Markov Decision Process (MDP), whose basic construction includes a state space. Action space Reward function.

[0102] The problem to be solved is the bidding decision of the alliance, and its core elements can be constructed as the following state vector:

[0103] (1)

[0104] In the above formula, This represents the total frequency regulation capacity (MW) that the alliance can provide during time period t. This represents the timestamp information of the data; The overall actual performance index of the alliance in time period t can be measured by a weighted average, specifically:

[0105] (2)

[0106] Where N represents the total number of members in the alliance; This represents the frequency modulation performance index of the i-th member; This represents the actual effective frequency modulation mileage of the i-th member during time period t.

[0107] In the state vector This represents a historical frequency-adjusted clearing price series over the past N periods, reflecting recent market volatility. Examples include:

[0108] (3)

[0109] In the state vector The exogenous policy constraints corresponding to time period t are as follows:

[0110] (4)

[0111] in, and These represent the price range for frequency modulation capacity and the price range for frequency modulation performance, respectively.

[0112] In the state vector This includes timestamp information, such as hours and weekdays / weekends, used to capture the periodic patterns of system frequency adjustment needs.

[0113] The set containing all state vectors mentioned above is called the state space, i.e. .

[0114] After completing the state space, we examine the actions corresponding to the decisions, i.e., the feasible set of decision variables. In this embodiment, the actions represent the two prices that the alliance submits to the trading center. Therefore, the actions can be represented as follows:

[0115] (5)

[0116] in, This represents the price declared by the alliance for frequency modulation capacity during time period t, i.e., the frequency modulation capacity quotation; This represents the price declared by the alliance for frequency modulation performance in time period t, i.e., the frequency modulation performance quotation; This represents the frequency regulation capacity declared by the alliance for time period t. Frequency regulation capacity is a specific technical parameter that directly determines how many megawatts of physical power will be used to respond to frequency fluctuations in the power grid. Actions based on this decision variable of frequency regulation capacity directly involve decisions regarding the output / capacity of physical equipment, enabling decision-making and control of power system resources. Actions must comply with bidding constraints, including the price limit for frequency regulation capacity, the price limit for frequency regulation performance, and frequency regulation capacity limitations, expressed by the formula:

[0117] (6)

[0118] The set of all actions is called the action space, that is...

[0119] Considering the alliance's revenue model, we define a reward function that includes two parts: the winning bid reward and the deviation penalty.

[0120] (7)

[0121] in, The revenue function representing the alliance's participation in frequency modulation (FM) reporting is as follows:

[0122] (8)

[0123] In the above formula, and These represent the frequency regulation capacity clearing price and the performance clearing price for time period t, respectively. This represents the alliance's winning bid capacity during time period t; Indicates the duration of FM service.

[0124] The deviation penalty portion is represented in the following form:

[0125] (9)

[0126] in, and The penalty coefficient is... This indicates a punishment for "inflated performance," and This indicates that "capacity breach" will be punished. This indicates the frequency modulation performance indicators submitted by the alliance; This indicates the actual frequency modulation capacity provided by the alliance during time period t.

[0127] The above steps complete the definition of the alliance bidding model under MDP.

[0128] 2. DRL-based alliance bidding decision optimization

[0129] Having completed the coalition bidding model under MDP in the previous step, we will now use DRL (deep reinforcement learning) to train the agent to solve the decision optimization problem.

[0130] Considering the application scenarios involved, the Soft Actor-Commentator (SAC) algorithm was chosen for training. SAC is an off-policy algorithm, which means that it can reuse old experience data for learning. This has higher sample utilization efficiency compared to the on-policy PPO (old data becomes invalid after each policy update). Given that historical data in the electricity market is abundant but simulation costs may be high, SAC can learn more fully from limited data and speed up the training process.

[0131] SAC typically contains three types of neural networks: policy networks (actors). Q-value network (commentator) and the corresponding target Q-value network (Used for stable training).

[0132] First, Equation (7) is used as the reward function of the agent. A simulation environment is constructed using historical market data to drive the agent to perform a large number of iterative learning processes. The training objective is to learn a policy. ,in The parameters of the neural network, That is to say, in a given state Give the corresponding strategy action The goal is to maximize the cumulative expected reward, thereby maximizing the alliance's total expected revenue. The specific form is as follows:

[0133] (10)

[0134] in, The objective function is represented by t; the iteration step is represented by t, and the maximum iteration step is represented by T. This represents the curve trajectory formed by all strategies; Indicates the discount factor; The reward function is expressed in the form of equation (7). same.

[0135] The above formula is the general form of the DRL network objective function. Next, the SAC algorithm will be used to complete the DRL training.

[0136] First, an experience replay pool needs to be constructed, which represents a huge data repository (usually a first-in, first-out queue) to store the experiences generated by the agent's interactions with the environment. Each experience is generated by extracting state and action vectors from historical data in the MDP process, specifically in the form of a tuple of the following form.

[0137] (11)

[0138] in, A Boolean flag indicating whether the round has ended (e.g., whether the daily bidding cycle has been completed).

[0139] If all experiences are stored in an experience replay pool and used as the dataset for training the agent, then:

[0140] (12)

[0141] Where M represents the total number of experiences.

[0142] Assuming the experience replay pool has a sufficiently large amount of data, a small batch of experiences (e.g., 256) is randomly sampled from it. For any one of these experiences, the objective Q-value is calculated as follows:

[0143] (13)

[0144] In the above formula, The soft-state value function of the target Q-network is defined as:

[0145] (14)

[0146] in, This represents the Q estimate of the target network; The parameters representing the policy network; This indicates adaptive adjustment of the temperature parameter, which can be adjusted by the following objective function:

[0147] (15)

[0148] in, This represents the target value of expected entropy, which is the "exploratory motivation" target set for the agent. Specifically, it specifies the average degree of randomness that the agent's strategy (i.e., decision-making method) should maintain during the learning process. Possible values ​​are:

[0149] (16)

[0150] In the above formula, This represents the number of dimensions in the action space.

[0151] After obtaining the target Q value from the small-batch experience using equation (13), the mean square error (MSE) between the current Q-network estimate and the target Q value is calculated. Then, the Q-network... The soft Bellman error loss function for parameter updates is:

[0152] (17)

[0153] in, This represents the estimated value of the Q-network; This represents the expectation of the soft state value function of the target Q-network.

[0154] Meanwhile, policy networks The parameter update method is as follows:

[0155] (18)

[0156] In addition, to ensure output action Within the price limit range, add a scaled activation function after the output layer of the policy network:

[0157] (19)

[0158] In the above formula, and Indicates the lower and upper limits of the price limit range; is the hyperbolic tangent function, and is the activation function; Indicates policy parameters The mean of the state vector under the given conditions; Indicates policy parameters Standard deviation of the state vector under the given conditions; It is noise sampled from a standard normal distribution.

[0159] By following the steps above, the agent can be trained to automatically provide the optimal strategy for alliance bidding.

[0160] III. Profit Distribution Calculation Engine Module: Intelligent Profit Distribution Based on Cooperative Game Theory

[0161] After winning a bid for FM service and reaping the revenue, the alliance must distribute it fairly among all members. Let the revenue distribution vector be denoted as... Where n represents the total number of alliance members, Let i represent the profit that the i-th member should receive. Then we have:

[0162] (20)

[0163] in, This represents the total revenue generated by the alliance's participation in FM service. The question then becomes: how to fairly and reasonably distribute this revenue among the participating members?

[0164] To solve this problem, we first define an arbitrary sub-alliance. (N represents the set of all members), assuming the previous DRL has been trained, the optimal bidding strategy is denoted as... .

[0165] Set the data to be associated with the sub-alliance Matching range, run Calculate the average expected return of sub-alliance C over the entire settlement period, i.e., the return of sub-alliance operating alone. This is its characteristic function, denoted as . .

[0166] To ensure the uniqueness and fairness of the optimal allocation scheme within the cooperative game framework, the following algorithm based on the nucleolus method is presented:

[0167] Step 1: For any allocation scheme Alliance with any non-empty true child The complaint value is defined as:

[0168] (twenty one)

[0169] This value represents the alliance. If I leave the major league N and go solo, how much more revenue can I earn compared to the current distribution scheme? The larger, the more it means The more dissatisfied one is with the allocation plan.

[0170] Step 2: Construct the sorted complaint value vector. That is, for each possible allocation scheme X, calculate the complaint values ​​of all non-empty sub-alliances, and sort them in descending order, denoted as:

[0171] (twenty two)

[0172] Step 3: Construct the following linear programming problem to minimize the first largest complaint value in the complaint value vector:

[0173] (twenty three)

[0174] Solving this linear programming problem LP will yield the result that makes The one with the lowest complaint value ranked first. and the corresponding allocation scheme However, the solution may not be unique at this point, so we proceed to step 4.

[0175] Step 4: In the fixed position Given this premise, construct the following LP to minimize the second largest complaint value in the complaint value vector:

[0176] (twenty four)

[0177] This LP guarantees Maximum complaint value not exceeding Under the premise of minimizing the second complaint value.

[0178] Step 5: Repeat steps 1-4 above until only one allocation scheme remains, i.e., a unique allocation scheme is determined. .

[0179] It can be observed that for a league with n participants, at most a certain number of solutions are needed. The computational cost is on the same order of magnitude as the Shapley value allocation method.

[0180] However, the direct goal of this method is to minimize the maximum complaint, which ensures that no sub-alliance (whether individual or group) has a strong incentive to leave the main alliance; this is crucial for maintaining a long-term, stable business alliance. Although the Shapley value is based on a rigorous axiomatic system, the allocation scheme it calculates may sometimes be outside the "core", that is, there may be a sub-alliance with a positive complaint value (e(C, x)>0), which means that the sub-alliance has an incentive to leave.

[0181] Furthermore, in the implementation system designed in this embodiment, the members of the alliance are heterogeneous (thermal power, energy storage, VPP), and their costs and contribution methods are very different; the "minimize maximum complaint" principle of the nucleus can better handle the fairness concerns brought about by this heterogeneity.

[0182] Example 2

[0183] One embodiment of the present invention provides a multi-agent collaborative game decision-making system, comprising:

[0184] The perception data acquisition module is configured to acquire the perception data required for decision-making, including market data, entity data, and external environment data.

[0185] The bidding strategy optimization module is configured to: construct a multi-entity cooperative alliance; based on perception data, under given bidding constraints, with the goal of maximizing the expected total revenue of the multi-entity cooperative alliance, use a deep reinforcement learning agent to find the best bidding strategy to participate in the FM market bidding.

[0186] The revenue distribution optimization module is configured as follows: After a multi-entity cooperative alliance wins a bid for frequency modulation services and obtains revenue, it determines the optimal revenue distribution scheme based on the value of each member in the multi-entity cooperative alliance and the revenue obtained by the alliance from winning the bid for frequency modulation services, combined with a cooperative game algorithm, and distributes the benefits to each member according to the optimal revenue distribution scheme.

[0187] Furthermore, the process of finding the optimal bidding strategy using a deep reinforcement learning agent specifically involves:

[0188] A coalition bidding model is defined using Markov decision processes;

[0189] Based on the alliance bidding model, the agent is trained through deep reinforcement learning. Under bidding constraints including FM capacity price limit, FM performance price limit and FM capacity limit, the optimal bidding strategy is sought with the goal of maximizing the expected total revenue constructed based on the winning bid revenue and deviation penalty.

[0190] The bidding strategy mentioned above involves the alliance offering FM capacity, FM performance, and FM capacity quotes for specific time periods.

[0191] Furthermore, the alliance bidding model is defined, including defining the state space, action space, and reward function;

[0192] The states in the state space This can be expressed as a formula:

[0193]

[0194] in, This represents the total frequency modulation capacity that the alliance can provide during time period t; This represents the overall actual performance index of the alliance in time period t. This represents the historical frequency-adjusted clearing price series over the past N time periods. This indicates the exogenous policy constraints corresponding to time period t. This is timestamp information;

[0195] Actions in the action space This can be expressed as a formula:

[0196]

[0197] in, This indicates the frequency modulation capacity quotation for the alliance during time period t; This indicates the alliance's frequency modulation performance quote for time period t; This represents the frequency modulation capacity of the alliance during time period t;

[0198] The reward function This can be expressed as a formula:

[0199]

[0200] in, This indicates the portion of the revenue the alliance receives from participating in frequency modulation. This indicates the penalty for deviation.

[0201] Furthermore, the agent is trained using deep reinforcement learning, employing a soft actor-commentator algorithm for training, with the training objective being to maximize the expected total revenue of the multi-agent cooperative alliance, expressed by the formula:

[0202]

[0203] in, The objective function is represented by t; the iteration step is represented by t, and the maximum iteration step is represented by T. This represents the curve trajectory formed by all strategies; Indicates the discount factor; Represents the reward function, This is the current state. This is the current action.

[0204] Furthermore, the optimal profit distribution scheme determined by combining cooperative game theory algorithms is used to assist in distributing alliance profits among alliance members, specifically as follows:

[0205] Initialize the current profit distribution plan;

[0206] Based on the trained agent, the individual earnings of each member are predicted to quantify the member's value. Then, based on this quantified value, the complaint value of each member under the current earnings distribution scheme is calculated. This constitutes the current profit distribution plan. Complaint value vector ;

[0207] To minimize the maximum complaint value in the complaint value vector With the goal of improving the current profit distribution plan... Perform iterative optimization until the condition for stopping the iteration is met.

[0208] Furthermore, the calculation of each member's complaint value under the current profit distribution scheme, whereby the complaint value represents how much more profit the sub-alliance could gain by acting alone compared to the current distribution scheme, is used to calculate the profit of the sub-alliance acting alone. Alliance with children Under the current allocation scheme Total revenue The difference is used to calculate, and the formula is as follows:

[0209]

[0210] in, For the Alliance Under the current allocation scheme The complaint value below, For the current profit distribution plan, Sub-alliances can be composed of several members, including special sub-alliances composed of a single member. To predict the profits of a sub-alliance acting alone, Indicates the current profit distribution plan The revenue to be allocated to the i-th member of the neutron alliance.

[0211] Example 3

[0212] One embodiment of the present invention provides a computer program product, including a computer program that, when executed by a processor, implements the aforementioned multi-agent collaborative game decision-making method.

[0213] Example 4

[0214] In one embodiment of the present invention, a non-transitory computer-readable storage medium is provided for storing computer instructions, which, when executed by a processor, implement the aforementioned multi-agent collaborative game decision-making method.

[0215] Example 5

[0216] One embodiment of the present invention provides an electronic device, including: a processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to enable the electronic device to implement the multi-agent collaborative game decision-making method.

[0217] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0218] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0219] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.

Claims

1. A multi-agent collaborative game decision-making method, characterized in that, include: Acquire the perceived data needed for decision-making; perceived data includes market data, entity data, and external environment data. Construct a multi-entity cooperative alliance. Based on perception data, under given bidding constraints, with the goal of maximizing the expected total revenue of the multi-entity cooperative alliance, use a deep reinforcement learning agent to find the best bidding strategy to participate in the FM market bidding. After a multi-entity cooperative alliance wins a bid for FM service and obtains revenue, the optimal revenue distribution scheme is determined based on the value of each member in the multi-entity cooperative alliance and the revenue obtained by the alliance from winning the bid for FM service, combined with a cooperative game algorithm, and the revenue is distributed to each member according to the optimal revenue distribution scheme. Specifically, the process of finding the optimal bidding strategy using a deep reinforcement learning agent involves: A coalition bidding model is defined using Markov decision processes; Based on the alliance bidding model, the agent is trained through deep reinforcement learning. Under bidding constraints including FM capacity price limit, FM performance price limit and FM capacity limit, the optimal bidding strategy is sought with the goal of maximizing the expected total revenue constructed based on the winning bid revenue and deviation penalty. The bidding strategy mentioned above refers to the alliance's pricing for frequency modulation capacity, frequency modulation performance, and frequency modulation capacity for specific time periods. Define the alliance bidding model, including defining the state space, action space, and reward function; The states in the state space This can be expressed as a formula: in, This represents the total frequency modulation capacity that the alliance can provide during time period t; This represents the alliance's overall actual performance index during time period t, obtained by weighted averaging of frequency modulation performance index and actual effective frequency modulation mileage. This represents the historical frequency-adjusted clearing price series over the past N time periods. This indicates the exogenous policy restrictions corresponding to time period t, consisting of the frequency modulation capacity price limit range and the frequency modulation performance price limit range. This is timestamp information; Actions in the action space This can be expressed as a formula: in, This indicates the frequency modulation capacity quotation for the alliance during time period t; This indicates the alliance's frequency modulation performance quote for time period t; This represents the frequency modulation capacity of the alliance during time period t; The reward function This can be expressed as a formula: in, This indicates the portion of the revenue the alliance receives from participating in frequency modulation. This indicates the penalty for deviation.

2. The multi-agent collaborative game decision-making method as described in claim 1, characterized in that, The process involves training the agent using deep reinforcement learning, employing a soft actor-commentator algorithm to maximize the expected total reward of the multi-agent cooperative alliance. This is expressed by the formula: in, The objective function is represented by t; the iteration step is represented by t, and the maximum iteration step is represented by T. This represents the curve trajectory formed by all strategies; Indicates the discount factor; Represents the reward function, This is the current state. This is the current action.

3. The multi-agent collaborative game decision-making method as described in claim 1, characterized in that, The optimal profit distribution scheme determined by combining cooperative game theory algorithms is used to assist in distributing alliance profits among alliance members, specifically as follows: Initialize the current profit distribution plan; Based on the trained agent, the individual earnings of each member are predicted to quantify the member's value. Then, based on this quantified value, the complaint value of each member under the current earnings distribution scheme is calculated. This constitutes the current profit distribution plan. Complaint value vector ; To minimize the maximum complaint value in the complaint value vector With the goal of improving the current profit distribution plan... Perform iterative optimization until the condition for stopping the iteration is met.

4. The multi-agent collaborative game decision-making method as described in claim 3, characterized in that, The calculation of each member's complaint value under the current profit distribution scheme indicates how much more profit the sub-alliance could gain by acting alone compared to the current distribution scheme. Alliance with children Under the current allocation scheme Total revenue The difference is used to calculate, and the formula is as follows: in, For the Alliance Under the current allocation scheme The complaint value below, For the current profit distribution plan, Sub-alliances can be composed of several members, including special sub-alliances composed of a single member. To predict the profits of a sub-alliance acting alone, Indicates the current profit distribution plan The revenue to be allocated to the i-th member of the neutron alliance.

5. A multi-agent collaborative game decision-making system, characterized in that, The multi-agent collaborative game decision-making method as described in any one of claims 1-4 includes: The perception data acquisition module is configured to acquire the perception data required for decision-making, including market data, entity data, and external environment data. The bidding strategy optimization module is configured to: construct a multi-entity cooperative alliance; based on perception data, under given bidding constraints, with the goal of maximizing the expected total revenue of the multi-entity cooperative alliance, use a deep reinforcement learning agent to find the best bidding strategy to participate in the FM market bidding. The revenue distribution optimization module is configured as follows: After a multi-entity cooperative alliance wins a bid for frequency modulation services and obtains revenue, it determines the optimal revenue distribution scheme based on the value of each member in the multi-entity cooperative alliance and the revenue obtained by the alliance from winning the bid for frequency modulation services, combined with a cooperative game algorithm, and distributes the benefits to each member according to the optimal revenue distribution scheme.

6. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements a multi-agent collaborative game decision-making method as described in any one of claims 1-4.

7. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium is used to store computer instructions, which, when executed by a processor, implement a multi-agent collaborative game decision-making method as described in any one of claims 1-4.

8. An electronic device, characterized in that, include: The device includes a processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to enable the electronic device to implement a multi-agent collaborative game decision-making method as described in any one of claims 1-4.