Large-scale power dispatching optimization method and system based on agent collaboration
By using an intelligent agent-based collaborative power dispatch optimization method, information on power generation, transmission, consumption, and the market is integrated to optimize unit output and line power flow. This solves the problem of coordinating market dynamics and the interests of multiple stakeholders in existing technologies, and achieves efficient and stable operation of the power system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-03-24
AI Technical Summary
Existing power dispatching methods are unable to combine dynamic market changes with the interests of multiple stakeholders, cannot accurately capture market price fluctuations and user electricity consumption behavior, and do not fully consider the dynamic changes in the real-time operating status of the power system, resulting in deviations in the execution of dispatching plans and making it difficult to ensure the safe and stable operation of large-scale power systems.
A large-scale power dispatch optimization method based on intelligent agent collaboration is adopted. Multi-dimensional parameters are collected through the intelligent decision-making platform of the energy internet to construct a dynamic market response strategy network. Combined with an asymmetric Nash bargaining model and a dynamic game deep reinforcement learning model, the unit output and line power flow are optimized to generate dispatch plans. The final plan is verified through simulation.
It has improved the operating efficiency and market adaptability of the power system, reduced operating costs, enhanced the ability to respond to market dynamics and user needs, and ensured the safe and stable operation of the power system.
Smart Images

Figure CN121727136A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power dispatch optimization technology, and in particular to a large-scale power dispatch optimization method and system based on agent collaboration. Background Technology
[0002] Current technological innovations in the energy sector are driving the development of power systems towards larger scale and greater complexity. On the generation side, this includes generating units with varying output characteristics; transmission networks need to cover wider geographical areas to achieve cross-regional power distribution; and on the consumption side, the proliferation of various electrical devices results in diverse load types and increased fluctuation frequencies. Simultaneously, the increasing marketization of the electricity market and the growing influence of price fluctuations and trading rules on dispatch decisions make traditional dispatching methods insufficient to balance the multifaceted requirements of cost control, safety assurance, and market demand response during system operation. In this context, it is necessary to leverage intelligent technologies to integrate various key information from generation, transmission, consumption, and market segments, promoting efficient cooperation among all participants to address dispatching challenges in complex scenarios and improve the overall operational quality and dispatching flexibility of the power system.
[0003] Existing technologies have two significant shortcomings in large-scale power dispatch optimization: First, existing dispatch methods fail to fully consider the dynamic changes in the market and the interests of different stakeholders when making decisions. They cannot accurately capture the impact of market price fluctuations and user electricity consumption behavior adjustments on dispatch schemes, nor can they effectively coordinate the conflicts of interest among stakeholders. As a result, the dispatch schemes formed are deficient in adapting to market changes and meeting the collaborative needs of multiple stakeholders. Second, in the process of dispatch-related calculations and optimizations, existing technologies do not comprehensively consider the dynamic changes of various parameters under the real-time operating state of the power system. When designing the analysis framework related to decision-making, they do not fully incorporate key constraints such as the unit output adjustment rate and the power transmission capacity of the lines. Furthermore, they lack a process for multi-dimensional effect verification and dynamic adjustment of the initially formed dispatch schemes. This makes it easy for the dispatch schemes to have execution deviations in practical applications, making it difficult to ensure the long-term safe and stable operation of large-scale power systems. Summary of the Invention
[0004] To overcome the shortcomings and deficiencies of existing technologies, this invention provides a method and system for large-scale power dispatch optimization based on agent collaboration.
[0005] The technical solution adopted in this invention is a large-scale power dispatch optimization method based on intelligent agent collaboration, comprising the following steps: S1, collecting different parameters under large-scale power dispatch scenarios through an energy internet intelligent decision-making platform, the parameters including the upper limit of generator output, the lower limit of generator output, and the generator ramp rate on the generation side; the line transmission capacity and line impedance on the transmission side; the load forecast and load type on the consumption side; and the electricity price fluctuation coefficient and transaction volume constraints on the market side; S2, constructing a dynamic market response strategy network, inputting the market-side parameters and consumption-side parameters collected in S1 into the network, and extracting features and assigning weights to the parameters through a multilayer perceptron and attention mechanism within the network to obtain the market response coefficients of different electricity consumption entities; S3, based on the market response coefficients obtained in S2, and combining the generation-side and transmission-side parameters, establishing an asymmetric Nash bargaining model, with the revenue function and bargaining weight of different agents as the core, to determine the initial bargaining scheme among multiple agents; S4, using the initial bargaining scheme determined in S3 as input, introducing dynamic... The game-theoretic deep reinforcement learning model sets the agent's state space as the real-time operating parameters of the power system, the action space as the unit output adjustment and line power flow allocation, and the reward function as the weighted sum of the power system operating cost reduction and power supply reliability improvement. Through iterative training, the agent's policy network and value network are updated to obtain an optimized game strategy. In step S5, based on the optimized game strategy obtained in step S4, different parameters for large-scale power dispatch are adjusted to generate dispatch plans including unit start-up and shutdown plans, unit output allocation schemes, and line power flow control schemes. In step S6, the dispatch plans generated in step S5 are fed back to the energy internet intelligent decision-making platform. The platform's simulation module simulates the operation of the dispatch plans, outputting the corresponding power system network loss rate, load fulfillment rate, and unit coal consumption rate. If all indicators meet the preset thresholds, the dispatch plan is adopted as the final large-scale power dispatch optimization scheme; otherwise, the process returns to step S2 to readjust the parameters of the dynamic market response strategy network and repeats steps S3, S4, S5, and S6.
[0006] Furthermore, the expression for the dynamic market response strategy network is: middle, Let be the market response coefficient of the i-th electricity consumer at time t. It is the Sigmoid activation function. , , These are the weight matrices for the first, second, and third layers of the network, respectively. These are the bias vectors for the first, second, and third layers of the network, respectively. Let be the feature vector of the i-th electricity consumer at time t, including the load forecast, electricity price fluctuation coefficient, and traded electricity volume constraint. Let be the response sensitivity coefficient of the i-th electricity consumer at time t. Let be the market participation coefficient of the i-th electricity consumer at time t. Let be the association weight between the i-th electricity consumer and the k-th market node. Let be the transaction volume of the k-th market node at time t. This represents the total number of market nodes.
[0007] Furthermore, the expression for the asymmetric Nash bargaining model is: , ,in, This is a vector of negotiation variables, including unit output adjustments and load allocation. Let this be the feasible region of the bargaining variable. Let j be the payoff function of the j-th agent. For the retention benefit of the j-th agent, Let the bargaining weight of the j-th agent satisfy the following conditions: The total number of agents. These are the lower and upper limits of the output of the j-th unit, respectively. The j-th unit at time t and Constant effort Let be the ramp rate of the j-th unit. The transmission power of the I-th line. These are the lower and upper limits of the transmission power for line I, respectively.
[0008] Furthermore, the policy network expression of the dynamic game deep reinforcement learning model is as follows: ,in, For parameters The policy network, Let t be the action of the agent. Let t be the state of the agent at time t. The output function of the policy network includes the state. Feature mapping and action Matching calculation, where A is the action space; the value network expression is: ,in, For parameters The value network Let be the weight coefficient of the k-th value calculation unit, p be the total number of value calculation units, and ReLU be the linear rectified activation function. This is the weight matrix for the k-th value calculation unit. This is the bias vector for the k-th value calculation unit.
[0009] Furthermore, the comprehensive index calculation model used in the energy internet intelligent decision-making platform for evaluating dispatch plans is as follows: ,in, For comprehensive evaluation indicators, , The weighting coefficients for network loss rate, load fulfillment rate, and unit coal consumption rate are respectively, and satisfy the following conditions: , For power system network loss rate, For load sufficiency, This refers to the unit's coal consumption rate; simultaneously, the real-time performance guarantee model for parameter acquisition in the platform is as follows: ,in, q represents the total time spent on parameter acquisition, and q represents the total number of parameter acquisition nodes. Let be the time taken for a single data collection at the q-th data collection node. Let be the sampling frequency coefficient of the q-th sampling node.
[0010] Furthermore, the accuracy calibration model for unit output allocation in the large-scale power dispatch optimization scheme is as follows: ,in, For the output calibration value of the j-th unit, Let be the real-time calibration coefficient for the j-th unit. These are the optimized output and actual output of the j-th unit, respectively. Let T be the historical deviation correction factor for the j-th unit, and T be the historical data statistical duration. Let be the historical optimized output and historical actual output of the j-th generating unit at time t, respectively; the stability adjustment model for line power flow control is: ,in, This represents the power flow regulation amount for the I-th line. Let be the power flow margin coefficient for the I-th line. This represents the upper limit of the transmission capacity of the I-th line. This represents the actual power flow of line I. This is the power flow rate adjustment coefficient for the I-th line. Let be the power flow rate of the I-th line.
[0011] Further, S3 includes the following sub-steps: S31, extracting the market response coefficients corresponding to each agent from the output results of S2, and constructing a revenue function for each agent by combining the output cost function of the generating units and the loss function of the transmission lines. The revenue function incorporates the influence factor of the market response coefficient on the revenue, which is calculated by multiplying the market response coefficient by the transaction scale of the corresponding agent; S32, determining the bargaining weight of each agent based on its role and influence in the power system. The bargaining weight of the generating agents refers to the generating unit capacity and the contribution to power supply reliability, while the bargaining weight of the transmission agents refers to the transmission capacity of the transmission lines. The importance of quantity and network connectivity is considered, and the bargaining weight of the electricity-consuming agents is referenced to the load scale and load priority; S33, the constraints of the asymmetric Nash bargaining model are set, including the upper and lower limits of the output of the generating units, the ramp rate constraint, the transmission power constraint of the transmission lines, the load demand constraint of the electricity-consuming side, and the electricity quantity and price constraint in the market transaction. Each constraint is transformed into a mathematical inequality and incorporated into the model; S34, the asymmetric Nash bargaining model is solved using the Lagrange multiplier method. By constructing the Lagrange function, the optimal solution of the bargaining variables that satisfies all constraints is obtained by solving the extreme value of the function, and then the initial bargaining scheme among multiple agents is determined.
[0012] Further, S4 includes the following sub-steps: S41, decompose the initial negotiation scheme obtained in S3 into a set of initial actions for each intelligent agent, and construct the state space of the intelligent agents by combining the real-time operating data of the power system, such as unit output, line power flow, and load. Each state vector in the state space includes fused information of real-time operating parameters and initial action parameters; S42, define the action space of the intelligent agents. The action space of the generation-side intelligent agents includes the adjustment range and direction of unit output; the action space of the transmission-side intelligent agents includes the distribution ratio and control method of line power flow; the action space of the consumption-side intelligent agents includes the load transfer period and transfer amount. The boundary of the action space is determined by the safety operation constraints of the power system; S43, set... The reward function is calculated, with positive reward items including the reduction in power system operating costs, the improvement in power supply reliability, and the increase in market transaction revenue, and negative reward items including penalties for unit over-limit operation, penalties for line power flow exceeding limits, and penalties for load deficit. The final reward function is obtained by weighted summation. In S44, the dynamic game deep reinforcement learning model is trained using experience replay and target network technology. The state-action-reward-next state data generated by the interaction between the agent and the environment are stored in the experience pool. The data in the experience pool is randomly sampled to update the policy network and value network. The stability of model training is improved by delaying the update of the target network. The model is iteratively trained until the reward function value converges, and the optimized game strategy is obtained.
[0013] Further, S5 includes the following sub-steps: S51, parsing the optimized game strategy output by S4, extracting the optimal actions of each agent, converting the optimal actions of the generation-side agents into unit start-up and shutdown commands and output setpoints, converting the optimal actions of the transmission-side agents into line power flow control commands and transmission power allocation values, and converting the optimal actions of the consumption-side agents into load adjustment plans and electricity consumption time arrangements; S52, combining the power system parameters collected in S1, performing feasibility verification on the converted commands and values, including whether the unit start-up and shutdown commands comply with the minimum start-up and shutdown time constraints, whether the output setpoints are within the upper and lower limits of unit output, and whether the line power flow control commands comply with the minimum start-up and shutdown time constraints. S53: Integrate the verified instructions and values to generate a unit start-up and shutdown plan, which specifies the start-up and shutdown times and continuous operating duration of each unit; generate a unit output allocation scheme, which specifies the output value of each unit at different time periods; and generate a line power flow control scheme, which specifies the power flow control value of each line at different time periods. S54: Combine the generated unit start-up and shutdown plan, unit output allocation scheme, and line power flow control scheme to form a complete dispatch plan. The dispatch plan includes the execution time period, execution subject, and monitoring indicators of each scheme. The monitoring indicators are used to track the operation status of the dispatch plan.
[0014] A large-scale power dispatch optimization system based on intelligent agent collaboration includes: a multi-dimensional power parameter acquisition and integration unit, which is connected to the perception layer of the energy internet intelligent decision-making platform, receives various parameters from the generation side, transmission side, consumption side, and market side, and integrates parameters from different sources into a unified parameter set through data cleaning and format conversion. The output of this unit is connected to a dynamic market response strategy network construction unit; and a dynamic market response strategy network construction and calculation unit, which receives the parameter set output by the multi-dimensional power parameter acquisition and integration unit, and constructs a dynamic market response strategy network based on a preset network structure and parameter initialization rules. The market response strategy network performs feature extraction and weight allocation calculations on the parameter set, outputting the market response coefficients of each electricity consumer. The output of this unit is connected to the asymmetric Nash bargaining model establishment unit. The asymmetric Nash bargaining model establishment and solution unit receives the market response coefficients output by the dynamic market response strategy network construction and calculation unit, combines them with the generation-side and transmission-side parameters output by the multi-dimensional power parameter acquisition and integration unit, constructs the asymmetric Nash bargaining model, sets model constraints and objective functions, and uses an optimization algorithm to solve the model to obtain the initial bargaining scheme. The output of this unit is connected to the dynamic game deep reinforcement learning model. The system includes a training unit; a dynamic game deep reinforcement learning model training and optimization unit, which receives the initial bargaining scheme output by the asymmetric Nash bargaining model establishment and solution unit, constructs the agent's state space, action space, and reward function, initializes the policy network and value network parameters, iteratively trains the model through experience replay and target network techniques, and outputs the optimized game strategy. The output of this unit is connected to the scheduling plan generation unit; and a scheduling plan generation and verification unit, which receives the optimized game strategy output by the dynamic game deep reinforcement learning model training and optimization unit, and transforms it into unit start-up and shutdown plans and unit output allocation. The system includes a power distribution scheme and a line power flow control scheme. Feasibility checks are performed on each scheme, and schemes that pass the combined checks are used to generate a dispatch plan. The output of this unit is connected to the dispatch plan simulation evaluation and determination unit. The dispatch plan simulation evaluation and determination unit receives the dispatch plan output from the dispatch plan generation and verification unit, calls the simulation module of the energy internet intelligent decision-making platform to simulate the operation of the dispatch plan, calculates the power system network loss rate, load fulfillment rate, and unit coal consumption rate, compares them with preset thresholds to determine whether the requirements are met, and outputs the final large-scale power dispatch optimization scheme. The output of this unit is connected to the external power dispatch execution system.
[0015] Beneficial Effects: This invention proposes a large-scale power dispatch optimization method and system based on intelligent agent collaboration. This method and system effectively integrate multi-dimensional parameters from the generation, transmission, consumption, and market sides. Through multi-stage collaboration, it achieves efficient optimization of large-scale power dispatch, not only improving system operating efficiency and reducing operating costs, but also enhancing responsiveness to market dynamics and user demands, ensuring the safe and stable operation of the power system. Addressing the difficulty of existing dispatch methods in combining market dynamics and the demands of multiple stakeholders, this system, through multi-unit collaboration, fully considers market price fluctuations, changes in user electricity consumption behavior, and the interests of various participating entities, coordinating the interests between different stakeholders to make the generated dispatch scheme more aligned with market realities and the collaborative needs of multiple stakeholders, thus improving market adaptability. Addressing the shortcomings of existing technologies in comprehensively considering real-time dynamic changes in system parameters and lacking multi-indicator verification and adjustment mechanisms, this method fully integrates key constraints such as unit output adjustment rates and line transmission capacity into the decision-making process. Furthermore, it verifies the dispatch plan through simulation evaluation, iteratively adjusting it in a timely manner when requirements are not met, reducing actual execution deviations and ensuring the long-term stable operation of the large-scale power system. Attached Figure Description
[0016] Figure 1 This is a flowchart of the method steps of the present invention; Figure 2 This is a diagram showing the system unit composition of the present invention. Detailed Implementation
[0017] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0018] like Figure 1 As shown, the large-scale power dispatch optimization method based on agent cooperation includes the following steps: S1, collects different parameters under large-scale power dispatch scenarios through the energy internet intelligent decision-making platform. The parameters include the upper limit of unit output, the lower limit of unit output, and the unit ramp rate on the generation side; the line transmission capacity and line impedance on the transmission side; the load forecast and load type on the power consumption side; and the electricity price fluctuation coefficient and trading volume constraints on the market side. Specifically, the implementation process of step S1 is as follows: Through the perception layer module of the energy internet intelligent decision-making platform, comprehensive collection of multi-dimensional parameters under large-scale power dispatch scenarios is conducted. Specifically, the generation-side parameters include setting the upper limit of thermal power unit output to 1000 MW to 1500 MW and the lower limit to 300 MW to 500 MW; the upper limit of wind turbine unit output is set to 150 MW to 300 MW based on installed capacity, and the lower limit is 0 MW; the ramp rate of all units is controlled at 100 MW / hour to 200 MW / hour. The transmission-side parameters include setting the 220 kV line transmission capacity to 200 MW to 400 MW and the 500 kV line transmission capacity to 200 MW to 400 MW. The transmission capacity of the power lines is set at 800 MW to 1200 MW, and the line impedance ranges from 0.01 ohms / km to 0.05 ohms / km depending on the conductor type. Electricity-side parameters include residential load forecasts divided by region, with each region receiving 50 MW to 200 MW, and industrial load forecasts ranging from 300 MW to 800 MW per plant area, with the load type indicated as adjustable or non-adjustable. Market-side parameters include day-ahead price fluctuation coefficients ranging from 0.8 to 1.5, real-time price fluctuation coefficients ranging from 0.5 to 2.0, and transaction volume constraints set at 100 MWh to 5000 MWh depending on the trading entity. This step ensures comprehensive and accurate basic data support for subsequent dispatch decisions through multi-source data collection, providing data assurance for subsequent model construction and strategy optimization.
[0019] S2, construct a dynamic market response strategy network, input the market-side parameters and electricity-side parameters collected in S1 into the network, and extract features and assign weights to the parameters through the multilayer perceptron and attention mechanism inside the network to obtain the market response coefficients of different electricity-consuming entities; Specifically, the implementation process of step S2 is as follows: Based on the parameters collected in step S1, a dynamic market response strategy network is constructed in the algorithm module of the energy internet intelligent decision-making platform. The network structure adopts a three-layer perceptron architecture. The number of input layer nodes is determined to be 12 to 18 based on the parameter dimensions of the market side and the electricity consumption side. The number of hidden layer nodes is set to 32 to 64. The number of output layer nodes is consistent with the number of electricity consumption entities, usually 20 to 50. The parameters collected in step S1, such as the market-side electricity price fluctuation coefficient, transaction electricity constraints, and electricity consumption side load forecast values and load types, are converted into network output parameters. The input vector is transmitted to the hidden layer through the input layer. The hidden layer uses the ReLU activation function to extract nonlinear features from the parameters. An attention mechanism is also introduced, assigning weights of 0.6 to 0.8 to the load forecast and electricity price fluctuation coefficient, and weights of 0.2 to 0.4 to the traded electricity constraints and load type. The feature vector processed by the hidden layer is then transmitted to the output layer. The output layer uses the Sigmoid activation function to map the results to the range of 0 to 1, obtaining the market response coefficient for each electricity consumer. The higher the coefficient value, the stronger the sensitivity of the electricity consumer to market price changes. This step achieves in-depth processing of market and electricity consumption parameters through the network model, providing key coefficient support for subsequent multi-agent bargaining and improving the adaptability of the scheduling scheme to market changes.
[0020] S3, based on the market response coefficient obtained from S2, combined with the parameters of the power generation side and the transmission side, establishes an asymmetric Nash bargaining model, with the revenue function and bargaining weight of different agents as the core, to determine the initial bargaining scheme among multiple agents; Specifically, the implementation process of step S3 is as follows: Based on the market response coefficients of each electricity consumer obtained in step S2, and combined with the parameters such as the upper and lower limits of the power output of the generating units, the ramp rate, and the transmission capacity and impedance of the transmission lines collected in step S1, an asymmetric Nash bargaining model is established in the game decision module of the platform. When constructing the model, the market response coefficients are first multiplied by the load scale of the electricity consumer to obtain the revenue impact factor of each electricity consumer. Then, the unit output cost of the generating units (300 yuan / MWh to 500 yuan / MWh for thermal power units and 0 yuan / MWh for wind power units) and the network loss cost of the transmission lines (valued at 5 yuan / MWh to 15 yuan / MWh) are combined to construct the intelligent agents of the generating side, transmission side, and electricity consumer side. The payoff function is defined; subsequently, bargaining weights are assigned based on the importance of each agent in the system. The weights for generation-side agents are allocated according to their installed capacity percentage, with a total weight percentage of 0.4 to 0.5; the weights for transmission-side agents are allocated according to their line transmission capacity percentage, with a total weight percentage of 0.2 to 0.3; and the weights for consumption-side agents are allocated according to their load size percentage, with a total weight percentage of 0.2 to 0.3. Simultaneously, model constraints are set, including that generator output does not exceed upper or lower limits, output variation does not exceed the ramp rate, and line transmission power does not exceed transmission capacity. The model is solved using a gradient descent algorithm to obtain the initial bargaining scheme among the agents. The scheme specifies the initial output allocation values for each generator, the initial power flow allocation values for each line, and the initial load allocation values for each consumption entity. This step achieves multi-agent interest coordination by establishing a bargaining model, providing an initial scheme for subsequent game strategy optimization and ensuring a balance of interests among the agents.
[0021] S4 takes the initial bargaining scheme determined in S3 as input, introduces a dynamic game deep reinforcement learning model, sets the state space of the agent as the real-time operating parameters of the power system, the action space as the unit output adjustment and line power flow allocation, and the reward function as the weighted sum of the power system operating cost reduction and power supply reliability improvement. The agent's policy network and value network are updated through iterative training of the model to obtain the optimized game strategy. Specifically, the implementation process of step S4 is as follows: The unit output allocation value, line power flow allocation value, and load allocation value in the initial negotiation scheme determined in step S3 are used as initial inputs. A dynamic game deep reinforcement learning model is introduced into the platform's reinforcement learning module. During model initialization, the state space of the agent is set to include real-time unit output (the value range is consistent with the upper and lower limits of output), real-time line power flow (the value range is consistent with the transmission capacity), real-time load consumption (the value range is consistent with the load forecast value), and real-time market electricity price (the value range corresponds to the electricity price fluctuation coefficient), totaling 30 to 50 state variables. The action space is set as unit output adjustment (value range is -50 MW to 50 MW), line power flow adjustment (value range is -30 MW to 30 MW), and load transfer (value range is -20 MW to 20 MW). The reward function is designed as a weighted average of the system operating cost reduction value and the power supply reliability improvement value. The system assigns a reward of 10 to 20 points to each reduction in operating costs (10,000 yuan) and a reward of 5 to 10 points to each increase in load satisfaction rate (1%). Penalties of -50 to -100 points are assigned to generator overload, line overload, and load deficit. The model training employs an experience replay mechanism, storing the state-action-reward-next state data generated by the agent's interaction with the environment in an experience pool of 10,000 to 50,000 records. During each training iteration, 32 to 64 data points are randomly sampled to update the policy network and value network. The learning rate for the policy network is set to 0.001 to 0.005, and the learning rate for the value network is set to 0.0001 to 0.001. After 1,000 to 5,000 rounds of iterative training, training stops and the optimized game strategy is output when the reward function value fluctuates by no more than 5% for 50 consecutive rounds. This strategy clarifies the optimal action selection rules for each agent under different states. This step uses a reinforcement learning model to dynamically optimize the strategy, improve the scheduling scheme's adaptability to real-time system changes, and ensure the optimality of scheduling decisions.
[0022] S5, based on the optimized game strategy obtained in S4, adjusts different parameters of large-scale power dispatch to generate a dispatch plan including unit start-up and shutdown plan, unit output allocation scheme and line power flow control scheme. Specifically, the implementation process of step S5 is as follows: Based on the optimized game strategy obtained in step S4, the parameters of large-scale power dispatch are systematically adjusted in the platform's plan generation module; firstly, according to the optimal action rules of the units in the strategy, combined with the upper and lower limits of unit output and ramp rate in step S1, a unit start-up and shutdown plan is generated, specifying the start-up and shutdown times (accurate to the hour) and start-up and shutdown sequence (sorted from low to high by unit coal consumption) of thermal power units, and the operating period of wind turbine units is determined according to the predicted wind speed (ranging from 3 m / s to 25 m / s); then, a unit output allocation scheme is generated, and the hourly load demand is allocated to each operating unit according to the load demand and unit cost characteristics of each time period, ensuring... The policy ensures that the output of each generating unit remains within its upper and lower limits and that output fluctuations do not exceed the ramp rate. For example, if the total load demand during a certain period is 1500 MWh, 500 MWh will be allocated to thermal power units with a unit cost of 300 yuan / MWh, 800 MWh to thermal power units with a unit cost of 350 yuan / MWh, and 200 MWh to wind power units. Simultaneously, a power flow control scheme is generated. Based on the transmission capacity of each line and the unit output allocation results, the power flow value for each line is calculated. By adjusting the transformer tap changers (adjustment range ±5%) or the control quantity of the flexible DC converter valve (control accuracy 0.1 MW), the power flow of the lines is controlled between 80% and 90% of the transmission capacity, avoiding power flow exceeding limits. This step, through parameter adjustment and scheme generation, transforms the optimization strategy into an executable scheduling plan, providing concrete scheme support for subsequent simulation verification.
[0023] S6 feeds the dispatch plan generated in S5 back to the Energy Internet Intelligent Decision Platform. The platform's simulation module simulates the operation of the dispatch plan and outputs the power system network loss rate, load fulfillment rate, and unit coal consumption rate corresponding to the dispatch plan. If all indicators meet the preset thresholds, the dispatch plan is taken as the final large-scale power dispatch optimization scheme. If not, it returns to S2 to readjust the parameters of the dynamic market response strategy network and repeats S3, S4, S5, and S6.
[0024] Specifically, the implementation process of step S6 is as follows: the dispatch plan, consisting of the unit start-up and shutdown plan, unit output allocation scheme, and line power flow control scheme generated in step S5, is fed back to the simulation evaluation module of the energy internet intelligent decision-making platform through a data interface; the simulation module uses digital twin technology to construct a power system simulation model, which includes component parameters such as units, lines, and loads consistent with the actual system. The simulation duration is set to 24 to 72 hours, and the time step is set to 15 to 60 minutes; during the simulation, the system network loss rate (calculation accuracy of 0.1%), load satisfaction rate (calculation accuracy of 0.1%), and unit coal consumption rate (calculation accuracy of 0.1 g / kWh) are calculated in real time at each time step. The system network loss rate is calculated through line impedance and power flow value, and the load satisfaction rate is calculated through actual power supply and power flow value. The ratio of load demand is calculated, and the unit coal consumption rate is calculated using the pre-fitted curves of unit output and coal consumption. After simulation, the average value of the above three indicators is output. The preset thresholds are set as follows: system network loss rate not exceeding 5%, load fulfillment rate not less than 99%, and unit coal consumption rate not exceeding 300 g / kWh. If all three indicators meet the threshold requirements, the dispatch plan is marked as the final large-scale power dispatch optimization scheme and transmitted to the dispatch execution system. If any indicator does not meet the threshold, for example, if the load fulfillment rate is 98.5% lower than the threshold, the backtracking mechanism is automatically triggered, returning to step S2, adjusting the attention weight of the dynamic market response strategy network (increasing the load forecast weight to 0.8 to 0.9), and re-executing steps S2 to S6 until a dispatch scheme that meets the threshold requirements is generated. This step, through simulation verification and iterative adjustment, ensures the feasibility and optimality of the dispatch scheme, guaranteeing the safe and stable operation of the large-scale power system.
[0025] Preferably, the expression for the dynamic market response strategy network is: middle, Let be the market response coefficient of the i-th electricity consumer at time t. It is the Sigmoid activation function. , , These are the weight matrices for the first, second, and third layers of the network, respectively. These are the bias vectors for the first, second, and third layers of the network, respectively. Let be the feature vector of the i-th electricity consumer at time t, including the load forecast, electricity price fluctuation coefficient, and traded electricity volume constraint. Let be the response sensitivity coefficient of the i-th electricity consumer at time t. Let be the market participation coefficient of the i-th electricity consumer at time t. Let be the association weight between the i-th electricity consumer and the k-th market node. Let be the transaction volume of the k-th market node at time t. This represents the total number of market nodes.
[0026] Specifically, based on the market-side and electricity-side parameters collected in step S1, a dynamic market response strategy network is constructed and the market response coefficient of each electricity consumer is calculated. This network adopts a three-layer computing architecture. The number of computing units in the first layer is determined according to the parameter dimensions, ranging from 12 to 18; the number in the second layer is 32 to 64; and the number in the third layer is matched with the number of electricity consumers, ranging from 20 to 50. During the calculation, the electricity price fluctuation coefficient (0.8 to 1.5 day-ahead, 0.5 to 2.0 real-time), transaction volume constraints (100 MWh to 5000 MWh), and load forecasts (50 MW to 200 MW / area for residential, 300 MW to 800 MW / plant for industrial), load type, and other parameters on the market side are first converted into input data in a unified format and transmitted to the first-level calculation unit. Then, the second-level calculation unit uses a specific nonlinear processing method to extract parameter features, while assigning a calculation weight of 0.6 to 0.8 to the load forecast and electricity price fluctuation coefficient, and a calculation weight of 0.2 to 0.4 to the transaction volume constraints and load type, to strengthen the influence of key parameters on the results. Finally, the third-level calculation unit maps the processing results to a numerical range of 0 to 1 to obtain the market response coefficient for each electricity user. By clarifying the network operation structure and parameter weights, the accuracy of market response coefficient calculation is improved, providing more reliable basic data for subsequent multi-agent bargaining, and making the dispatching scheme more suitable for the different sensitivities of different electricity users to market changes.
[0027] Preferably, the expression for the asymmetric Nash bargaining model is: , ,in, This is a vector of negotiation variables, including unit output adjustments and load allocation. Let this be the feasible region of the bargaining variable. Let j be the payoff function of the j-th agent. For the retention benefit of the j-th agent, Let the bargaining weight of the j-th agent satisfy the following conditions: The total number of agents. These are the lower and upper limits of the output of the j-th unit, respectively. The j-th unit at time t and Constant effort Let be the ramp rate of the j-th unit. The transmission power of the I-th line. These are the lower and upper limits of the transmission power for line I, respectively.
[0028] Specifically, using the market response coefficient obtained in step S2 as the core, and combining it with the generation and transmission side parameters from step S1, an asymmetric Nash bargaining model is constructed and the initial bargaining scheme is solved. During model construction, the market response coefficient of each electricity consumer is first multiplied by its corresponding load scale (50 MW to 200 MW for residential, 300 MW to 800 MW for industrial) to obtain the key factors affecting the revenue of each agent. Then, combining the unit output cost of the generation side units (300 RMB / MWh to 500 RMB / MWh for thermal power units, 0 RMB / MWh for wind power units) and the network loss cost of the transmission line (5 RMB / MWh to 15 RMB / MWh), revenue calculation rules for each agent on the generation, transmission, and consumption sides are established. Subsequently, bargaining weights were assigned based on the roles of each agent in the system. Agents on the power generation side were allocated weights according to their installed capacity, with a total weight of 0.4 to 0.5. Agents on the transmission side were allocated weights according to their line transmission capacity (220 kV 200 MW to 400 MW, 500 kV 800 MW to 1200 MW), with a total weight of 0.2 to 0.3. Agents on the power consumption side were allocated weights according to their load size, with a total weight of 0.2 to 0.3. Constraints were also set, including that generator output did not exceed upper or lower limits (300 MW to 1500 MW for thermal power units, 0 MW to 300 MW for wind power units), output variation did not exceed the ramp rate (100 MW / h to 200 MW / h), and line transmission power did not exceed transmission capacity. The model was solved using a gradient descent calculation method to obtain the initial output allocation values for each generator unit, the initial power flow allocation values for the lines, and the initial load allocation values for the main power consumption entities. By clearly defining the elements of model construction and the solution method, we can ensure the rationality of the coordination of interests among multiple agents and provide a scientific initial plan for subsequent game strategy optimization.
[0029] Preferably, the policy network expression of the dynamic game deep reinforcement learning model is: ,in, For parameters The policy network, Let t be the action of the agent. Let t be the state of the agent at time t. The output function of the policy network includes the state. Feature mapping and action Matching calculation, where A is the action space; the value network expression is: ,in, For parameters The value network Let be the weight coefficient of the k-th value calculation unit, p be the total number of value calculation units, and ReLU be the linear rectified activation function. This is the weight matrix for the k-th value calculation unit. This is the bias vector for the k-th value calculation unit.
[0030] Specifically, using the initial negotiation scheme from step S3 as input, a dynamic game deep reinforcement learning model is constructed and trained to optimize the game strategy. During model initialization, the agent's state space is set to include real-time power output of generating units (300 MW to 1500 MW for thermal power units, 0 MW to 300 MW for wind power units), real-time power flow of transmission lines (200 MW to 400 MW for 220 kV, 800 MW to 1200 MW for 500 kV), real-time load consumption (50 MW to 200 MW for residential, 300 MW to 800 MW for industrial), and real-time market electricity price (0.5 to 2.0 times the benchmark price), totaling 30 to 50 state dimensions; actions... The space is set as the unit output adjustment amount (-50 MW to 50 MW), the line power flow adjustment amount (-30 MW to 30 MW), and the load transfer amount (-20 MW to 20 MW). The reward calculation rule is designed as a weighted sum of the reduction in system operating costs and the improvement in power supply reliability. For every 10,000 yuan reduction in operating costs, a reward score of 10 to 20 is given. For every 1% increase in load satisfaction rate, a reward score of 5 to 10 is given. Unit over-limit, line overload, and load deficit are respectively penalized with a penalty score of -50 to -100. During model training, the states, actions, rewards, and next states generated by the agent's interaction with the system are stored in a database with a capacity of 10,000 to 50,000 records. In each training iteration, 32 to 64 records are randomly selected to update the policy calculation rules and value calculation rules. The learning rate for the policy calculation rules is set to 0.001 to 0.005, and the learning rate for the value calculation rules is set to 0.0001 to 0.001. Training is iterated for 1000 to 5000 rounds. Training stops when the reward score fluctuates by no more than 5% for 50 consecutive rounds, and the optimized game strategy is output. By clearly defining the model training parameters and process, the adaptability of the game strategy to dynamic changes in the system is improved, ensuring the optimality of scheduling decisions.
[0031] Preferably, the comprehensive index calculation model used in the energy internet intelligent decision-making platform for evaluating dispatch plans is as follows: ,in, For comprehensive evaluation indicators, , The weighting coefficients for network loss rate, load fulfillment rate, and unit coal consumption rate are respectively, and satisfy the following conditions: , For power system network loss rate, For load sufficiency, This refers to the unit's coal consumption rate; simultaneously, the real-time performance guarantee model for parameter acquisition in the platform is as follows: ,in, q represents the total time spent on parameter acquisition, and q represents the total number of parameter acquisition nodes. Let be the time taken for a single data collection at the q-th data collection node. Let be the sampling frequency coefficient of the q-th sampling node.
[0032] Specifically, comprehensive indicator calculation rules and real-time parameter acquisition calculation rules are constructed for evaluating dispatch plans. When calculating the comprehensive indicator, system network loss rate, load fulfillment rate, and unit coal consumption rate are used as core evaluation items. The network loss rate is assigned a weight of 0.3 to 0.4, the load fulfillment rate a weight of 0.4 to 0.5, and the unit coal consumption rate a weight of 0.2 to 0.3, with a total weight of 1. The comprehensive evaluation indicator is obtained through weighted summation. The calculation accuracy of the network loss rate is 0.1%, the load fulfillment rate is 0.1%, and the unit coal consumption rate is 0.1 g / kWh. Lower indicator values indicate better overall performance of the dispatch plan. When calculating the real-time performance guarantee for parameter acquisition, the time taken for each acquisition (0.1 to 1 second) and the acquisition frequency coefficient (1 to 5, with higher values indicating higher acquisition frequency) of all parameter acquisition nodes (50 to 100 nodes in total, including generation, transmission, consumption, and market sides) are statistically analyzed. The time taken for each acquisition at each node is multiplied by the acquisition frequency coefficient, and then summed to obtain the total parameter acquisition time. The total time must be controlled within 5 to 30 seconds to ensure that the acquired data reflects the system's operating status in real time. By clarifying the calculation rules for comprehensive indicators and real-time performance guarantees, quantitative standards are provided for evaluating dispatch plans and assessing the effectiveness of parameter acquisition, thereby improving the reliability of dispatch schemes and the stability of system operation.
[0033] Preferably, the accuracy calibration model for unit output allocation in the large-scale power dispatch optimization scheme is as follows: ,in, For the output calibration value of the j-th unit, Let be the real-time calibration coefficient for the j-th unit. These are the optimized output and actual output of the j-th unit, respectively. Let T be the historical deviation correction factor for the j-th unit, and T be the historical data statistical duration. Let be the historical optimized output and historical actual output of the j-th generating unit at time t, respectively; the stability adjustment model for line power flow control is: ,in, This represents the power flow regulation amount for the I-th line. Let be the power flow margin coefficient for the I-th line. This represents the upper limit of the transmission capacity of the I-th line. This represents the actual power flow of line I. This is the power flow rate adjustment coefficient for the I-th line. Let be the power flow rate of the I-th line.
[0034] Specifically, the calculation rules for the accuracy calibration of unit output allocation and the calculation rules for the stability adjustment of line power flow control are constructed. During unit output accuracy calibration, the deviation between the optimized output value and the actual output value of each unit is first obtained. This deviation is then combined with the unit's real-time calibration coefficient (0.1 to 0.3, with higher values indicating greater calibration strength) to calculate the real-time calibration amount. Simultaneously, the deviation between the optimized output and the actual output of the unit is statistically analyzed hourly over a historical 12-24 hour period. This deviation is then combined with the historical deviation correction coefficient (0.05 to 0.2) to calculate the historical correction amount. The real-time calibration amount and the historical correction amount are added together to obtain the final unit output calibration amount, ensuring that the deviation between the actual and optimized output of the unit is controlled within 5 MW to 10 MW. When adjusting power flow stability on a power line, the difference between the upper limit of the line's transmission capacity and the actual power flow value (i.e., the power flow margin) is first calculated. This margin adjustment is then calculated using the line's power flow margin coefficient (0.2 to 0.4). Next, the rate of change of the power flow is calculated (from -50 MW to 50 MW per hour), and the rate of change adjustment is calculated using the power flow rate adjustment coefficient (0.1 to 0.3). The margin adjustment and the rate of change adjustment are then added together to obtain the total power flow adjustment, ensuring that power flow fluctuations are controlled within 10 to 20 MW, thus preventing sudden power flow changes from affecting system stability. By clarifying the calculation rules for accuracy calibration and stability adjustment, the accuracy of unit output and power flow control on the power line is improved, ensuring the safety and stability of large-scale power system operation.
[0035] Preferably, step S3 includes the following sub-steps: S31, extracting the market response coefficients corresponding to each agent from the output results of S2, and constructing a revenue function for each agent by combining the output cost function of the generating units and the loss function of the transmission lines. The revenue function incorporates the influence factor of the market response coefficient on the revenue, which is calculated by multiplying the market response coefficient by the transaction scale of the corresponding agent; S32, determining the bargaining weight of each agent based on its role and influence in the power system. The bargaining weight of the generating agents refers to the generating unit capacity and the contribution to power supply reliability, and the bargaining weight of the transmission agents refers to the transmission capacity of the transmission lines. The importance of quantity and network connectivity is considered, and the bargaining weight of the electricity-consuming agents is referenced to the load scale and load priority; S33, the constraints of the asymmetric Nash bargaining model are set, including the upper and lower limits of the output of the generating units, the ramp rate constraint, the transmission power constraint of the transmission lines, the load demand constraint of the electricity-consuming side, and the electricity quantity and price constraint in the market transaction. Each constraint is transformed into a mathematical inequality and incorporated into the model; S34, the asymmetric Nash bargaining model is solved using the Lagrange multiplier method. By constructing the Lagrange function, the optimal solution of the bargaining variables that satisfies all constraints is obtained by solving the extreme value of the function, and then the initial bargaining scheme among multiple agents is determined.
[0036] Specifically, step S3 includes four sub-steps, S31 to S34: In S31, the market response coefficient of each agent is extracted from the output of step S2. This is combined with the output cost of the generating units (300-500 RMB / MWh for thermal power units, 0 RMB / MWh for wind power units) and the loss cost of the transmission lines (5-15 RMB / MWh) to construct a revenue function for each agent. The function incorporates the product of the market response coefficient and the corresponding agent's transaction scale (100-5000 MWh) as an influencing factor, strengthening the impact of market dynamics on revenue. In S32, the bargaining weight is determined based on the agent's role and influence. Weights are allocated to the generating unit's reference installed capacity (100-1500 MW per unit) and its contribution to power supply reliability (quantified based on historical power supply compliance rates of 95%-99%). The transmission capacity of the reference transmission lines (220 kV, 200-400 MW) is also considered. Weights are assigned to the importance of network connectivity (quantified by the number of connection nodes, from 3 to 10) for 00 MW and 500 kV (800 MW to 1200 MW), and weights are assigned to the reference load scale of the electricity consumption side (50 MW to 200 MW for residential and 300 MW to 800 MW for industrial) and load priority (0.3 to 0.5 for adjustable loads and 0.6 to 0.8 for non-adjustable loads). In S33, constraints are set, including the upper and lower limits of unit output (300 MW to 1500 MW for thermal power units and 0 MW to 300 MW for wind power units), ramp rate (100 MW / h to 200 MW / h), upper and lower limits of line transmission power (consistent with transmission capacity), and market transaction constraints, which are transformed into mathematical inequalities and incorporated into the model. In S34, the Lagrange multiplier method is used to solve the model, construct the Lagrange function and find the extreme value to obtain the optimal solution of the bargaining variables that satisfy all constraints, and determine the initial bargaining scheme. By refining the steps and parameter settings of S3, the scientific nature of multi-agent bargaining is ensured, providing a reliable foundation for subsequent optimization.
[0037] Preferably, step S4 includes the following sub-steps: S41, decompose the initial negotiation scheme obtained in S3 into a set of initial actions for each intelligent agent, and construct the state space of the intelligent agents by combining the real-time operating data of the power system, such as unit output, line power flow, and load. Each state vector in the state space includes fused information of real-time operating parameters and initial action parameters; S42, define the action space of the intelligent agents. The action space of the generation-side intelligent agents includes the adjustment range and direction of unit output; the action space of the transmission-side intelligent agents includes the distribution ratio and control method of line power flow; and the action space of the consumption-side intelligent agents includes the load transfer period and transfer amount. The boundary of the action space is determined by the safety operation constraints of the power system; S43, design... The reward function includes positive rewards such as reduced power system operating costs, improved power supply reliability, and increased market transaction revenue, and negative rewards such as penalties for generator over-limit operation, excessive line power flow, and load deficit. The final reward function is obtained by weighted summation. In step S44, the dynamic game deep reinforcement learning model is trained using experience replay and target network techniques. The state-action-reward-next state data generated by the agent's interaction with the environment is stored in an experience pool. Data from the experience pool is randomly sampled to update the policy network and value network. The stability of model training is improved by delaying updates to the target network. Iterative training continues until the reward function value converges, resulting in the optimized game strategy.
[0038] Specifically, step S4 includes four sub-steps from S41 to S44: In S41, the initial negotiation scheme in step S3 is broken down into a set of initial actions for each intelligent agent. Combined with real-time parameters of the power system (generator output 300 MW to 1500 MW, line power flow 200 MW to 1200 MW, load 50 MW to 800 MW), a state space is constructed. Each state vector includes the fusion information of the above real-time parameters and initial action parameters, with a total of 30 to 50 dimensions. In S42, the action space is defined. The generator side actions include the output adjustment range (-50 MW to 50 MW) and direction. The transmission side includes the power flow allocation ratio (0.1 to 0.9) and control mode. The power consumption side includes the load transfer period (1 hour to 4 hours) and transfer amount (-20 MW to 20 MW). The action boundaries are determined by system safety constraints (such as output not exceeding the upper and lower limits and power flow not exceeding the capacity). In S43, a reward function is designed. Positive reward items include operating cost reduction (10 to 20 points for every 10,000 yuan reduction), power supply reliability improvement (5 to 10 points for every 1% increase in load satisfaction rate), and market revenue increase (8 to 15 points for every 5,000 yuan increase). Negative reward items include unit over-limit penalties (-50 to -100 points for every 10 MW over-limit), line overload penalties (-40 to -80 points for every 5 MW overload), and load deficit penalties (-60 to -120 points for every 5 MW deficit). The weighted sum is used to obtain the final reward. In S44, experience replay (experience pool capacity of 10,000 to 50,000 records) and target network technology are used to train the model. 32 to 64 data records are randomly sampled to update the strategy and value network. The target network is updated with a delay of 50 to 100 rounds. Iterative training is performed for 1,000 to 5,000 rounds until the reward converges, resulting in an optimized game strategy. By refining the S4 sub-steps, the accuracy of model training and the effectiveness of strategy optimization can be improved.
[0039] Preferably, step S5 includes the following sub-steps: S51, parsing the optimized game strategy output by S4, extracting the optimal actions of each agent, converting the optimal actions of the generation-side agents into unit start-stop commands and output setpoints, converting the optimal actions of the transmission-side agents into line power flow control commands and transmission power allocation values, and converting the optimal actions of the consumption-side agents into load adjustment plans and power consumption time arrangements; S52, combining the power system parameters collected in S1, performing feasibility verification on the converted commands and values, including whether the unit start-stop commands meet the minimum start-stop time constraints, whether the output setpoints are within the upper and lower limits of unit output, and whether the line power flow control commands meet the requirements. S53: Integrate the verified instructions and values to generate a unit start-up and shutdown plan, which specifies the start-up and shutdown times and continuous operating duration of each unit; generate a unit output allocation scheme, which specifies the output value of each unit at different time periods; and generate a line power flow control scheme, which specifies the power flow control value of each line at different time periods. S54: Combine the generated unit start-up and shutdown plan, unit output allocation scheme, and line power flow control scheme to form a complete dispatch plan. The dispatch plan includes the execution time period, execution subject, and monitoring indicators of each scheme. The monitoring indicators are used to track the operation status of the dispatch plan.
[0040] Specifically, step S5 includes four sub-steps, S51 to S54: In S51, the optimization game strategy from step S4 is analyzed, and the optimal actions of each agent are extracted. The optimal actions on the power generation side are converted into unit start-up and shutdown commands (accurate to the hour) and output setpoints (300 MW to 1500 MW). On the transmission side, they are converted into line power flow control commands (200 MW to 1200 MW) and transmission power allocation values (allocated according to the line capacity ratio of 0.1 to 0.9). On the power consumption side, they are converted into load adjustment plans (adjustable load transfer of 1 to 4 hours) and power consumption time arrangements (peak and valley time division: peak hours 8:00-22:00, valley hours 22:00-8:00 the next day). In S52, feasibility is verified by combining the system parameters from step S1 to verify whether the unit start-up and shutdown commands meet the minimum start-up and shutdown time (4 to 8 hours for thermal power units and 0.5 to 1 hour for wind power units) and whether the output setpoints are within the upper and lower limits. In S53, the verified instructions and values are integrated to generate unit start-up and shutdown plans (specifying the start-up and shutdown times of each unit and the continuous operating duration of 4 to 20 hours), unit output allocation schemes (allocated according to time-period load demand of 500 MW to 2000 MW, with a single unit output share of 0.05 to 0.3), and line power flow control schemes (specifying the time-period power flow control values for each line of 200 MW to 1200 MW, with a deviation allowable of ±5 MW). In S54, the above schemes are combined to form a dispatch plan, which includes the execution period (24 to 72 hours), the executing entities (power plant, transmission company, user-side management platform), and monitoring indicators (unit output deviation ±5 MW, line power flow deviation ±3 MW, load satisfaction rate ≥99%), for subsequent simulation tracking. By refining the steps in S5, the feasibility and completeness of the dispatch plan are ensured, laying the foundation for subsequent simulation evaluation.
[0041] The dynamic market response strategy network in this invention is an intelligent computing architecture used to capture the sensitivity of electricity users to market changes. Through multi-layered computation and weight allocation, it transforms market-side and electricity-side parameters into quantified market response coefficients. The implementation process first involves collecting market-side parameters such as electricity price fluctuation coefficients (0.8 to 1.5 day-ahead, 0.5 to 2.0 real-time), transaction volume constraints (100 MWh to 5000 MWh), and electricity-side load forecasts (50 to 200 MW / area for residential use, 300 to 800 MW / factory for industrial use), and load types in step S1. Then, in step S2, a three-layer computational structure is constructed (12 to 18 nodes in the input layer, 32 to 64 nodes in the hidden layer, and 20 to 50 nodes in the output layer). Through nonlinear feature extraction and an attention mechanism (assigning weights of 0.6 to 0.8 to load forecasts and electricity price fluctuation coefficients, and 0.2 to 0.4 to transaction volume constraints and load types), the parameters are mapped to market response coefficients in the range of 0 to 1. The network serves to provide crucial data support for subsequent multi-agent bargaining, accurately distinguishing the differences in market sensitivity among different electricity consumers. It breaks down the disconnect between market factors and electricity consumption behavior in traditional dispatching, making dispatching decisions more aligned with market dynamics, enhancing the power system's adaptability to market changes, and laying a data foundation for multi-agent collaborative optimization.
[0042] The asymmetric Nash bargaining model in this invention is a game-theoretic computation model that coordinates the interests of multiple agents on the power generation, transmission, and consumption sides. It aims to determine the initial bargaining scheme by quantifying the revenue and bargaining weight of each agent. In step S3, the implementation process first extracts the market response coefficient from step S2. Combined with the power generation unit output cost (300-500 RMB / MWh for thermal power units, 0 RMB / MWh for wind power units) and the transmission line network loss cost (5-15 RMB / MWh), a revenue function for each agent is constructed (incorporating the product of the market response coefficient and transaction size as an influencing factor). Then, bargaining weights are set according to the agent's role (0.4-0.5 for the total weight on the power generation side, 0.2-0.3 for the transmission side, and 0.2-0.3 for the power consumption side), and constraints are set (upper and lower limits of unit output, ramp rate of 100-200 MW / h, and line transmission capacity). Finally, the Lagrange multiplier method is used to solve for the initial unit output, initial line power flow, and initial load allocation values for the main power consumption entities. The purpose of this model is to balance the interests of multiple agents and avoid a decrease in overall system efficiency due to the maximization of the interests of a single agent. This addresses the difficulty in quantifying conflicts of interest among multiple stakeholders in traditional dispatching, achieving a balance of interests through a scientific bargaining mechanism. This provides a reasonable initial plan for subsequent optimization of game strategies, ensuring the fairness and effectiveness of collaborative operation among multiple stakeholders in the power system.
[0043] The dynamic game deep reinforcement learning model in this invention is an intelligent algorithm model that optimizes the multi-agent scheduling strategy through iterative training. Its core is to enable agents to learn the optimal action selection rules in the process of interacting with the power system environment. In step S4, the initial negotiation scheme from step S3 is used as input. A state space (containing 30 to 50 dimensions including generator output, line power flow, load consumption, and market electricity price), an action space (generator output adjustment from -50 MW to 50 MW, line power flow adjustment from -30 MW to 30 MW, and load transfer from -20 MW to 20 MW), and a reward function (positive rewards include reduced operating costs and improved power supply reliability; negative rewards include penalties for generator overruns and line overloads) are set. Then, experience replay (experience pool capacity of 10,000 to 50,000 records) and target network technology are used. The strategy network learning rate is 0.001 to 0.005, and the value network learning rate is 0.0001 to 0.001. The system is iteratively trained for 1000 to 5000 rounds until the reward value fluctuates by no more than 5% for 50 consecutive rounds. The optimized game strategy is then output. The model's function is to dynamically optimize the scheduling strategy, enabling scheduling decisions to adapt to real-time changes in the power system's operation. Breaking through the limitations of traditional static scheduling, this method achieves autonomous iterative optimization of strategies through reinforcement learning, thereby enhancing the adaptability of scheduling schemes to dynamic changes in the system and ensuring the safety and optimal operation of large-scale power systems.
[0044] The intelligent decision-making platform for the energy internet in this invention is an integrated technology platform supporting the entire process of large-scale power dispatch optimization, including core functions such as data acquisition, model calculation, contingency plan generation, and simulation evaluation. The implementation process spans steps S1 to S6. In S1, multiple parameters are collected through the perception layer and integrated into a unified format. In S2 to S4, a computational environment is provided for the dynamic market response strategy network, the asymmetric Nash bargaining model, and the dynamic game deep reinforcement learning model. In S5, dispatch contingency plan generation and verification are supported. In S6, the simulation module (simulation duration 24 to 72 hours, time step 15 to 60 minutes) calculates the system network loss rate (accuracy 0.1%), load fulfillment rate (accuracy 0.1%), and unit coal consumption rate (accuracy 0.1 g / kWh), and determines whether iterative adjustments are needed based on thresholds. The platform's role is to connect the entire dispatch optimization process, providing data interaction and computational support for each model and step. Breaking away from the data silos and process fragmentation issues in traditional dispatching, a closed-loop dispatching optimization system is constructed to achieve multi-side data integration, multi-model collaborative computation, and dynamic adjustment of multiple links. This provides an integrated technical carrier for large-scale power dispatching optimization, ensuring the scientific, efficient, and reliable nature of dispatching decisions.
[0045] like Figure 2The aforementioned large-scale power dispatch optimization system based on intelligent agent collaboration includes: a multi-dimensional power parameter acquisition and integration unit, which is connected to the perception layer of the energy internet intelligent decision-making platform, receives various parameters from the generation side, transmission side, consumption side, and market side, and integrates parameters from different sources into a unified parameter set through data cleaning and format conversion. The output of this unit is connected to a dynamic market response strategy network construction unit; and a dynamic market response strategy network construction and calculation unit, which receives the parameter set output by the multi-dimensional power parameter acquisition and integration unit, and constructs a dynamic market response strategy network based on a preset network structure and parameter initialization rules. The dynamic market response strategy network performs feature extraction and weight allocation calculations on the parameter set, outputting the market response coefficients of each electricity consumer. The output of this unit is connected to the asymmetric Nash bargaining model establishment unit. The asymmetric Nash bargaining model establishment and solution unit receives the market response coefficients output by the dynamic market response strategy network construction and calculation unit, combines them with the generation-side and transmission-side parameters output by the multi-dimensional power parameter acquisition and integration unit, constructs the asymmetric Nash bargaining model, sets model constraints and objective functions, and uses an optimization algorithm to solve the model to obtain the initial bargaining scheme. The output of this unit is connected to the dynamic game deep reinforcement learning unit. The model training unit is connected to the dynamic game deep reinforcement learning model training and optimization unit. This unit receives the initial bargaining scheme output by the asymmetric Nash bargaining model establishment and solution unit, constructs the agent's state space, action space, and reward function, initializes the policy network and value network parameters, iteratively trains the model through experience replay and target network techniques, and outputs the optimized game strategy. The output of this unit is connected to the scheduling plan generation unit. The scheduling plan generation and verification unit receives the optimized game strategy output by the dynamic game deep reinforcement learning model training and optimization unit and transforms it into unit start-up and shutdown plans and unit output. The system includes allocation schemes and line power flow control schemes. Feasibility checks are performed on each scheme, and schemes that pass the combined checks are used to generate a dispatch plan. The output of this unit is connected to the dispatch plan simulation evaluation and determination unit. The dispatch plan simulation evaluation and determination unit receives the dispatch plan output by the dispatch plan generation and verification unit, calls the simulation module of the energy internet intelligent decision-making platform to simulate the operation of the dispatch plan, calculates the power system network loss rate, load fulfillment rate, and unit coal consumption rate, compares them with preset thresholds to determine whether the requirements are met, and outputs the final large-scale power dispatch optimization scheme. The output of this unit is connected to the external power dispatch execution system.
[0046] A method and system for large-scale power dispatch optimization based on intelligent agent collaboration is proposed. This method and system can comprehensively collect and integrate various key parameters from the generation, transmission, consumption, and market sides, breaking down data barriers between different links and providing complete and accurate basic information support for dispatch decisions. Through the collaborative operation of multiple links, parameter processing, scheme construction, model training, contingency plan generation, and simulation evaluation are closely linked to form a closed-loop dispatch optimization process. This not only improves the scientificity and rationality of dispatch schemes but also enhances the ability to control the operating status of complex power systems. In the dispatch optimization process, the system fully considers the needs of system operating efficiency, cost control, market response, and safety and stability, achieving balanced optimization of multiple objectives and avoiding performance degradation in other aspects caused by optimization of a single objective.
[0047] To overcome the shortcomings of the prior art, addressing the difficulty of existing dispatching methods in integrating market dynamics and the demands of multiple stakeholders, this method and system, through multi-unit collaboration, fully incorporates market dynamics such as market price fluctuations and changes in user electricity consumption behavior during the decision-making process. It also considers the interests of different stakeholders, including the generation, transmission, and consumption sides, balancing their interests through a reasonable coordination mechanism. This allows the generated dispatching scheme to better adapt to market changes, meet the needs of multi-stakeholder collaborative operation, and effectively improve market adaptability. Furthermore, addressing the shortcomings of existing technologies in comprehensively considering real-time dynamic changes in system parameters and lacking multi-indicator verification and adjustment mechanisms, this method fully integrates real-time parameter constraints such as unit output adjustment rates and line transmission capacity during the scheme design phase. Through a dedicated simulation evaluation process, it verifies multiple indicators of the dispatching plan, including system network losses, load fulfillment, and unit coal consumption. If requirements are not met, timely backtracking adjustments are made, significantly reducing actual execution deviations and ensuring the long-term stable operation of large-scale power systems.
[0048] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," "connect," "link," and "fix" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal communication between two components. Those skilled in the art will understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0049] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various equivalent changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A large-scale power dispatch optimization method based on agent-based cooperation, characterized in that, Includes the following steps: S1. Collect various parameters for large-scale power dispatch scenarios through the energy internet intelligent decision-making platform. These parameters include the upper and lower limits of generator output and the unit ramp-up rate on the generation side; the line transmission capacity and line impedance on the transmission side; the load forecast and load type on the consumption side; and the electricity price fluctuation coefficient and transaction volume constraints on the market side. S2. Construct a dynamic market response strategy network. Input the market-side parameters and consumption-side parameters collected in S1 into this network. Use a multilayer perceptron and attention mechanism within the network to extract features and assign weights to the parameters, obtaining the market response coefficients for different electricity consumers. S3. Based on S2... Based on the obtained market response coefficients and parameters from the generation and transmission sides, an asymmetric Nash bargaining model is established. Taking the revenue functions and bargaining weights of different agents as the core, the initial bargaining scheme among multiple agents is determined. In S4, the initial bargaining scheme determined in S3 is used as input, and a dynamic game deep reinforcement learning model is introduced. The state space of the agents is set as the real-time operating parameters of the power system, the action space is the unit output adjustment and the line power flow allocation, and the reward function is the weighted sum of the power system operating cost reduction and the power supply reliability improvement. The strategy network and value network of the agents are updated through iterative training of the model to obtain the optimized game strategy. S5, based on the optimized game strategy obtained in S4, adjusts different parameters of large-scale power dispatch to generate a dispatch plan including unit start-up and shutdown plans, unit output allocation schemes, and line power flow control schemes; S6, feeds the dispatch plan generated in S5 back to the energy internet intelligent decision-making platform, and simulates the operation of the dispatch plan through the platform's simulation module, outputting the power system network loss rate, load fulfillment rate, and unit coal consumption rate corresponding to the dispatch plan. If all indicators meet the preset thresholds, the dispatch plan is taken as the final large-scale power dispatch optimization scheme. If not, it returns to S2 to readjust the parameters of the dynamic market response strategy network and repeats S3, S4, S5, and S6.
2. The large-scale power dispatch optimization method based on agent collaboration according to claim 1, characterized in that, The expression for the dynamic market response strategy network is: middle, Let be the market response coefficient of the i-th electricity consumer at time t. It is the Sigmoid activation function. , , These are the weight matrices for the first, second, and third layers of the network, respectively. These are the bias vectors for the first, second, and third layers of the network, respectively. Let be the feature vector of the i-th electricity consumer at time t, including the load forecast, electricity price fluctuation coefficient, and traded electricity volume constraint. Let be the response sensitivity coefficient of the i-th electricity consumer at time t. Let be the market participation coefficient of the i-th electricity consumer at time t. Let be the association weight between the i-th electricity consumer and the k-th market node. Let be the transaction volume of the k-th market node at time t. This represents the total number of market nodes.
3. The large-scale power dispatch optimization method based on agent collaboration according to claim 1, characterized in that, The expression for the asymmetric Nash bargaining model is: , ,in, This is a vector of negotiation variables, including unit output adjustments and load allocation. Let this be the feasible region of the bargaining variable. Let j be the payoff function of the j-th agent. For the retention benefit of the j-th agent, Let the bargaining weight of the j-th agent satisfy the following conditions: The total number of agents. These are the lower and upper limits of the output of the j-th unit, respectively. The j-th unit at time t and Constant effort Let be the ramp rate of the j-th unit. The transmission power of the I-th line, These are the lower and upper limits of the transmission power for line I, respectively.
4. The large-scale power dispatch optimization method based on agent collaboration according to claim 1, characterized in that, The policy network expression of the dynamic game deep reinforcement learning model is: ,in, For parameters The policy network, Let t be the action of the agent. Let t be the state of the agent at time t. The output function of the policy network includes the state Feature mapping and action Matching calculation, where A is the action space; the value network expression is: ,in, For parameters The value network Let be the weight coefficient of the k-th value calculation unit, p be the total number of value calculation units, and ReLU be the linear rectified activation function. This is the weight matrix for the k-th value calculation unit. This is the bias vector for the k-th value calculation unit.
5. The large-scale power dispatch optimization method based on agent collaboration according to claim 1, characterized in that, The comprehensive index calculation model used to evaluate dispatch plans in the energy internet intelligent decision-making platform is as follows: ,in, For comprehensive evaluation indicators, , The weighting coefficients for network loss rate, load fulfillment rate, and unit coal consumption rate are respectively, and satisfy the following conditions: , For power system network loss rate, For load sufficiency, This refers to the unit's coal consumption rate; simultaneously, the real-time performance guarantee model for parameter acquisition in the platform is as follows: ,in, q represents the total time spent on parameter acquisition, and q represents the total number of parameter acquisition nodes. Let be the time taken for a single data collection at the q-th data collection node. Let be the sampling frequency coefficient of the q-th sampling node.
6. The large-scale power dispatch optimization method based on agent cooperation according to claim 1, characterized in that, The accuracy calibration model for unit output allocation in the aforementioned large-scale power dispatch optimization scheme is as follows: ,in, For the output calibration value of the j-th unit, Let be the real-time calibration coefficient for the j-th unit. These are the optimized output and actual output of the j-th unit, respectively. Let T be the historical deviation correction factor for the j-th unit, and T be the historical data statistical duration. Let be the historical optimized output and historical actual output of the j-th generating unit at time t, respectively; the stability adjustment model for line power flow control is: ,in, This represents the power flow regulation amount for the I-th line. Let be the power flow margin coefficient for the I-th line. This represents the upper limit of the transmission capacity of the I-th line. This represents the actual power flow of line I. This is the power flow rate adjustment coefficient for the I-th line. Let be the power flow rate of the I-th line.
7. The large-scale power dispatch optimization method based on agent collaboration according to claim 1, characterized in that, S3 includes the following steps: S31, extracting the market response coefficients corresponding to each agent from the output results of S2, and constructing a revenue function for each agent by combining the output cost function of the generating units and the loss function of the transmission lines. The revenue function incorporates the influence factor of the market response coefficient on the revenue, which is calculated by multiplying the market response coefficient by the transaction scale of the corresponding agent; S32, determining the bargaining weight of each agent based on its role and influence in the power system. The bargaining weight of the generating agents refers to the generating unit capacity and the contribution to power supply reliability, while the bargaining weight of the transmission agents refers to the transmission capacity of the transmission lines and the... The importance of network connectivity is considered, and the bargaining weight of the electricity-consuming agents is referenced to the load scale and load priority; S33, the constraints of the asymmetric Nash bargaining model are set, including the upper and lower limits of the output of the generating units, the ramp rate constraint, the transmission power constraint of the transmission lines, the load demand constraint of the electricity-consuming side, and the electricity and price constraints in market transactions. Each constraint is transformed into a mathematical inequality and incorporated into the model; S34, the asymmetric Nash bargaining model is solved using the Lagrange multiplier method. By constructing the Lagrange function, the optimal solution of the bargaining variables that satisfies all constraints is obtained by solving the extreme value of the function, thereby determining the initial bargaining scheme among multiple agents.
8. The large-scale power dispatch optimization method based on agent collaboration according to claim 1, characterized in that, S4 includes the following steps: S41, decompose the initial negotiation scheme obtained in S3 into a set of initial actions for each agent, and construct the state space of the agents by combining the real-time power system output, line flow, and load data. Each state vector in the state space includes fused information of real-time operating parameters and initial action parameters; S42, define the action space of the agents. The action space of the generation-side agents includes the adjustment range and direction of the unit output; the action space of the transmission-side agents includes the distribution ratio and control method of the line flow; and the action space of the consumption-side agents includes the load transfer period and transfer amount. The boundary of the action space is determined by the safety operation constraints of the power system; S43, design rewards. The reward function includes positive rewards such as reduced power system operating costs, improved power supply reliability, and increased market transaction revenue, and negative rewards such as penalties for generator over-limit operation, excessive line power flow, and load deficit. The final reward function is obtained by weighted summation. S44: The dynamic game deep reinforcement learning model is trained using experience replay and target network techniques. The state-action-reward-next state data generated by the agent's interaction with the environment is stored in the experience pool. Data from the experience pool is randomly sampled to update the policy network and value network. The stability of model training is improved by delaying the update of the target network. Iterative training continues until the reward function value converges, resulting in the optimized game strategy.
9. The large-scale power dispatch optimization method based on agent collaboration according to claim 1, characterized in that, S5 includes the following steps: S51, analyzing the optimized game strategy output by S4, extracting the optimal actions of each agent, converting the optimal actions of the generator-side agents into generator start-up and shutdown commands and output setpoints, converting the optimal actions of the transmission-side agents into line power flow control commands and transmission power allocation values, and converting the optimal actions of the consumer-side agents into load adjustment plans and electricity consumption schedules; S52, combining the power system parameters collected in S1, performing feasibility verification on the converted commands and values, including whether the generator start-up and shutdown commands comply with the minimum start-up and shutdown time constraints, whether the output setpoints are within the upper and lower limits of generator output, and whether the line power flow control commands meet the line power flow requirements. S53: Integrate the verified instructions and values to generate a unit start-up and shutdown plan, which specifies the start-up and shutdown times and continuous operating duration of each unit; generate a unit output allocation scheme, which specifies the output value of each unit at different time periods; and generate a line power flow control scheme, which specifies the power flow control value of each line at different time periods. S54: Combine the generated unit start-up and shutdown plan, unit output allocation scheme, and line power flow control scheme to form a complete dispatch plan. The dispatch plan includes the execution time period, execution subject, and monitoring indicators of each scheme. The monitoring indicators are used to track the operation status of the dispatch plan.
10. A large-scale power dispatch optimization system based on agent-based collaboration, characterized in that, include: The power parameter multi-dimensional acquisition and integration unit connects to the perception layer of the energy internet intelligent decision-making platform. It receives various parameters from the generation, transmission, consumption, and market sides, and integrates these parameters into a unified parameter set through data cleaning and format conversion. The output of this unit connects to the dynamic market response strategy network construction unit. The dynamic market response strategy network construction and calculation unit receives the parameter set output from the power parameter multi-dimensional acquisition and integration unit, constructs a dynamic market response strategy network based on a preset network structure and parameter initialization rules, and performs feature extraction on the parameter set. The first unit performs weighted allocation calculations and outputs the market response coefficients for each electricity consumer. The output of this unit is connected to the asymmetric Nash bargaining model establishment unit. The second unit, the asymmetric Nash bargaining model establishment and solution unit, receives the market response coefficients output from the dynamic market response strategy network construction and calculation unit. It combines these with generation-side and transmission-side parameters output from the multi-dimensional power parameter acquisition and integration unit to construct an asymmetric Nash bargaining model. It sets model constraints and objective functions, and uses an optimization algorithm to solve the model to obtain the initial bargaining scheme. The output of this unit is connected to the dynamic game deep reinforcement learning model training unit. The deep reinforcement learning model training and optimization unit receives the initial bargaining scheme output by the asymmetric Nash bargaining model establishment and solution unit, constructs the agent's state space, action space, and reward function, initializes the policy network and value network parameters, iteratively trains the model through experience replay and target network techniques, and outputs the optimized game strategy. The output of this unit is connected to the scheduling plan generation unit. The scheduling plan generation and verification unit receives the optimized game strategy output by the dynamic game deep reinforcement learning model training and optimization unit, and transforms it into unit start-up and shutdown plans, unit output allocation schemes, and line scheduling plans. The power flow control scheme verifies the feasibility of each scheme, and generates a scheduling plan from the schemes that pass the combined verification. The output of this unit is connected to the scheduling scheme simulation evaluation and determination unit. The scheduling scheme simulation evaluation and determination unit receives the scheduling plan output by the scheduling plan generation and verification unit, calls the simulation module of the energy internet intelligent decision-making platform to simulate the operation of the scheduling plan, calculates the power system network loss rate, load fulfillment rate, and unit coal consumption rate, compares them with preset thresholds to determine whether the requirements are met, and outputs the final large-scale power dispatch optimization scheme. The output of this unit is connected to the external power dispatch execution system.
Citation Information
Cited By
Autonomous controllable environment-oriented power business system tuning method and device, and medium
CN120782193A