Distributed energy storage and power scheduling method and device based on master-slave game
By employing a master-slave game-theoretic distributed energy storage and power scheduling method, the problem of dynamic power fluctuations in energy storage systems is solved, achieving efficient power allocation matrix updates and real-time feedback control, thereby improving the stability and responsiveness of energy storage systems.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- LINXIA COUNTY ELECTRIC POWER CO
- Filing Date
- 2026-01-12
- Publication Date
- 2026-04-21
AI Technical Summary
Existing energy storage power dispatch methods are difficult to adapt to the complex fluctuations of dynamic power curves, leading to the risk of overcharging or over-discharging. Furthermore, they lack dynamic game-theoretic mechanisms across multiple time scales, making it impossible to respond in real time to grid dispatch commands or changes in the state of energy storage units.
A distributed energy storage and power dispatching method based on master-slave game theory is adopted. By incrementally updating the power allocation matrix and dynamic weight parameters in the short term, combined with power deviation detection, closed-loop control is achieved to optimize the economy, safety and grid interaction requirements.
It achieves discrete approximation of continuous power demand, reduces errors, ensures energy conservation of energy storage units, improves the consistency between dispatch commands and actual output, and enhances the robustness and real-time response capability of the system.
Smart Images

Figure CN121906558A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of distributed energy storage management technology, specifically to a distributed energy storage and power scheduling method and device based on master-slave game theory. Background Technology
[0002] Existing energy storage power dispatch methods typically rely on continuous function modeling or fixed time-scale discretization strategies, which are difficult to effectively adapt to the complex fluctuation characteristics of dynamic power curves. Static discretization strategies lead to the accumulation of deviations between power allocation values and actual demand in short-term periods. Especially when wind and solar power output fluctuates drastically or load changes suddenly, they are prone to overcharging or over-discharging risks. Existing optimization frameworks are mostly based on offline prediction and lack dynamic game mechanisms with multiple time scales, making it impossible to respond in real time to grid dispatch commands or changes in the state of energy storage units.
[0003] In the prior art, CN116031948A discloses a master-slave game scheduling model constructed based on the game relationship between the distribution network and the microgrid. The model is solved using a particle swarm optimization algorithm to determine the optimal and efficient operation strategy for the distribution network under the condition of maximizing the benefits for both the distribution network and the microgrid. However, this method does not achieve discrete approximation of continuous power demand through incremental updates of the power allocation matrix over short periods, nor does it support minute-level real-time feedback correction to significantly improve the consistency between scheduling instructions and actual output. Furthermore, it lacks a target benefit function based on dynamic weight parameters to simultaneously optimize economic efficiency, security, and grid interaction requirements, and it does not trigger matrix reset when power deviation exceeds a threshold, nor does it incorporate dynamic scaling based on capacity continuity constraints. Therefore, an efficient and reliable distributed energy storage and power scheduling method based on master-slave game theory is urgently needed.
[0004] The information disclosed in the background section is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0005] The purpose of this invention is to provide a distributed energy storage and power scheduling method and apparatus based on master-slave game theory, so as to solve the problems mentioned in the background art.
[0006] To achieve the above objectives, the present invention provides the following technical solution: The distributed energy storage and power scheduling method based on master-slave game theory includes the following steps: S1: Obtain the power distribution curve of the energy storage unit, and for the energy storage unit, divide it into equal lengths within a set time window. Each long-term period is divided into several long-term periods, and the power data allocated to each long-term period is determined accordingly. The system extracts power data allocated within each short-term period, calculates the average power of each short-term period as the power allocation value for that short-term period, and constructs a power allocation matrix by iterating through the power allocation values allocated in all short-term periods. S2: Adjust the power allocation matrix according to the power continuity constraints to obtain the adjusted power allocation matrix; based on the adjusted power allocation matrix, use differential game theory to set the energy storage dynamic equation, and set the target payoff function according to the energy storage dynamic equation, and use a dynamic weight allocation mechanism to calculate the weight parameters of the target payoff function. S3: Calculate the deviation between the actual discharge power of the energy storage unit and the adjusted power allocation matrix in real time for each long period. If the deviation is not less than the set power difference threshold, reset the power allocation matrix and update the weight parameters of the target revenue function. When the power allocation matrix is less than the minimum allowable power or the rate of change of the target revenue function is greater than the preset threshold, output the power allocation matrix in real time.
[0007] Further, the power allocation matrix is constructed, and the specific steps are as follows: Define the time window length as Divide the time window into equal parts There are several long-term periods, and each long-term period is numbered as follows: ,in, The index number represents the long-term time period, and each long-term time period is further divided into... There are several short-term periods, and all short-term periods are numbered as follows: ,in, ,in, Indicates the index number for a short-term period; The specific logic for calculating the average power of each short-term period is as follows: determine the start and end time points of each short-term period, perform dense sampling on them using sensors, calculate the corresponding area under the power curve of the short-term period using the rectangular area method, take the area value as the total power of the short-term period, divide the total power by the length of the start time period to obtain the average power of the corresponding short-term period, and take it as the power allocation value of the short-term period. This process is repeated to obtain the power allocation value of all short-term periods. Divide all power allocation values into OK Power allocation matrix of columns ,in, This represents the product of the total number of long-term periods and the total number of short-term periods. Indicates the energy storage unit in the first Within a long-term period, the first Power allocation values for short-term periods.
[0008] Further, the adjusted power allocation matrix is obtained, and the specific steps are as follows: Iterate through each element of the power allocation matrix for each long time period, and adjust the power allocation matrix based on constraints set according to power continuity: in, Indicates a long period of time The actual total charging energy; This indicates the adjusted energy storage unit during the long-term period. Intra-short period The allocated power distribution value.
[0009] Furthermore, differential game theory is used to set the dynamic equations for energy storage. The specific steps are as follows: Based on the adjusted Differential game theory is used to set the dynamic equations for energy storage: in, in, express Real-time energy storage capacity of the energy storage unit; This represents the overall efficiency coefficient of the energy storage unit; express The instantaneous power applied to the energy storage unit at all times; This represents the equivalent self-attenuation coefficient of the energy storage unit; Represents the Dirac impulse function; Indicates the start time of the current window; This indicates the energy storage capacity of the initial energy storage unit.
[0010] Furthermore, a dynamic weight allocation mechanism is used to calculate the weight parameters of the objective return function. The specific steps are as follows: based on and The target return function is set according to the energy storage dynamic equation: in, This represents the total revenue of the energy storage unit; Indicates time Real-time electricity price weighting parameters; This represents the capacity fluctuation penalty coefficient; This represents the power fluctuation penalty coefficient; The weight parameters of the objective return function are calculated using a dynamic weight allocation mechanism. in, Indicates short period of time Actual dispatch power; This indicates the maximum power limit allowed within a short period of time.
[0011] Furthermore, the deviation between the actual power of the energy storage unit and the total capacity over the current long-term period is calculated in real time. The specific steps are as follows: Real-time calculation of the deviation between the actual power of the energy storage unit and the adjusted power allocation matrix for each long-term period: in, Indicates the first The absolute value of the deviation between the actual discharge power and the adjusted power distribution matrix over a long period; This indicates the current actual discharge power of the energy storage unit. Furthermore, the weight parameters of the target return function are reset, specifically through the following steps: when Greater than or equal to When this happens, the weight parameters of the objective return function are updated as follows: in, Indicates the number of consecutive resets; Indicates the first The adaptive attenuation coefficient for the next reset; Indicates the first The power sensitivity factor of the second reset; Indicates the first The time decay coefficient of the next reset; Indicates the first The threshold for the actual discharge power difference over a long period of time.
[0012] Further, the power allocation matrix is reset, specifically through the following steps: in, This indicates a reset over a long period of time. Intra-short period The allocated power distribution value; The power allocation matrix is output in real time when the power allocation matrix is less than the minimum allowable power or the rate of change of the target revenue function is greater than the preset threshold.
[0013] The present invention also provides a distributed energy storage and power scheduling device based on master-slave game theory, wherein the scheduling device is used to execute the above-described scheduling method, comprising: The energy storage planning module is used to obtain the power allocation curve of the energy storage unit. For the energy storage unit, the time window is divided into N long-term periods of equal length. The power data allocated to each long-term period is determined. Each long-term period is divided into M short-term periods, and the power data allocated to each short-term period is extracted. The average power of each short-term period is calculated as the power allocation value of that short-term period. The power allocation matrix is constructed by traversing the power data allocated to all short-term periods. The game optimization module is used to adjust the power allocation matrix based on the power allocation matrix and set constraints according to power continuity to obtain the adjusted power allocation matrix; based on the optimized power allocation matrix, the energy storage dynamic equation is set using differential game theory, and the target payoff function is set according to the energy storage dynamic equation. The weight parameters of the target payoff function are calculated using a dynamic weight allocation mechanism. The dynamic feedback module is used to calculate in real time the deviation between the actual discharge power of the energy storage unit and the adjusted power allocation matrix for each long-term period. If the deviation is not less than the set power difference threshold, the power allocation matrix is reset and the weight parameters of the target revenue function are updated. When the power allocation matrix is less than the minimum allowable power or the rate of change of the target revenue function is greater than the preset threshold, the power allocation matrix is output in real time.
[0014] Compared with the prior art, the beneficial effects of the present invention are: By incrementally updating the power allocation matrix for short-term periods, discrete approximation of continuous power demand and error reduction are achieved. Energy conservation issues in the discretization process are resolved through capacity constraint scaling, ensuring that the total charging or discharging amount strictly matches the physical limits of the energy storage unit. Dynamic weight parameters are introduced to adaptively adjust the priorities of economy, safety, and grid service based on real-time electricity price signals and grid demand, achieving dynamic tracking of the multi-objective Pareto front. Closed-loop control from prediction to execution to correction is achieved through power deviation detection, suppressing scheduling deviations caused by prediction errors or equipment aging. Attached Figure Description
[0015] Figure 1 This is a schematic diagram of the overall method flow of the present invention; Figure 2 This is a graph showing the relationship between the absolute value of the power deviation and the power distribution value. Figure 3 This is a schematic diagram of the overall device of the present invention. Detailed Implementation
[0016] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments.
[0017] It should be noted that, unless otherwise defined, the technical or scientific terms used in this invention should have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in this invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0018] Example: Please see Figures 1-2 The present invention provides a technical solution: The distributed energy storage and power scheduling method based on master-slave game theory includes the following steps: S1: Obtain the power distribution curve of the energy storage unit, and for the energy storage unit, divide it into equal lengths within a set time window. Each long-term period is divided into several long-term periods, and the power data allocated to each long-term period is determined accordingly. The system extracts power data allocated within each short-term period, calculates the average power of each short-term period as the power allocation value for that short-term period, and constructs a power allocation matrix by iterating through the power allocation values allocated in all short-term periods. The specific steps for constructing the capacity allocation matrix are as follows: Define the time window length as Divide the time window into equal parts There are several long-term periods, and each long-term period is numbered as follows: ,in, The index number represents the long-term time period, and each long-term time period is further divided into... There are several short-term periods, and all short-term periods are numbered as follows: ,in, ,in, Indicates the index number for a short-term period; The specific logic for calculating the average power of each short-term period is as follows: determine the start and end time points of each short-term period, perform dense sampling on them using sensors, calculate the corresponding area under the power curve of the short-term period using the rectangular area method, take the area value as the total power of the short-term period, divide the total power by the length of the start time period to obtain the average power of the corresponding short-term period, and take it as the power allocation value of the short-term period. This process is repeated to obtain the power allocation value of all short-term periods. Divide all power allocation values into OK Power allocation matrix of columns ,in, This represents the product of the total number of long-term periods and the total number of short-term periods. Indicates the energy storage unit in the first Within a long-term period, the first Power allocation values for short-term periods.
[0019] In the above process, the rectangular area method is used to calculate the area below the power curve for the short term as follows: Based on a pre-defined short time period, the start and end times are determined. Within this period, the sensor samples data at a time difference equal to the defined short time period. The power values are densely sampled at a high frequency to ensure that the sampling points can fully capture the power changes within the short time period, thereby obtaining a series of discrete time-power data points. These sampling points are then connected in chronological order to approximate the power change curve for that period. The sensor uses a time difference equal to the segmented short duration to measure the time between two adjacent samples. The reason for densely sampling the power values at a frequency of 1 / 1000 is that this means collecting 1000 data points within a short period. This ensures that even if the power fluctuates rapidly within that period, it can be fully captured by the dense sampling points. This results in highly accurate total power (area under the power change curve) and average power, far exceeding the typical engineering accuracy requirement of 1% error. Too few sampling points may not be able to fully capture some rapid fluctuations, leading to a decrease in integration accuracy and affecting optimization quality. Too many sampling points, while theoretically more accurate, will generate massive amounts of data, placing a huge burden on real-time data transmission, storage, and processing, and placing excessive demands on sensor, communication, and computing resources. The rectangular method from numerical integration is used for area estimation: the entire short-term period is divided into multiple smaller micro-periods, the length of which is the sampling interval. For each micro-period, the power value sampled at its starting time is used as an approximate constant power value within that period, and the product of this power value and the length of the micro-period is used as the rectangular area corresponding to that micro-period. Finally, the rectangular areas corresponding to all micro-periods are summed, and the total area obtained is the approximate total area under the power curve within that short-term period, which is also the total power of that period. Power allocation matrix The format is designed to cover the entire scheduling timeline, i.e. A complete framework for a short period of time, This represents the product of the total number of long-term periods and the total number of short-term periods. Its physical meaning is the total number of short-term periods within the entire scheduling window. ,in, This represents the total number of short-term periods, therefore one The matrix has exactly the following... Each element is then indexed by a short-term time period. This accurately "anchors" each power value to its long-term time period and its specific position in the global time series.
[0020] S2: Adjust the power allocation matrix according to the power continuity constraints to obtain the adjusted power allocation matrix; based on the adjusted power allocation matrix, use differential game theory to set the energy storage dynamic equation, and set the target payoff function according to the energy storage dynamic equation, and use a dynamic weight allocation mechanism to calculate the weight parameters of the target payoff function. The specific steps for obtaining the adjusted capacity allocation matrix are as follows: Iterate through each element of the power allocation matrix for each long time period, and adjust the power allocation matrix based on constraints set according to power continuity: in, Indicates a long period of time The actual total charging energy; This indicates the adjusted energy storage unit during the long-term period. Intra-short period The allocated power distribution value.
[0021] In the above process, through and By comparing these two thresholds, the system aims to ensure that, within each long-term period, the sum of all short-term power allocation values closely matches the total amount of energy actually available to the energy storage unit during that period. ,set up Right now to The system uses a tolerance range of several times, rather than requiring strict equality, to avoid frequent and unnecessary adjustments to the plan due to overly rigid plans or minor errors in sensor measurements, thus enhancing the system's robustness. At the same time, it allows for a certain degree of autonomous fluctuation in power allocation in the short term. When the total planned power exceeds this reasonable range, the system proportionally scales all short-term power values within that period. This is a fast, fair, and corrective method that maintains the shape of the power curve and the relative proportions of each period, ensuring that the corrected plan meets the total energy constraint without disrupting the overall coordination of the original multi-period power allocation strategy.
[0022] The specific steps for setting the energy storage dynamic equation using differential game theory are as follows: Based on the adjusted Differential game theory is used to set the dynamic equations for energy storage: in, in, express Real-time energy storage capacity of the energy storage unit; This represents the overall efficiency coefficient of the energy storage unit; express The instantaneous power applied to the energy storage unit at all times; This represents the equivalent self-attenuation coefficient of the energy storage unit; Represents the Dirac impulse function; Indicates the start time of the current window; This indicates the energy storage capacity of the initial energy storage unit.
[0023] In the above process, the power scheduling plan under discrete time scale is transformed into a precise mathematical model of a continuous-time, dynamically evolving energy storage system, thus providing a rigorous time-domain analysis basis for master-slave game optimization; First, by introducing differential equations, the continuous change of energy storage capacity over time can be accurately described, while simultaneously characterizing energy storage efficiency. and self-decay The influence of these two key dynamic characteristics on the state makes the model more closely resemble the actual system. Secondly, the Dirac impulse function is utilized. The adjusted energy storage unit will be used in the long term. Intra-short period Allocated power allocation value Transformed into a series of instantaneous power pulses acting at specific moments. By embedding discrete scheduling instructions into a continuous time frame, the connection between planning and dynamics is achieved; finally, the state is given. The analytical solution fully characterizes the explicit relationship between capacity and initial state, historical power command and system parameters at any time, providing a direct analytical tool for subsequent integral calculation and optimization of the revenue function. This process elevates the scheduling problem from static allocation to the dynamic system control level. exist In the calculation, the entire scheduling window Divided into Each of the three equal short-term periods corresponds to a discrete power allocation value. Each power value It is modeled as a Dirac pulse with a duration of That is, the moment when this short period begins, therefore Indicates any consecutive time points Only when A value of [value] will only occur when it is exactly equal to the start time of a certain short period of time. Instantaneous power; at other times The mathematical expression of 0 essentially transforms the average power scheduling plan on the discrete time scale into an idealized pulse sequence on the continuous time scale. exist In the calculation, the first part Indicates the initial time. Energy storage capacity As the remaining amount decays naturally over time, the decay rate changes from the equivalent self-decay coefficient. The decision exhibits exponential decay characteristics; Part Two This represents the cumulative contribution of all applied discrete power pulses to the current capacity during the scheduling period; each power pulse applied at a specific time... It will go through energy storage efficiency The conversion of the pulse, and similarly following the exponential decay law, affects the capacity at subsequent times. The magnitude of this effect depends on the time elapsed since the pulse's occurrence. The time difference formula clearly reveals the dynamic process of the continuous evolution of energy storage capacity over time. It is the result of the combined effects of the initial state, historical power dispatch commands, energy storage efficiency, and the system's self-degradation characteristics.
[0024] The specific steps for calculating the weight parameters of the target return function using a dynamic weight allocation mechanism are as follows: based on and The target return function is set according to the energy storage dynamic equation: in, This represents the target revenue function of the energy storage unit; The objective return function at time t represents The weighting parameters of the real-time electricity price, This represents the capacity fluctuation penalty coefficient; This represents the power fluctuation penalty coefficient; In the above process, by setting the target return function at time... Weighting parameters of real-time electricity price Designed with dynamic parameters, the energy storage dispatch system can respond sensitively to real-time electricity price signals, automatically favoring discharge for profit during peak electricity price periods and charging for energy storage during off-peak periods, thereby maximizing arbitrage opportunities. A quadratic penalty term is introduced for fluctuations in energy storage capacity and power, effectively suppressing drastic changes in state of charge and sharp increases and decreases in charging and discharging power. This not only ensures that the energy storage equipment operates within a safe operating range and delays aging, but also mitigates its power impact on the grid, improving grid-friendliness. By adjusting the coefficient... and The weight between economy and stability can be flexibly adjusted according to actual operational needs; When calculating the total revenue of an energy storage unit, its revenue function consists of two terms integrated over time. The first term... Represents electricity revenue, of which It is a real-time electricity price weight that changes over time, serving as an adjustable weight parameter for the objective revenue function. The first term is the instantaneous charging or discharging power of the energy storage unit; the product of the two reflects the direct economic benefits or costs gained through participation in the electricity market or response to electricity price signals. These are system operation penalty items, among which, To mitigate large fluctuations in energy storage capacity in order to maintain a stable state of charge and extend equipment lifespan, The penalty is to mitigate drastic changes in power to ensure a smooth charging and discharging process and reduce equipment stress; coefficient Control the overall intensity of punishment. The function adjusts the penalty weight of power fluctuation relative to capacity fluctuation. The purpose of this function is to achieve a trade-off between maximizing economic benefits and operational stability, guiding the dispatch strategy to pursue electricity price arbitrage while taking into account the safe, stable and durable operation of the energy storage system.
[0025] The weight parameters of the objective return function are calculated using a dynamic weight allocation mechanism. In the above process, an exponential decay term based on time distance is introduced. Assigning higher scheduling weights to recent time periods allows energy storage units to focus their power response more on current and nearby electricity price signals or grid demand, enhancing the real-time nature and responsiveness of the scheduling strategy. This is achieved by incorporating adjustment terms related to actual power state. When the power margin, i.e., the difference between the maximum allowable power and the actual power, is large, the weight is proactively increased to incentivize the system to participate more actively in power regulation when it is capable, thereby improving capacity utilization and market responsiveness; while when the power is close to the limit, the impact of this factor is reduced to avoid the risk of exceeding the limit due to over-scheduling; this mechanism together achieves agile focus in the time dimension and adaptive incentive in the power dimension. Dependent variable Directly reflects the scheduling decision at time. The intensity of economic incentives refers to the dynamically adjusted benefit weights that take into account both time urgency and power margin. Determined by multiple independent variables, the independent variable in the first term The distance between the current time and the starting time reflects the time decay effect; the closer the distance, the stronger the effect. The independent variable in the second term... The difference between the upper limit of power and the actual power reflects the power margin status; the larger the difference, the greater the available space. Distance in time They are negatively correlated, meaning the further away from the start time, the smaller the contribution of the time decay term; at the same time With power margin When the margin is positive, it shows a positive correlation, meaning that the larger the available power space, the greater the contribution of the power regulation term, thereby encouraging the system to schedule more actively within the safety margin. constitute The coefficients 0.7 and 0.3 in the polynomial are essentially weighted allocations for two influencing factors: the time decay term and the power margin term. Giving the time decay term a higher weight of 0.7 ensures the dispatch strategy has a clear time orientation, prioritizing responses to near-real-time demands, consistent with the general principle in power dispatch that "near-real-time decisions are more important." The power margin term, weighted at 0.3, serves as an important regulatory supplement, providing additional incentives when the system has sufficient regulation capacity, while preventing its over-dominance from causing the strategy to deviate from time optimality. The 0.15 in the denominator is a normalized design parameter; choosing 0.15... of As a benchmark, it is based on engineering experience with the safe operating range of typical energy storage systems or power equipment: it usually represents a reasonable power margin percentage threshold that is both regulatory and not too small to cause oversensitivity or too large to cause insufficient incentive. It ensures that when the actual power is close to the maximum limit, the contribution of this term can be smoothly reduced to zero, thereby achieving a soft constraint on the safety boundary; while when the margin is large, this term can provide significant but not excessive incentive.
[0026] S3: Calculate the deviation between the actual discharge power of the energy storage unit and the adjusted power allocation matrix in real time for each long period. If the deviation is not less than the set power difference threshold, reset the power allocation matrix and update the weight parameters of the target revenue function. When the power allocation matrix is less than the minimum allowable power or the rate of change of the target revenue function is greater than the preset threshold, output the power allocation matrix in real time.
[0027] The specific steps for calculating the deviation between the actual power of the energy storage unit and the total capacity over the current long-term period are as follows: Real-time calculation of the deviation between the actual power of the energy storage unit and the adjusted power allocation matrix for each long-term period: in, Indicates the first The absolute value of the deviation between the actual discharge power and the adjusted power distribution matrix over a long period; This indicates the current actual discharge power of the energy storage unit. In the above process, The calculation establishes a crucial closed-loop feedback monitoring mechanism. By comparing the real-time measured actual discharge power with the sum of the pre-optimized short-term power plan, the execution accuracy of the scheduling plan can be continuously and quantitatively evaluated, thereby promptly identifying discrepancies between the plan and reality caused by model errors, external disturbances, or changes in equipment status. Secondly, this deviation value... It directly serves as the trigger and adjustment basis for subsequent adaptive reset logic. When the deviation exceeds the threshold, the system will automatically start updating the weight parameters and power allocation matrix, enabling the scheduling strategy to dynamically respond to the actual operating status, thereby effectively improving the system's self-correction capability and overall robustness in the face of uncertainty. The specific logic for setting the power difference threshold is as follows: set the power difference threshold to the total power of the long-term period. When the power difference is not less than this power difference threshold, the system will automatically start updating the weight parameters and power allocation matrix. When the power difference is less than this power difference threshold, it is determined that the current scheduling strategy does not need further adjustment, and the current power allocation matrix is output in real time as the final scheduling instruction, and the optimization process of the current window ends. Set the power difference threshold to the total power over that long-term period. This is because in power system operation, many key performance indicators, such as frequency deviation and power control accuracy, are often set within a few percent of their rated values. Setting the deviation threshold to 5% of the planned value aligns with industry expectations for power control accuracy, ensuring that the performance objectives of this method are consistent with the actual operating standards of the power grid. The bias tolerance provides an "absence" for the optimization algorithm to fine-tune locally without immediately triggering a global reset, which helps the algorithm converge more smoothly to a robust solution.
[0028] Calculated deviation value This means that in actual operation, the energy storage unit, at the [number]th [time period], [is able to achieve its intended purpose]. The absolute difference between the total discharge power observed over a long period and the sum of the power allocation plans for that period generated in advance through optimized scheduling is determined by two core variables: one is the actual total discharge power obtained through real-time monitoring. It reflects the actual operating status of the energy storage unit; secondly, in the adjusted power allocation matrix, the corresponding long-term time period The sum of all short-term power allocation values represents the optimal scheduling plan target for that period; the deviation is the absolute value of the difference between these two variables; the dependent variable... Difference between the two independent variables The absolute values of the differences show a strict positive correlation, meaning the larger the absolute value of the difference, the larger the calculated power deviation. The specific steps for resetting the weight parameters of the target return function are as follows: when Greater than or equal to When this happens, the weight parameters of the objective return function are updated as follows: in, Indicates the number of consecutive resets; Indicates the first The adaptive attenuation coefficient for the next reset; Indicates the first The power sensitivity factor of the second reset; Indicates the first The time decay coefficient of the next reset; Indicates the first The threshold for the actual discharge power difference over a long period of time.
[0029] In the above process, the dependent variable This is the new weighting parameter obtained after resetting following the detection of a significant power deviation. Its meaning is that it increases with the number of consecutive resets. The adaptive dynamic incentive coefficient is used to adjust the strength of the economic benefit term in the benefit function. By making the structure and coefficient of the weight parameters change dynamically with the number of resets, the scheduling strategy can intelligently adjust its emphasis on time proximity and power margin after multiple deviations from the plan, thereby enhancing the system's learning and adaptive capabilities in a continuous uncertain environment. After resetting It is mainly affected by three independent variables: first, time distance. It represents the time difference between the current assessment time and the scheduling start time; secondly, it represents the power margin. This reflects the gap between the current actual operating power and the limit value; thirdly, the number of consecutive resets. It is a cumulative process state variable, and the influence of the independent variable on the dependent variable is specifically manifested in the time distance. Through coefficients The effect of the adjusted exponential decay term Power margin through coefficient Effect of the linear term of adjustment And the number of resets Then by dynamically changing the coefficients , and , Distance in time There is a negative correlation, meaning that the further the evaluation time is from the start time, the stronger the time decay term becomes. The smaller the contribution, With power margin A positive correlation means that the larger the available power space, the greater the effect of the power regulation term. The greater the contribution, the greater the initial amplitude coefficient of the time decay term. and The power sensitivity coefficient exhibits a negative correlation, meaning it decreases geometrically with increasing n. and time decay rate coefficient All with The positive correlation means that after repeated resets, the strategy gradually reduces its focus on recent time periods, while increasing its sensitivity to power margin and the rate of time decay.
[0030] With the number of resets The increase decreases geometrically by 0.8, and when the system repeatedly deviates from the original plan, that is... When the time is increased, it will actively reduce its dependence on the initial principle of "the most recent time has the highest priority". This will cause the system to reduce its path dependence on the recent time period after the original time-oriented strategy has failed multiple times. Instead, it will allocate more attention to a wider time range or other factors such as power status in order to seek a more robust scheduling scheme. When the value increases, it means that after experiencing multiple deviations, the system will pay more attention to the current real-time power margin information, increase its sensitivity to "available adjustability," and encourage or constrain its actions within the safety boundary. When this value increases, it means that the time decay term decays faster with distance, further weakening the weight of the distant time period; exist hour, , , This ensures the adjusted weight parameters During the initial correction, the time decay term and power margin term still have comparable magnitudes, avoiding abrupt policy changes or bias towards a certain extreme due to improper initial coefficient settings, thus maintaining the smoothness of the adjustment. The geometric decay base of 0.8 is a key convergence adjustment parameter. Choosing 0.8, which is less than 1 and close to 1, is to achieve a "mild but continuous" decay effect. It affects the amplitude coefficient of the time decay term. With the number of resets The influence of time factors decreases steadily as time increases, so that the influence does not disappear too early due to excessively rapid decay, nor does it cause the strategy adjustment to be too slow due to excessively slow decay. Each reset increases the value by 0.1, and Each reset increment increases the value by 0.2; these two linearly increasing steps control the "learning rate" of policy evolution. The 0.1 increment makes the policy more sensitive to power margins. The increase is made at a relatively gradual rate to ensure that the system can gradually improve its response to real-time conditions without overreacting. An increment of 0.2, which is greater than 0.1, will slow down the rate of time decay. Accelerating at a faster pace and more decisively shifting the focus of decision-making towards real-time conditions and near-term periods, this differentiated growth rate reflects design priorities: in the face of ongoing uncertainty, accelerating the move away from reliance on long-term fixed plans is more urgent than linearly increasing the focus on power margin.
[0031] The specific steps to reset the capacity allocation matrix are as follows: in, This indicates a reset over a long period of time. Intra-short period The allocated power distribution value.
[0032] The power allocation matrix is output in real time when the power allocation matrix is less than the minimum allowable power or the rate of change of the target revenue function is greater than the preset threshold.
[0033] The technical effect of resetting the capacity allocation matrix in the above process is to achieve a proportional dynamic scaling correction mechanism based on real-time feedback. When a significant deviation between the actual total discharge power and the planned value is detected over a long period, this method does not simply amortize the difference or crudely replace the original plan, but rather... The ratio of the planned total power for that period to the actual power allocation is used as a uniform scaling factor and linearly applied to each short-term power allocation value within that period. This maintains the relative proportion and shape of the original power allocation curve across different short periods. In other words, the correction is performed without altering the internal structure of the original optimized scheduling scheme, avoiding the introduction of new inconsistencies due to local adjustments. Secondly, by scaling the entire system proportionally, the planned total power can be quickly and smoothly aligned with the actual execution result, ensuring consistency between the planned and actual total energy. Finally, this correction method is simple to calculate and responds quickly, providing the system with an instantly updated and internally self-consistent power reference trajectory for the next scheduling cycle or real-time control, thereby effectively improving the tracking accuracy of the scheduling strategy and the stability of closed-loop control.
[0034] Dependent variable Specifically, this reflects the first During a long-term period, when a significant power deviation is detected, the original short-term power allocation plan will be adjusted. The final power allocation value obtained after reset and correction. It uses a proportional scaling mechanism to quickly and smoothly calibrate the planned total power to the actual execution result while maintaining the relative proportions of power values in each short period of the original plan (i.e., the shape of the power curve). This ensures consistency between the scheduling plan and the actual total energy output and provides an instantly updated and internally coordinated power reference trajectory for the next cycle; dependent variable It is directly determined by two independent variables: one is the original adjusted power allocation value for this short period of time. It represents the initial plan based on the model and optimization calculations; secondly, it represents the total planned power for that period. Deviation from actual execution The relative ratio quantitatively reflects the degree of deviation between the planned total and the actual total; the dependent variable Compared with the original planned value There is a positive correlation, meaning that the larger the original planned value, the larger the reset value tends to be. Ratio of deviation The scaling factor exhibits a negative correlation, meaning that when the actual total discharge power is less than the planned total, the scaling factor is less than 1, leading to... Less than All short-term planned values will be adjusted downwards proportionally; conversely, they will be adjusted upwards proportionally.
[0035] The rate of change of the objective return function is calculated by comparing the total return function calculated in two adjacent short-term decision periods. The value of the change rate is determined by taking the ratio of its difference to the time interval. This change rate essentially characterizes the convergence speed of the optimization process or the marginal benefit of the profit improvement. The determination of the minimum allowable power is usually based on the technical characteristics of the energy storage unit itself, such as the minimum stable charging and discharging power, the lower limit of sensor accuracy, and the stable operation requirements of the power grid, such as avoiding oscillations caused by frequent fine-tuning. It is a pre-configured fixed parameter or a safety lower limit that is dynamically adjusted according to the operating status. The setting of the preset threshold depends on the specific optimization goal and engineering experience. The threshold used for the profit change rate is intended to determine whether the strategy has become stable, i.e., the change rate is small enough, or whether it has fallen into local optimization, i.e., the change rate has not improved significantly in a long period of time. This requires calibrating an empirical value that can balance the convergence speed and optimization accuracy through historical data simulation or system trial operation.
[0036] In the above embodiments, the first... The absolute value of the deviation between the actual discharge power and the adjusted power distribution matrix over a long period and the long-term discharge power. Intra-short period The 20 sets of power allocation values are used to reflect the changes in power allocation values as the absolute value of the deviation changes, as shown in Table 1: Table 1: Relationship between absolute value of deviation and power distribution value In Table 1 above, the first preset is... , By changing the power deviation value, you can see It decreases as the absolute value of the deviation increases, which is consistent with the characteristics of actual power distribution.
[0037] Please see Figure 3 The present invention also provides a distributed energy storage and power scheduling device based on master-slave game theory, wherein the scheduling device is used to execute the above-mentioned scheduling method, including: The energy storage planning module is used to obtain the power allocation curve of the energy storage unit. For the energy storage unit, the time window is divided into N long-term periods of equal length. The power data allocated to each long-term period is determined. Each long-term period is divided into M short-term periods, and the power data allocated to each short-term period is extracted. The average power of each short-term period is calculated as the power allocation value of that short-term period. The power allocation matrix is constructed by traversing the power data allocated to all short-term periods. The game optimization module is used to adjust the power allocation matrix based on the power allocation matrix and set constraints according to power continuity to obtain the adjusted power allocation matrix; based on the optimized power allocation matrix, the energy storage dynamic equation is set using differential game theory, and the target payoff function is set according to the energy storage dynamic equation. The weight parameters of the target payoff function are calculated using a dynamic weight allocation mechanism. The dynamic feedback module is used to calculate in real time the deviation between the actual discharge power of the energy storage unit and the adjusted power allocation matrix for each long-term period. If the deviation is not less than the set power difference threshold, the power allocation matrix is reset and the weight parameters of the target revenue function are updated. When the power allocation matrix is less than the minimum allowable power or the rate of change of the target revenue function is greater than the preset threshold, the power allocation matrix is output in real time.
[0038] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0039] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented in software, the above embodiments can be implemented, in whole or in part, as a computer program product. Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution.
[0040] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0041] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A distributed energy storage and power scheduling method based on master-slave game theory, characterized by the following steps: include: S1: Obtain the power distribution curve of the energy storage unit, and for the energy storage unit, divide it into equal lengths within a set time window. Each long-term period is divided into several long-term periods, and the power data allocated to each long-term period is determined accordingly. The system extracts power data allocated within each short-term period, calculates the average power of each short-term period as the power allocation value for that short-term period, and constructs a power allocation matrix by iterating through the power allocation values allocated in all short-term periods. S2: Adjust the power allocation matrix according to the power continuity constraints to obtain the adjusted power allocation matrix; based on the adjusted power allocation matrix, use differential game theory to set the energy storage dynamic equation, and set the target payoff function according to the energy storage dynamic equation, and use a dynamic weight allocation mechanism to calculate the weight parameters of the target payoff function. S3: Calculate the deviation between the actual discharge power of the energy storage unit and the adjusted power allocation matrix in real time for each long period. If the deviation is not less than the set power difference threshold, reset the power allocation matrix and update the weight parameters of the target revenue function. When the power allocation matrix is less than the minimum allowable power or the rate of change of the target revenue function is greater than the preset threshold, output the power allocation matrix in real time.
2. The distributed energy storage and power scheduling method based on master-slave game theory according to claim 1, characterized in that, The specific steps for constructing the power allocation matrix are as follows: Define the time window length as Divide the time window into equal parts There are several long-term periods, and each long-term period is numbered as follows: ,in, The index number represents the long-term time period, and each long-term time period is further divided into... There are several short-term periods, and all short-term periods are numbered as follows: ,in, ,in, Indicates the index number for a short-term period; The specific logic for calculating the average power of each short-term period is as follows: determine the start and end time points of each short-term period, perform dense sampling on them using sensors, calculate the corresponding area under the power curve of the short-term period using the rectangular area method, take the area value as the total power of the short-term period, divide the total power by the length of the start time period to obtain the average power of the corresponding short-term period, and take it as the power allocation value of the short-term period. This process is repeated to obtain the power allocation value of all short-term periods. Divide all power allocation values into OK Power allocation matrix of columns ,in, This represents the product of the total number of long-term periods and the total number of short-term periods. Indicates the energy storage unit in the first Within a long-term period, the first Power allocation values for short-term periods Indicates the long-term time period index number. Indicates the index number for a short period of time.
3. The distributed energy storage and power scheduling method based on master-slave game theory according to claim 2, characterized in that, The specific steps for obtaining the adjusted power allocation matrix are as follows: Iterate through each element of the power allocation matrix for each long time period, and adjust the power allocation matrix based on constraints set according to power continuity: in, Indicates a long period of time The actual total charging energy; This indicates the adjusted energy storage unit during the long-term period. Intra-short period The allocated power distribution value, Indicates the energy storage unit in the first The sum of the power allocation values for all short periods within a long-term period.
4. The distributed energy storage and power scheduling method based on master-slave game theory according to claim 1, characterized in that, The specific steps for setting the energy storage dynamic equation using differential game theory are as follows: Based on the adjusted Differential game theory is used to set the dynamic equations for energy storage: in, in, express Real-time energy storage capacity of the energy storage unit; This represents the overall efficiency coefficient of the energy storage unit; express The instantaneous power applied to the energy storage unit at all times; This represents the equivalent self-attenuation coefficient of the energy storage unit; Represents the Dirac impulse function; Indicates the start time of the current window; This indicates the energy storage capacity of the initial energy storage unit; Indicates the length of the time window; Indicates the total number of items in a long-term period; This represents the product of the total number of long-term periods and the total number of short-term periods. This indicates the index number for a short-term period.
5. The distributed energy storage and power scheduling method based on master-slave game theory according to claim 1, characterized in that, The specific steps for calculating the weight parameters of the target return function using a dynamic weight allocation mechanism are as follows: based on and The target return function is set according to the energy storage dynamic equation: in, This represents the target revenue function of the energy storage unit; The objective return function at time t represents The weighting parameters of the real-time electricity price; This represents the capacity fluctuation penalty coefficient; This represents the power fluctuation penalty coefficient; The weight parameters of the objective return function are calculated using a dynamic weight allocation mechanism. in, Indicates short period of time Actual dispatch power; This indicates the maximum power limit allowed within a short period of time.
6. The distributed energy storage and power scheduling method based on master-slave game theory according to claim 1, characterized in that, The specific steps for calculating the deviation between the actual discharge power of the energy storage unit and the adjusted power allocation matrix for each long-term period in real time are as follows: in, Indicates the first The absolute value of the deviation between the actual discharge power and the adjusted power distribution matrix over a long period; This indicates the current actual discharge power of the energy storage unit.
7. The distributed energy storage and power scheduling method based on master-slave game theory according to claim 6, characterized in that, The specific steps to update the weight parameters of the target return function are as follows: when Greater than or equal to When this happens, the weight parameters of the objective return function are updated as follows: in, Indicates the number of consecutive resets; Indicates the first The adaptive attenuation coefficient for the next reset; Indicates the first The power sensitivity factor of the second reset; Indicates the first The time decay coefficient of the next reset; Indicates the first The threshold for the actual discharge power difference over a long period of time.
8. The distributed energy storage and power scheduling method based on master-slave game theory according to claim 6, characterized in that, To reset the power allocation matrix, the specific steps are as follows: in, This indicates a reset over a long period of time. Intra-short period The allocated power distribution value; The power allocation matrix is output in real time when the power allocation matrix is less than the minimum allowable power or the rate of change of the target revenue function is greater than the preset threshold.
9. A distributed energy storage and power dispatching device based on master-slave game theory, characterized in that: The distributed energy storage and power scheduling device based on master-slave game theory is used to execute the scheduling method according to any one of claims 1-8, including: The energy storage planning module is used to obtain the power allocation curve of the energy storage unit. For the energy storage unit, the time window is divided into N long-term periods of equal length. The power data allocated to each long-term period is determined. Each long-term period is divided into M short-term periods, and the power data allocated to each short-term period is extracted. The average power of each short-term period is calculated as the power allocation value of that short-term period. The power allocation matrix is constructed by traversing the power data allocated to all short-term periods. The game optimization module is used to adjust the power allocation matrix based on the power allocation matrix and set constraints according to power continuity to obtain the adjusted power allocation matrix; based on the optimized power allocation matrix, the energy storage dynamic equation is set using differential game theory, and the target payoff function is set according to the energy storage dynamic equation. The weight parameters of the target payoff function are calculated using a dynamic weight allocation mechanism. The dynamic feedback module is used to calculate in real time the deviation between the actual discharge power of the energy storage unit and the adjusted power allocation matrix for each long-term period. If the deviation is not less than the set power difference threshold, the power allocation matrix is reset and the weight parameters of the target revenue function are updated. When the power allocation matrix is less than the minimum allowable power or the rate of change of the target revenue function is greater than the preset threshold, the power allocation matrix is output in real time.
Citation Information
Patent Citations
Distributed energy storage and power scheduling method and device based on master-slave game
CN116031948A