Virtual power plant backup capacity optimal configuration management method and system
By using clustering and grouping based on response delay and energy state, and a two-layer optimization framework, combined with a multi-agent non-cooperative game model within the cluster, the complexity of virtual power plant reserve capacity configuration is solved, achieving efficient, economical and reliable resource management and improving the safety and stability of the power system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING TRUTH WISDOM POWER TECH CO LTD
- Filing Date
- 2026-02-03
- Publication Date
- 2026-07-07
AI Technical Summary
Traditional methods for configuring reserve capacity in virtual power plants are difficult to optimize effectively when dealing with large-scale and diverse distributed energy units. This leads to difficulties in managing response characteristics and energy state differences caused by resource heterogeneity, which affects the safe and stable operation of the power system.
By clustering based on response delay and energy state, a state clustering mapping table is constructed. A two-layer optimization framework and a non-cooperative game model of multiple agents within the cluster are adopted to optimize the reserve capacity configuration of distributed energy units. The KKT multiplier feedback mechanism is combined to achieve coordination between the upper and lower layer optimizations.
It enables efficient management of a large number of heterogeneous resources within a virtual power plant, reduces system computational complexity, improves the economy of resource allocation and the accuracy and reliability of allocation schemes, and balances overall benefits with local cost control.
Smart Images

Figure CN122026478B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power system management technology, and in particular to a method and system for optimizing the configuration and management of virtual power plant reserve capacity. Background Technology
[0002] With the large-scale integration of renewable energy and the widespread application of distributed energy, the operating characteristics of power systems have undergone profound changes. Virtual power plants, as a new form of energy aggregation, integrate and dispatch a large number of dispersed distributed energy units through information and communication technologies, and have become an important means to promote renewable energy consumption and improve the flexibility of power systems. Especially against the backdrop of ever-increasing grid-side reserve capacity demands, virtual power plants can effectively support the safe and stable operation of power systems by aggregating different types of distributed energy units to provide backup services.
[0003] Traditional methods for configuring reserve capacity in virtual power plants typically employ a centralized optimization approach. This involves determining the reserve capacity configuration for each distributed energy unit by solving an overall optimization problem, based on the unit's technical characteristics and cost parameters. However, with the expansion of virtual power plant scale and the diversification of distributed energy types, these resources exhibit significant heterogeneity in response characteristics and energy state, posing new challenges to the optimal configuration of virtual power plant reserve capacity. Summary of the Invention
[0004] This invention provides a method and system for optimizing the configuration and management of virtual power plant reserve capacity, which can solve the problems in the prior art.
[0005] A first aspect of the present invention provides a method for optimizing the configuration and management of virtual power plant reserve capacity, comprising:
[0006] Obtain the adjustable capacity, response latency, and energy status of each distributed energy unit within the virtual power plant, as well as the standby capacity demand information on the grid side;
[0007] Based on the response delay and energy state, each distributed energy unit is clustered in the two-dimensional space of response delay-energy state to obtain multiple state clusters. The centroid coordinates and the number of resources within each state cluster are calculated to construct a state cluster mapping table.
[0008] Based on the state cluster mapping table, a two-layer optimization framework is constructed. The upper layer optimization aims to maximize the operating revenue of the virtual power plant and optimizes the allocation target amount of the reserve capacity of each state cluster. The lower layer optimization aims to minimize the intra-cluster response cost of each state cluster and optimizes the actual configuration amount.
[0009] For each state cluster, based on the allocation target quantity and the centroid coordinates, a multi-agent non-cooperative game model within the cluster is constructed. The Nash equilibrium solution is obtained by solving the optimal response function of each game agent, and the initial configuration quantity of each distributed energy unit is calculated based on the equilibrium solution.
[0010] Substitute the initial configuration amount into the lower-level optimization, extract the KKT multipliers as feedback signals to adjust the allocation target amount of the upper-level optimization, and re-execute the intra-cluster game decomposition until the deviation between the upper and lower-level configuration amounts converges to obtain the final spare capacity configuration scheme.
[0011] Based on the final backup capacity configuration scheme, backup capacity configuration instructions are issued to each distributed energy unit.
[0012] Based on the response delay and energy state, distributed energy units are clustered in the two-dimensional space of response delay-energy state to obtain multiple state clusters. The centroid coordinates and the number of resources within each state cluster are calculated, and a state cluster mapping table is constructed, including:
[0013] Using the response delay of each distributed energy unit as the first dimension coordinate and the energy state as the second dimension coordinate, the spatial location of each distributed energy unit is determined in the two-dimensional space of response delay-energy state.
[0014] A density- and distance-based dual-constraint clustering algorithm is used to iteratively group the spatial location points. In each iteration, the weighted distance from each spatial location point to the current centroid of each cluster is calculated. The weighted distance includes time similarity weight and capacity similarity weight.
[0015] Based on the weighted distance, each distributed energy unit is assigned to a corresponding state cluster, and the centroid coordinates of each state cluster are recalculated. When the change in the centroid coordinates of each state cluster in continuous iterations is less than the convergence threshold, the iteration is terminated, resulting in the final multiple state clusters.
[0016] For each state cluster, the mean coordinates of the distributed energy units within the cluster are calculated and used as the centroid coordinates. The number of units within the cluster is counted, and the association between the centroid coordinates, the number of units within the cluster, and the identifiers of each distributed energy unit within the cluster is established to construct a state cluster mapping table.
[0017] A two-layer optimization framework is constructed based on the aforementioned state cluster mapping table. The upper-layer optimization aims to maximize the operating revenue of the virtual power plant by optimizing the target allocation of reserve capacity for each state cluster. The lower-layer optimization aims to minimize the intra-cluster response cost of each state cluster by optimizing the actual configuration, including:
[0018] Based on the state cluster mapping table, a response delay-energy state coupling matrix is constructed. The centroid coordinates are processed by Cartesian product to obtain the coupling feature vector. The inner product of the feature vector and the reserve capacity requirement is calculated to obtain the demand matching degree.
[0019] Using the allocation target amount of the reserve capacity of each state cluster as the decision variable, an upper-level optimization model is constructed. The upper-level optimization objective function includes a deviation measure term of the ratio between the allocation target amount and the demand matching degree, and a boundary jump constraint term of the allocation target amount between adjacent clusters. The constraint condition is that the total allocation amount meets the reserve demand and the allocation amount of each cluster is non-negative.
[0020] For each state cluster, based on its allocation target quantity and centroid coordinates, a cluster content allocation benchmark function is constructed to map the response delay characteristics and energy state of each distributed energy unit into a capacity allocation benchmark value.
[0021] A lower-level optimization model is constructed, with the actual configuration amount of each distributed energy unit within the cluster as the decision variable. The lower-level optimization objective function includes a deviation term between the actual configuration amount and the capacity allocation benchmark value, as well as a coupling term for response cost, and ensures that the total configuration amount within the cluster is equal to the allocation target amount and does not exceed the adjustable capacity.
[0022] For each state cluster, based on the allocation target quantity and the centroid coordinates, a multi-agent non-cooperative game model within the cluster is constructed. The Nash equilibrium solution is obtained by solving the optimal response function of each player, and the initial configuration quantity of each distributed energy unit is calculated based on the equilibrium solution, including:
[0023] For each state cluster, each distributed energy unit within the cluster is taken as the main player in the game. Based on the allocation target quantity and the centroid coordinates, the delay competition factor and capacity competition factor are calculated, and tensor product operation is performed to obtain the two-dimensional competition intensity matrix of each state cluster.
[0024] A payoff function is established for each game player. The payoff function includes a product term of allocation amount and unit capacity payoff, and a bilinear competition loss term based on a two-dimensional competition intensity matrix. The partial derivative of the payoff function with respect to the allocation amount of each game player is calculated and set to zero to obtain the optimal response function of each game player relative to the allocation amount of other game players.
[0025] The optimal response functions of each game player are combined to form a system of nonlinear equations. The allocation target quantity is introduced as a total constraint. By iteratively solving the system of nonlinear equations, the Nash equilibrium solution that satisfies the simultaneous validity of the optimal response functions of all game players is obtained.
[0026] Extract each component from the Nash equilibrium solution and use each component as the initial configuration of each distributed energy unit within the state cluster.
[0027] The optimal response functions of each player are combined to form a system of nonlinear equations. The allocation objective is introduced as a total constraint. By iteratively solving this system of nonlinear equations, a Nash equilibrium solution satisfying the simultaneous validity of the optimal response functions of all players is obtained, including:
[0028] The optimal response functions of each game subject are combined to form a nonlinear equation system. The coupling coefficient matrix of the configuration quantity is extracted. Based on the spectral radius and condition number of the two-dimensional competition intensity matrix, a precondition transformation is performed to obtain a positive definite symmetric preprocessing matrix.
[0029] An augmented Lagrangian function is constructed using the allocation target quantity as a constraint. The augmented Lagrangian function includes a residual term of the nonlinear equation system, a quadratic penalty term weighted by the preprocessing matrix, and a linear coupling term between the Lagrange multiplier and the constraint deviation.
[0030] The augmented Lagrangian function is decomposed into a configuration quantum problem and a multiplier update subproblem using the alternating direction multiplier method. The inverse of the preprocessing matrix is introduced as the descent direction correction operator, and the uniform distribution is set as the initial value for iteration.
[0031] During the iteration process, the configuration quantum problem is solved sequentially to obtain updated values. The Lagrange multipliers are updated based on the deviation between the updated values and the target quantity, and the residual norm of the nonlinear equation system is calculated. When it is less than the residual threshold and the updated configuration quantity satisfies the total quantity constraint, the current configuration quantity is output as the Nash equilibrium solution.
[0032] Substituting the initial configuration amount into the lower-level optimization, extracting the KKT multipliers as feedback signals to adjust the allocation target amount of the upper-level optimization, and re-executing the intra-cluster game decomposition until the deviation between the upper and lower-level configuration amounts converges, the final spare capacity configuration scheme is obtained, including:
[0033] Substitute the initial allocation amount into the lower-level optimization objective function, construct the KKT optimality condition equation system in combination with the constraint conditions, solve to obtain the KKT multipliers corresponding to the allocation objective amount constraint, and construct a sensitivity vector characterizing the marginal response strength of the actual allocation amount based on the sign and value of the KKT multipliers.
[0034] The sensitivity vector and the gradient vector of the upper-level optimization objective function are multiplied to obtain the adjustment direction factor. The adaptive step size is determined based on its magnitude. The allocation objective quantity of the upper-level optimization is corrected to obtain the corrected allocation objective quantity. This is then substituted into the multi-agent non-cooperative game model within the cluster, and game decomposition is performed to obtain the updated initial configuration quantity. This is then substituted into the lower-level optimization to calculate the new actual configuration quantity.
[0035] Calculate the deviation vector norm between the updated initial configuration and the new actual configuration. When the deviation is less than a preset threshold, output the new actual configuration as the final solution. Otherwise, repeat the iterative process of KKT multiplier extraction, sensitivity calculation and target quantity correction based on the new actual configuration until convergence.
[0036] Substituting the initial allocation amount into the lower-level optimization objective function, and combining it with the constraints, a set of KKT optimality condition equations is constructed. Solving these equations yields the KKT multipliers corresponding to the allocation objective constraints. Based on the signs and values of these KKT multipliers, a sensitivity vector characterizing the marginal response strength of the actual allocation amount is constructed, including:
[0037] Substitute the initial configuration into the lower-level optimization objective function, construct an augmented objective functional through a penalty term to soften the inequality constraints, solve the first-order necessary conditions and combine them with complementary relaxation conditions to construct the KKT optimality equation system.
[0038] The KKT optimality condition equations are linearized and decomposed to extract the sensitive submatrix corresponding to the allocation objective constraint. This submatrix is then decomposed into a product of an orthogonal matrix and an upper triangular matrix. The KKT multipliers of each cluster are obtained by back substitution.
[0039] Sign discrimination is performed on the KKT multipliers to obtain the indicator vector. Logarithmic transformation is performed on its absolute value to obtain the magnitude vector. Element-wise multiplication is performed to obtain the logarithmic domain sensitivity vector. The logarithmic domain sensitivity vector is nonlinearly scaled and mapped through inverse exponential transformation to obtain the sensitivity vector characterizing the marginal response strength of the actual configuration quantity.
[0040] A second aspect of the present invention provides a virtual power plant reserve capacity optimization configuration management system, comprising:
[0041] The resource information acquisition unit is used to acquire the adjustable capacity, response delay and energy status of each distributed energy unit in the virtual power plant, as well as the standby capacity demand information on the grid side.
[0042] The state space clustering unit is used to cluster and group each distributed energy unit in the two-dimensional space of response delay and energy state based on the response delay and energy state, to obtain multiple state clusters, and to calculate the centroid coordinates and the number of resources in each state cluster, and to construct a state cluster mapping table.
[0043] The two-layer optimization decision unit is used to construct a two-layer optimization framework based on the state cluster mapping table. The upper layer optimization aims to maximize the operating revenue of the virtual power plant and optimizes the allocation target amount of the reserve capacity of each state cluster. The lower layer optimization aims to minimize the intra-cluster response cost of each state cluster and optimizes the actual configuration amount.
[0044] The intra-cluster game equilibrium unit is used to construct an intra-cluster multi-agent non-cooperative game model for each state cluster based on the allocation target quantity and the centroid coordinates. The Nash equilibrium solution is obtained by solving the optimal response function of each game agent, and the initial configuration quantity of each distributed energy unit is calculated based on the equilibrium solution.
[0045] The iterative convergence coordination unit is used to substitute the initial configuration amount into the lower-level optimization, extract the KKT multipliers as feedback signals to adjust the allocation target amount of the upper-level optimization, and re-execute the intra-cluster game decomposition until the deviation between the upper and lower-level configuration amounts converges to obtain the final spare capacity configuration scheme.
[0046] The instruction distribution and execution unit is used to issue backup capacity configuration instructions to each distributed energy unit based on the final backup capacity configuration scheme.
[0047] A third aspect of the embodiments of the present invention,
[0048] An electronic device is provided, comprising:
[0049] processor;
[0050] Memory used to store processor-executable instructions;
[0051] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0052] Fourth aspect of the present invention,
[0053] A computer-readable storage medium is provided, having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0054] The beneficial effects of this application are as follows:
[0055] The virtual power plant standby capacity optimization configuration management method provided by this invention achieves efficient management of a large number of heterogeneous resources in the virtual power plant by clustering distributed energy units based on response delay and energy status and constructing a state cluster mapping table, which effectively reduces the computational complexity of the system.
[0056] Based on the two-layer optimization framework, the upper layer focuses on maximizing the operating revenue of the virtual power plant, while the lower layer minimizes the response cost within each cluster. This makes the optimization process more in line with the actual operating logic, taking into account both overall revenue and local cost control, and improving the economy of resource allocation.
[0057] By introducing a multi-agent non-cooperative game model within the cluster and solving for Nash equilibrium, the autonomous decision-making characteristics of each distributed energy unit are fully considered, making the configuration scheme more fair and reasonable. At the same time, the KKT multiplier feedback mechanism ensures the coordination and consistency of optimization between upper and lower layers, improving the accuracy and reliability of reserve capacity configuration. Attached Figure Description
[0058] Figure 1 This is a flowchart illustrating the virtual power plant standby capacity optimization configuration management method according to an embodiment of the present invention;
[0059] Figure 2 This is a schematic diagram of the process for constructing a state clustering mapping table according to an embodiment of the present invention. Detailed Implementation
[0060] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0061] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0062] Figure 1 This is a flowchart illustrating the virtual power plant reserve capacity optimization configuration management method according to an embodiment of the present invention, as shown below. Figure 1 As shown, the method includes:
[0063] Obtain the adjustable capacity, response latency, and energy status of each distributed energy unit within the virtual power plant, as well as the standby capacity demand information on the grid side;
[0064] Based on the response delay and energy state, each distributed energy unit is clustered in the two-dimensional space of response delay-energy state to obtain multiple state clusters. The centroid coordinates and the number of resources within each state cluster are calculated to construct a state cluster mapping table.
[0065] Based on the state cluster mapping table, a two-layer optimization framework is constructed. The upper layer optimization aims to maximize the operating revenue of the virtual power plant and optimizes the allocation target amount of the reserve capacity of each state cluster. The lower layer optimization aims to minimize the intra-cluster response cost of each state cluster and optimizes the actual configuration amount.
[0066] For each state cluster, based on the allocation target quantity and the centroid coordinates, a multi-agent non-cooperative game model within the cluster is constructed. The Nash equilibrium solution is obtained by solving the optimal response function of each game agent, and the initial configuration quantity of each distributed energy unit is calculated based on the equilibrium solution.
[0067] Substitute the initial configuration amount into the lower-level optimization, extract the KKT multipliers as feedback signals to adjust the allocation target amount of the upper-level optimization, and re-execute the intra-cluster game decomposition until the deviation between the upper and lower-level configuration amounts converges to obtain the final spare capacity configuration scheme.
[0068] Based on the final backup capacity configuration scheme, backup capacity configuration instructions are issued to each distributed energy unit.
[0069] Figure 2 This is a schematic diagram of the process for constructing a state cluster mapping table according to an embodiment of the present invention. In an optional implementation, based on the response delay and energy state, each distributed energy unit is clustered in a two-dimensional space of response delay-energy state to obtain multiple state clusters. The centroid coordinates and the number of resources within each state cluster are calculated to construct a state cluster mapping table, including:
[0070] Using the response delay of each distributed energy unit as the first dimension coordinate and the energy state as the second dimension coordinate, the spatial location of each distributed energy unit is determined in the two-dimensional space of response delay-energy state.
[0071] A density- and distance-based dual-constraint clustering algorithm is used to iteratively group the spatial location points. In each iteration, the weighted distance from each spatial location point to the current centroid of each cluster is calculated. The weighted distance includes time similarity weight and capacity similarity weight.
[0072] Based on the weighted distance, each distributed energy unit is assigned to a corresponding state cluster, and the centroid coordinates of each state cluster are recalculated. When the change in the centroid coordinates of each state cluster in continuous iterations is less than the convergence threshold, the iteration is terminated, resulting in the final multiple state clusters.
[0073] For each state cluster, the mean coordinates of the distributed energy units within the cluster are calculated and used as the centroid coordinates. The number of units within the cluster is counted, and the association between the centroid coordinates, the number of units within the cluster, and the identifiers of each distributed energy unit within the cluster is established to construct a state cluster mapping table.
[0074] In this specific implementation, real-time data of each distributed energy unit needs to be collected, including response latency and energy status. Response latency refers to the time from when an energy unit receives a dispatch command to when it actually responds, in seconds. Energy status refers to the percentage of the energy unit's current available capacity relative to its total capacity, in percentage. For example, for a photovoltaic power generation unit DER-001, its response latency is 2.5 seconds and its current energy status is 85%; for an energy storage unit DER-002, its response latency is 0.8 seconds and its energy status is 60%; for a controllable load DER-003, its response latency is 3.2 seconds and its energy status is 70%.
[0075] After data collection, these data are mapped onto a two-dimensional space. Each distributed energy unit is represented as a spatial location point in this two-dimensional space, with the horizontal axis representing the response delay value and the vertical axis representing the energy state value. Taking the three energy units mentioned above as examples, the spatial location point coordinates of DER-001 are (2.5, 85), those of DER-002 are (0.8, 60), and those of DER-003 are (3.2, 70). Assuming there are a total of 30 distributed energy units, 30 spatial location points will be formed in the two-dimensional space.
[0076] A density- and distance-based dual-constraint clustering algorithm is used to group these spatial locations. The algorithm randomly selects K initial cluster centers (the value of K can be preset according to the system size and distribution characteristics; in this example, K=4) and performs iterative clustering. In each iteration, the weighted distance from each spatial location point to the current centroid of each cluster is calculated.
[0077] The weighted distance calculation takes into account both temporal similarity weight and capacity similarity weight. Specifically, for spatial location point i and cluster centroid j, the weighted distance is the product of the square of the time dimension difference and the temporal similarity weight, plus the product of the square of the energy state dimension difference and the capacity similarity weight, and then the square root of the result. In this embodiment, the temporal similarity weight is set to 0.6 and the capacity similarity weight is set to 0.4 to reflect the relative importance of response delay to scheduling decisions.
[0078] For example, suppose the coordinates of the four cluster centroids in the current iteration are C1(1.5, 75), C2(4.0, 65), C3(2.0, 90), and C4(0.7, 55). For the energy unit DER-001, the weighted distances to each centroid are calculated as follows: Weighted distance to C1: (2.5-1.5) 2 ×0.6+(85-75) 2 ×0.4 square root, approximately 6.37; weighted distance to C2: (2.5-4.0) 2 ×0.6+(85-65) 2×0.4 square root, approximately 12.70; weighted distance to C3: (2.5-2.0) 2 ×0.6+(85-90) 2 ×0.4 square root, approximately 3.18; weighted distance to C4: (2.5-0.7) 2 ×0.6+(85-55) 2 The result is approximately 19.02 (×0.4 square root). Based on the calculation results, DER-001 was assigned to the C3 cluster with the smallest distance, and the same calculation and assignment process was performed on all energy units.
[0079] After completing one round of allocation, the centroid coordinates of each cluster are recalculated. For a cluster containing multiple energy units, the new centroid coordinates are the arithmetic mean of all units in the cluster in each dimension. For example, if after allocation, cluster C3 contains three units: DER-001 (2.5, 85), DER-004 (1.8, 92), and DER-007 (2.2, 88), then the new centroid coordinates of C3 are ((2.5+1.8+2.2) / 3, (85+92+88) / 3), which is (2.17, 88.33).
[0080] Continue the iterative process, and check the change in centroid coordinates after each iteration. When the change in centroid coordinates of all clusters in two consecutive iterations is less than the preset convergence threshold (set to 0.01 in this example), terminate the iteration. Assume that after 12 iterations, the algorithm converges and obtains the final 4 state clusters.
[0081] For each final state cluster, calculate and record the following information: the centroid coordinates of the cluster, i.e., the average value of all energy units in the cluster in the two dimensions of response delay and energy state, the number of energy units contained in the cluster, and the list of identifiers of each energy unit in the cluster.
[0082] For example, the final clustering results are as follows: Cluster 1: Centroid coordinates (1.2, 80), containing 8 units, identifier list [DER-005, DER-008, DER-012, DER-015, DER-019, DER-022, DER-026, DER-029]; Cluster 2: Centroid coordinates (3.8, 65), containing 7 units, identifier list [DER-003, DER-006, DER-010, DER-014, DER-020, DER-024, DE...]. [R-028]; Cluster 3: Centroid coordinates (2.1, 88), containing 9 units, identifier list [DER-001, DER-004, DER-007, DER-011, DER-016, DER-021, DER-025, DER-027, DER-030]; Cluster 4: Centroid coordinates (0.7, 58), containing 6 units, identifier list [DER-002, DER-009, DER-013, DER-017, DER-018, DER-023]
[0083] This information is organized into a state cluster mapping table for subsequent energy scheduling and optimization decisions. This table contains the centroid coordinates (mean response delay and mean energy state) of each cluster, the number of energy units within the cluster, and a detailed list of unit identifiers. In this way, the energy management system can quickly identify groups of energy units with similar response characteristics and energy states, enabling more efficient hierarchical scheduling and optimization control.
[0084] In one optional implementation, a two-layer optimization framework is constructed based on the state cluster mapping table. The upper-layer optimization aims to maximize the operating revenue of the virtual power plant by optimizing the target allocation of reserve capacity for each state cluster. The lower-layer optimization aims to minimize the intra-cluster response cost of each state cluster by optimizing the actual configuration, including:
[0085] Based on the state cluster mapping table, a response delay-energy state coupling matrix is constructed. The centroid coordinates are processed by Cartesian product to obtain the coupling feature vector. The inner product of the feature vector and the reserve capacity requirement is calculated to obtain the demand matching degree.
[0086] Using the allocation target amount of the reserve capacity of each state cluster as the decision variable, an upper-level optimization model is constructed. The upper-level optimization objective function includes a deviation measure term of the ratio between the allocation target amount and the demand matching degree, and a boundary jump constraint term of the allocation target amount between adjacent clusters. The constraint condition is that the total allocation amount meets the reserve demand and the allocation amount of each cluster is non-negative.
[0087] For each state cluster, based on its allocation target quantity and centroid coordinates, a cluster content allocation benchmark function is constructed to map the response delay characteristics and energy state of each distributed energy unit into a capacity allocation benchmark value.
[0088] A lower-level optimization model is constructed, with the actual configuration amount of each distributed energy unit within the cluster as the decision variable. The lower-level optimization objective function includes a deviation term between the actual configuration amount and the capacity allocation benchmark value, as well as a coupling term for response cost, and ensures that the total configuration amount within the cluster is equal to the allocation target amount and does not exceed the adjustable capacity.
[0089] In practical applications, the virtual power plant management system needs to acquire historical operating data of distributed energy units, including response latency and energy status data. Based on the historical operating data, the distributed energy units are divided into different state clusters. Assuming the clustering results formed in this embodiment are as follows: Cluster 1 contains energy storage units with fast response and sufficient capacity; Cluster 2 contains controllable power generation units with relatively fast response and moderate capacity; Cluster 3 contains demand response resources with moderate response speed; Cluster 4 contains traditional adjustable power generation units with slower response but larger capacity; Cluster 5 contains edge resources with slow response and limited capacity.
[0090] Based on the clustering results, a state cluster mapping table is constructed to record the centroid coordinates of each cluster and the list of distributed energy units it contains. For example, the centroid coordinates of cluster 1 are (15 seconds, 85%), and it contains battery energy storage units A, B, and C, etc.; the centroid coordinates of cluster 2 are (40 seconds, 65%), and it contains controllable photovoltaic units D and E, etc. The mapping table also records the adjustable capacity limit of each unit, such as 2.5MW for unit A and 1.8MW for unit B, etc.
[0091] Based on the state clustering mapping table, a response delay-energy state coupling matrix is constructed. By performing a Cartesian product operation on the centroid coordinates of each cluster, the coupling feature vector is obtained. For example, the inner product of the centroid coordinates (15 seconds, 85%) of cluster 1 and the reserve capacity demand feature (such as 3MW of reserve capacity requiring a response within 10 minutes) yields a demand matching degree of 0.92; the demand matching degree of cluster 2 is 0.78; that of cluster 3 is 0.65; that of cluster 4 is 0.45; and that of cluster 5 is 0.32. These matching degree values represent the degree of suitability of each cluster for responding to a specific reserve capacity demand.
[0092] A higher-level optimization model was constructed, with the allocation target amount of the reserve capacity of each cluster as the decision variable. The objective function of the higher-level optimization consists of two parts: the first part is a deviation measure term of the ratio of the allocation target amount to the demand matching degree, which ensures that the capacity allocation matches the response capability of each cluster; the second part is a boundary jump constraint term of the allocation target amount between adjacent clusters, which avoids excessive differences in the allocation of clusters with similar characteristics. The constraint condition is set as the total allocation amount meeting the reserve demand and the allocation amount of each cluster being non-negative. The optimization problem is solved by a genetic algorithm. After 100 generations of population iteration, the optimal allocation target amount is obtained: 1.2MW for cluster 1, 0.9MW for cluster 2, 0.5MW for cluster 3, 0.3MW for cluster 4, and 0.1MW for cluster 5, which together meet the reserve capacity demand of 3MW.
[0093] For each state cluster, a capacity allocation benchmark function is constructed based on its allocation target quantity and centroid coordinates. This function maps the response delay characteristics and energy state of each distributed energy unit to a capacity allocation benchmark value. Specifically, the deviation of each unit within the cluster from the cluster centroid is calculated. The smaller the deviation, the higher the benchmark allocation ratio. For example, energy storage unit A in cluster 1 has a response delay of 12 seconds and a state of charge (SOC) of 87%, with a deviation of 0.05 from the cluster centroid (15 seconds, 85%), corresponding to a benchmark allocation value of 0.5MW; unit B has a response delay of 16 seconds and an SOC of 84%, with a deviation of 0.03, corresponding to a benchmark allocation value of 0.4MW; unit C has a response delay of 18 seconds and an SOC of 82%, with a deviation of 0.08, corresponding to a benchmark allocation value of 0.3MW.
[0094] A lower-level optimization model is constructed, with the actual configuration of each distributed energy unit within the cluster as the decision variable. The lower-level optimization objective function includes two parts: a deviation term between the actual configuration and the capacity allocation benchmark value, and a coupling term for response cost. Considering the actual response cost of the unit, such as the response cost of energy storage unit A being 50 yuan / MW, unit B being 55 yuan / MW, and unit C being 45 yuan / MW, the constraint is that the total configuration within the cluster is equal to the allocation target and does not exceed the adjustable capacity. The interior point method is used to solve this optimization problem, and the optimal configuration scheme within cluster 1 is obtained: unit A is configured with 0.6MW, unit B with 0.4MW, and unit C with 0.2MW, for a total of 1.2MW, which is consistent with the upper-level allocation target.
[0095] Using a similar method, the optimal configuration schemes for other clusters are calculated and summarized to form a complete virtual power plant operation scheme. In practical applications, this two-layer optimization framework effectively improves the virtual power plant's response capability to grid dispatch and reduces overall operating costs.
[0096] In one optional implementation, for each state cluster, based on the allocation target quantity and the centroid coordinates, a multi-agent non-cooperative game model within the cluster is constructed. The Nash equilibrium solution is obtained by solving the optimal response function of each player, and the initial configuration quantity of each distributed energy unit is calculated based on the equilibrium solution, including:
[0097] For each state cluster, each distributed energy unit within the cluster is taken as the main player in the game. Based on the allocation target quantity and the centroid coordinates, the delay competition factor and capacity competition factor are calculated, and tensor product operation is performed to obtain the two-dimensional competition intensity matrix of each state cluster.
[0098] A payoff function is established for each game player. The payoff function includes a product term of allocation amount and unit capacity payoff, and a bilinear competition loss term based on a two-dimensional competition intensity matrix. The partial derivative of the payoff function with respect to the allocation amount of each game player is calculated and set to zero to obtain the optimal response function of each game player relative to the allocation amount of other game players.
[0099] The optimal response functions of each game player are combined to form a system of nonlinear equations. The allocation target quantity is introduced as a total constraint. By iteratively solving the system of nonlinear equations, the Nash equilibrium solution that satisfies the simultaneous validity of the optimal response functions of all game players is obtained.
[0100] Extract each component from the Nash equilibrium solution and use each component as the initial configuration of each distributed energy unit within the state cluster.
[0101] In the implementation of constructing a multi-agent non-cooperative game model within each state cluster, this invention obtains the Nash equilibrium solution by solving the optimal response function of each game agent, and calculates the initial configuration quantity of the distributed energy unit accordingly.
[0102] In the specific implementation process, for each state cluster, the distributed energy units within the cluster are identified as the game subjects. Assume that a certain state cluster contains 5 distributed energy units, namely energy units A, B, C, D and E, with corresponding unit capacity revenues of 12 yuan / kWh, 15 yuan / kWh, 10 yuan / kWh, 18 yuan / kWh and 14 yuan / kWh, respectively. The allocation target of this cluster is 100 kWh, and the centroid coordinates are 120.5 degrees longitude and 30.2 degrees latitude.
[0103] When calculating the delay contention factor, the distance from each distributed energy unit to the cluster centroid is used as the basis. For example, the distance from energy unit A to the centroid is 0.8 km, B is 1.2 km, C is 0.5 km, D is 1.5 km, and E is 0.9 km. These distances are converted into delay contention factors through an exponential decay function, resulting in delay contention factors of 0.85 for A, 0.70 for B, 0.92 for C, 0.65 for D, and 0.82 for E.
[0104] The capacity competition factor is calculated based on the maximum available capacity of each energy unit. Assuming that the maximum available capacity of A is 30 kWh, B is 25 kWh, C is 40 kWh, D is 20 kWh, and E is 35 kWh, after normalization, the capacity competition factor of A is 0.75, B is 0.62, C is 1.00, D is 0.50, and E is 0.88.
[0105] Perform tensor product operation to multiply the delay competition factor and the capacity competition factor to form a two-dimensional matrix. For example, the two-dimensional competition strength between A and B is 0.85×0.70=0.595, and between A and C it is 0.85×0.92=0.782. And so on to construct a complete 5×5 two-dimensional competition strength matrix.
[0106] A payoff function is established for each game player, including the product of the allocation amount and the payoff per unit capacity, minus the competition loss term based on the two-dimensional competition intensity matrix. Taking energy unit A as an example, its payoff function is 12 × allocation amount A minus the competition loss with other units. The competition loss is calculated as the product of A's allocation amount and the allocation amounts of other units, multiplied by the corresponding two-dimensional competition intensity coefficient. For example, the competition loss with B is allocation amount A × allocation amount B × 0.595.
[0107] By taking the partial derivative of the payoff function and setting it to zero, we obtain the optimal response function. For example, the optimal response function for A is A = 6 - 0.297 × B - 0.391 × C - 0.276 × D - 0.348 × E. Similarly, the optimal response functions for B, C, D, and E can be derived.
[0108] These optimal response functions are combined into a system of nonlinear equations, with a total constraint: configuration quantity A + configuration quantity B + configuration quantity C + configuration quantity D + configuration quantity E = 100. An iterative solution method is used, starting with random initial values, for example, the initial configuration quantity of each unit is 20. In each iteration, the configuration quantity of each unit is updated according to the current configuration quantity of other units. For example, after the first iteration, the configuration quantity of A is updated to 6 - 0.297 × 20 - 0.391 × 20 - 0.276 × 20 - 0.348 × 20 = 6 - 26.24 = -20.24, but since the configuration quantity cannot be negative, it is adjusted to 0.
[0109] After multiple iterations (50 in this example), when the change in the allocation amount of each unit is less than the preset threshold of 0.001, convergence is considered to have been achieved, and the Nash equilibrium solution is finally obtained: allocation amount A = 25.3 kWh, allocation amount B = 18.7 kWh, allocation amount C = 21.5 kWh, allocation amount D = 19.8 kWh, allocation amount E = 14.7 kWh, and the total is 100 kWh, which meets the allocation target amount constraint.
[0110] The components of these equilibrium solutions serve as the initial configuration values for each distributed energy unit within the state cluster. In practical applications, these configuration values can be adjusted based on real-time data, but the initial configuration values provide a starting point for satisfying Nash equilibrium, ensuring a reasonable allocation of resources considering competition factors.
[0111] The above process is repeated for all state clusters to obtain the initial configuration of all distributed energy units in the entire system. This game theory-based configuration method takes into account the competitive relationship and characteristics between energy units, so that resource allocation can achieve overall optimization while satisfying the total amount constraint.
[0112] In one optional implementation, the optimal response functions of each game player are simultaneously formed into a system of nonlinear equations, and the allocation objective is introduced as a total constraint. By iteratively solving this system of nonlinear equations, a Nash equilibrium solution satisfying the simultaneous validity of the optimal response functions of all game players is obtained, including:
[0113] The optimal response functions of each game subject are combined to form a nonlinear equation system. The coupling coefficient matrix of the configuration quantity is extracted. Based on the spectral radius and condition number of the two-dimensional competition intensity matrix, a precondition transformation is performed to obtain a positive definite symmetric preprocessing matrix.
[0114] An augmented Lagrangian function is constructed using the allocation target quantity as a constraint. The augmented Lagrangian function includes a residual term of the nonlinear equation system, a quadratic penalty term weighted by the preprocessing matrix, and a linear coupling term between the Lagrange multiplier and the constraint deviation.
[0115] The augmented Lagrangian function is decomposed into a configuration quantum problem and a multiplier update subproblem using the alternating direction multiplier method. The inverse of the preprocessing matrix is introduced as the descent direction correction operator, and the uniform distribution is set as the initial value for iteration.
[0116] During the iteration process, the configuration quantum problem is solved sequentially to obtain updated values. The Lagrange multipliers are updated based on the deviation between the updated values and the target quantity, and the residual norm of the nonlinear equation system is calculated. When it is less than the residual threshold and the updated configuration quantity satisfies the total quantity constraint, the current configuration quantity is output as the Nash equilibrium solution.
[0117] In this specific embodiment, the utility functions of each game player and their corresponding optimal response functions are collected. Taking the allocation of electricity resources as an example, it is assumed that there are 3 users participating in the allocation, and their respective optimal response functions are: x1=10-0.5x2-0.3x3, x2=12-0.4x1-0.6x3, x3=8-0.2x1-0.3x2, where x1, x2, and x3 represent the amount of resources obtained by the three users, and the total allocation target is 25 units.
[0118] The three optimal response functions are combined into a system of nonlinear equations, from which the collocation coupling coefficient matrix A is extracted. In this example, the collocation coupling coefficient matrix is a 3×3 matrix, where all diagonal elements are 1, and the off-diagonal elements are a, b, and c respectively. 12 =0.5, a 13 =0.3, a 21 =0.4, a 23 =0.6, a 31 =0.2, a 32 =0.3. By analyzing the spectral radius (modulus of the largest eigenvalue) and condition number (ratio of the largest eigenvalue to the smallest eigenvalue) of the matrix, the preconditioning transformation strategy is determined.
[0119] For the above matrix, the calculated spectral radius is approximately 1.2 and the condition number is approximately 3.8, indicating that the equation system exhibits a moderate degree of ill-conditionedness. Based on these characteristics, a preprocessing matrix P is constructed to be a positive definite symmetric matrix. In practice, the preprocessing matrix can be obtained by regularizing the original coupling coefficient matrix, such as P=A T ·A+βI, where β is the regularization parameter with a value of 0.1, and I is the identity matrix.
[0120] Construct an augmented Lagrangian function L(x, λ), which contains three terms: the residual term of the nonlinear equation system, the quadratic penalty term weighted by the preprocessing matrix, and the linear coupling term of the Lagrange multipliers and constraint bias. The purpose of the augmented Lagrangian function is to transform the original constrained optimization problem into an unconstrained optimization problem while maintaining the accuracy of the solution. In this example, the residual vector r = [x1 - (10 - 0.5x2 - 0.3x3), x2 - (12 - 0.4x1 - 0.6x3), x3 - (8 - 0.2x1 - 0.3x2)], and the total constraint is x1 + x2 + x3 = 25.
[0121] The alternating direction multiplier method is employed to solve the problem, decomposing the augmented Lagrangian function into a collocation quantum problem and a multiplier update subproblem. For the collocation quantum problem, a modified Newton method is used, introducing the inverse of the preprocessing matrix. As a descent direction correction operator, to improve iteration efficiency, the initial value of the iteration is set to a uniformly distributed amount, i.e., x1. 0 =x2 0 =x3 0 =25 / 3≈8.33, initial value of Lagrange multiplier λ 0 =0.
[0122] During the iteration process, the following steps are performed: For the k-th iteration, first solve the collocation quantum problem to obtain x. k +1Specifically, the augmented Lagrangian function with respect to x is minimized using the preprocessed conjugate gradient method, and the Lagrange multipliers λ are updated based on the new collocation and constraint bias. k+1 =λ k +ρ(x1 k+1 +x2 k+1 +x3 k+1 -25), where ρ is the penalty factor, set to 0.5, and the residual norm ||r(x) of the nonlinear equation system is calculated. k+1 The constraint is checked and the residual threshold is set to 0.001. When the residual norm is less than this threshold and |x1 k+1 +x2 k+1 +x3 k+1 When -25| < 0.001, the iteration terminates, and the current configuration value is output as the Nash equilibrium solution.
[0123] In this example, after 7 iterations, the obtained Nash equilibrium solution is x1=8.72, x2=9.45, x3=6.83, with a total of 25 and a residual norm of 0.0008, which satisfies the termination condition. This solution shows the optimal resource allocation obtained by the three users, considering the optimal response strategies of each game player and the total amount constraint.
[0124] The key advantage of this method lies in its effective handling of the ill-conditioned problems caused by the coupling coefficient matrix through preprocessing techniques and augmented Lagrange methods, thereby improving the stability and efficiency of iterative convergence. In practical applications, this method can be extended to situations involving more game participants, such as smart grid demand-side response, computing resource allocation, and communication network bandwidth allocation. It can effectively solve the equilibrium resource allocation scheme of each participant under mutual game, providing a theoretical basis and practical tools for system-level resource optimization.
[0125] In one optional implementation, the initial configuration amount is substituted into the lower-level optimization, KKT multipliers are extracted as feedback signals to adjust the allocation target amount of the upper-level optimization, and the intra-cluster game decomposition is re-executed until the deviation between the upper and lower-level configuration amounts converges, resulting in the final spare capacity configuration scheme, including:
[0126] Substitute the initial allocation amount into the lower-level optimization objective function, construct the KKT optimality condition equation system in combination with the constraint conditions, solve to obtain the KKT multipliers corresponding to the allocation objective amount constraint, and construct a sensitivity vector characterizing the marginal response strength of the actual allocation amount based on the sign and value of the KKT multipliers.
[0127] The sensitivity vector and the gradient vector of the upper-level optimization objective function are multiplied to obtain the adjustment direction factor. The adaptive step size is determined based on its magnitude. The allocation objective quantity of the upper-level optimization is corrected to obtain the corrected allocation objective quantity. This is then substituted into the multi-agent non-cooperative game model within the cluster, and game decomposition is performed to obtain the updated initial configuration quantity. This is then substituted into the lower-level optimization to calculate the new actual configuration quantity.
[0128] Calculate the deviation vector norm between the updated initial configuration and the new actual configuration. When the deviation is less than a preset threshold, output the new actual configuration as the final solution. Otherwise, repeat the iterative process of KKT multiplier extraction, sensitivity calculation and target quantity correction based on the new actual configuration until convergence.
[0129] In practical applications, after obtaining the initial allocation, it needs to be substituted into the lower-level optimization objective function to construct a complete set of KKT optimality condition equations. This set of equations includes the original problem constraints, complementary relaxation conditions, and the condition that the partial derivative of the Lagrange function with respect to the original variables is zero. Solving this set of equations yields a set of KKT multiplier λ values corresponding to the allocation objective constraints. The sign and value of the KKT multipliers directly reflect the degree of influence of the constraints on the objective function. Positive values indicate that the corresponding constraint is effective, and the larger the value, the greater the influence of the constraint on the objective function; negative values indicate that the current constraint has no effect; and zero values indicate that the constraint is exactly at a critical state. Taking a four-unit power system as an example, the initial allocation is [12MW, 15MW, 10MW, 8MW]. By solving the KKT equations, the KKT multipliers corresponding to the four constraints are [0.8, 0, -0.2, 0.5], indicating that the first and fourth constraints are effective and have a significant impact, the second constraint is critical, and the third constraint has no effect.
[0130] When constructing the sensitivity vector S based on KKT multipliers, the constraint coefficients of each effective constraint (with positive KKT multipliers) are extracted to form a sensitivity vector with the same dimension as the target quantity vector. The numerical values of each element in this vector represent the marginal response strength of the actual allocation quantity to the objective function in the corresponding dimension. Continuing with the previous example, the extracted sensitivity vector is [0.8, 0, 0, 0.5], indicating that the first and fourth allocation quantities have a greater impact on the objective function, while the second and third allocation quantities currently have a smaller impact.
[0131] The sensitivity vector S is multiplied by the gradient vector G of the upper-level objective function to obtain the adjustment direction factor D = S·G. If D is positive, it indicates that the current allocation of the objective quantity is aligned with the system sensitivity direction and should be adjusted in the positive direction; if it is negative, it should be adjusted in the opposite direction. The adjustment step size α is calculated adaptively and dynamically based on the magnitude of |D|. It can be set as α = β / |D|, where β is a preset constant, typically ranging from [0.01, 0.1]. Continuing with the previous example, assuming the gradient vector G of the upper-level objective function is [0.6, 0.3, -0.2, 0.4], then D = 0.8 × 0.6 + 0 × 0.3 + 0 × (-0.2) + 0.5 × 0.4 = 0.68. Taking β = 0.05, the step size α = 0.05 / 0.68 = 0.0735.
[0132] When correcting the allocation target amount of the upper-level optimization, the gradient descent method is used. The correction formula is: Corrected allocation target amount = original allocation target amount - α × G. For the example above, if the original allocation target amount is [13MW, 16MW, 9MW, 7MW], then the corrected allocation target amount is [13 - 0.0735 × 0.6, 16 - 0.0735 × 0.3, 9 - 0.0735 × (-0.2), 7 - 0.0735 × 0.4] = [12.96MW, 15.98MW, 9.01MW, 6.97MW].
[0133] Substituting the revised allocation target into the multi-agent non-cooperative game model within the cluster, the game decomposition calculation is re-executed. During the game, each unit adjusts its strategy based on its own cost characteristics and the revised allocation target, eventually reaching a Nash equilibrium state and obtaining the updated initial configuration. Continuing with the above example, the updated initial configuration obtained after the game is [12.8MW, 15.5MW, 9.5MW, 7.2MW].
[0134] The updated initial configuration is substituted into the next layer of optimization to calculate the new actual configuration as [12.75MW, 15.45MW, 9.55MW, 7.25MW]. The deviation vector norm between the updated initial configuration and the new actual configuration is calculated as ||[12.8-12.75, 15.5-15.45, 9.5-9.55, 7.2-7.25]||=||[0.05, 0.05, -0.05, -0.05]||=0.1. If the preset convergence threshold is 0.2, then 0.1<0.2, satisfying the convergence condition. The output [12.75MW, 15.45MW, 9.55MW, 7.25MW] is used as the final reserve capacity configuration scheme.
[0135] If the convergence condition is not met, the iterative process of KKT multiplier extraction, sensitivity calculation, and target quantity correction is repeated based on the new actual configuration. During the iteration, the relative rate of change of the solutions in two consecutive iterations is checked to determine whether a local optimum has been reached. If the relative rate of change is less than the preset accuracy requirement (e.g., 0.001), the iteration can be terminated early. To prevent oscillation, the momentum method can be used, that is, a certain proportion (e.g., 0.3) of the previous update direction is retained in each update and added to the current update direction.
[0136] In practical applications, the computational complexity of the entire method is mainly determined by two parts: intra-cluster game decomposition and solving the KKT equations. For a system with n units, if game convergence requires m iterations, the time complexity of solving the KKT equations is O(n^m). 3 The overall time complexity is approximately O(k×m×n+k×n). 3 ), where k is the number of outer loop iterations. For large-scale systems, distributed computing can be used to accelerate the game process, while sparse matrix techniques can be used to improve the efficiency of solving the KKT equations.
[0137] In one optional implementation, the initial allocation is substituted into the lower-level optimization objective function, and a set of KKT optimality condition equations is constructed in conjunction with the constraints. The KKT multipliers corresponding to the allocation objective constraints are then solved. Based on the signs and values of the KKT multipliers, a sensitivity vector characterizing the marginal response strength of the actual allocation is constructed, including:
[0138] Substitute the initial configuration into the lower-level optimization objective function, construct an augmented objective functional through a penalty term to soften the inequality constraints, solve the first-order necessary conditions and combine them with complementary relaxation conditions to construct the KKT optimality equation system.
[0139] The KKT optimality condition equations are linearized and decomposed to extract the sensitive submatrix corresponding to the allocation objective constraint. This submatrix is then decomposed into a product of an orthogonal matrix and an upper triangular matrix. The KKT multipliers of each cluster are obtained by back substitution.
[0140] Sign discrimination is performed on the KKT multipliers to obtain the indicator vector. Logarithmic transformation is performed on its absolute value to obtain the magnitude vector. Element-wise multiplication is performed to obtain the logarithmic domain sensitivity vector. The logarithmic domain sensitivity vector is nonlinearly scaled and mapped through inverse exponential transformation to obtain the sensitivity vector characterizing the marginal response strength of the actual configuration quantity.
[0141] In a resource allocation optimization system, the sensitivity analysis process based on the initial allocation and the lower-level optimization objective function can be divided into three key steps. First, the initial allocation needs to be substituted into the lower-level optimization objective function. Assuming the system has N clusters, the initial allocation is represented as an N-dimensional vector [a1, a2, ..., a...]. NThe initial resource allocation is represented by the array F(a), where each element represents the initial resource amount allocated to the corresponding cluster. The lower-level optimization objective function can be expressed as a nonlinear function F(a) with respect to the allocation amount, and is subject to a series of constraints, including the equality constraint h(a) = 0 and the inequality constraint g(a) ≤ 0. To handle the inequality constraints, a penalty function method is used to introduce a penalty term, and an augmented objective functional F is constructed. aug (a, μ) = F(a) + μ·P(g(a)), where P(·) is the penalty function and μ is the penalty factor. Consider a practical case involving resource allocation for 5 clusters with initial allocation values of [100, 150, 200, 120, 180]. The lower-level objective function is the weighted sum of the resource utilities of each cluster, with constraints that the total resources do not exceed 800 and the allocation value for each cluster is not less than 80.
[0142] By solving for the first-order necessary condition of the augmented objective functional and combining it with the complementary relaxation condition λ·g(a)=0 (where λ is a Lagrange multiplier) of the inequality constraints, a system of KKT optimality condition equations can be constructed. This system includes the gradient condition, expressed as... The equations h(a) = 0, the inequality constraints g(a) ≤ 0, the complementary relaxation condition λ·g(a) = 0, and the nonnegativity condition λ ≥ 0 are given. Substituting the initial configuration values from the above example, we can obtain the specific KKT equations.
[0143] The KKT optimality condition equations are linearized by performing a Taylor expansion of the nonlinear KKT equations around the initial allocation, retaining first-order terms to obtain the linearized KKT equations. From this linearized equations, a sensitivity submatrix S, of size m×n, is extracted, corresponding to the allocation target constraints. Here, m is the number of active constraints and n is the number of allocation variables. In the example, assuming two active constraints (total resource cap and minimum allocation for the first cluster), a 2×5 sensitivity submatrix is extracted. This matrix is then decomposed using QR decomposition, expressed as S = Q·R, where Q is an orthogonal matrix and R is an upper triangular matrix. The solution R·x = Q is obtained through back substitution. T ·b (where b is the perturbation vector) can be used to obtain the KKT multipliers corresponding to each active constraint. In the example, it is assumed that the calculated KKT multipliers are λ1=0.85 (total resource constraint) and λ2=-0.25 (minimum allocation constraint).
[0144] A sensitivity vector is constructed based on the obtained KKT multipliers. The KKT multipliers are then sign-determined to obtain an indicator vector I = [1, -1], indicating that the total resource constraint is positive and the minimum allocation constraint is negative. A logarithmic transformation is performed on the absolute values of the KKT multipliers to obtain an amplitude vector M = [ln(0.85), ln(0.25)] ≈ [-0.163, -1.386]. Element-wise multiplication of the indicator vector and the amplitude vector yields a logarithmic domain sensitivity vector D = [-0.163, 1.386]. A nonlinear scaling function S(D) = sign(D)·|D| is then applied to this vector. α (where α is the scaling parameter, usually between 0.5 and 0.8). With α=0.6, the scaled vector [-0.107, 1.037] is obtained, and the sensitivity vector [0.899, 2.821], which characterizes the marginal response strength of the actual configuration quantity, is obtained by inverse exponential transformation exp(S(D)).
[0145] Each element of the sensitivity vector represents the magnitude of change in the objective function when the corresponding constraint changes by one unit. In the example, the sensitivity vector [0.899, 2.821] indicates that relaxing the total resource constraint by one unit will increase the objective function by 0.899 units, while relaxing the minimum allocation constraint of the first cluster by one unit will increase the objective function by 2.821 units. This result provides a quantitative basis for resource allocation decisions and indicates the priority direction of constraint adjustment.
[0146] In practical applications, this method can dynamically adjust the stress level of each constraint based on the calculated sensitivity vector, thereby optimizing resource allocation strategies. For example, in the above example, the minimum allocation requirement for the first cluster can be appropriately reduced because this constraint has high sensitivity, and adjusting it can bring a more significant improvement in system performance. This technique is particularly suitable for large-scale, multi-constraint resource allocation scenarios, effectively identifying system bottlenecks and providing targeted optimization suggestions.
[0147] The virtual power plant standby capacity optimization configuration management system of this invention includes:
[0148] The resource information acquisition unit is used to acquire the adjustable capacity, response delay and energy status of each distributed energy unit in the virtual power plant, as well as the standby capacity demand information on the grid side.
[0149] The state space clustering unit is used to cluster and group each distributed energy unit in the two-dimensional space of response delay and energy state based on the response delay and energy state, to obtain multiple state clusters, and to calculate the centroid coordinates and the number of resources in each state cluster, and to construct a state cluster mapping table.
[0150] The two-layer optimization decision unit is used to construct a two-layer optimization framework based on the state cluster mapping table. The upper layer optimization aims to maximize the operating revenue of the virtual power plant and optimizes the allocation target amount of the reserve capacity of each state cluster. The lower layer optimization aims to minimize the intra-cluster response cost of each state cluster and optimizes the actual configuration amount.
[0151] The intra-cluster game equilibrium unit is used to construct an intra-cluster multi-agent non-cooperative game model for each state cluster based on the allocation target quantity and the centroid coordinates. The Nash equilibrium solution is obtained by solving the optimal response function of each game agent, and the initial configuration quantity of each distributed energy unit is calculated based on the equilibrium solution.
[0152] The iterative convergence coordination unit is used to substitute the initial configuration amount into the lower-level optimization, extract the KKT multipliers as feedback signals to adjust the allocation target amount of the upper-level optimization, and re-execute the intra-cluster game decomposition until the deviation between the upper and lower-level configuration amounts converges to obtain the final spare capacity configuration scheme.
[0153] The instruction distribution and execution unit is used to issue backup capacity configuration instructions to each distributed energy unit based on the final backup capacity configuration scheme.
[0154] A third aspect of the present invention provides an electronic device, comprising:
[0155] processor;
[0156] Memory used to store processor-executable instructions;
[0157] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0158] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0159] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.
[0160] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for optimizing and managing the standby capacity of a virtual power plant, characterized in that: include: Obtain the adjustable capacity, response latency, and energy status of each distributed energy unit within the virtual power plant, as well as the standby capacity demand information on the grid side; Based on the response delay and energy state, each distributed energy unit is clustered in the two-dimensional space of response delay-energy state to obtain multiple state clusters. The centroid coordinates and the number of resources within each state cluster are calculated to construct a state cluster mapping table. Based on the state cluster mapping table, a two-layer optimization framework is constructed. The upper layer optimization aims to maximize the operating revenue of the virtual power plant and optimizes the allocation target amount of the reserve capacity of each state cluster. The lower layer optimization aims to minimize the intra-cluster response cost of each state cluster and optimizes the actual configuration amount. For each state cluster, based on the allocation target quantity and the centroid coordinates, a multi-agent non-cooperative game model within the cluster is constructed. The Nash equilibrium solution is obtained by solving the optimal response function of each game agent, and the initial configuration quantity of each distributed energy unit is calculated based on the equilibrium solution. Substitute the initial configuration amount into the lower-level optimization, extract the KKT multipliers as feedback signals to adjust the allocation target amount of the upper-level optimization, and re-execute the intra-cluster game decomposition until the deviation between the upper and lower-level configuration amounts converges to obtain the final spare capacity configuration scheme. Based on the final backup capacity configuration scheme, backup capacity configuration instructions are issued to each distributed energy unit.
2. The method according to claim 1, characterized in that, Based on the response delay and energy state, distributed energy units are clustered in the two-dimensional space of response delay-energy state to obtain multiple state clusters. The centroid coordinates and the number of resources within each state cluster are calculated, and a state cluster mapping table is constructed, including: Using the response delay of each distributed energy unit as the first dimension coordinate and the energy state as the second dimension coordinate, the spatial location of each distributed energy unit is determined in the two-dimensional space of response delay-energy state. A density- and distance-based dual-constraint clustering algorithm is used to iteratively group the spatial location points. In each iteration, the weighted distance from each spatial location point to the current centroid of each cluster is calculated. The weighted distance includes time similarity weight and capacity similarity weight. Based on the weighted distance, each distributed energy unit is assigned to a corresponding state cluster, and the centroid coordinates of each state cluster are recalculated. When the change in the centroid coordinates of each state cluster in continuous iterations is less than the convergence threshold, the iteration is terminated, resulting in the final multiple state clusters. For each state cluster, the mean coordinates of the distributed energy units within the cluster are calculated and used as the centroid coordinates. The number of units within the cluster is counted, and the association between the centroid coordinates, the number of units within the cluster, and the identifiers of each distributed energy unit within the cluster is established to construct a state cluster mapping table.
3. The method according to claim 1, characterized in that, A two-layer optimization framework is constructed based on the aforementioned state cluster mapping table. The upper-layer optimization aims to maximize the operating revenue of the virtual power plant by optimizing the target allocation of reserve capacity for each state cluster. The lower-layer optimization aims to minimize the intra-cluster response cost of each state cluster by optimizing the actual configuration, including: Based on the state cluster mapping table, a response delay-energy state coupling matrix is constructed. The centroid coordinates are processed by Cartesian product to obtain the coupling feature vector. The inner product of the feature vector and the reserve capacity requirement is calculated to obtain the demand matching degree. Using the allocation target amount of the reserve capacity of each state cluster as the decision variable, an upper-level optimization model is constructed. The upper-level optimization objective function includes a deviation measure term of the ratio between the allocation target amount and the demand matching degree, and a boundary jump constraint term of the allocation target amount between adjacent clusters. The constraint condition is that the total allocation amount meets the reserve demand and the allocation amount of each cluster is non-negative. For each state cluster, based on its allocation target quantity and centroid coordinates, a cluster content allocation benchmark function is constructed to map the response delay characteristics and energy state of each distributed energy unit into a capacity allocation benchmark value. A lower-level optimization model is constructed, with the actual configuration amount of each distributed energy unit within the cluster as the decision variable. The lower-level optimization objective function includes a deviation term between the actual configuration amount and the capacity allocation benchmark value, as well as a coupling term for response cost, and ensures that the total configuration amount within the cluster is equal to the allocation target amount and does not exceed the adjustable capacity.
4. The method according to claim 1, characterized in that, For each state cluster, based on the allocation target quantity and the centroid coordinates, a multi-agent non-cooperative game model within the cluster is constructed. The Nash equilibrium solution is obtained by solving the optimal response function of each player, and the initial configuration quantity of each distributed energy unit is calculated based on the equilibrium solution, including: For each state cluster, each distributed energy unit within the cluster is taken as the main player in the game. Based on the allocation target quantity and the centroid coordinates, the delay competition factor and capacity competition factor are calculated, and tensor product operation is performed to obtain the two-dimensional competition intensity matrix of each state cluster. A payoff function is established for each game player. The payoff function includes a product term of allocation amount and unit capacity payoff, and a bilinear competition loss term based on a two-dimensional competition intensity matrix. The partial derivative of the payoff function with respect to the allocation amount of each game player is calculated and set to zero to obtain the optimal response function of each game player relative to the allocation amount of other game players. The optimal response functions of each game player are combined to form a system of nonlinear equations. The allocation target quantity is introduced as a total constraint. By iteratively solving the system of nonlinear equations, the Nash equilibrium solution that satisfies the simultaneous validity of the optimal response functions of all game players is obtained. Extract each component from the Nash equilibrium solution and use each component as the initial configuration of each distributed energy unit within the state cluster.
5. The method according to claim 4, characterized in that, The optimal response functions of each player are combined to form a system of nonlinear equations. The allocation objective is introduced as a total constraint. By iteratively solving this system of nonlinear equations, a Nash equilibrium solution satisfying the simultaneous validity of the optimal response functions of all players is obtained, including: The optimal response functions of each game subject are combined to form a nonlinear equation system. The coupling coefficient matrix of the configuration quantity is extracted. Based on the spectral radius and condition number of the two-dimensional competition intensity matrix, a precondition transformation is performed to obtain a positive definite symmetric preprocessing matrix. An augmented Lagrangian function is constructed using the allocation target quantity as a constraint. The augmented Lagrangian function includes a residual term of the nonlinear equation system, a quadratic penalty term weighted by the preprocessing matrix, and a linear coupling term between the Lagrange multiplier and the constraint deviation. The augmented Lagrangian function is decomposed into a configuration quantum problem and a multiplier update subproblem using the alternating direction multiplier method. The inverse of the preprocessing matrix is introduced as the descent direction correction operator, and the uniform distribution is set as the initial value for iteration. During the iteration process, the configuration quantum problem is solved sequentially to obtain updated values. The Lagrange multipliers are updated based on the deviation between the updated values and the target quantity, and the residual norm of the nonlinear equation system is calculated. When it is less than the residual threshold and the updated configuration quantity satisfies the total quantity constraint, the current configuration quantity is output as the Nash equilibrium solution.
6. The method according to claim 1, characterized in that, Substituting the initial configuration amount into the lower-level optimization, extracting the KKT multipliers as feedback signals to adjust the allocation target amount of the upper-level optimization, and re-executing the intra-cluster game decomposition until the deviation between the upper and lower-level configuration amounts converges, the final spare capacity configuration scheme is obtained, including: Substitute the initial allocation amount into the lower-level optimization objective function, construct the KKT optimality condition equation system in combination with the constraint conditions, solve to obtain the KKT multipliers corresponding to the allocation objective amount constraint, and construct a sensitivity vector characterizing the marginal response strength of the actual allocation amount based on the sign and value of the KKT multipliers. The sensitivity vector and the gradient vector of the upper-level optimization objective function are multiplied to obtain the adjustment direction factor. The adaptive step size is determined based on its magnitude. The allocation objective quantity of the upper-level optimization is corrected to obtain the corrected allocation objective quantity. This is then substituted into the multi-agent non-cooperative game model within the cluster, and game decomposition is performed to obtain the updated initial configuration quantity. This is then substituted into the lower-level optimization to calculate the new actual configuration quantity. Calculate the deviation vector norm between the updated initial configuration and the new actual configuration. When the deviation is less than a preset threshold, output the new actual configuration as the final solution. Otherwise, repeat the iterative process of KKT multiplier extraction, sensitivity calculation and target quantity correction based on the new actual configuration until convergence.
7. The method according to claim 6, characterized in that, Substituting the initial allocation amount into the lower-level optimization objective function, and combining it with the constraints, a set of KKT optimality condition equations is constructed. Solving these equations yields the KKT multipliers corresponding to the allocation objective constraints. Based on the signs and values of these KKT multipliers, a sensitivity vector characterizing the marginal response strength of the actual allocation amount is constructed, including: Substitute the initial configuration into the lower-level optimization objective function, construct an augmented objective functional through a penalty term to soften the inequality constraints, solve the first-order necessary conditions and combine them with complementary relaxation conditions to construct the KKT optimality equation system. The KKT optimality condition equations are linearized and decomposed to extract the sensitive submatrix corresponding to the allocation objective constraint. This submatrix is then decomposed into a product of an orthogonal matrix and an upper triangular matrix. The KKT multipliers of each cluster are obtained by back substitution. Sign discrimination is performed on the KKT multipliers to obtain the indicator vector. Logarithmic transformation is performed on its absolute value to obtain the magnitude vector. Element-wise multiplication is performed to obtain the logarithmic domain sensitivity vector. The logarithmic domain sensitivity vector is nonlinearly scaled and mapped through inverse exponential transformation to obtain the sensitivity vector characterizing the marginal response strength of the actual configuration quantity.
8. A virtual power plant standby capacity optimization and configuration management system, used to implement the method as described in any one of claims 1-7, characterized in that, include: The resource information acquisition unit is used to acquire the adjustable capacity, response delay and energy status of each distributed energy unit in the virtual power plant, as well as the standby capacity demand information on the grid side. The state space clustering unit is used to cluster and group each distributed energy unit in the two-dimensional space of response delay and energy state based on the response delay and energy state, to obtain multiple state clusters, and to calculate the centroid coordinates and the number of resources in each state cluster, and to construct a state cluster mapping table. The two-layer optimization decision unit is used to construct a two-layer optimization framework based on the state cluster mapping table. The upper layer optimization aims to maximize the operating revenue of the virtual power plant and optimizes the allocation target amount of the reserve capacity of each state cluster. The lower layer optimization aims to minimize the intra-cluster response cost of each state cluster and optimizes the actual configuration amount. The intra-cluster game equilibrium unit is used to construct an intra-cluster multi-agent non-cooperative game model for each state cluster based on the allocation target quantity and the centroid coordinates. The Nash equilibrium solution is obtained by solving the optimal response function of each game agent, and the initial configuration quantity of each distributed energy unit is calculated based on the equilibrium solution. The iterative convergence coordination unit is used to substitute the initial configuration amount into the lower-level optimization, extract the KKT multipliers as feedback signals to adjust the allocation target amount of the upper-level optimization, and re-execute the intra-cluster game decomposition until the deviation between the upper and lower-level configuration amounts converges to obtain the final spare capacity configuration scheme. The instruction distribution and execution unit is used to issue backup capacity configuration instructions to each distributed energy unit based on the final backup capacity configuration scheme.
9. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 7.
Citation Information
Patent Citations
Virtual power plant scheduling and clearing method based on conditional value-at-risk in market mode
CN117691681A
Virtual power plant control scheduling method and system
CN117893350A