Multi-agent cooperative game-based virtual power plant configuration optimization method and system

By analyzing the collaborative value and dynamic responsiveness of the entities in the virtual power plant, and combining the risk compensation coefficient with the optimized configuration scheme, the problems of unfair distribution of benefits and incentive failure in the virtual power plant were solved, thereby improving the overall operational efficiency and alliance stability.

CN120875183BActive Publication Date: 2025-12-09SHAANXI HANSHUNENG TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511383870.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-26
Publication Date
2025-12-09
Estimated Expiration
2045-09-26

AI Technical Summary

Technical Problem

Existing virtual power plant configuration optimization methods fail to effectively incorporate individual differences in dispatch frequency, risk-bearing capacity, and actual output capacity, leading to unfair revenue distribution and ineffective incentive mechanisms, thus undermining the stability of the alliance.

Method used

By analyzing the collaborative value, dynamic responsiveness, risk compensation coefficient, and contribution-benefit matching degree of the main entities, a distributed optimization algorithm is adopted to optimize the configuration scheme, ensuring that the overall benefit is maximized while the individual benefit is reasonable.

Benefits of technology

It has achieved a fair and effective profit distribution mechanism, improved the operational efficiency of virtual power plants and the stability of multi-entity alliances, and reduced the structural disconnect between configuration optimization and profit distribution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120875183B_ABST
    Figure CN120875183B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of virtual power plant operation, and particularly relates to a multi-agent cooperative game virtual power plant configuration optimization method and system. The method evaluates the cooperation value of a single agent by analyzing the difference between the individual operation income and the cooperative operation income of each agent in the virtual power plant; determines the contribution ability of the agent according to the dynamic response degree determined based on the current power of the agent and the power fluctuation deviation of the agent in the historical operation, and in combination with the loss risk that can be borne by each agent in the operation of the virtual power plant; and optimizes the configuration operation based on the matching of the dynamic contribution ability and the cooperation value. The present application introduces the contribution perception to optimize the overall configuration scheme by the matching relationship between the contribution and the obtained income of each agent in the cooperative process, improves the operation efficiency of the virtual power plant, the incentive effectiveness of the income distribution and the stability of the multi-agent alliance, and reduces the structural disconnection problem between the configuration optimization and the income distribution.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of virtual power plant operation, in particular to a multi-agent cooperative game virtual power plant configuration optimization method and system. BACKGROUND

[0002] With the rapid development of new energy and distributed energy, the traditional centralized power system is gradually transforming into a distributed and intelligent system. As a new energy management method, the virtual power plant (VPP) uses information communication and control technology to virtually aggregate various energy resources such as photovoltaic, wind power, energy storage, and electric vehicles distributed in different locations for unified scheduling and control, achieving the function of "looking like a large power plant". There are multiple interest subjects in the virtual power plant, such as energy producers, energy storage operators, and demand response users, each having different objective functions and operation preferences, and there are conflicts in resource competition, operation, and income distribution. Game theory can be used to simulate and solve strategic interaction problems among multiple agents. Under the framework of cooperative game theory, a cooperation mechanism between agents is established through alliance modeling, marginal contribution calculation, and income distribution strategies, and energy resource configuration is optimized to improve the overall efficiency of the system and protect individual income.

[0003] In the actual operation of the current virtual power plant, the participating agents show significant dynamic changes and inequality in resource capacity, regulation and response behavior, and cost structure. The adjustable resource size, load response frequency, equipment availability, and ability to withstand market fluctuations of different agents are different. However, the existing mainstream configuration optimization methods generally focus on maximizing the total system revenue as the core objective, and fail to effectively incorporate the differences in dispatching frequency, risk tolerance, and actual output capacity of individuals, resulting in an uneven distribution of benefits and insufficient returns for some agents in the configuration scheme, causing a discrepancy between income distribution and actual contribution, and undermining the incentive mechanism and alliance stability. SUMMARY

[0004] To solve the technical problem of uneven income distribution and ineffective incentives caused by traditional configuration optimization methods focusing on maximizing system revenue and ignoring individual participation behavior and contribution differences in the prior art, the purpose of the present application is to provide a multi-agent cooperative game virtual power plant configuration optimization method and system, and the technical solution adopted is as follows:

[0005] The present application provides a multi-agent cooperative game virtual power plant configuration optimization method, which comprises:

[0006] Obtaining the cooperative income and actual power of each resource scheduling period of each agent during the cooperative operation period in the virtual power plant;

[0007] The cooperation value of the single subject is analyzed based on independent operation income and cooperation income of the single subject in the current resource scheduling period;

[0008] The fluctuation deviation coefficient of the single subject is determined by the change degree between the actual power fluctuation and the expected power supply deviation in the continuous adjacent resource scheduling periods in the cooperation operation period of the single subject; the dynamic response degree of the single subject in the current resource scheduling period is obtained in combination with the fluctuation deviation coefficient based on the actual power fluctuation degree and the expected power in the current resource scheduling period of the single subject;

[0009] The risk compensation coefficient of the single subject is analyzed based on the equipment loss of the single subject device in the current resource scheduling period and other subjects;

[0010] The dynamic contribution ability of the single subject is analyzed based on the dynamic response degree and the risk compensation coefficient of the single subject; and the contribution income matching degree of the single subject is obtained according to the matching relationship between the dynamic contribution ability and the cooperation value of the single subject in the current resource scheduling period.

[0011] The configuration scheme is obtained by distributed optimization based on the contribution income matching degree.

[0012] Further, the cooperation value obtaining method comprises:

[0013] The independent income of each subject in the current resource scheduling period is obtained;

[0014] The difference between the cooperation income and the independent income of each subject in the current resource scheduling period is normalized to obtain the cooperation value of each subject in the current resource scheduling period.

[0015] Further, the fluctuation deviation coefficient obtaining method comprises:

[0016] For any subject, the resource fluctuation of the subject in each resource scheduling period is obtained according to the fluctuation confusion degree of the actual power of the subject in each resource scheduling period and the actual power fluctuation deviation of the resource scheduling period in the cooperation operation period;

[0017] The difference between the actual power and the expected power of the subject at each time point in each resource scheduling period is calculated, and the mean of all differences is normalized to obtain the resource supply difference of the subject in each resource scheduling period.

[0018] In the cooperation operation period of the subject, the ratio of the difference between the resource fluctuation of each adjacent two resource scheduling periods and the resource supply difference is taken as the adjacent change ratio of each adjacent two resource scheduling periods; and the mean of the adjacent change ratios of all adjacent two resource scheduling periods in the cooperation operation period is taken as the fluctuation deviation coefficient of the subject.

[0019] Further, the resource fluctuation acquisition method comprises:

[0020] The variance of the actual power of the subject in each resource scheduling period is taken as the chaos variation degree of each resource scheduling period, and the average of the chaos variation degrees of all resource scheduling periods in the cooperative operation period is taken as the variation average of the subject;

[0021] The difference between the chaos variation degree and the variation average of each resource scheduling period is calculated, and the ratio of the difference to the variation average is taken as the fluctuation deviation of each resource scheduling period;

[0022] The product of the fluctuation deviation and the chaos variation degree of the subject in each resource scheduling period is normalized to obtain the resource fluctuation of the subject.

[0023] Further, the dynamic response degree acquisition method comprises:

[0024] For any subject, the variance of the actual power of the subject in the current resource scheduling period is taken as the current power fluctuation degree of the subject, and the product of the current fluctuation degree and the fluctuation deviation coefficient is taken as the current response deviation degree of the subject;

[0025] The difference between the expected power and the current response deviation degree of the subject in the current resource scheduling period is normalized to obtain the current dynamic response degree of the subject.

[0026] Further, the risk compensation coefficient acquisition method comprises:

[0027] The device loss of each subject under the expected power in the current resource scheduling period is obtained;

[0028] The difference between the device loss of each subject and the average of the device losses of all subjects is normalized to obtain the risk compensation coefficient of each subject.

[0029] Further, the dynamic contribution ability acquisition method comprises:

[0030] For any subject, the product of the risk compensation coefficient and the dynamic response degree of the subject is taken as the risk adjustment degree of the subject, and the sum of the risk adjustment degree and the dynamic response degree of the subject is taken as the dynamic contribution ability of the subject.

[0031] Further, the contribution income matching degree acquisition method comprises:

[0032] The difference between the current dynamic contribution ability of a single subject and the cooperative value is negatively correlated and normalized to obtain the contribution income matching degree of the subject.

[0033] Further, the configuration scheme is obtained by distributed optimization based on the contribution benefit matching degree, comprising:

[0034] The contribution benefit matching degree of each subject is taken as a weight, and an optimized configuration scheme is obtained by collaborative solving through an ADMM algorithm.

[0035] The application further provides a virtual power plant configuration optimization system for multi-subject cooperative game, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the above multi-subject cooperative game virtual power plant configuration optimization method when executing the computer program.

[0036] The application has the following beneficial effects:

[0037] The application analyzes the difference between the individual operation benefit and the cooperative operation benefit of each subject in the virtual power plant, evaluates the cooperative value of the single subject, quantifies the real benefit improvement of each subject caused by cooperation, and clearly defines the value expectation of each subject participating in the virtual power plant. Then, the dynamic response of each subject in the actual configuration execution is evaluated, the degree of task undertaken by the subject is represented by the power fluctuation deviation in the actual operation, the effort of each subject is truly reflected, the dynamic response degree is determined in combination with the current power of the subject, and a quantitative basis is provided for dynamically constructing a fair and effective benefit distribution mechanism. The risk compensation coefficient is calculated by analyzing the possible loss risk of each subject in the operation of the virtual power plant, the fairness and strategy stability of the incentive mechanism are improved, and the subject with strong resource and capacity gradually exits or responds passively due to long-term high load is avoided. The current contribution is determined in combination with the risk compensation and the dynamic response, the benefit matching is based on the dynamic contribution ability and the cooperative value, and the configuration decision is optimized. The application introduces the contribution perception to optimize the overall configuration scheme through the matching relationship between the contribution and the obtained benefit of each subject in the cooperative process, ensures the maximization of the overall benefit, takes into account the individual benefit rationality and the incentive compatibility, improves the operation efficiency of the virtual power plant, the incentive effectiveness of the benefit distribution, and the stability of the multi-subject alliance, and reduces the structural disconnection problem between the configuration optimization and the benefit distribution. BRIEF DESCRIPTION OF DRAWINGS

[0038] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creating any creative labor.

[0039] Figure 1 A multi-subject cooperative game virtual power plant configuration optimization method flowchart is provided for an embodiment of the present application.

[0040] Figure 2 A virtual power plant configuration scheme flow diagram provided by one embodiment of the present application;

[0041] Figure 3 A distribution diagram of actual power and expected power provided by one embodiment of the present application;

[0042] Figure 4 A flow diagram of an ADMM algorithm system for solving an optimal matching scheme provided by one embodiment of the present application. DETAILED DESCRIPTION

[0043] In order to further illustrate the technical means and effects taken by the present application to achieve the predetermined purposes, the specific embodiments, structures, features and effects of the multi-agent cooperative game virtual power plant configuration optimization method and system according to the present application are described in detail as follows. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.

[0044] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs.

[0045] The specific scheme of the multi-agent cooperative game virtual power plant configuration optimization method and system provided by the present application is described in detail below with reference to the accompanying drawings.

[0046] Please refer to Figure 1 , which shows a multi-agent cooperative game virtual power plant configuration optimization method flowchart provided by one embodiment of the present application. The method includes the following steps:

[0047] S1: Obtain the cooperative income and actual power of each agent in each resource scheduling period during the cooperative operation period in the virtual power plant.

[0048] In a virtual power plant (VPP), a group of agents composed of multiple types of resources such as photovoltaic, wind power, energy storage, and adjustable load participate in scheduling and service together, and achieve resource complementation and income improvement through cooperative operation. In order to realize efficient access and optimal scheduling of resources in the process of multi-agent cooperation, a game model considering agent differences needs to be constructed, and a scientific resource configuration strategy needs to be formulated from the system level. Please refer to Figure 2 , which shows a virtual power plant configuration scheme flow diagram provided by one embodiment of the present application.

[0049] In the embodiment of the present application, the subject actual participation record is called from the virtual power plant platform, and the focus includes the expected power issued by the platform, the resource scheduling arrangement of each subject formulated by the system, and the actual power of the subject, which is the active power injected by the subject into the system or absorbed from the system at a certain time, measured by the intelligent meter in real time in kilowatts (kW).

[0050] In the embodiment of the present application, the operation time of the virtual power plant in the past week is taken as the analysis cooperation operation period, when the real-time optimization configuration decision is made in minute units, each 15 minutes is taken as a time unit as the resource scheduling period, and the Shapley value is used to obtain the cooperation benefit of each subject in a certain resource scheduling period based on the fair allocation method of the cooperative game theory. It should be noted that the implementer of the setting of collecting operation data can adjust according to the specific implementation condition, which is not limited here, and the Shapley value calculation is a well-known technical means familiar to those skilled in the art, which is not described here.

[0051] S2: Based on the independent operation benefit and the cooperation benefit of a single subject in the current resource scheduling period, the cooperation value of the subject is analyzed.

[0052] The cooperation value reflects the value judgment and benefit expectation of the subject participating in the cooperation. If the cooperation value deviates seriously from the benefit obtained by independent operation, it is easy to lead to the decrease of the participation willingness and the weakening of the response enthusiasm, and even to cause the instability of the resource exit and the cooperation alliance. Therefore, the individual benefit level needs to be reasonably evaluated under the premise of maximizing the overall benefit of the system, so that the optimization configuration scheme has the incentive compatibility and the collaborative sustainability.

[0053] The purpose of each subject participating in the cooperation in the virtual power plant is to benefit from the alliance, and the marginal benefit brought by resource sharing, unified scheduling and complementary cooperation is more than the benefit under the condition of independent operation. The greater the difference between the benefit allocated to the subject in the cooperative operation and the benefit of independent operation, the higher the cooperation value. According to the benefit of each subject in the cooperation and the independent operation, the cooperation value of each subject is obtained.

[0054] Preferably, in the embodiment of the present application, the cooperation value obtaining method comprises:

[0055] Firstly, the independent income of each subject in the current resource scheduling period is obtained. In the embodiment of the present application, the historical operation data of the subject before joining the virtual power plant is obtained, including the power output or load consumption curve, the autonomous scheduling strategy and the operation cost information, and combined with external factors such as weather conditions and electricity price curve, a dynamic programming (DP) model of the subject "separated from the virtual power plant" is constructed to simulate and estimate the independent operation income of the subject in the same period, which is used as a reference for measuring the value change of the subject before and after participating in the cooperation. It should be noted that the dynamic programming model is a well-known technical means for those skilled in the art, and will not be described further here.

[0056] Further, the difference between the cooperation income and the independent income of each subject in the current resource scheduling period is normalized to obtain the current cooperation value of each subject. It should be noted that normalization is a well-known technical means for those skilled in the art, and the selection of the normalization function can be linear normalization or standard normalization, and the specific normalization method is not limited here.

[0057] S3: In the cooperation running period of a single subject, the fluctuation deviation coefficient of the subject is determined by the change degree between the actual power fluctuation and the expected power supply deviation in the continuous adjacent resource scheduling periods; the dynamic response degree of the subject is obtained by combining the actual power fluctuation degree and the expected power in the current resource scheduling period of the single subject with the fluctuation deviation coefficient.

[0058] In the virtual power plant cooperation operation, the actual execution of the configuration of each subject is obviously different due to the differences in resource characteristics, response ability and task division. For example, photovoltaic may not be able to complete the scheduling task due to output fluctuation, and energy storage or adjustable load needs to bear additional system balancing responsibility. Therefore, the dynamic response contribution of each subject can be comprehensively analyzed, and the income situation can be optimized in combination with the response contribution of each subject.

[0059] The virtual power plant integrates various types of distributed resources, including fluctuating resources such as photovoltaic and wind power that are significantly affected by external environment, and subjects such as energy storage system and adjustable load that have relatively stable response. In actual operation, external factors such as extreme weather, equipment failure and power grid disturbance often cause some subjects to be unable to complete the expected output according to the established configuration scheme, resulting in an inequality between the actual contribution and the planned configuration. In order to improve the accuracy of resource scheduling and the executability of the configuration scheme, based on the historical operation data, the availability and response ability of various types of subject resources under external disturbance conditions are identified and quantified, providing a reliable basis for subsequent prediction of the resource dynamic response degree of each subject in the current resource scheduling period.

[0060] The fluctuation deviation coefficient of the subject is obtained by the time sequence characteristics of the change fluctuation of the actual power and the deviation from the expected planned power in the historical cooperation runtime period, and reflects the change relationship between the actual output and the deviation in the current historical time sequence.

[0061] For any subject, the resource fluctuation of the subject in each resource scheduling period is obtained according to the fluctuation confusion degree of the actual power of the subject in each resource scheduling period and the actual power fluctuation deviation of the subject in the resource scheduling period in the cooperation runtime period. In the virtual power plant cooperation runtime period, the actual power change of each subject output to the system in each resource scheduling period is analyzed. If the actual power change in the period is large and exceeds the general level, external interference may be encountered, and the resource fluctuation of the subject in the period is greater.

[0062] In the embodiment of the application, the variance of the actual power of the subject in each resource scheduling period is taken as the confusion change degree of each resource scheduling period, and reflects the power change degree in each period. The mean value of the confusion change degrees of the subject in all resource scheduling periods in the cooperation runtime period is taken as the change mean value of the subject, and as the general power change degree of the subject in the recent cooperation runtime period.

[0063] Further, the difference between the confusion change degree and the change mean value of each resource scheduling period is calculated, and the ratio of the difference to the change mean value is taken as the fluctuation deviation degree of each resource scheduling period, which quantifies the degree of power change deviation from the general level in a single resource scheduling period. Finally, the product of the fluctuation deviation degree and the confusion change degree of the subject in each resource scheduling period is normalized to obtain the resource fluctuation of the subject.

[0064] Further, the difference between the actual power and the expected power of the subject at each time in each resource scheduling period is calculated, and the mean value of all the differences is normalized to obtain the resource supply difference of the subject in each resource scheduling period. By comparing and analyzing the actual output power and the expected output power of each resource scheduling period, the scheduling uncertainty of the subject due to power fluctuation can be measured. Please refer to Figure 3 , which shows a distribution diagram of actual power and expected power provided by an embodiment of the application, wherein curve A represents the expected power required by the virtual power plant scheduling resource instruction, and curve B represents the actual power. In the subsequent scheduling, the subject provides a large deviation in the output power of the resource due to some factors.

[0065] The relationship between the resource fluctuation and the resource supply difference of each subject in the recent operation process is analyzed to help predict the current possible output power deviation. If the change range of the resource supply difference and the range of the resource fluctuation tend to be consistent in the adjacent period, it is indicated that the resource fluctuation can be used to predict the dynamic response degree of the subject in the current period.

[0066] Therefore, in the cooperative operation period of the subject, the ratio of the difference of the resource fluctuation of each adjacent two resource scheduling periods to the resource supply difference is taken as the adjacent change ratio of each adjacent two resource scheduling periods, which reflects the proportional relationship of the relative change.

[0067] The average of the adjacent change ratios of all adjacent two resource scheduling periods in the cooperative operation period is taken as the fluctuation deviation coefficient of the subject based on the proportional relationship in all continuous periods. Based on the fluctuation deviation coefficient, the actual contribution deviation of the subject in the current resource scheduling period can be predicted and analyzed.

[0068] When each subject receives the resource scheduling task from the virtual power plant, the actual output may deviate due to the fluctuation. The deviation between the expected power and the fluctuation situation predicted power is analyzed. The greater the fluctuation is, the greater the resource supply deviation is, and the smaller the actual possible dynamic response of the subject is.

[0069] Preferably, in the embodiment of the present application, the method for obtaining the dynamic response degree comprises:

[0070] For any subject, the actual power variance of the subject in the current resource scheduling period is taken as the current power fluctuation degree of the subject. The product of the current fluctuation degree and the fluctuation deviation coefficient of the subject is taken as the current response deviation degree of the subject. The expected deviation degree of the actual contribution of the subject is obtained through the rule of the overall deviation change in the historical time sequence.

[0071] The difference between the expected power of the subject in the current resource scheduling period and the current response deviation degree is further normalized to obtain the current dynamic response degree of the subject. The current dynamic response is estimated through the influence reference of the historical scheduling situation.

[0072] S4: Based on the device loss of the single subject device in the current resource scheduling period and other subjects, the risk compensation coefficient of the single subject is analyzed.

[0073] When the subjects in the virtual power plant need to provide more resources, the single subject bears more risk, loss or maintenance cost. If the loss is more than that of other subjects, more risk compensation income is needed. Therefore, according to the resource scheduling arrangement of each subject in the current period through the configuration scheme, the loss of each subject under the expected power output in the current resource scheduling period is obtained to obtain the risk compensation coefficient of each subject.

[0074] In the embodiment of the present application, the equipment loss of each subject under the expected power in the current resource scheduling period is obtained, the difference between the equipment loss of each subject and the average loss of all subject equipment is normalized to obtain the risk compensation coefficient of each subject, and the greater the risk compensation coefficient, the higher the risk borne by the subject in providing resources, and the higher the contribution ability evaluation in the subsequent period.

[0075] In one specific embodiment of the present application, the equipment loss under the expected power can be obtained from the attenuation model provided by the subject corresponding manufacturer, which is not described herein.

[0076] S5: Analyze the current dynamic contribution ability of the subject through the current dynamic response degree and the risk compensation coefficient of the subject, and obtain the contribution benefit matching degree of the subject according to the matching relationship between the current dynamic contribution ability of the single subject and the cooperation value.

[0077] Since the greater the risk compensation coefficient, the more the risk borne by the subject in providing resources, the actual dynamic contribution should be amplified, and the more the contribution of the subject which mainly provides resources in the long term, the dynamic contribution ability is obtained through the current dynamic response degree and the risk compensation coefficient of the subject.

[0078] In the embodiment of the present application, for any subject, the product of the risk compensation coefficient and the dynamic response degree of the subject is taken as the risk adjustment degree of the subject as an amplification adjustment degree. The sum of the risk adjustment degree and the dynamic response degree of the subject is taken as the dynamic contribution ability of the subject.

[0079] By comprehensively analyzing the dynamic contribution behavior and the cooperation value of each subject, a direct corresponding relationship between the actual contribution of the subject and the obtained benefit can be established, the linkage optimization of the resource allocation decision and the benefit distribution mechanism is realized, the virtual power plant ensures the maximization of the total system benefit, and the rationality of the benefit of each subject and the fairness of the incentive are also ensured, thereby improving the overall cooperation efficiency and the continuous stability of system operation.

[0080] When the subject cooperates in the virtual power plant, the cooperation value represents the benefit of the subject, and the greater the dynamic contribution ability of the subject in the current period, the higher the cooperation value, the more reasonable the benefit of the subject, and the greater the matching degree of contribution and benefit.

[0081] In the embodiment of the present application, the difference between the current dynamic contribution ability of the single subject and the cooperation value is negatively correlated and normalized to obtain the contribution benefit matching degree of the subject, and the greater the contribution benefit matching degree, the greater the matching degree of the contribution ability and the benefit of the subject. It should be noted that the negative correlation mapping is a well-known technical means for those skilled in the art, such as using inverse proportion or negative exponential power form, which is not described and limited herein.

[0082] As an example, the expression of the contribution benefit matching degree is: ; in the formula, represent the contribution benefit matching degree of the first subject at present; represent the cooperative value of the first subject in the current resource scheduling period; represent the dynamic contribution ability of the first subject in the current resource scheduling period. represent the absolute value extraction function, represent the exponential function with a natural constant as the base.

[0083] S6: Obtain the configuration scheme based on the contribution benefit matching degree through distributed optimization.

[0084] In the process of multi-agent collaborative operation of the virtual power plant, the core goal of configuration optimization is to maximize the overall benefit of the system, which may ignore the actual interests of individual participants. In order to prevent free riding, incentive failure and resource idling, it is necessary to ensure that the input and output of each participant are matched, maintain the stability of the alliance, cooperate with the distributed optimization algorithm, realize the collaborative solution mode of “computing and negotiating”, and introduce the benefit distribution mechanism into the configuration optimization problem as part of the solution.

[0085] The configuration optimization model of the virtual power plant needs to comprehensively consider the economic benefits at the system level, that is, under the premise of meeting the power balance, equipment constraints and market rules, the overall operation benefit, such as the power benefit minus the operation cost, is maximized as the operation benefit. At the same time, the fair distribution of each subject is realized, that is, the high matching between the cooperative contribution made by each participating subject in the operation and the benefit obtained by it is ensured.

[0086] Therefore, in the embodiment of the present application, the contribution benefit matching degree of each subject is taken as the weight, and the ADMM algorithm is used to cooperatively solve and obtain the optimal configuration scheme. In the virtual power plant with multiple participants, in order to realize efficient optimization and fair cooperation of resource configuration, a distributed optimization algorithm such as the alternating direction multiplier method (ADMM algorithm) is used, the contribution benefit matching degree of each subject is introduced into the local objective function, a collaborative solution mechanism of “computing and negotiating” is constructed, and the process of obtaining the best configuration scheme is obtained. The goal is to maximize the overall benefit of the virtual power plant, reduce the cost of each subject and maximize the benefit. Please refer to Figure 4 , which shows a process schematic diagram of an ADMM algorithm system for solving an optimal configuration scheme according to an embodiment of the present application.

[0087] It should be noted that the method of solving the cooperation and coordination task among multiple agents by a distributed optimization algorithm is a well-known technical means familiar to those skilled in the art, and will not be described here.

[0088] In one specific embodiment of the present application, the obtaining process of the optimal configuration scheme includes: firstly, decomposing the resource configuration optimization problem of the entire virtual power plant into multiple local sub-problems, and independently solving the optimal strategy of each resource subject. Secondly, introducing the contribution income matching degree of each subject into its local objective function to guide the optimization direction of the subject in the form of a weight factor. Then in each iteration, each subject locally performs optimization to solve its objective function and reports the results (such as power allocation, response capability, etc.). Finally, the virtual power plant system platform aggregates all subject results, calculates the global consistent variables (such as system scheduling total or price, etc.), and broadcasts them to each subject. Each subject updates its multiplier variable according to the global feedback information and adjusts the optimization direction of the next round.

[0089] Finally, the resource optimization configuration scheme not only meets the global goal of maximizing the total income of the virtual power plant, but also fully considers the dynamic contribution, cooperation value and income fairness of each subject, and finally forms a stable, efficient and incentive-compatible collaborative configuration scheme.

[0090] In summary, the present application analyzes the difference between the individual operation income and the collaborative operation income of each subject in the virtual power plant, evaluates the cooperation value of a single subject, quantifies the real income improvement of each subject due to cooperation, and clearly defines the value expectation of each subject participating in the virtual power plant. Further, the dynamic response of each subject in actual configuration execution is evaluated, the degree of task assumption is represented by the power fluctuation deviation of the subject in actual operation, the effort of each subject is truly reflected, the dynamic response degree is determined based on the current power of the subject, and a quantitative basis is provided for dynamically building a fair and effective income distribution mechanism. The risk of loss that each subject may bear in the operation of the virtual power plant is analyzed, the risk compensation coefficient is calculated, the fairness and strategy stability of the incentive mechanism are improved, and the subject with strong resource capacity is prevented from gradually withdrawing or responding negatively due to long-term high load. The current contribution is determined in combination with the risk compensation and dynamic response, the income matching based on the dynamic contribution ability and cooperation value is used to optimize the configuration decision. The present application introduces the contribution perception to optimize the overall configuration scheme through the matching relationship between the contribution and the obtained income of each subject in the process of participating in cooperation, ensures the maximization of the overall income, takes into account the individual income rationality and incentive compatibility, improves the operation efficiency of the virtual power plant, the incentive effectiveness of the income distribution, and the stability of the multi-agent alliance, and reduces the structural disconnection problem between the configuration optimization and the income distribution.

[0091] The application further provides a virtual power plant configuration optimization system for multi-agent cooperative game, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the multi-agent cooperative game virtual power plant configuration optimization method when executing the computer program.

[0092] It should be noted that the above-mentioned embodiment sequence is only for description, and does not represent the advantages and disadvantages of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or can be advantageous.

[0093] Each of the embodiments in the specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each embodiment mainly describes the difference from other embodiments.

Claims

1. A method for virtual power plant configuration optimization of a multi-agent cooperative game, characterized in that, The method comprises: obtaining the cooperative benefit and actual power of each subject in each resource scheduling period during the cooperative operation period in the virtual power plant; analyzing the current cooperative value of the subject based on the independent operation benefit and the cooperative benefit of the single subject in the current resource scheduling period; during the cooperative operation period of the single subject, determining the fluctuation deviation coefficient of the subject by the change degree between the actual power fluctuation and the expected power supply deviation in the continuous adjacent resource scheduling periods; obtaining the dynamic response degree of the subject in the current resource scheduling period based on the actual power fluctuation degree and the expected power, and combining the fluctuation deviation coefficient; analyzing the risk compensation coefficient of the single subject in the current resource scheduling period based on the equipment loss of the single subject device and other subjects; analyzing the dynamic contribution ability of the subject in the current resource scheduling period through the dynamic response degree and the risk compensation coefficient of the subject; obtaining the contribution benefit matching degree of the subject according to the matching relationship between the dynamic contribution ability and the cooperative value of the single subject in the current resource scheduling period; obtaining the configuration scheme through distributed optimization based on the contribution benefit matching degree; the method for obtaining the fluctuation deviation coefficient comprises: for any subject, obtaining the resource fluctuation of the subject in each resource scheduling period according to the fluctuation confusion degree of the actual power of the subject in each resource scheduling period and the actual power fluctuation deviation of the resource scheduling period in the cooperative operation period; calculating the difference between the actual power and the expected power of the subject at each time in each resource scheduling period, normalizing the mean value of all differences, and obtaining the resource supply difference of the subject in each resource scheduling period; during the cooperative operation period of the subject, taking the ratio of the difference of the resource fluctuation and the resource supply difference of each adjacent two resource scheduling periods as the adjacent change ratio of each adjacent two resource scheduling periods; taking the mean value of the adjacent change ratios of all adjacent two resource scheduling periods in the cooperative operation period as the fluctuation deviation coefficient of the subject; the method for obtaining the dynamic response degree comprises: for any subject, taking the actual power variance of the subject in the current resource scheduling period as the current power fluctuation degree of the subject; taking the product of the current fluctuation degree and the fluctuation deviation coefficient of the subject as the current response deviation degree of the subject; normalizing the difference between the expected power and the current response deviation degree of the subject in the current resource scheduling period to obtain the dynamic response degree of the subject in the current resource scheduling period; the method for obtaining the dynamic contribution ability comprises: for any subject, taking the product of the risk compensation coefficient and the dynamic response degree of the subject as the risk adjustment degree of the subject; taking the sum of the risk adjustment degree and the dynamic response degree of the subject as the dynamic contribution ability of the subject; obtaining the configuration scheme through distributed optimization based on the contribution benefit matching degree, comprising: taking the contribution benefit matching degree of each subject as the weight, and obtaining the optimized configuration scheme through ADMM algorithm.

2. The method of claim 1, wherein the method for obtaining the cooperative value comprises: obtaining the independent benefit of each subject in the independent operation in the current resource scheduling period; The difference between the cooperation benefit and the independent benefit of each subject in the current resource scheduling period is normalized to obtain the cooperation value of each subject at present. 3.The method of claim 1, wherein The method for obtaining the resource volatility comprises: The variance of the actual power of the subject in each resource scheduling period is taken as the chaotic variation degree of each resource scheduling period, and the average of the chaotic variation degrees of all resource scheduling periods in the cooperation operation period is taken as the variation average of the subject; The difference between the chaotic variation degree and the variation average of each resource scheduling period is calculated, and the ratio of the difference to the variation average is taken as the fluctuation deviation of each resource scheduling period; The product of the fluctuation deviation and the chaotic variation degree of the subject in each resource scheduling period is normalized to obtain the resource volatility of the subject. 4.The method of claim 1, wherein The method for obtaining the risk compensation coefficient comprises: The equipment loss of each subject under the expected power in the current resource scheduling period is obtained; The difference between the equipment loss of each subject and the average of the equipment losses of all subjects is normalized to obtain the risk compensation coefficient of each subject.

5. The method of claim 1, wherein The method for obtaining the contribution benefit matching degree comprises: The difference between the dynamic contribution ability and the cooperation value of each subject at present is negatively correlated and normalized to obtain the contribution benefit matching degree of the subject. 6.A virtual power plant configuration optimization system for multi-agent cooperative game, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein, The processor executes the computer program to realize the steps of the virtual power plant configuration optimization method of the multi-subject cooperation game according to any one of claims 1-5.

Citation Information

Patent Citations

  • Virtual power plant group cooperative operation method, electronic equipment and storage medium

    CN117175539A

  • Virtual power plant income distribution method with flexible resources considering external contribution

    CN119204302A