Distributed flexible resource cluster regulation method and system

By employing a distributed flexible resource cluster control method, utilizing a hierarchical architecture of local control and inter-cluster collaborative control, and combining global reinforcement learning and federated learning algorithms, the efficiency and real-time performance issues of large-scale flexible resource online control are resolved, achieving efficient flexible resource collaborative control.

CN120728748BActive Publication Date: 2025-11-28STATE GRID ZHEJIANG ELECTRIC POWER CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511172194.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-21
Publication Date
2025-11-28
Estimated Expiration
2045-08-21

AI Technical Summary

Technical Problem

Existing technologies are insufficient to meet the online control requirements of large-scale flexible resources. In particular, the computational scale of centralized control methods increases significantly with the expansion of flexible resource scale, resulting in limited participation of flexible resources in scheduling and a single mode of operation.

Method used

A distributed flexible resource cluster regulation method is adopted. By obtaining the power adjustment rate of local flexible resources, local regulation is carried out based on the control law and the maximum power regulation capacity. When the power adjustment rate is unbalanced, a collaborative regulation request between clusters is triggered. The power adjustment amount is optimized by using a global reinforcement learning model. The collaborative regulation of flexible resources is achieved by combining a hierarchical distributed architecture and federated learning algorithm.

Benefits of technology

It improves the efficiency and real-time performance of online control in large-scale flexible resource scenarios, meets the online control needs of large-scale flexible resources, and reduces processing complexity and privacy leakage risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120728748B_ABST
    Figure CN120728748B_ABST
Patent Text Reader

Abstract

The application discloses a distributed flexible resource cluster regulation method and system. Through a hierarchical distributed flexible resource cluster regulation architecture, the computing task of flexible resource power regulation is decoupled into cluster power regulation and inter-cluster collaborative regulation. When the inter-cluster collaborative regulation is performed, the regulation decision end only needs to be responsible for the distribution of the power transfer matrix between various distributed flexible resource clusters, and does not need to process the operation data of each flexible resource. The power regulation amount distribution of each flexible resource is responsible for each distributed flexible resource cluster, thereby effectively improving the regulation efficiency in the online regulation scene of large-scale flexible resources, ensuring the real-time performance in the online regulation scene of large-scale flexible resources, and meeting the online regulation demand of large-scale flexible resources.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of power systems, in particular to a distributed flexible resource cluster regulation method and system. BACKGROUND

[0002] With the continuous advancement of new power system construction, the number of distributed power sources, controllable loads and energy storage and other flexible resources is showing an explosive growth trend. Compared with traditional loads and power generation side resources, flexible resources have the characteristics of small individual capacity, large number and geographical dispersion, and have stronger temporal and spatial uncertainty. In addition, due to the limitation of communication and data processing capacity, a large number of terminal devices cannot realize timely interconnection and intercommunication with the dispatching center, directly leading to the limitation of the scale and the single mode of flexible resources participating in dispatching. Under this background, how to realize the coordinated regulation of a large number of distributed flexible resources is a problem to be solved.

[0003] The existing technology mainly adopts a centralized control mode for the coordinated regulation of flexible resources. This mode usually collects the power consumption demand information of flexible resources within the jurisdiction of the load aggregator, and makes real-time centralized decision on the power consumption of each flexible resource. However, the calculation scale of this mode will significantly increase with the expansion of the scale of the regulated flexible resources, resulting in that the existing technology cannot meet the online regulation needs of large-scale flexible resources. SUMMARY

[0004] The present application provides a distributed flexible resource cluster regulation method and system to solve the technical problem that the existing technology cannot meet the online regulation needs of large-scale flexible resources.

[0005] To solve the above technical problems, the first aspect of the embodiment of the present application provides a distributed flexible resource cluster regulation method, comprising:

[0006] obtaining the power adjustment rate of each local flexible resource belonging to the local distributed flexible resource cluster at each time, and determining the control law of the local flexible resource based on the power adjustment rate;

[0007] if it is detected that each power adjustment rate is not imbalanced, then performing power regulation on the local flexible resource based on the power adjustment rate of the local flexible resource, the control law and the preset maximum power regulation capacity;

[0008] if it is detected that any one power adjustment rate is imbalanced, then sending a cluster intercoordination regulation request to a regulation decision end; wherein the cluster intercoordination regulation request is used to instruct the regulation decision end to feed back the power transfer matrix between each distributed flexible resource cluster;

[0009] Based on the power transfer matrix, the power adjustment amount of each of the local flexible resources is optimized using a global reinforcement learning model to determine the target power adjustment amount of each of the local flexible resources; wherein, the global reinforcement learning model is generated based on the weights of the local reinforcement learning models of the control decision-making terminal and each of the distributed flexible resource clusters.

[0010] As a preferred embodiment, the step of obtaining the power adjustment rate of each local flexible resource belonging to the local distributed flexible resource cluster at each time moment, and determining the control law of the local flexible resource based on the power adjustment rate, specifically includes:

[0011] The power regulation rate of each local flexible resource at each time moment is obtained based on the ratio between the power regulation amount of each local flexible resource at each time moment and the maximum power regulation capacity.

[0012] Based on the power adjustment rate of each local flexible resource at each time step, the control law of the local flexible resource is determined by the following expression:

[0013] ;

[0014] in, Indicates local flexible resources exist The control law at time t; Indicates local flexible resources exist The power regulation rate at that time; Indicates local flexible resources exist The power regulation rate at that time; Indicates local flexible resources With local flexible resources Between Connection weights at specific times; Indicates local flexible resources The neighbor set, which includes Time and local flexible resources All local flexible resources with physical connections.

[0015] As a preferred embodiment, the method further includes:

[0016] The remaining regulation capacity of each local flexible resource at each time is determined based on the difference between the maximum power regulation capacity and the power regulation amount of each local flexible resource at each time.

[0017] According to the residual adjustment capacity of each local flexible resource at each time, the connection weight between each local flexible resource is adjusted by the following expression:

[0018] ;

[0019] wherein, denotes the residual adjustment capacity of local flexible resource at the time ; denotes the residual adjustment capacity of local flexible resource at the time .

[0020] As a preferred solution, the power adjustment of the local flexible resource based on the power adjustment rate, the control law and the preset maximum power adjustment capacity of the local flexible resource specifically comprises:

[0021] Based on the power adjustment rate, the control law and the maximum power adjustment capacity of the local flexible resource, the target power adjustment amount of the local flexible resource is determined by the following expression:

[0022] ;

[0023] According to the target power adjustment amount of the local flexible resource, the power of the local flexible resource is adjusted;

[0024] wherein, denotes the target power adjustment amount of local flexible resource at the time ; denotes the power adjustment rate of local flexible resource at the time ; denotes the control law of local flexible resource at the time ; denotes the maximum power adjustment capacity of local flexible resource at the time ; denotes a preset proportion coefficient.

[0025] As a preferred solution, the power transfer matrix is generated by the global reinforcement learning model based on the state space of the total power deviation of the distributed flexible resource cluster, the upper limit and the lower limit of the power adjustment capacity of each distributed flexible resource cluster, which is utilized by the regulation decision end;

[0026] The power transfer matrix is used to represent the power transfer relationship and the power transfer amount between each of the distributed flexible resource clusters; and the total power deviation of the distributed flexible resource clusters is the sum of the difference between the actual power and the expected power of each of the distributed flexible resource clusters at the current time.

[0027] As a preferred solution, the target power adjustment amount of each local flexible resource is determined by optimizing the power adjustment amount of each local flexible resource using a global reinforcement learning model according to the power transfer matrix, and specifically includes:

[0028] The power transfer matrix is decomposed into the target power adjustment amount of each local flexible resource based on a consensus algorithm using the global reinforcement learning model, taking the injection power, voltage amplitude of the cluster grid connection point, upper limit of the power adjustment capacity and lower limit of the power adjustment capacity of the local distributed flexible resource cluster as the state space.

[0029] As a preferred solution, the target power adjustment amount of each local flexible resource is determined by decomposing the power transfer matrix based on a consensus algorithm using the global reinforcement learning model, and specifically includes:

[0030] According to the power transfer matrix, the total power adjustment demand of the local distributed flexible resource cluster at the current time is determined;

[0031] Based on a consensus algorithm, the power adjustment rate of each local flexible resource is iteratively adjusted with the same power adjustment rate of each local flexible resource as the optimization target, until the power adjustment rate of each local flexible resource converges, and the initial power adjustment rate of each local flexible resource is obtained;

[0032] Based on the product between the initial power adjustment rate of each local flexible resource and the maximum power adjustment capacity, the initial power adjustment amount of each local flexible resource is determined;

[0033] According to the total power adjustment demand and the initial power adjustment amount of each local flexible resource, a power adjustment proportion coefficient is determined, and based on the product between the power adjustment proportion coefficient and the initial power adjustment amount, the target power adjustment amount of each local flexible resource is determined; wherein the sum of each target power adjustment amount is equal to the total power adjustment demand.

[0034] As a preferred solution, the global reinforcement learning model is generated by the following steps:

[0035] Based on a horizontal federated learning strategy, the parameter server aggregates the local reinforcement learning model weights of the control decision-making end and each of the distributed flexible resource clusters according to the number of local training samples in the current round of training to obtain updated global model weights. The updated global model weights are then used as the initial local reinforcement learning model weights for the control decision-making end and each of the distributed flexible resource clusters to perform the next round of training until the reward value converges. Finally, the global reinforcement learning model is generated based on the current global model weights.

[0036] The weights of the local reinforcement learning model at the control decision-making end are generated by the control decision-making end training the local reinforcement learning model using local control data.

[0037] The weights of the local reinforcement learning model of the distributed flexible resource cluster are generated by the distributed flexible resource cluster training the local reinforcement learning model using local cluster operation data;

[0038] The reward value is calculated based on power regulation cost, power regulation limit penalty, and branch power flow limit penalty;

[0039] The specific expression for the power regulation cost is as follows:

[0040] ;

[0041] The specific expression for the power over-limit penalty is as follows:

[0042] ;

[0043] The specific expression for the branch power flow exceeding the limit penalty is as follows:

[0044] ;

[0045] in, This indicates the cost of the power regulation; This indicates the number of the distributed flexible resource clusters; Indicates the first The control cost of a distributed flexible resource cluster; Indicates the first Power regulation of a distributed flexible resource cluster; and They represent the first The upper limit and lower limit of the power regulation capacity of a distributed flexible resource cluster; This indicates the penalty for exceeding the power regulation limit; This indicates the preset penalty factor for exceeding the adjustment power limit; represents the branch flow over-limit punishment; represents a preset branch flow over-limit punishment factor; represents the power flow of the branch . represents the power flow upper limit of the branch . represents a branch set.

[0046] As a preferred solution, the method specifically judges whether each power adjustment rate is unbalanced through the following steps:

[0047] According to the power adjustment rate of each local flexible resource, an average power adjustment rate at the current time is determined;

[0048] According to the difference between the power adjustment rate of each local flexible resource and the average power adjustment rate, a power adjustment rate deviation of each local flexible resource is determined;

[0049] Obtain the communication load rate and the node residual energy at the current time;

[0050] Map the power adjustment rate deviation, the communication load rate and the node residual energy into a fuzzy set through a preset event detector;

[0051] Based on the fuzzy rule corresponding to the fuzzy set, fuzzy reasoning is performed on the power adjustment rate deviation, the communication load rate and the node residual energy to determine the power adjustment rate imbalance triggering threshold corresponding to the power adjustment rate deviation, the communication load rate and the node residual energy; wherein the fuzzy rule is used to define the fuzzy output result corresponding to each fuzzy set for representing the power adjustment rate imbalance triggering threshold;

[0052] If it is detected that the power adjustment rate deviation of each local flexible resource is less than or equal to the power adjustment rate imbalance triggering threshold, it is determined that the power adjustment rate of each local flexible resource is not unbalanced;

[0053] If it is detected that the power adjustment rate deviation of any one local flexible resource is greater than the power adjustment rate imbalance triggering threshold, it is determined that the power adjustment rate of the any one local flexible resource is unbalanced.

[0054] The second aspect of the embodiment of the application provides a distributed flexible resource cluster regulation system, comprising:

[0055] A data acquisition module is configured to:

[0056] Obtain the power adjustment rate of each local flexible resource belonging to a local distributed flexible resource cluster at each time, and determine the control law of the local flexible resource based on the power adjustment rate.

[0057] an intra-cluster power regulation module, configured to:

[0058] if it is detected that none of the power adjustment rates is imbalanced, performing power regulation on the local flexible resource based on the power adjustment rate of the local flexible resource, the control law and a preset maximum power regulation capacity;

[0059] an inter-cluster cooperative regulation module, configured to:

[0060] if it is detected that any one of the power adjustment rates is imbalanced, sending an inter-cluster cooperative regulation request to a regulation decision end, wherein the inter-cluster cooperative regulation request is used to instruct the regulation decision end to feed back a power transfer matrix between the distributed flexible resource clusters;

[0061] optimizing the power regulation amount of each local flexible resource by using a global reinforcement learning model based on the power transfer matrix, to determine a target power regulation amount of each local flexible resource, wherein the global reinforcement learning model is generated based on local reinforcement learning model weights of the regulation decision end and each distributed flexible resource cluster.

[0062] Compared with the prior art, the embodiment of the present application has the beneficial effects that, by using the hierarchical distributed flexible resource cluster regulation architecture, the computing task of flexible resource power regulation is decoupled into intra-cluster power regulation and inter-cluster cooperative regulation, when performing inter-cluster cooperative regulation, the regulation decision end only needs to be responsible for the distribution of the power transfer matrix between the distributed flexible resource clusters, without processing the operation data of each flexible resource, and the power regulation amount of each flexible resource is distributed by each distributed flexible resource cluster, thereby effectively improving the regulation efficiency in the online regulation scenario of large-scale flexible resources, ensuring the real-time performance in the online regulation scenario of large-scale flexible resources, and meeting the online regulation requirements of large-scale flexible resources. BRIEF DESCRIPTION OF DRAWINGS

[0063] Figure 1 is a flowchart of the distributed flexible resource cluster regulation method in the embodiment of the present application;

[0064] Figure 2 is a schematic diagram of the distributed flexible resource cluster regulation architecture in the embodiment of the present application;

[0065] Figure 3 is a schematic diagram of the event-triggered communication mechanism based on sampling data in the embodiment of the present application;

[0066] Figure 4 is a structural schematic diagram of the distributed flexible resource cluster regulation system in the embodiment of the present application. DETAILED DESCRIPTION

[0067] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.

[0068] Please refer to Figure 1 The first aspect of the embodiments of the present application provides a distributed flexible resource cluster regulation method, comprising the following steps S1 to S4:

[0069] S1, obtaining the power adjustment rate of each local flexible resource belonging to the local distributed flexible resource cluster at each time, and determining the control law of the local flexible resource based on the power adjustment rate;

[0070] S2, if it is detected that each power adjustment rate is not imbalanced, then based on the power adjustment rate of the local flexible resource, the control law and the preset maximum power regulation capacity, the power of the local flexible resource is regulated;

[0071] S3, if it is detected that any one power adjustment rate is imbalanced, then a cluster inter-collaborative regulation request is sent to a regulation decision end; wherein the cluster inter-collaborative regulation request is used to indicate the regulation decision end to feed back the power transfer matrix between each distributed flexible resource cluster;

[0072] S4, according to the power transfer matrix, the power adjustment amount of each local flexible resource is optimized by using a global reinforcement learning model to determine the target power adjustment amount of each local flexible resource; wherein the global reinforcement learning model is generated based on the local reinforcement learning model weight of the regulation decision end and each distributed flexible resource cluster.

[0073] It is worth noting that the distributed flexible resource cluster regulation architecture diagram in the present embodiment is as shown in Figure 2As shown, the distributed flexible resource cluster regulation architecture is divided into two layers, a regulation layer and a cluster layer. The regulation decision end is located in the upper regulation layer, which takes each distributed flexible resource cluster as a whole for regulation, and realizes the collaborative regulation between clusters in the region. Each distributed flexible resource cluster constitutes the lower cluster layer. The power adjustment amount of each flexible resource in the distributed flexible resource cluster is allocated by the cluster corresponding sub-group scheduling control unit, so that each distributed flexible resource cluster responds to the regulation instruction of the regulation layer while realizing intelligent self-adaptation within the cluster. Each terminal in the distributed flexible resource cluster is an independent flexible resource terminal. To realize the hierarchical collaborative distributed flexible resource cluster regulation architecture, first, the flexible resources in the region need to be clustered according to the type of flexible resource, geographical location distribution, or affiliation, etc., so as to form each distributed flexible resource cluster. For example, the flexible resources within the power supply range are divided into a local cluster with the power supply radius of the substation as the boundary. This embodiment will not be described in more detail.

[0074] Further, the main task of power regulation in the cluster is the rapid allocation of power adjustment amount of each local flexible resource. In this embodiment, the power regulation of each local flexible resource is realized based on the consensus algorithm. The power adjustment rate of each local flexible resource at each time is selected as the consensus state variable, and the control law of each local flexible resource is determined. It can be understood that the control law of each local flexible resource is used to guide each local flexible resource to dynamically adjust its control output according to the difference between its power adjustment rate and that of the neighbor flexible resource, so as to realize the convergence and balance of power allocation in the cluster. If the current power adjustment rate is not imbalanced, i.e., the power adjustment rate in the cluster can be balanced through cluster power regulation, the cluster power regulation is continued, i.e., based on the power adjustment rate, control law and maximum power adjustment capacity of each local flexible resource, the power regulation of each local flexible resource is performed, so as to realize the convergence and balance of power allocation in the cluster while meeting the power adjustment capacity constraint of each local flexible resource.

[0075] Further, if the power adjustment rate of any one of the clusters is unbalanced, i.e., the power adjustment rates of the clusters in the cluster cannot be balanced, interaction between adjacent clusters needs to be triggered. First, a cluster coordination request is sent to the regulation and decision end to indicate the regulation and decision end to feed back the power transfer matrix between the distributed flexible resource clusters. It can be understood that the main task of the cluster coordination is the dynamic exchange of power between the clusters. In this embodiment, a multi-agent hierarchical reinforcement learning algorithm based on federated learning is used to realize the interaction between the multiple distributed flexible resource clusters. This algorithm process divides the cluster coordination architecture into a regional coordination layer (macro power distribution) and a cluster execution layer (micro power distribution), thereby significantly reducing the processing complexity of the distributed flexible resource power regulation and avoiding the processing of the operation data of a large number of flexible resources by a single control unit at the same time. Specifically, the regulation and decision end is the regional coordination layer in the cluster coordination architecture, and the feedback of the power transfer matrix between the distributed flexible resource clusters is the process of executing macro power distribution. After receiving the power transfer matrix issued by the regulation and decision end, the global reinforcement learning model is used to optimize the power adjustment amount of each local flexible resource to determine the target power adjustment amount of each local flexible resource, thereby realizing micro power distribution. The global reinforcement learning model in this embodiment is generated by aggregating the local reinforcement learning model weights of the regulation and decision end and the distributed flexible resource clusters based on the federated learning strategy.

[0076] The distributed flexible resource cluster regulation method provided by the embodiment of the application decouples the calculation task of flexible resource power regulation into cluster internal power regulation and cluster coordination by using a hierarchical distributed flexible resource cluster regulation architecture. When the cluster coordination is executed, the regulation and decision end only needs to be responsible for the issuance of the power transfer matrix between the distributed flexible resource clusters, and does not need to process the operation data of each flexible resource. The power adjustment amount of each flexible resource is distributed by each distributed flexible resource cluster, thereby effectively improving the regulation efficiency in the online regulation scenario of a large number of flexible resources, ensuring the real-time performance in the online regulation scenario of a large number of flexible resources, and meeting the online regulation requirements of a large number of flexible resources.

[0077] As a preferred solution, the power adjustment rate of each local flexible resource belonging to the local distributed flexible resource cluster at each time is obtained, and the control law of the local flexible resource is determined based on the power adjustment rate. Specifically, the method comprises:

[0078] The power adjustment rate of each local flexible resource at each time is obtained according to the ratio between the power adjustment amount of each local flexible resource at each time and the maximum power adjustment capacity.

[0079] Based on the power adjustment rate of each local flexible resource at each time step, the control law of the local flexible resource is determined by the following expression:

[0080] ;

[0081] in, Indicates local flexible resources exist The control law at time t; Indicates local flexible resources exist The power regulation rate at that time; Indicates local flexible resources exist The power regulation rate at that time; Indicates local flexible resources with local flexible resources Between Connection weights at specific times; Indicates local flexible resources The neighbor set, which includes Time and local flexible resources All local flexible resources with physical connections.

[0082] Specifically, in this embodiment, the power regulation rate of each local flexible resource at each time moment is obtained by the following expression based on the ratio between the power regulation amount of each local flexible resource at each time moment and the maximum power regulation capacity:

[0083] ;

[0084] in, Indicates local flexible resources exist Power regulation rate at any given time; Indicates local flexible resources exist The amount of power adjustment at any given time; Indicates local flexible resources exist Maximum power regulation capacity at any given time.

[0085] As a preferred embodiment, the method further includes:

[0086] The remaining regulation capacity of each local flexible resource at each time is determined based on the difference between the maximum power regulation capacity and the power regulation amount of each local flexible resource at each time.

[0087] According to the residual adjustment capacity of each local flexible resource at each time, the connection weight between each local flexible resource is adjusted by the following expression:

[0088] ;

[0089] wherein, represents the residual adjustment capacity of local flexible resource at the time; represents the residual adjustment capacity of local flexible resource at the time.

[0090] Specifically, the connection weight between each local flexible resource in the embodiment is an adaptive weight coefficient, which is dynamically adjusted according to the residual adjustment capacity of the local flexible resource, so that the connection weight between each local flexible resource considers the real-time adjustment capacity of the local flexible resource, and ensures that the power distribution is more fair and effective.

[0091] As a preferred solution, the power of the local flexible resource is regulated based on the power adjustment rate of the local flexible resource, the control law and the preset maximum power adjustment capacity, and specifically includes:

[0092] Based on the power adjustment rate of the local flexible resource, the control law and the maximum power adjustment capacity, the target power adjustment amount of the local flexible resource is determined by the following expression:

[0093] ;

[0094] The power of the local flexible resource is regulated according to the target power adjustment amount of the local flexible resource;

[0095] wherein, represents the target power adjustment amount of local flexible resource at the time; represents the power adjustment rate of local flexible resource at the time; represents the control law of local flexible resource at the time; represents the maximum power adjustment capacity of local flexible resource at the time; represents a preset proportion coefficient.

[0096] ​​​​​​Specifically, for any one local flexible resource, according to the calculated control law, the increment or decrement of the power adjustment of the local flexible resource can be determined through the above expression, wherein the size of the increment or decrement of the power adjustment can be in a proportional relationship with the control law, so that the embodiment introduces a proportional coefficient The target power adjustment amount of each local flexible resource is calculated by using the above expression.

[0097] As a preferred solution, the power transfer matrix is generated by the regulation and decision end based on a state space of a total power deviation of the distributed flexible resource cluster, an upper limit of a power adjustment capacity and a lower limit of the power adjustment capacity of each distributed flexible resource cluster, and by using the global reinforcement learning model generated based on the state space;

[0098] The power transfer matrix is used to represent the power transfer relationship and the power transfer amount between each distributed flexible resource cluster. The total power deviation of the distributed flexible resource cluster is the sum of the difference between the actual power and the expected power of each distributed flexible resource cluster at the current time.

[0099] Specifically, the regulation and decision end in the embodiment also generates the power transfer matrix of the current state space based on the hierarchical reinforcement learning algorithm. First, the total power deviation of the distributed flexible resource cluster, the upper limit of the power adjustment capacity of each distributed flexible resource cluster, and the lower limit of the power adjustment capacity are taken as the state space. Then, the corresponding power transfer matrix is generated as the action space by using the trained global reinforcement learning model under the constraint of the state space. It can be understood that the power transfer matrix in the embodiment is used to represent the power transfer relationship and the power transfer amount between each distributed flexible resource cluster. Assuming that an element in the power transfer matrix is represented as , the element represents the power transfer amount from the distributed flexible resource cluster to the distributed flexible resource cluster . A positive value represents output power, and a negative value represents absorbed power. In addition, since the power transfer matrix has determined the power transfer relationship and the power transfer amount between each distributed flexible resource cluster, in combination with the current power of each distributed flexible resource cluster, the expected power of each distributed flexible resource cluster at the next time can be determined. Therefore, in combination with the actual power of the distributed flexible resource cluster at each time, the total power deviation of the distributed flexible resource cluster can be detected.

[0100] As a preferred solution, the power adjustment amount of each local flexible resource is optimized by using the global reinforcement learning model according to the power transfer matrix, and the target power adjustment amount of each local flexible resource is determined, specifically including:

[0101] The injection power of the cluster grid connection point, the voltage amplitude, the upper limit of the power regulation capacity and the lower limit of the power regulation capacity of the local distributed flexible resource cluster are taken as the state space, and the power transfer matrix is decomposed into the target power regulation amount of each local flexible resource based on a consistency algorithm by using the global reinforcement learning model.

[0102] Specifically, the local distributed flexible resource cluster serves as a cluster execution layer, and needs to complete power distribution within the cluster according to the power transfer matrix. First, the injection power of the cluster grid connection point, the voltage amplitude, the upper limit of the power regulation capacity and the lower limit of the power regulation capacity of the local distributed flexible resource cluster are taken as the state space, and then under the constraint of the state space, the power transfer matrix is decomposed into the target power regulation amount of each local flexible resource based on a consistency algorithm by using the global reinforcement learning model, so that the sum of the target power regulation amounts of each local flexible resource meets the total power regulation demand required by the current power transfer matrix.

[0103] As a preferred solution, the power transfer matrix is decomposed into the target power regulation amount of each local flexible resource based on a consistency algorithm by using the global reinforcement learning model, which specifically includes:

[0104] According to the power transfer matrix, the total power regulation demand of the local distributed flexible resource cluster is determined;

[0105] Based on a consistency algorithm, the power adjustment rate of each local flexible resource is iteratively adjusted with the same power adjustment rate of each local flexible resource as the optimization target, until the power adjustment rate of each local flexible resource converges, and the initial power adjustment rate of each local flexible resource is obtained;

[0106] Based on the product between the initial power adjustment rate of each local flexible resource and the maximum power regulation capacity, the initial power regulation amount of each local flexible resource is determined;

[0107] According to the total power regulation demand and the initial power regulation amount of each local flexible resource, a power adjustment proportion coefficient is determined, and based on the product between the power adjustment proportion coefficient and the initial power regulation amount, the target power regulation amount of each local flexible resource is determined; wherein the sum of each target power regulation amount is equal to the total power regulation demand.

[0108] Specifically, the total power regulation demand of the local distributed flexible resource cluster is first calculated according to the power transfer matrix issued by the regulation and control decision end , which is specifically shown in the following expression:

[0109] ;

[0110] wherein, represents the power transfer amount from the local distributed flexible resource cluster to the rest of the distributed flexible resource clusters; while the total power adjustment demand is the total power transfer amount from the local distributed flexible resource cluster to all the rest of the distributed flexible resource clusters.

[0111] Further, based on a consensus algorithm, taking the same power adjustment rate of each local flexible resource as the optimization goal, the expression: is taken as the iterative formula, and the power adjustment rate of each local flexible resource is gradually adjusted until the power adjustment rate of each local flexible resource converges. It is worth noting that since the consensus algorithm takes the convergence of the power adjustment rate as the goal, it does not guarantee that the sum of the power adjustment amounts of each local flexible resource under the converged power adjustment rate is equal to the total power adjustment demand described above, so the initial power adjustment amount after convergence needs to be adjusted. Assuming that the power adjustment rate of each local flexible resource after iteration of the consensus algorithm is determined, the initial power adjustment amount is shown in the following expression:

[0112] ;

[0113] wherein, represents the initial power adjustment amount of the local flexible resource ; represents the power adjustment rate of the local flexible resource after iteration of the consensus algorithm.

[0114] Then, the initial power adjustment amount is proportionally adjusted to satisfy the sum of each target power adjustment amount equal to the total power adjustment demand , i.e. The proportional adjustment of the initial power adjustment amount is shown in the following expression:

[0115] ;

[0116] wherein, represents the local distributed flexible resource cluster; represents the target power adjustment amount of the local flexible resource ; is the power adjustment proportionality coefficient determined according to the total power adjustment demand and each initial power adjustment amount.

[0117] As a preferred solution, the method specifically generates the global reinforcement learning model through the following steps: ​

[0118] Based on a horizontal federated learning strategy, the parameter server aggregates the local reinforcement learning model weights of the control decision-making end and each of the distributed flexible resource clusters according to the number of local training samples in the current round of training to obtain updated global model weights. The updated global model weights are then used as the initial local reinforcement learning model weights for the control decision-making end and each of the distributed flexible resource clusters to perform the next round of training until the reward value converges. Finally, the global reinforcement learning model is generated based on the current global model weights.

[0119] The weights of the local reinforcement learning model at the control decision-making end are generated by the control decision-making end training the local reinforcement learning model using local control data.

[0120] The weights of the local reinforcement learning model of the distributed flexible resource cluster are generated by the distributed flexible resource cluster training the local reinforcement learning model using local cluster operation data;

[0121] The reward value is calculated based on power regulation cost, power regulation limit penalty, and branch power flow limit penalty;

[0122] The specific expression for the power regulation cost is as follows:

[0123] ;

[0124] The specific expression for the power over-limit penalty is as follows:

[0125] ;

[0126] The specific expression for the branch power flow exceeding the limit penalty is as follows:

[0127] ;

[0128] in, This indicates the cost of the power regulation; This indicates the number of the distributed flexible resource clusters; Indicates the first The control cost of a distributed flexible resource cluster; Indicates the first Power regulation of a distributed flexible resource cluster; and They represent the first The upper limit and lower limit of the power regulation capacity of a distributed flexible resource cluster; This indicates the penalty for exceeding the power regulation limit; This indicates the preset penalty factor for exceeding the adjustment power limit; represents the branch flow out-of-limit punishment; represents a preset branch flow out-of-limit punishment factor; represents the flow of the branch ; represents the flow upper limit of the branch ; represents the branch set.

[0129] It is worth noting that the traditional centralized multi-agent reinforcement learning needs to collect the original operation data (such as power, voltage, and regulation capacity) of each cluster to the central server to train the model, which may leak the business secrets (such as user power consumption mode and equipment running state) or sensitive information (such as geographic location load characteristics) of the cluster owner. Therefore, the present embodiment adopts a horizontal federated learning strategy to train the multi-agent reinforcement learning model, and the regulation and decision-making end and each distributed flexible resource cluster only upload the local reinforcement learning model weight instead of the original data, so as to solve the privacy leakage problem of the centralized multi-agent reinforcement learning algorithm. The step flow of the horizontal federated learning strategy in the present embodiment is as follows:

[0130] (1) First, the regulation and decision-making end trains the local reinforcement learning model using the local regulation data (such as the total power deviation of the distributed flexible resource cluster, the upper limit and lower limit of the power regulation capacity of each distributed flexible resource cluster, etc.) to generate the local reinforcement learning model weight of the regulation and decision-making end; each distributed flexible resource cluster trains the local reinforcement learning model using the local cluster operation data (such as the injection power of the cluster grid connection point, the voltage amplitude, the upper limit and lower limit of the power regulation capacity of the local distributed flexible resource cluster, and the power transfer matrix received from the regulation and decision-making end, etc.) to generate the local reinforcement learning model weight of the distributed flexible resource cluster.

[0131] (2) The regulation and decision-making end and each distributed flexible resource cluster regularly upload the latest local reinforcement learning model weight generated by the training to the parameter server without uploading the specific local training samples.

[0132] (3) The parameter server aggregates the local reinforcement learning model weights of the regulation and decision-making end and each distributed flexible resource cluster according to the number of local training samples of the regulation and decision-making end and each distributed flexible resource cluster in the current round of training, to obtain the updated global model weight. Assuming that in the round of federated training, a total of clusters participate, each cluster trains to obtain a set of updated local reinforcement learning model weights . The aggregation process of the global model weight is shown in the following expression:

[0133] ;

[0134] wherein, denotes the updated global model weight; denotes the number of local training samples of the i-th distributed flexible resource cluster; denotes the total number of local training samples of all distributed flexible resource clusters participating in the current round of training. Thus, the weight of each local reinforcement learning model weight in aggregation is determined by the proportion of the number of local training samples of the corresponding distributed flexible resource cluster to the total number of local training samples. The distributed flexible resource cluster with more local training samples contributes more in aggregating the global model weight. (4) The parameter server distributes the updated global model weight to the regulation and decision end and each distributed flexible resource cluster as the initial local reinforcement learning model weight for the regulation and decision end and each distributed flexible resource cluster to perform the next round of training.

[0135] (5) Repeat the above process until the reward value converges, i.e., the global reinforcement learning model converges, and generate the global reinforcement learning model according to the current global model weight.

[0136] Specifically, at each time step, the regulation and decision end as the upper decision unit first observes the total power deviation of the distributed flexible resource clusters, the upper limit and lower limit of the power regulation capacity of each distributed flexible resource cluster as the state space, generates a power transfer matrix as the action space according to the state space, and issues the power transfer matrix to the distributed flexible resource clusters; the distributed flexible resource clusters as the lower execution unit observe the injected power, voltage amplitude, upper limit and lower limit of the power regulation capacity of the local distributed flexible resource cluster at the grid connection point as the state space after receiving the power transfer matrix, decompose the power transfer matrix into the target power regulation amount of each local flexible resource based on a consensus algorithm and execute the action; after executing the action, the environment calculates the reward value according to the power regulation cost, the regulation power out-of-limit penalty and the branch power flow out-of-limit penalty, and the regulation and decision end and each distributed flexible resource cluster obtain the reward value and observe the new state. Through continuous interaction with the environment, the strategy of the regulation and decision end and each distributed flexible resource cluster is constantly optimized, the cumulative reward gradually increases and tends to be stable, and finally the model training is completed.

[0137] The global reinforcement learning model training method provided in this embodiment eliminates the risk of commercial secret leakage while maintaining cross-cluster coordination performance, and meets the privacy protection needs of multi-stakeholder scenarios.

[0138] As a preferred solution, the method specifically determines whether the power adjustment rates are unbalanced by the following steps:

[0139]

[0140] ​determine an average power adjustment rate at the current time according to the power adjustment rates of the respective local flexible resources;

[0141] determine a power adjustment rate deviation of each of the local flexible resources according to a difference between the power adjustment rate of each of the local flexible resources and the average power adjustment rate;

[0142] obtain a communication load rate and a node residual energy at the current time;

[0143] map the power adjustment rate deviation, the communication load rate and the node residual energy into a fuzzy set through a preset event detector;

[0144] perform fuzzy inference on the power adjustment rate deviation, the communication load rate and the node residual energy based on a fuzzy rule corresponding to the fuzzy set to determine a power adjustment rate imbalance triggering threshold corresponding to the power adjustment rate deviation, the communication load rate and the node residual energy, wherein the fuzzy rule is used to define a fuzzy output result corresponding to each fuzzy set and used to represent the power adjustment rate imbalance triggering threshold;

[0145] if the power adjustment rate deviation of each of the local flexible resources is less than or equal to the power adjustment rate imbalance triggering threshold, determine that the power adjustment rate of each of the local flexible resources is not imbalanced;

[0146] if the power adjustment rate deviation of any one of the local flexible resources is greater than the power adjustment rate imbalance triggering threshold, determine that the power adjustment rate of the any one of the local flexible resources is imbalanced.

[0147] Specifically, since the hierarchical distributed flexible resource cluster regulation system in the embodiment does not need frequent information exchange when in steady state operation, i.e., each distributed flexible resource cluster performs power regulation within the cluster without interaction with other distributed flexible resource clusters, the traditional time-triggered communication mechanism has serious waste of communication resources and computing resources. To reduce the communication amount and the computing amount of the regulation system, the embodiment proposes an event-triggered communication mechanism based on sampling data, such as Figure 3As shown, each distributed flexible resource cluster only triggers interaction with adjacent distributed flexible resource clusters when it meets the predetermined power adjustment rate imbalance event trigger condition, i.e., the power change rate within the cluster cannot be balanced, and does not participate in interaction at other times. Among them, the sampler is responsible for measuring the state of the distributed flexible resource cluster, including the power adjustment rate deviation of each local flexible resource, the communication load rate and the node residual energy, and sends it to the event detector, which judges whether the power adjustment rate imbalance event trigger condition is met. If it is met, the communication network between the local distributed flexible resource cluster and its adjacent distributed flexible resource cluster is triggered and the controller is updated, and the inter-cluster collaborative regulation is carried out.

[0148] Further, the embodiment introduces fuzzy logic to handle nonlinear and uncertain multi-parameter coupling problems, and converts the qualitative description of power adjustment rate deviation, communication load rate and node residual energy into quantitative power adjustment rate imbalance trigger threshold, realizing dynamic adaptive adjustment. The design logic of the event detector is as follows:

[0149] (1) Input variables: including the power adjustment rate deviation of each local flexible resource , the communication load rate and the node residual energy , the specific expressions are as follows:

[0150] ;

[0151] ;

[0152] ;

[0153] Among them, represents the average power adjustment rate of the distributed flexible resource cluster; represents the data transmission rate of the current link; represents the maximum available bandwidth of the link; the communication load rate reflects the degree of communication link congestion, indicates that the link is saturated, and unnecessary communication should be avoided at this time; represents the current residual capacity of the node; represents the total capacity of the node battery; the node residual energy reflects the energy state of the edge node or terminal device, and the low power needs to reduce communication to prolong the device life.

[0154] (2) Output variable: power adjustment rate imbalance trigger threshold , ranging from 0 to 1, when , communication is triggered.

[0155] (3) Fuzzy inference: Firstly, the input parameters (including power adjustment rate deviation, communication load rate and node residual energy) are mapped to fuzzy sets (such as three triangular fuzzy membership functions corresponding to "large", "medium" and "small"); secondly, the corresponding fuzzy rules are activated according to the input membership; finally, the fuzzy output result is converted into a definite value by the barycenter method. Among them, the logical relationship between multiple parameters is described by "if-then" fuzzy rules, and a typical fuzzy rule is shown in Table 1:

[0156] Table 1 Fuzzy rule definition

[0157]

[0158] The embodiment utilizes the event-triggered communication mechanism based on sampling data and fuzzy processing, can prolong the life of the edge device, and avoid information storm under network congestion condition.

[0159] Referring to Figure 4 , the second aspect of the embodiment of the present application provides a distributed flexible resource cluster regulation system, comprising:

[0160] The data acquisition module 101 is configured to:

[0161] acquire the power adjustment rate of each local flexible resource belonging to the local distributed flexible resource cluster at each time point, and determine the control law of the local flexible resource based on the power adjustment rate;

[0162] The cluster power regulation module 102 is configured to:

[0163] If it is detected that each power adjustment rate is not imbalanced, the local flexible resource is subjected to power regulation based on the power adjustment rate of the local flexible resource, the control law and the preset maximum power regulation capacity;

[0164] The inter-cluster cooperative regulation module 103 is configured to:

[0165] If it is detected that any one power adjustment rate is imbalanced, an inter-cluster cooperative regulation request is sent to the regulation decision end; wherein the inter-cluster cooperative regulation request is used to instruct the regulation decision end to feed back the power transfer matrix between each distributed flexible resource cluster;

[0166] According to the power transfer matrix, the power regulation amount of each local flexible resource is optimized by using a global reinforcement learning model to determine the target power regulation amount of each local flexible resource; wherein the global reinforcement learning model is generated based on the local reinforcement learning model weight of the regulation decision end and each distributed flexible resource cluster.

[0167] As a preferred solution, the data acquisition module 101 is configured to acquire a power adjustment rate of each local flexible resource in the local distributed flexible resource cluster at each time instant, and determine a control law of the local flexible resource based on the power adjustment rate, and specifically comprises:

[0168] obtain the power adjustment rate of each local flexible resource at each time instant according to a ratio between the power adjustment amount of each local flexible resource at each time instant and the maximum power adjustment capacity;

[0169] determine the control law of the local flexible resource according to the power adjustment rate of each local flexible resource at each time instant by the following expression:

[0170] ;

[0171] wherein, denotes the control law of the local flexible resource at the time instant; denotes the power adjustment rate of the local flexible resource at the time instant; denotes the power adjustment rate of the local flexible resource at the time instant; denotes a connection weight between the local flexible resource and the local flexible resource at the time instant; denotes a neighbor set of the local flexible resource , the neighbor set comprising all local flexible resources having a physical connection relationship with the local flexible resource at the time instant. denotes a neighbor set of the local flexible resource , the neighbor set comprising all local flexible resources having a physical connection relationship with the local flexible resource at the time instant.

[0172] As a preferred solution, the data acquisition module 101 is further configured to:

[0173] determine a residual adjustment capacity of each local flexible resource at each time instant according to a difference between the maximum power adjustment capacity and the power adjustment amount of each local flexible resource at each time instant;

[0174] adjust a connection weight between each local flexible resource according to the residual adjustment capacity of each local flexible resource at each time instant by the following expression:

[0175] ;

[0176] wherein, denotes the control law of the local flexible resource ​​At the remaining regulation capacity at the time instant; denotes a local flexible resource At the remaining regulation capacity at the time instant.

[0177] As a preferred solution, the in-cluster power regulation module 102 is configured to regulate power of the local flexible resource based on the power adjustment rate of the local flexible resource, the control law and a preset maximum power regulation capacity, and specifically includes:

[0178] The target power regulation amount of the local flexible resource is determined based on the power adjustment rate of the local flexible resource, the control law and the maximum power regulation capacity through the following expression:

[0179] ;

[0180] The power of the local flexible resource is regulated according to the target power regulation amount of the local flexible resource;

[0181] wherein, denotes a local flexible resource At the target power regulation amount at the time instant; denotes a local flexible resource At the power adjustment rate at the time instant; denotes a local flexible resource At the control law at the time instant; denotes a local flexible resource At the maximum power regulation capacity at the time instant; denotes a preset proportion coefficient.

[0182] As a preferred solution, the power transfer matrix is generated by the regulation decision end based on a state space of a total power deviation of the distributed flexible resource cluster, an upper limit and a lower limit of the power regulation capacity of each distributed flexible resource cluster, by using the global reinforcement learning model;

[0183] wherein, the power transfer matrix is used to represent the power transfer relationship and the power transfer amount between each distributed flexible resource cluster; and the total power deviation of the distributed flexible resource cluster is the sum of the difference between the actual power and the expected power of each distributed flexible resource cluster at the current time instant.

[0184] As a preferred solution, the inter-cluster cooperative regulation module 103 is configured to optimize, according to the power transfer matrix, a power adjustment amount of each local flexible resource by using a global reinforcement learning model, and determine a target power adjustment amount of each local flexible resource, specifically including:

[0185] Taking the injection power of the cluster grid-connected point, the voltage amplitude, the upper limit of the power adjustment capacity and the lower limit of the power adjustment capacity of the local distributed flexible resource cluster as the state space, the global reinforcement learning model is used to decompose the power transfer matrix into the target power adjustment amount of each local flexible resource based on a consensus algorithm.

[0186] As a preferred solution, the inter-cluster cooperative regulation module 103 is configured to decompose the power transfer matrix into the target power adjustment amount of each local flexible resource by using the global reinforcement learning model based on a consensus algorithm, specifically including:

[0187] According to the power transfer matrix, determine the current total power adjustment demand of the local distributed flexible resource cluster;

[0188] Based on a consensus algorithm, taking the same power adjustment rate of each local flexible resource as an optimization target, iteratively adjusting the power adjustment rate of each local flexible resource until the power adjustment rate of each local flexible resource converges, and obtaining an initial power adjustment rate of each local flexible resource;

[0189] Based on the product between the initial power adjustment rate of each local flexible resource and the maximum power adjustment capacity, determine an initial power adjustment amount of each local flexible resource;

[0190] According to the total power adjustment demand and each initial power adjustment amount, determine a power adjustment proportion coefficient, and based on the product between the power adjustment proportion coefficient and the initial power adjustment amount, determine the target power adjustment amount of each local flexible resource; wherein the sum of each target power adjustment amount is equal to the total power adjustment demand.

[0191] As a preferred solution, the system further comprises a global reinforcement learning model generation module, configured to:

[0192] Based on a horizontal federated learning strategy, the parameter server aggregates the local reinforcement learning model weights of the control decision-making end and each of the distributed flexible resource clusters according to the number of local training samples in the current round of training, to obtain updated global model weights. The updated global model weights are then used as the initial local reinforcement learning model weights for the control decision-making end and each of the distributed flexible resource clusters to perform the next round of training, until the reward value converges. Finally, the global reinforcement learning model is generated based on the current global model weights.

[0193] The weights of the local reinforcement learning model at the control decision-making end are generated by the control decision-making end training the local reinforcement learning model using local control data.

[0194] The weights of the local reinforcement learning model of the distributed flexible resource cluster are generated by training the local reinforcement learning model using local cluster operation data.

[0195] The reward value is calculated based on power regulation cost, power regulation over-limit penalty, and branch power flow over-limit penalty;

[0196] The specific expression for the power regulation cost is as follows:

[0197] ;

[0198] The specific expression for the power over-limit penalty is as follows:

[0199] ;

[0200] The specific expression for the branch flow over-limit penalty is as follows:

[0201] ;

[0202] in, This indicates the cost of the power regulation; This indicates the number of the distributed flexible resource clusters; Indicates the first The control cost of a distributed flexible resource cluster; Indicates the first Power regulation of a distributed flexible resource cluster; and They represent the first The upper limit and lower limit of the power regulation capacity of a distributed flexible resource cluster; This indicates the penalty for exceeding the power regulation limit; This indicates the preset penalty factor for exceeding the adjustment power limit; represents a branch flow out-of-limit punishment; represents a preset branch flow out-of-limit punishment factor; represents a branch flow; represents an upper limit of a branch flow; represents a branch set.

[0203] As a preferred solution, the system further comprises a power adjustment rate imbalance detection module, configured to:

[0204] determine an average power adjustment rate at a current time according to the power adjustment rate of each local flexible resource;

[0205] determine a power adjustment rate deviation of each local flexible resource according to a difference between the power adjustment rate of each local flexible resource and the average power adjustment rate;

[0206] obtain a communication load rate and a node residual energy at the current time;

[0207] map the power adjustment rate deviation, the communication load rate and the node residual energy into a fuzzy set through a preset event detector;

[0208] perform fuzzy reasoning on the power adjustment rate deviation, the communication load rate and the node residual energy based on a fuzzy rule corresponding to the fuzzy set, to determine a power adjustment rate imbalance triggering threshold corresponding to the power adjustment rate deviation, the communication load rate and the node residual energy; wherein the fuzzy rule is used to define a fuzzy output result corresponding to each fuzzy set for representing the power adjustment rate imbalance triggering threshold;

[0209] if the power adjustment rate deviation of each local flexible resource is detected to be less than or equal to the power adjustment rate imbalance triggering threshold, it is determined that the power adjustment rate of each local flexible resource is not imbalanced;

[0210] if the power adjustment rate deviation of any one local flexible resource is detected to be greater than the power adjustment rate imbalance triggering threshold, it is determined that the power adjustment rate of the any one local flexible resource is imbalanced.

[0211] The distributed flexible resource cluster regulation system provided by the embodiment of the present application decouples the computing task of flexible resource power regulation into cluster internal power regulation and inter-cluster collaborative regulation through a hierarchical distributed flexible resource cluster regulation architecture. When performing inter-cluster collaborative regulation, the regulation decision end only needs to be responsible for the distribution of the power transfer matrix between various distributed flexible resource clusters, and does not need to process the operation data of each flexible resource. The power regulation amount distribution of each flexible resource is responsible for each distributed flexible resource cluster, thereby effectively improving the regulation efficiency in the online regulation scenario of large-scale flexible resources, ensuring the real-time performance in the online regulation scenario of large-scale flexible resources, and meeting the online regulation requirements of large-scale flexible resources.

[0212] The above is the preferred embodiment of the present application. It should be noted that, for those skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, which are also considered within the scope of protection of the present application.

Claims

1. A distributed flexible resource cluster regulation method, characterized in that, The method comprises: obtaining the power adjustment rate of each local flexible resource belonging to a local distributed flexible resource cluster at each time point, and determining the control law of the local flexible resource based on the power adjustment rate; wherein the power adjustment rate is determined based on the ratio between the power adjustment amount of the local flexible resource at each time point and the maximum power adjustment capacity; if it is detected that none of the power adjustment rates is unbalanced, then power regulation is performed on the local flexible resource based on the power adjustment rate, the control law and the preset maximum power adjustment capacity of the local flexible resource; if it is detected that any one of the power adjustment rates is unbalanced, then a cluster intercoordination regulation request is sent to a regulation decision end; wherein the cluster intercoordination regulation request is used to instruct the regulation decision end to feed back the power transfer matrix between each distributed flexible resource cluster; according to the power transfer matrix, the power adjustment amount of each local flexible resource is optimized by using a global reinforcement learning model to determine the target power adjustment amount of each local flexible resource; wherein the global reinforcement learning model is generated based on the local reinforcement learning model weight of the regulation decision end and the local reinforcement learning model weight of each distributed flexible resource cluster.

2. The distributed flexible resource cluster regulation method of claim 1, wherein, The method further comprises: obtaining the power adjustment rate of each local flexible resource belonging to a local distributed flexible resource cluster at each time point, and determining the control law of the local flexible resource based on the power adjustment rate; wherein the power adjustment rate is determined based on the ratio between the power adjustment amount of the local flexible resource at each time point and the maximum power adjustment capacity; according to the difference between the maximum power adjustment capacity and the power adjustment amount of each local flexible resource at each time point, the residual adjustment capacity of each local flexible resource at each time point is determined; ; wherein, denotes a local flexible resource at a time instant the control law; denotes a local flexible resource at a time instant the power adjustment rate; denotes a local flexible resource at a time instant the power adjustment rate; denotes a local flexible resource a connection weight between a local flexible resource at a time instant ; denotes a neighbor set of a local flexible resource , the neighbor set comprising all local flexible resources having a physical connection relationship with the local flexible resource at a time instant .

3. The distributed flexible resource cluster regulation method of claim 2, wherein, according to the residual adjustment capacity of each local flexible resource at each time point, the connection weight between each local flexible resource is adjusted by the following expression: The method further comprises: obtaining the power adjustment rate of each local flexible resource belonging to a local distributed flexible resource cluster at each time point, and determining the control law of the local flexible resource based on the power adjustment rate; wherein the power adjustment rate is determined based on the ratio between the power adjustment amount of the local flexible resource at each time point and the maximum power adjustment capacity; ; wherein, represents a local flexible resource At the remaining regulation capacity at the time instant; represents a local flexible resource At the remaining regulation capacity at the time instant.

4. The distributed flexible resource cluster regulation method of claim 1, wherein, according to the difference between the maximum power adjustment capacity and the power adjustment amount of each local flexible resource at each time point, the residual adjustment capacity of each local flexible resource at each time point is determined; according to the residual adjustment capacity of each local flexible resource at each time point, the connection weight between each local flexible resource is adjusted by the following expression: ; The method further comprises: wherein, represents a local flexible resource at the target power adjustment amount at the time instant; represents a local flexible resource at the power adjustment rate at the time instant; represents a local flexible resource at the control law at the time instant; represents a local flexible resource at the maximum power adjustment capacity at the time instant; represents a local flexible resource at the target power adjustment amount at the time instant; represents a local flexible resource at the maximum power adjustment capacity at the time instant; represents a preset proportion coefficient.

5. The distributed flexible resource cluster regulation method of claim 1, wherein, obtaining the power adjustment rate of each local flexible resource belonging to a local distributed flexible resource cluster at each time point, and determining the control law of the local flexible resource based on the power adjustment rate; wherein the power adjustment rate is determined based on the ratio between the power adjustment amount of the local flexible resource at each time point and the maximum power adjustment capacity; according to the difference between the maximum power adjustment capacity and the power adjustment amount of each local flexible resource at each time point, the residual adjustment capacity of each local flexible resource at each time point is determined; according to the residual adjustment capacity of each local flexible resource at each time point, the connection weight between each local flexible resource is adjusted by the following expression: The power transfer matrix is generated by the regulation decision end based on the state space of the total power deviation of the distributed flexible resource cluster, the upper limit and the lower limit of the power adjustment capacity of each distributed flexible resource cluster by using the global reinforcement learning model; The power transfer matrix is used to represent the power transfer relationship and the power transfer amount between each of the distributed flexible resource clusters; and the total power deviation of the distributed flexible resource clusters is the sum of the difference between the actual power and the expected power of each of the distributed flexible resource clusters at the current time.

6. The distributed flexible resource cluster regulation method of claim 1 or 5, wherein, The target power adjustment amount of each of the local flexible resources is determined by optimizing the power adjustment amount of each of the local flexible resources using a global reinforcement learning model based on the power transfer matrix, and specifically includes: The power transfer matrix is decomposed into the target power adjustment amount of each of the local flexible resources based on a consensus algorithm using the global reinforcement learning model, with the injection power, voltage amplitude of the cluster grid-connected point, upper limit of the power adjustment capacity and lower limit of the power adjustment capacity of the local distributed flexible resource cluster as the state space.

7. The distributed flexible resource cluster regulation method of claim 6, wherein, The target power adjustment amount of each of the local flexible resources is determined by decomposing the power transfer matrix based on a consensus algorithm using the global reinforcement learning model, and specifically includes: The total power adjustment demand of the local distributed flexible resource cluster is determined based on the power transfer matrix. The power adjustment rate of each of the local flexible resources is iteratively adjusted based on a consensus algorithm, with the same power adjustment rate of each of the local flexible resources as the optimization target, until the power adjustment rate of each of the local flexible resources converges, to obtain the initial power adjustment rate of each of the local flexible resources. The initial power adjustment amount of each of the local flexible resources is determined based on the product between the initial power adjustment rate of each of the local flexible resources and the maximum power adjustment capacity. The target power adjustment amount of each of the local flexible resources is determined based on the product between the power adjustment proportion coefficient and the initial power adjustment amount, and the sum of the target power adjustment amount of each of the local flexible resources is equal to the total power adjustment demand.

8. The distributed flexible resource cluster regulation method of claim 6, wherein, The global reinforcement learning model is generated by the following steps: The local reinforcement learning model weight of the regulation and decision end and the local reinforcement learning model weight of each of the distributed flexible resource clusters are aggregated by the parameter server based on a horizontal federated learning strategy, according to the number of local training samples of the regulation and decision end and each of the distributed flexible resource clusters in the current round of training, to obtain updated global model weight, and the updated global model weight is used as the initial local reinforcement learning model weight for the regulation and decision end and each of the distributed flexible resource clusters to perform the next round of training until the reward value converges, and the global reinforcement learning model is generated based on the current global model weight; The local reinforcement learning model weight of the regulation and decision end is generated by training the local reinforcement learning model using local regulation and decision data; and The local reinforcement learning model weight of each of the distributed flexible resource clusters is generated by training the local reinforcement learning model using local regulation and decision data. The local reinforcement learning model weight of the distributed flexible resource cluster is generated by training a local reinforcement learning model using local cluster operation data; The reward value is calculated based on a power regulation cost, a regulation power out-of-limit penalty, and a branch power flow out-of-limit penalty; The expression of the power regulation cost is specifically: ; The expression of the regulation power out-of-limit penalty is specifically: ; The expression of the branch power flow out-of-limit penalty is specifically: ; wherein, denotes the power regulation cost; denotes the number of distributed flexible resource clusters; denotes the regulation cost of the th distributed flexible resource cluster; denotes the power adjustment amount of the th distributed flexible resource cluster; and denote the upper limit of power adjustment capacity and the lower limit of power adjustment capacity of the th distributed flexible resource cluster, respectively; denotes the regulation power out-of-limit penalty; denotes a preset regulation power out-of-limit penalty factor; denotes the branch power flow out-of-limit penalty; denotes a preset branch power flow out-of-limit penalty factor; denotes the power flow of the branch ; denotes the upper limit of power flow of the branch ; denotes the branch set.

9. The distributed flexible resource cluster regulation method of claim 1, wherein, The method specifically determines whether the power adjustment rates of each local flexible resource are unbalanced through the following steps: According to the power adjustment rate of each local flexible resource, an average power adjustment rate at the current time is determined; According to the difference between the power adjustment rate of each local flexible resource and the average power adjustment rate, a power adjustment rate deviation of each local flexible resource is determined; A communication load rate and a node residual energy at the current time are obtained; The power adjustment rate deviation, the communication load rate, and the node residual energy are mapped into a fuzzy set through a preset event detector; Based on the fuzzy rule corresponding to the fuzzy set, the power adjustment rate deviation, the communication load rate, and the node residual energy are subjected to fuzzy reasoning to determine a power adjustment rate imbalance triggering threshold corresponding to the power adjustment rate deviation, the communication load rate, and the node residual energy; wherein the fuzzy rule is used to define the fuzzy output result corresponding to each fuzzy set, which is used to represent the power adjustment rate imbalance triggering threshold; If it is detected that the power adjustment rate deviation of each local flexible resource is less than or equal to the power adjustment rate imbalance triggering threshold, it is determined that the power adjustment rate of each local flexible resource is not unbalanced; If it is detected that the power adjustment rate deviation of any one local flexible resource is greater than the power adjustment rate imbalance triggering threshold, it is determined that the power adjustment rate of the any one local flexible resource is unbalanced.

10. A distributed flexible resource cluster regulation system, characterized in that, The method comprises: The data acquisition module is configured to: acquire a power adjustment rate of each local flexible resource belonging to a local distributed flexible resource cluster at each time, and determine a control law of the local flexible resource based on the power adjustment rate; wherein the power adjustment rate is determined based on a ratio between a power regulation amount of the local flexible resource at each time and a maximum power regulation capacity; The intra-cluster power regulation module is configured to: If it is detected that each power adjustment rate is not unbalanced, the local flexible resource is subjected to power regulation based on the power adjustment rate of the local flexible resource, the control law, and a preset maximum power regulation capacity; The inter-cluster collaborative regulation module is configured to: If it is detected that any one power adjustment rate is unbalanced, an inter-cluster collaborative regulation request is sent to a regulation decision end; wherein the inter-cluster collaborative regulation request is used to instruct the regulation decision end to feed back a power transfer matrix between each distributed flexible resource cluster. According to the power transfer matrix, a global reinforcement learning model is used to optimize the power adjustment amount of each local flexible resource, and a target power adjustment amount of each local flexible resource is determined; wherein the global reinforcement learning model is generated based on the local reinforcement learning model weight of the regulation and control decision end and the local reinforcement learning model weight of each distributed flexible resource cluster.

Citation Information

Patent Citations

  • Park multi-flexible resource collaborative optimization control method based on consistency algorithm

    CN117578447A

  • Multi-region integrated energy system scheduling method, device, and storage medium

    WO2024087319A1