Distributed flexible resource cluster regulation and control method and system
Through the distributed flexible resource cluster control method, local power adjustment rate and control law are used to perform intra-cluster control, and inter-cluster collaborative control is performed when necessary. This solves the problem of low efficiency of large-scale flexible resource online control and realizes efficient flexible resource collaborative control.
Patent Information
- Application Number
- CN202511172194.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-08-21
AI Technical Summary
Existing technologies are unable to meet the needs of online regulation of large-scale flexible resources, and the increase in computing scale in centralized control methods leads to low efficiency.
A distributed flexible resource cluster control method is adopted to perform intra-cluster power control by obtaining the power adjustment rate and control law of local flexible resources. When necessary, inter-cluster collaborative control is requested from the control decision-making end. The power transfer matrix is optimized using a global reinforcement learning model to achieve collaborative control of flexible resources.
It improves the control efficiency and real-time performance in the online control scenario of large-scale flexible resources, and meets the online control needs of large-scale flexible resources.
Smart Images

Figure CN120728748A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power systems, and in particular to a distributed flexible resource cluster control method and system. Background Art
[0002] With the continuous advancement of new power system construction, the number of flexible resources such as distributed power sources, controllable loads, and energy storage has experienced explosive growth. Compared with traditional loads and generation-side resources, flexible resources are characterized by smaller individual capacity, larger numbers, and geographical dispersion, resulting in greater temporal and spatial uncertainty. Furthermore, due to limitations in communication and data processing capabilities, massive amounts of terminal devices cannot achieve timely interconnection with the dispatch center, directly limiting the scale and limited methods of flexible resource participation in dispatch. Against this backdrop, achieving coordinated control of massive distributed flexible resources is an urgent issue.
[0003] In the existing technology, the coordinated regulation of flexible resources mainly adopts a centralized control method. This method usually collects electricity demand information of flexible resources under its jurisdiction through load aggregators, and makes real-time centralized decisions on the power consumption of each flexible resource. However, the computing scale of this method will increase significantly with the expansion of the scale of the flexible resources being regulated, making it difficult for existing technologies to meet the online regulation needs of large-scale flexible resources. Summary of the Invention
[0004] The present invention provides a distributed flexible resource cluster control method and system to solve the technical problem that the existing technology is difficult to meet the online control requirements of large-scale flexible resources.
[0005] In order to solve the above technical problems, a first aspect of an embodiment of the present invention provides a distributed flexible resource cluster control method, including: Obtaining a power adjustment rate of each local flexible resource belonging to the local distributed flexible resource cluster at each moment, and determining a control law for the local flexible resource based on the power adjustment rate; If it is detected that none of the power adjustment rates are unbalanced, power regulation is performed on the local flexible resource based on the power adjustment rate of the local flexible resource, the control law, and a preset maximum power regulation capacity; If any power adjustment rate imbalance is detected, an inter-cluster collaborative control request is sent to the control decision end; wherein the inter-cluster collaborative control request is used to instruct the control decision end to feedback the power transfer matrix between each distributed flexible resource cluster; According to the power transfer matrix, the power regulation amount of each of the local flexible resources is optimized using a global reinforcement learning model to determine the target power regulation amount of each of the local flexible resources; wherein, the global reinforcement learning model is generated based on the local reinforcement learning model weights of the control decision end and each of the distributed flexible resource clusters.
[0006] As a preferred solution, obtaining the power adjustment rate of each local flexible resource belonging to the local distributed flexible resource cluster at each moment and determining the control law of the local flexible resource based on the power adjustment rate specifically includes: Obtaining the power adjustment rate of each of the local flexible resources at each moment according to a ratio between the power adjustment amount of each of the local flexible resources at each moment and the maximum power adjustment capacity; According to the power adjustment rate of each local flexible resource at each moment, the control law of the local flexible resource is determined by the following expression: ; in, Represents local flexible resources exist The control law at time t; Represents local flexible resources exist The power regulation rate at the time; Represents local flexible resources exist The power regulation rate at the time; Represents local flexible resources With local flexible resources between The connection weight at the moment; Represents local flexible resources The neighbor set of Time and local flexible resources All local flexible resources that have a physical connection relationship.
[0007] As a preferred solution, the method further comprises: determining a remaining adjustment capacity of each of the local flexible resources at each moment according to a difference between the maximum power adjustment capacity and the power adjustment amount of each of the local flexible resources at each moment; According to the remaining adjustment capacity of each local flexible resource at each moment, the connection weights between the local flexible resources are adjusted by the following expression: ; in, Represents local flexible resources exist The remaining regulating capacity at the time; Represents local flexible resources exist The remaining regulation capacity at the moment.
[0008] As a preferred solution, the power regulation of the local flexible resource based on the power regulation rate of the local flexible resource, the control law, and the preset maximum power regulation capacity specifically includes: Based on the power regulation rate of the local flexible resource, the control law, and the maximum power regulation capacity, a target power regulation amount of the local flexible resource is determined by the following expression: ; performing power regulation on the local flexible resource according to the target power regulation amount of the local flexible resource; in, Represents local flexible resources exist The target power adjustment amount at the time; Represents local flexible resources exist The power regulation rate at the time; Represents local flexible resources exist The control law at time t; Represents local flexible resources exist The maximum power regulation capacity at the time; Indicates the preset scale factor.
[0009] As a preferred solution, the power transfer matrix is generated by the control decision end using the total power deviation of the distributed flexible resource cluster, the upper limit and lower limit of the power regulation capacity of each distributed flexible resource cluster as the state space, using the global reinforcement learning model based on the state space; Among them, the power transfer matrix is used to represent the power transfer relationship and power transfer amount between each of the distributed flexible resource clusters; the total power deviation of the distributed flexible resource cluster is the sum of the differences between the actual power and the expected power of each of the distributed flexible resource clusters at the current moment.
[0010] As a preferred solution, optimizing the power adjustment amount of each of the local flexible resources using a global reinforcement learning model according to the power transfer matrix to determine the target power adjustment amount of each of the local flexible resources specifically includes: Taking the injected power, voltage amplitude of the cluster grid connection point, the upper limit and lower limit of the power regulation capacity of the local distributed flexible resource cluster as the state space, the global reinforcement learning model is used to decompose the power transfer matrix into the target power regulation amount of each local flexible resource based on the consistency algorithm.
[0011] As a preferred solution, decomposing the power transfer matrix into target power adjustment amounts of each of the local flexible resources based on a consistency algorithm using the global reinforcement learning model specifically includes: Determining a current total power regulation requirement of the local distributed flexible resource cluster according to the power transfer matrix; Based on a consistency algorithm, taking the same power adjustment rate of each local flexible resource as an optimization goal, iteratively adjusting the power adjustment rate of each local flexible resource until the power adjustment rate of each local flexible resource converges, thereby obtaining an initial power adjustment rate of each local flexible resource; determining an initial power adjustment amount of each of the local flexible resources based on a product of the initial power adjustment rate of each of the local flexible resources and the maximum power adjustment capacity; According to the total power regulation requirement and each of the initial power regulation amounts, a power adjustment proportional coefficient is determined, and based on the product of the power adjustment proportional coefficient and the initial power regulation amount, the target power adjustment amount of each of the local flexible resources is determined; wherein the sum of each of the target power adjustment amounts is equal to the total power regulation requirement.
[0012] As a preferred solution, the method specifically generates the global reinforcement learning model through the following steps: Based on the horizontal federated learning strategy, the parameter server aggregates the local reinforcement learning model weights of the control decision-making end and each of the distributed flexible resource clusters according to the number of local training samples of the control decision-making end and each of the distributed flexible resource clusters in the current round of training to obtain an updated global model weight, and uses the updated global model weight as the initial local reinforcement learning model weight for the control decision-making end and each of the distributed flexible resource clusters to perform the next round of training until the reward value converges, and generates the global reinforcement learning model according to the current global model weight; The weight of the local reinforcement learning model of the control decision end is generated by the control decision end training the local reinforcement learning model using the local control data; The local reinforcement learning model weights of the distributed flexible resource cluster are generated by the distributed flexible resource cluster training the local reinforcement learning model using local cluster operation data; The reward value is calculated based on the power control cost, the power over-limit penalty and the branch power over-limit penalty; The power control cost is specifically expressed as: ; The expression for adjusting the power limit penalty is specifically: ; The expression of the branch power flow exceeding limit penalty is specifically: ; in, represents the power regulation cost; Indicates the number of distributed flexible resource clusters; Indicates the The cost of regulating a distributed flexible resource cluster; Indicates the The power regulation of a distributed flexible resource cluster; and Respectively represent The upper and lower limits of the power regulation capacity of a distributed flexible resource cluster; Indicates the penalty for exceeding the adjustment power limit; Indicates the preset penalty factor for regulating power exceeding the limit; Indicates the penalty for exceeding the limit of the branch power flow; Indicates the preset branch power flow over-limit penalty factor; Indicates a branch The trend; Indicates a branch The upper limit of the trend; Indicates a branch set.
[0013] As a preferred solution, the method specifically determines whether each power regulation rate is unbalanced by the following steps: determining an average power regulation rate at a current moment according to the power regulation rates of the respective local flexible resources; determining a power regulation rate deviation of each of the local flexible resources according to a difference between the power regulation rate of each of the local flexible resources and the average power regulation rate; Obtain the current communication load rate and node remaining energy; Mapping the power regulation rate deviation, the communication load rate, and the node residual energy into a fuzzy set through a preset event detector; Based on the fuzzy rules corresponding to the fuzzy sets, fuzzy reasoning is performed on the power regulation rate deviation, the communication load rate, and the node residual energy to determine a power regulation rate imbalance trigger threshold corresponding to the power regulation rate deviation, the communication load rate, and the node residual energy; wherein the fuzzy rules are used to define a fuzzy output result corresponding to each fuzzy set for representing the power regulation rate imbalance trigger threshold; If it is detected that the power regulation rate deviations of the local flexible resources are all less than or equal to the power regulation rate imbalance trigger threshold, it is determined that the power regulation rates of the local flexible resources are not unbalanced; If it is detected that the power regulation rate deviation of any one of the local flexible resources is greater than the power regulation rate imbalance trigger threshold, it is determined that the power regulation rate of the any one of the local flexible resources is unbalanced.
[0014] A second aspect of an embodiment of the present invention provides a distributed flexible resource cluster control system, including: Data acquisition module, used to: Obtaining a power adjustment rate of each local flexible resource belonging to the local distributed flexible resource cluster at each moment, and determining a control law for the local flexible resource based on the power adjustment rate; The power control module within the cluster is used to: If it is detected that none of the power adjustment rates are unbalanced, power regulation is performed on the local flexible resource based on the power adjustment rate of the local flexible resource, the control law, and a preset maximum power regulation capacity; Inter-cluster collaborative control module, used for: If any power adjustment rate imbalance is detected, an inter-cluster collaborative control request is sent to the control decision end; wherein the inter-cluster collaborative control request is used to instruct the control decision end to feedback the power transfer matrix between each distributed flexible resource cluster; According to the power transfer matrix, the power regulation amount of each of the local flexible resources is optimized using a global reinforcement learning model to determine the target power regulation amount of each of the local flexible resources; wherein, the global reinforcement learning model is generated based on the local reinforcement learning model weights of the control decision end and each of the distributed flexible resource clusters.
[0015] Compared with the existing technology, the beneficial effect of the embodiments of the present invention is that, through the hierarchical distributed flexible resource cluster control architecture, the computing task of flexible resource power control is decoupled into intra-cluster power control and inter-cluster collaborative control. When performing inter-cluster collaborative control, the control decision-making end is only responsible for the distribution of the power transfer matrix between each distributed flexible resource cluster, without the need to process the operating data of each flexible resource. The power adjustment amount allocation of each flexible resource is the responsibility of each distributed flexible resource cluster, thereby effectively improving the control efficiency in the online control scenario of large-scale flexible resources, ensuring the real-time performance in the online control scenario of large-scale flexible resources, and being able to meet the online control needs of large-scale flexible resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 Schematic diagram of the distributed flexible resource cluster control method according to an embodiment of the present invention; Figure 2 Schematic diagram of a distributed flexible resource cluster control architecture in an embodiment of the present invention; Figure 3 Schematic diagram of an event-triggered communication mechanism based on sampled data in an embodiment of the present invention; Figure 4 It is a structural diagram of a distributed flexible resource cluster control system in an embodiment of the present invention. DETAILED DESCRIPTION
[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0018] See Figure 1 A first aspect of an embodiment of the present invention provides a distributed flexible resource cluster control method, comprising the following steps S1 to S4: Step S1, obtaining a power adjustment rate of each local flexible resource belonging to a local distributed flexible resource cluster at each moment, and determining a control law of the local flexible resource based on the power adjustment rate; Step S2: If it is detected that none of the power regulation rates are unbalanced, power regulation is performed on the local flexible resource based on the power regulation rate of the local flexible resource, the control law, and a preset maximum power regulation capacity; Step S3: If any power regulation rate imbalance is detected, an inter-cluster collaborative regulation request is sent to the regulation decision end; wherein the inter-cluster collaborative regulation request is used to instruct the regulation decision end to feed back the power transfer matrix between each distributed flexible resource cluster; Step S4, according to the power transfer matrix, use the global reinforcement learning model to optimize the power adjustment amount of each of the local flexible resources, and determine the target power adjustment amount of each of the local flexible resources; wherein, the global reinforcement learning model is generated based on the local reinforcement learning model weights of the control decision end and each of the distributed flexible resource clusters.
[0019] It is worth noting that the distributed flexible resource cluster control architecture diagram in this embodiment is as follows Figure 2 As shown, the distributed flexible resource cluster control architecture is divided into two layers: the control layer and the cluster layer. The control decision-making end is located in the upper control layer. It takes each distributed flexible resource cluster as a whole as the control object to achieve coordinated control between clusters in the region. Each distributed flexible resource cluster together constitutes the cluster layer located in the lower layer. The flexible resources within each distributed flexible resource cluster are allocated power adjustment by the sub-group scheduling control unit corresponding to the cluster. Therefore, each distributed flexible resource cluster responds to the control instructions of the control layer while achieving intelligent self-adaptation within the cluster. Each terminal within the distributed flexible resource cluster is an independent flexible resource terminal. To achieve this hierarchical and collaborative distributed flexible resource cluster control architecture, it is first necessary to cluster the flexible resources in the region according to the flexible resource type, geographical distribution, or the stakeholders and subordinate relationships, thereby forming each distributed flexible resource cluster. For example, the flexible resources within the power supply range are divided into a local cluster based on the power supply radius of the substation. This embodiment will not be described in detail here.
[0020] Furthermore, the main task of power regulation within the cluster is to quickly allocate the power regulation amount of each local flexible resource. This embodiment implements power regulation of each local flexible resource based on a consistency algorithm, selects the power regulation rate of each local flexible resource at each moment as a consistency state variable, and uses this to determine the control law of each local flexible resource. It can be understood that the control law of each local flexible resource is used to guide each local flexible resource to dynamically adjust its own control output based on the power regulation rate difference between it and its neighboring flexible resources, so as to achieve convergence and balance of power distribution within the cluster. If the current power regulation rates are not unbalanced, that is, the power regulation rates within the cluster can be balanced through power regulation within the cluster, then the power regulation within the cluster is maintained, that is, based on the power regulation rate, control law and maximum power regulation capacity of each local flexible resource, the power regulation of each local flexible resource is performed, so as to achieve convergence and balance of power distribution within the cluster while meeting the power regulation capacity constraints of each local flexible resource.
[0021] Furthermore, if any of the current power adjustment rates is unbalanced, that is, the power adjustment rates within the cluster cannot be balanced, it is necessary to trigger interaction with adjacent clusters. First, an inter-cluster collaborative control request is sent to the control decision-making end to instruct the control decision-making end to feedback the power transfer matrix between each distributed flexible resource cluster. It can be understood that the main task of inter-cluster collaborative control is the dynamic exchange of power between clusters. This embodiment adopts a multi-agent hierarchical reinforcement learning algorithm based on federated learning to realize the interaction between multiple distributed flexible resource clusters. This algorithm process divides the inter-cluster collaborative control architecture into a regional coordination layer (macro power allocation) and a cluster execution layer (micro power allocation), thereby significantly reducing the processing complexity of distributed flexible resource power control and avoiding a single control unit from processing the operating data of large-scale flexible resources at the same time. Specifically, the control decision-making end serves as the regional coordination layer in the above-mentioned inter-cluster collaborative control architecture. Its feedback of the power transfer matrix between each distributed flexible resource cluster is the process of executing macro power allocation. Each distributed flexible resource cluster serves as the cluster execution layer in the above-mentioned inter-cluster collaborative control architecture. After receiving the power transfer matrix sent by the control decision-making end, this embodiment uses a global reinforcement learning model to optimize the power adjustment amount of each local flexible resource to determine the target power adjustment amount of each local flexible resource and realize micro power allocation. Among them, the global reinforcement learning model in this embodiment is generated by aggregating the weights of the local reinforcement learning models of the control decision-making end and each distributed flexible resource cluster based on the federated learning strategy.
[0022] The distributed flexible resource cluster control method provided in an embodiment of the present invention decouples the computational task of flexible resource power control into intra-cluster power control and inter-cluster collaborative control through a hierarchical distributed flexible resource cluster control architecture. When performing inter-cluster collaborative control, the control decision-making end is only responsible for sending the power transfer matrix between each distributed flexible resource cluster, without having to process the operating data of each flexible resource. The power adjustment amount allocation of each flexible resource is the responsibility of each distributed flexible resource cluster, thereby effectively improving the control efficiency in the online control scenario of large-scale flexible resources, ensuring the real-time performance in the online control scenario of large-scale flexible resources, and being able to meet the online control needs of large-scale flexible resources.
[0023] As a preferred solution, obtaining the power adjustment rate of each local flexible resource belonging to the local distributed flexible resource cluster at each moment and determining the control law of the local flexible resource based on the power adjustment rate specifically includes: Obtaining the power adjustment rate of each of the local flexible resources at each moment according to a ratio between the power adjustment amount of each of the local flexible resources at each moment and the maximum power adjustment capacity; According to the power adjustment rate of each local flexible resource at each moment, the control law of the local flexible resource is determined by the following expression: ; in, Represents local flexible resources exist The control law at time t; Represents local flexible resources exist The power regulation rate at the time; Represents local flexible resources exist The power regulation rate at the time; Represents local flexible resources With local flexible resources between The connection weight at the moment; Represents local flexible resources The neighbor set of Time and local flexible resources All local flexible resources that have a physical connection relationship.
[0024] Specifically, this embodiment obtains the power adjustment rate of each local flexible resource at each moment through the following expression based on the ratio between the power adjustment amount of each local flexible resource at each moment and the maximum power adjustment capacity: ; in, Represents local flexible resources exist Power regulation rate at all times; Represents local flexible resources exist The power adjustment at each moment; Represents local flexible resources exist The maximum power regulation capacity at any moment.
[0025] As a preferred solution, the method further comprises: determining a remaining adjustment capacity of each of the local flexible resources at each moment according to a difference between the maximum power adjustment capacity and the power adjustment amount of each of the local flexible resources at each moment; According to the remaining adjustment capacity of each local flexible resource at each moment, the connection weights between the local flexible resources are adjusted by the following expression: ; in, Represents local flexible resources exist The remaining regulating capacity at the time; Represents local flexible resources exist The remaining regulation capacity at the moment.
[0026] Specifically, in this embodiment, the connection weight between each local flexible resource is an adaptive weight coefficient, which is dynamically adjusted according to the remaining regulation capacity of the local flexible resource, so that the connection weight between each local flexible resource takes into account the real-time regulation capability of the local flexible resource, ensuring fairer and more effective power distribution.
[0027] As a preferred solution, the power regulation of the local flexible resource based on the power regulation rate of the local flexible resource, the control law, and the preset maximum power regulation capacity specifically includes: Based on the power regulation rate of the local flexible resource, the control law, and the maximum power regulation capacity, a target power regulation amount of the local flexible resource is determined by the following expression: ; performing power regulation on the local flexible resource according to the target power regulation amount of the local flexible resource; in, Represents local flexible resources exist The target power adjustment amount at the time; Represents local flexible resources exist The power regulation rate at the time; Represents local flexible resources exist The control law at time t; Represents local flexible resources exist The maximum power regulation capacity at the time; Indicates the preset scale factor.
[0028] Specifically, for any local flexible resource, according to the calculated control law, the increment or decrement of its power adjustment can be determined by the above expression, wherein the increment or decrement of the power adjustment can be proportional to the control law, so this embodiment introduces a proportional coefficient , the above expression is used to calculate the target power regulation amount of each local flexible resource.
[0029] As a preferred solution, the power transfer matrix is generated by the control decision end using the total power deviation of the distributed flexible resource cluster, the upper limit and lower limit of the power regulation capacity of each distributed flexible resource cluster as the state space, using the global reinforcement learning model based on the state space; Among them, the power transfer matrix is used to represent the power transfer relationship and power transfer amount between each of the distributed flexible resource clusters; the total power deviation of the distributed flexible resource cluster is the sum of the differences between the actual power and the expected power of each of the distributed flexible resource clusters at the current moment.
[0030] Specifically, the control decision-making end in this embodiment also generates the power transfer matrix of the current state space based on the hierarchical reinforcement learning algorithm. First, the total power deviation of the distributed flexible resource cluster, the upper limit of the power regulation capacity of each distributed flexible resource cluster and the lower limit of the power regulation capacity are used as the state space. Then, under the constraints of this state space, the trained global reinforcement learning model is used to generate the corresponding power transfer matrix as the action space. It can be understood that the power transfer matrix in this embodiment is used to represent the power transfer relationship and power transfer amount between each distributed flexible resource cluster. Assume that an element in the power transfer matrix is represented as , then this element represents the distributed flexible resource cluster Towards distributed flexible resource clusters The power transfer amount is represented by a positive value, indicating output power, while a negative value indicates absorbed power. Furthermore, since the power transfer matrix clearly defines the power transfer relationship and power transfer amount between each distributed flexible resource cluster, the expected power of each distributed flexible resource cluster at the next moment can be determined by combining the current power of each distributed flexible resource cluster. Furthermore, the total power deviation of each distributed flexible resource cluster can be detected by combining the actual power of each distributed flexible resource cluster at each moment.
[0031] As a preferred solution, optimizing the power adjustment amount of each of the local flexible resources using a global reinforcement learning model according to the power transfer matrix to determine the target power adjustment amount of each of the local flexible resources specifically includes: Taking the injected power, voltage amplitude of the cluster grid connection point, the upper limit and lower limit of the power regulation capacity of the local distributed flexible resource cluster as the state space, the global reinforcement learning model is used to decompose the power transfer matrix into the target power regulation amount of each local flexible resource based on the consistency algorithm.
[0032] Specifically, the local distributed flexible resource cluster serves as the cluster execution layer and needs to complete the power allocation within the cluster according to the power transfer matrix. First, this embodiment uses the injected power, voltage amplitude, upper limit and lower limit of the power regulation capacity of the cluster grid connection point as the state space. Then, under the constraints of this state space, the trained global reinforcement learning model is used to decompose the power transfer matrix into the target power regulation amount of each local flexible resource based on the consistency algorithm, so that the sum of the target power regulation amount of each local flexible resource meets the total power regulation requirement required by the current power transfer matrix.
[0033] As a preferred solution, decomposing the power transfer matrix into target power adjustment amounts of each of the local flexible resources based on a consistency algorithm using the global reinforcement learning model specifically includes: Determining a current total power regulation requirement of the local distributed flexible resource cluster according to the power transfer matrix; Based on a consistency algorithm, taking the same power adjustment rate of each local flexible resource as an optimization goal, iteratively adjusting the power adjustment rate of each local flexible resource until the power adjustment rate of each local flexible resource converges, thereby obtaining an initial power adjustment rate of each local flexible resource; determining an initial power adjustment amount of each of the local flexible resources based on a product of the initial power adjustment rate of each of the local flexible resources and the maximum power adjustment capacity; According to the total power regulation requirement and each of the initial power regulation amounts, a power adjustment proportional coefficient is determined, and based on the product of the power adjustment proportional coefficient and the initial power regulation amount, the target power adjustment amount of each of the local flexible resources is determined; wherein the sum of each of the target power adjustment amounts is equal to the total power regulation requirement.
[0034] Specifically, this embodiment first calculates the current total power regulation demand of the local distributed flexible resource cluster based on the power transfer matrix issued by the control decision end. , as shown in the following expression: ; in, Indicates that the local distributed flexible resource cluster To the rest of the distributed flexible resource clusters The amount of power transfer; and the total power regulation requirement That is, from the local distributed flexible resource cluster To all other distributed flexible resource clusters The total power transfer.
[0035] Furthermore, based on the consistency algorithm, the optimization goal is to make the power adjustment rate of each local flexible resource the same, using the expression: As an iterative formula, the power adjustment rate of each local flexible resource is gradually adjusted until the power adjustment rate of each local flexible resource converges. It is worth noting that since the consistency algorithm aims to converge the power adjustment rate, it does not guarantee that the sum of the power adjustment amounts of each local flexible resource is equal to the above-mentioned total power adjustment demand under the power adjustment rate after convergence. Therefore, it is necessary to adjust the initial power adjustment amount after convergence. Assuming that the power adjustment rate of each local flexible resource has been determined after the consistency algorithm iteration, the initial power adjustment amount is shown in the following expression: ; in, Represents local flexible resources The initial power adjustment amount; Represents local flexible resources Power regulation rate after consistency algorithm iteration.
[0036] Then the initial power regulation amount is proportionally adjusted to meet the various target power regulation amounts. The sum of the power regulation requirements is equal to the total ,Right now , the proportional adjustment of the initial power regulation is shown in the following expression: ; in, Represents a local distributed flexible resource cluster; Represents local flexible resources The target power regulation amount; That is, the power adjustment proportional coefficient determined according to the total power adjustment demand and each initial power adjustment amount.
[0037] As a preferred solution, the method specifically generates the global reinforcement learning model through the following steps: Based on the horizontal federated learning strategy, the parameter server aggregates the local reinforcement learning model weights of the control decision-making end and each of the distributed flexible resource clusters according to the number of local training samples of the control decision-making end and each of the distributed flexible resource clusters in the current round of training to obtain an updated global model weight, and uses the updated global model weight as the initial local reinforcement learning model weight for the control decision-making end and each of the distributed flexible resource clusters to perform the next round of training until the reward value converges, and generates the global reinforcement learning model according to the current global model weight; The weight of the local reinforcement learning model of the control decision end is generated by the control decision end training the local reinforcement learning model using the local control data; The local reinforcement learning model weights of the distributed flexible resource cluster are generated by the distributed flexible resource cluster training the local reinforcement learning model using local cluster operation data; The reward value is calculated based on the power control cost, the power over-limit penalty and the branch power over-limit penalty; The power control cost is specifically expressed as: ; The expression for adjusting the power limit penalty is specifically: ; The expression of the branch power flow exceeding limit penalty is specifically: ; in, represents the power regulation cost; Indicates the number of distributed flexible resource clusters; Indicates the The cost of regulating a distributed flexible resource cluster; Indicates the The power regulation of a distributed flexible resource cluster; and Respectively represent The upper and lower limits of the power regulation capacity of a distributed flexible resource cluster; Indicates the penalty for exceeding the adjustment power limit; Indicates the preset penalty factor for regulating power exceeding the limit; Indicates the penalty for exceeding the limit of the branch power flow; Indicates the preset branch power flow over-limit penalty factor; Indicates a branch The trend; Indicates a branch The upper limit of the trend; Indicates a branch set.
[0038] It is worth noting that traditional centralized multi-agent reinforcement learning requires collecting the original operating data of each cluster (such as power, voltage, and regulation capacity) to the central server for model training. This may leak the commercial secrets of the cluster's entities (such as user power usage patterns and equipment operating status) or sensitive information (such as geographical location load characteristics). Therefore, this embodiment adopts a horizontal federated learning strategy to train the multi-agent reinforcement learning model, regulating the decision-making end and each distributed flexible resource cluster to only upload the local reinforcement learning model weights rather than the original data, thereby solving the privacy leakage problem of the centralized multi-agent reinforcement learning algorithm. The steps of the horizontal federated learning strategy in this embodiment are as follows: (1) First, the control decision-making end uses local control data (such as the total power deviation of the distributed flexible resource cluster, the upper limit and lower limit of the power regulation capacity of each distributed flexible resource cluster, etc.) to train the local reinforcement learning model and generate the weight of the local reinforcement learning model of the control decision-making end; each distributed flexible resource cluster uses local cluster operation data (such as the injected power and voltage amplitude of the cluster grid connection point, the upper limit and lower limit of the power regulation capacity of the local distributed flexible resource cluster, the power transfer matrix received from the control decision-making end, etc.) to train the local reinforcement learning model and generate the weight of the local reinforcement learning model of the distributed flexible resource cluster.
[0039] (2) The decision-making end and each distributed flexible resource cluster regularly upload the latest local reinforcement learning model weights generated by training to the parameter server instead of uploading specific local training samples.
[0040] (3) The parameter server aggregates the local reinforcement learning model weights of the control decision-making end and each distributed flexible resource cluster according to the number of local training samples of the control decision-making end and each distributed flexible resource cluster in the current round of training to obtain the updated global model weights. Assume that in the first round of federated training, there are clusters participate, each cluster Train to obtain a set of updated local reinforcement learning model weights The aggregation process of the global model weight is expressed as follows: ; in, represents the updated global model weight; Indicates the The number of local training samples in a distributed flexible resource cluster; Represents the total number of local training samples across all distributed flexible resource clusters participating in the current round of training. Therefore, the weight of each local reinforcement learning model during aggregation is determined by the ratio of the number of local training samples of the corresponding distributed flexible resource cluster to the total number of local training samples. Distributed flexible resource clusters with more local training samples contribute more to the aggregation of global model weights.
[0041] (4) The parameter server distributes the updated global model weights to the control decision-making end and each distributed flexible resource cluster to serve as the initial local reinforcement learning model weights for the control decision-making end and each distributed flexible resource cluster to perform the next round of training.
[0042] (5) Repeat the above process until the reward value converges, which means the global reinforcement learning model converges. Generate a global reinforcement learning model based on the current global model weights.
[0043] Specifically, at each time step, the control decision-making end, as the upper-level decision-making unit, first observes the total power deviation of the distributed flexible resource cluster, the upper and lower limits of the power regulation capacity of each distributed flexible resource cluster as the state space, and generates a power transfer matrix as the action space based on this, and sends the power transfer matrix to the distributed flexible resource cluster; the distributed flexible resource cluster, as the lower-level execution unit, observes the injected power, voltage amplitude, upper and lower limits of the power regulation capacity of the cluster grid connection point as the state space after receiving the power transfer matrix, and decomposes the power transfer matrix into the target power regulation amount of each local flexible resource based on the consistency algorithm and executes the action; after executing the action, the environment calculates the reward value based on the power regulation cost, the penalty for over-limit of regulation power and the penalty for over-limit of branch flow, the control decision-making end and each distributed flexible resource cluster obtain the reward value and observe the new state, and continuously optimize their own strategies through continuous interaction with the environment, so that the cumulative reward gradually increases and tends to be stable, and finally completes the model training.
[0044] The global reinforcement learning model training method provided in this embodiment maintains cross-cluster collaborative performance while eliminating the risk of commercial secret leakage, meeting the privacy protection requirements of multi-stakeholder scenarios.
[0045] As a preferred solution, the method specifically determines whether each power regulation rate is unbalanced by the following steps: determining an average power regulation rate at a current moment according to the power regulation rates of the respective local flexible resources; determining a power regulation rate deviation of each of the local flexible resources according to a difference between the power regulation rate of each of the local flexible resources and the average power regulation rate; Obtain the current communication load rate and node remaining energy; Mapping the power regulation rate deviation, the communication load rate, and the node residual energy into a fuzzy set through a preset event detector; Based on the fuzzy rules corresponding to the fuzzy sets, fuzzy reasoning is performed on the power regulation rate deviation, the communication load rate, and the node residual energy to determine a power regulation rate imbalance trigger threshold corresponding to the power regulation rate deviation, the communication load rate, and the node residual energy; wherein the fuzzy rules are used to define a fuzzy output result corresponding to each fuzzy set for representing the power regulation rate imbalance trigger threshold; If it is detected that the power regulation rate deviations of the local flexible resources are all less than or equal to the power regulation rate imbalance trigger threshold, it is determined that the power regulation rates of the local flexible resources are not unbalanced; If it is detected that the power regulation rate deviation of any one of the local flexible resources is greater than the power regulation rate imbalance trigger threshold, it is determined that the power regulation rate of the any one of the local flexible resources is unbalanced.
[0046] Specifically, since the hierarchical distributed flexible resource cluster control system in this embodiment does not require frequent information exchange during steady-state operation, that is, each distributed flexible resource cluster performs power control within the cluster without interacting with other distributed flexible resource clusters, the traditional time-triggered communication mechanism has serious waste of communication resources and computing resources. In order to reduce the communication and computing workload of the control system, this embodiment proposes an event-triggered communication mechanism based on sampled data, such as Figure 3 As shown, each distributed flexible resource cluster only triggers interaction with adjacent distributed flexible resource clusters when it meets the predetermined power regulation rate imbalance event triggering conditions, that is, when the power change rate within the cluster cannot be balanced, and does not participate in interaction at other times. Among them, the sampler is responsible for measuring the status of the distributed flexible resource cluster, including the power regulation rate deviation, communication load rate, and node residual energy of each local flexible resource, and sending it to the event detector. The event detector determines whether the power regulation rate imbalance event triggering conditions are met. If so, it triggers the communication network between the local distributed flexible resource cluster and its adjacent distributed flexible resource clusters and updates the controller for inter-cluster coordinated control.
[0047] Furthermore, this embodiment introduces fuzzy logic to handle nonlinear, uncertain multi-parameter coupling problems, converting qualitative descriptions of power regulation rate deviation, communication load rate, and node residual energy into quantitative power regulation rate imbalance trigger thresholds, thereby achieving dynamic adaptive regulation. The design logic of the event detector is as follows: (1) Input variables: including the power regulation rate deviation of each local flexible resource , communication load rate and node residual energy , the specific expression is as follows: ; ; ; in, Indicates the average power adjustment rate of the distributed flexible resource cluster; Indicates the data transmission rate of the current link; Indicates the maximum available bandwidth of the link; communication load rate reflects the degree of congestion of the communication link. Indicates link saturation. In this case, unnecessary communication should be avoided. Indicates the current remaining power of the node; Indicates the total battery capacity of the node; the remaining energy of the node It reflects the energy status of edge nodes or terminal devices. When the battery is low, communication needs to be reduced to extend the life of the device.
[0048] (2) Output variable: Power regulation rate imbalance trigger threshold , ranging from 0 to 1, when Communication is triggered.
[0049] (3) Fuzzy reasoning: First, the input parameters (including power regulation rate deviation, communication load rate, and node residual energy) are mapped to fuzzy sets (such as triangular fuzzy membership functions corresponding to "large", "medium", and "small"); second, the corresponding fuzzy rules are activated according to the input membership; finally, the fuzzy output results are converted into definite values through the center of gravity method. Among them, the logical relationship between multiple parameters is described by "if-then" fuzzy rules. Typical fuzzy rules are shown in Table 1 below: Table 1 Fuzzy rule definitions This embodiment utilizes an event-triggered communication mechanism based on sampled data and fuzzy processing, which can extend the life of edge devices and avoid information storms under network congestion conditions.
[0050] See Figure 4A second aspect of an embodiment of the present invention provides a distributed flexible resource cluster control system, including: The data acquisition module 101 is used to: Obtaining a power adjustment rate of each local flexible resource belonging to the local distributed flexible resource cluster at each moment, and determining a control law for the local flexible resource based on the power adjustment rate; The intra-cluster power control module 102 is configured to: If it is detected that none of the power adjustment rates are unbalanced, power regulation is performed on the local flexible resource based on the power adjustment rate of the local flexible resource, the control law, and a preset maximum power regulation capacity; The inter-cluster collaborative control module 103 is used to: If any power adjustment rate imbalance is detected, an inter-cluster collaborative control request is sent to the control decision end; wherein the inter-cluster collaborative control request is used to instruct the control decision end to feedback the power transfer matrix between each distributed flexible resource cluster; According to the power transfer matrix, the power regulation amount of each of the local flexible resources is optimized using a global reinforcement learning model to determine the target power regulation amount of each of the local flexible resources; wherein, the global reinforcement learning model is generated based on the local reinforcement learning model weights of the control decision end and each of the distributed flexible resource clusters.
[0051] As a preferred solution, the data acquisition module 101 is used to obtain the power adjustment rate of each local flexible resource belonging to the local distributed flexible resource cluster at each moment, and determine the control law of the local flexible resource based on the power adjustment rate, specifically including: Obtaining the power adjustment rate of each of the local flexible resources at each moment according to a ratio between the power adjustment amount of each of the local flexible resources at each moment and the maximum power adjustment capacity; According to the power adjustment rate of each local flexible resource at each moment, the control law of the local flexible resource is determined by the following expression: ; in, Represents local flexible resources exist The control law at time t; Represents local flexible resources exist The power regulation rate at the time; Represents local flexible resources exist The power regulation rate at the time; Represents local flexible resources With local flexible resources between The connection weight at the moment; Represents local flexible resources The neighbor set of Time and local flexible resources All local flexible resources that have a physical connection relationship.
[0052] As a preferred solution, the data acquisition module 101 is further configured to: determining a remaining adjustment capacity of each of the local flexible resources at each moment according to a difference between the maximum power adjustment capacity and the power adjustment amount of each of the local flexible resources at each moment; According to the remaining adjustment capacity of each local flexible resource at each moment, the connection weights between the local flexible resources are adjusted by the following expression: ; in, Represents local flexible resources exist The remaining regulating capacity at the time; Represents local flexible resources exist The remaining regulation capacity at the moment.
[0053] As a preferred solution, the intra-cluster power control module 102 is configured to perform power control on the local flexible resource based on the power adjustment rate of the local flexible resource, the control law, and the preset maximum power regulation capacity, specifically including: Based on the power regulation rate of the local flexible resource, the control law, and the maximum power regulation capacity, a target power regulation amount of the local flexible resource is determined by the following expression: ; performing power regulation on the local flexible resource according to the target power regulation amount of the local flexible resource; in, Represents local flexible resources exist The target power adjustment amount at the time; Represents local flexible resources exist The power regulation rate at the time; Represents local flexible resources exist The control law at time t; Represents local flexible resources exist The maximum power regulation capacity at the time; Indicates the preset scale factor.
[0054] As a preferred solution, the power transfer matrix is generated by the control decision end using the total power deviation of the distributed flexible resource cluster, the upper limit and lower limit of the power regulation capacity of each distributed flexible resource cluster as the state space, using the global reinforcement learning model based on the state space; Among them, the power transfer matrix is used to represent the power transfer relationship and power transfer amount between each of the distributed flexible resource clusters; the total power deviation of the distributed flexible resource cluster is the sum of the differences between the actual power and the expected power of each of the distributed flexible resource clusters at the current moment.
[0055] As a preferred solution, the inter-cluster collaborative control module 103 is configured to optimize the power adjustment amount of each of the local flexible resources according to the power transfer matrix using a global reinforcement learning model to determine the target power adjustment amount of each of the local flexible resources, specifically including: Taking the injected power, voltage amplitude of the cluster grid connection point, the upper limit and lower limit of the power regulation capacity of the local distributed flexible resource cluster as the state space, the global reinforcement learning model is used to decompose the power transfer matrix into the target power regulation amount of each local flexible resource based on the consistency algorithm.
[0056] As a preferred solution, the inter-cluster collaborative control module 103 is configured to decompose the power transfer matrix into target power adjustment amounts of each of the local flexible resources based on a consensus algorithm using the global reinforcement learning model, specifically including: Determining a current total power regulation requirement of the local distributed flexible resource cluster according to the power transfer matrix; Based on a consistency algorithm, taking the same power adjustment rate of each local flexible resource as an optimization goal, iteratively adjusting the power adjustment rate of each local flexible resource until the power adjustment rate of each local flexible resource converges, thereby obtaining an initial power adjustment rate of each local flexible resource; determining an initial power adjustment amount of each of the local flexible resources based on a product of the initial power adjustment rate of each of the local flexible resources and the maximum power adjustment capacity; According to the total power regulation requirement and each of the initial power regulation amounts, a power adjustment proportional coefficient is determined, and based on the product of the power adjustment proportional coefficient and the initial power regulation amount, the target power adjustment amount of each of the local flexible resources is determined; wherein the sum of each of the target power adjustment amounts is equal to the total power regulation requirement.
[0057] As a preferred solution, the system further includes a global reinforcement learning model generation module, which is used to: Based on the horizontal federated learning strategy, the parameter server aggregates the local reinforcement learning model weights of the control decision-making end and each of the distributed flexible resource clusters according to the number of local training samples of the control decision-making end and each of the distributed flexible resource clusters in the current round of training to obtain an updated global model weight, and uses the updated global model weight as the initial local reinforcement learning model weight for the control decision-making end and each of the distributed flexible resource clusters to perform the next round of training until the reward value converges, and generates the global reinforcement learning model according to the current global model weight; The weight of the local reinforcement learning model of the control decision end is generated by the control decision end training the local reinforcement learning model using the local control data; The local reinforcement learning model weights of the distributed flexible resource cluster are generated by the distributed flexible resource cluster training the local reinforcement learning model using local cluster operation data; The reward value is calculated based on the power control cost, the power over-limit penalty and the branch power over-limit penalty; The power control cost is specifically expressed as: ; The expression for adjusting the power limit penalty is specifically: ; The expression of the branch power flow exceeding limit penalty is specifically: ; in, represents the power regulation cost; Indicates the number of distributed flexible resource clusters; Indicates the The cost of regulating a distributed flexible resource cluster; Indicates the The power regulation of a distributed flexible resource cluster; and Respectively represent The upper and lower limits of the power regulation capacity of a distributed flexible resource cluster; Indicates the penalty for exceeding the adjustment power limit; Indicates the preset penalty factor for regulating power exceeding the limit; Indicates the penalty for exceeding the limit of the branch power flow; Indicates the preset branch power flow over-limit penalty factor; Indicates a branch The trend; Indicates a branch The upper limit of the trend; Indicates a branch set.
[0058] As a preferred solution, the system further includes a power regulation rate imbalance detection module, which is used to: determining an average power regulation rate at a current moment according to the power regulation rates of the respective local flexible resources; determining a power regulation rate deviation of each of the local flexible resources according to a difference between the power regulation rate of each of the local flexible resources and the average power regulation rate; Obtain the current communication load rate and node remaining energy; Mapping the power regulation rate deviation, the communication load rate, and the node residual energy into a fuzzy set through a preset event detector; Based on the fuzzy rules corresponding to the fuzzy sets, fuzzy reasoning is performed on the power regulation rate deviation, the communication load rate, and the node residual energy to determine a power regulation rate imbalance trigger threshold corresponding to the power regulation rate deviation, the communication load rate, and the node residual energy; wherein the fuzzy rules are used to define a fuzzy output result corresponding to each fuzzy set for representing the power regulation rate imbalance trigger threshold; If it is detected that the power regulation rate deviations of the local flexible resources are all less than or equal to the power regulation rate imbalance trigger threshold, it is determined that the power regulation rates of the local flexible resources are not unbalanced; If it is detected that the power regulation rate deviation of any one of the local flexible resources is greater than the power regulation rate imbalance trigger threshold, it is determined that the power regulation rate of the any one of the local flexible resources is unbalanced.
[0059] The distributed flexible resource cluster control system provided by the embodiment of the present invention decouples the computing task of flexible resource power control into intra-cluster power control and inter-cluster collaborative control through a hierarchical distributed flexible resource cluster control architecture. When performing inter-cluster collaborative control, the control decision-making end is only responsible for sending the power transfer matrix between each distributed flexible resource cluster, without having to process the operating data of each flexible resource. The power adjustment amount allocation of each flexible resource is the responsibility of each distributed flexible resource cluster, thereby effectively improving the control efficiency in the online control scenario of large-scale flexible resources, ensuring the real-time performance in the online control scenario of large-scale flexible resources, and being able to meet the online control needs of large-scale flexible resources.
[0060] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A distributed flexible resource cluster control method, characterized in that: include: Obtaining a power adjustment rate of each local flexible resource belonging to the local distributed flexible resource cluster at each moment, and determining a control law for the local flexible resource based on the power adjustment rate; If it is detected that none of the power adjustment rates are unbalanced, power regulation is performed on the local flexible resource based on the power adjustment rate of the local flexible resource, the control law, and a preset maximum power regulation capacity; If any power adjustment rate imbalance is detected, an inter-cluster collaborative control request is sent to the control decision end; wherein the inter-cluster collaborative control request is used to instruct the control decision end to feedback the power transfer matrix between each distributed flexible resource cluster; According to the power transfer matrix, the power regulation amount of each of the local flexible resources is optimized using a global reinforcement learning model to determine the target power regulation amount of each of the local flexible resources; wherein, the global reinforcement learning model is generated based on the local reinforcement learning model weights of the control decision end and each of the distributed flexible resource clusters.
2. The distributed flexible resource cluster control method according to claim 1, characterized in that: The obtaining of the power adjustment rate of each local flexible resource belonging to the local distributed flexible resource cluster at each moment, and determining the control law of the local flexible resource based on the power adjustment rate, specifically includes: Obtaining the power adjustment rate of each of the local flexible resources at each moment according to a ratio between the power adjustment amount of each of the local flexible resources at each moment and the maximum power adjustment capacity; According to the power adjustment rate of each local flexible resource at each moment, the control law of the local flexible resource is determined by the following expression: ; in, Represents local flexible resources exist The control law at time t; Represents local flexible resources exist The power regulation rate at the time; Represents local flexible resources exist The power regulation rate at the time; Represents local flexible resources With local flexible resources between The connection weight at the moment; Represents local flexible resources The neighbor set of Time and local flexible resources All local flexible resources that have a physical connection relationship.
3. The distributed flexible resource cluster control method according to claim 2, characterized in that: The method further comprises: determining a remaining adjustment capacity of each of the local flexible resources at each moment according to a difference between the maximum power adjustment capacity and the power adjustment amount of each of the local flexible resources at each moment; According to the remaining adjustment capacity of each local flexible resource at each moment, the connection weights between the local flexible resources are adjusted by the following expression: ; in, Represents local flexible resources exist The remaining regulating capacity at the time; Represents local flexible resources exist The remaining regulation capacity at the moment.
4. The distributed flexible resource cluster control method according to claim 1, characterized in that: The performing power regulation on the local flexible resource based on the power regulation rate of the local flexible resource, the control law, and a preset maximum power regulation capacity specifically includes: Based on the power regulation rate of the local flexible resource, the control law, and the maximum power regulation capacity, a target power regulation amount of the local flexible resource is determined by the following expression: ; performing power regulation on the local flexible resource according to the target power regulation amount of the local flexible resource; in, Represents local flexible resources exist The target power adjustment amount at the time; Represents local flexible resources exist The power regulation rate at the time; Represents local flexible resources exist The control law at time t; Represents local flexible resources exist The maximum power regulation capacity at the time; Indicates the preset scale factor.
5. The distributed flexible resource cluster control method according to claim 1, characterized in that: The power transfer matrix is generated by the control decision end using the total power deviation of the distributed flexible resource cluster, the upper limit of the power regulation capacity and the lower limit of the power regulation capacity of each distributed flexible resource cluster as the state space, using the global reinforcement learning model based on the state space; Among them, the power transfer matrix is used to represent the power transfer relationship and power transfer amount between each of the distributed flexible resource clusters; the total power deviation of the distributed flexible resource cluster is the sum of the differences between the actual power and the expected power of each of the distributed flexible resource clusters at the current moment.
6. The distributed flexible resource cluster control method according to claim 1 or 5, characterized in that: Optimizing the power adjustment amount of each of the local flexible resources using a global reinforcement learning model according to the power transfer matrix to determine a target power adjustment amount of each of the local flexible resources specifically includes: Taking the injected power, voltage amplitude of the cluster grid connection point, the upper limit and lower limit of the power regulation capacity of the local distributed flexible resource cluster as the state space, the global reinforcement learning model is used to decompose the power transfer matrix into the target power regulation amount of each local flexible resource based on the consistency algorithm.
7. The distributed flexible resource cluster control method according to claim 6, characterized in that: Decomposing the power transfer matrix into target power adjustment amounts of each of the local flexible resources based on a consistency algorithm using the global reinforcement learning model specifically includes: Determining a current total power regulation requirement of the local distributed flexible resource cluster according to the power transfer matrix; Based on a consistency algorithm, taking the same power adjustment rate of each local flexible resource as an optimization goal, iteratively adjusting the power adjustment rate of each local flexible resource until the power adjustment rate of each local flexible resource converges, thereby obtaining an initial power adjustment rate of each local flexible resource; determining an initial power adjustment amount of each of the local flexible resources based on a product of the initial power adjustment rate of each of the local flexible resources and the maximum power adjustment capacity; According to the total power regulation requirement and each of the initial power regulation amounts, a power adjustment proportional coefficient is determined, and based on the product of the power adjustment proportional coefficient and the initial power regulation amount, the target power adjustment amount of each of the local flexible resources is determined; wherein the sum of each of the target power adjustment amounts is equal to the total power regulation requirement.
8. The distributed flexible resource cluster control method according to claim 6, characterized in that: The method specifically generates the global reinforcement learning model through the following steps: Based on the horizontal federated learning strategy, the parameter server aggregates the local reinforcement learning model weights of the control decision-making end and each of the distributed flexible resource clusters according to the number of local training samples of the control decision-making end and each of the distributed flexible resource clusters in the current round of training to obtain an updated global model weight, and uses the updated global model weight as the initial local reinforcement learning model weight for the control decision-making end and each of the distributed flexible resource clusters to perform the next round of training until the reward value converges, and generates the global reinforcement learning model according to the current global model weight; The weight of the local reinforcement learning model of the control decision end is generated by the control decision end training the local reinforcement learning model using the local control data; The local reinforcement learning model weights of the distributed flexible resource cluster are generated by the distributed flexible resource cluster training the local reinforcement learning model using local cluster operation data; The reward value is calculated based on the power control cost, the power over-limit penalty and the branch power over-limit penalty; The power control cost is specifically expressed as: ; The expression for adjusting the power limit penalty is specifically: ; The expression of the branch power flow exceeding limit penalty is specifically: ; in, represents the power regulation cost; Indicates the number of distributed flexible resource clusters; Indicates the The cost of regulating a distributed flexible resource cluster; Indicates the The power regulation of a distributed flexible resource cluster; and Respectively represent The upper and lower limits of the power regulation capacity of a distributed flexible resource cluster; Indicates the penalty for exceeding the adjustment power limit; Indicates the preset penalty factor for regulating power exceeding the limit; Indicates the penalty for exceeding the limit of the branch power flow; Indicates the preset branch power flow over-limit penalty factor; Indicates a branch The trend; Indicates a branch The upper limit of the trend; Indicates a branch set.
9. The distributed flexible resource cluster control method according to claim 1, characterized in that: The method specifically determines whether each power regulation rate is unbalanced by the following steps: determining an average power regulation rate at a current moment according to the power regulation rates of the respective local flexible resources; determining a power regulation rate deviation of each of the local flexible resources according to a difference between the power regulation rate of each of the local flexible resources and the average power regulation rate; Obtain the current communication load rate and node remaining energy; Mapping the power regulation rate deviation, the communication load rate, and the node residual energy into a fuzzy set through a preset event detector; Based on the fuzzy rules corresponding to the fuzzy sets, fuzzy reasoning is performed on the power regulation rate deviation, the communication load rate, and the node residual energy to determine a power regulation rate imbalance trigger threshold corresponding to the power regulation rate deviation, the communication load rate, and the node residual energy; wherein the fuzzy rules are used to define a fuzzy output result corresponding to each fuzzy set for representing the power regulation rate imbalance trigger threshold; If it is detected that the power regulation rate deviations of the local flexible resources are all less than or equal to the power regulation rate imbalance trigger threshold, it is determined that the power regulation rates of the local flexible resources are not unbalanced; If it is detected that the power regulation rate deviation of any one of the local flexible resources is greater than the power regulation rate imbalance trigger threshold, it is determined that the power regulation rate of the any one of the local flexible resources is unbalanced.
10. A distributed flexible resource cluster control system, characterized in that: include: Data acquisition module, used to: Obtaining a power adjustment rate of each local flexible resource belonging to the local distributed flexible resource cluster at each moment, and determining a control law for the local flexible resource based on the power adjustment rate; The power control module within the cluster is used to: If it is detected that none of the power adjustment rates are unbalanced, power regulation is performed on the local flexible resource based on the power adjustment rate of the local flexible resource, the control law, and a preset maximum power regulation capacity; Inter-cluster collaborative control module, used for: If any power adjustment rate imbalance is detected, an inter-cluster collaborative control request is sent to the control decision end; wherein the inter-cluster collaborative control request is used to instruct the control decision end to feedback the power transfer matrix between each distributed flexible resource cluster; According to the power transfer matrix, the power regulation amount of each of the local flexible resources is optimized using a global reinforcement learning model to determine the target power regulation amount of each of the local flexible resources; wherein, the global reinforcement learning model is generated based on the local reinforcement learning model weights of the control decision end and each of the distributed flexible resource clusters.
Citation Information
Patent Citations
Power distribution system power distribution method and system, computer equipment and storage medium
CN116436013A
Power distribution network multi-type flexible resource cluster regulation and control method
CN116488156A
Power distribution network collaborative planning method and device considering multiple types of flexible resources
CN116683534A
Park multi-flexible resource collaborative optimization control method based on consistency algorithm
CN117578447A
Power distribution method based on BMS power supply control and related equipment thereof
CN120320450A
Cited By
Multi-station power coordinated regulation method and system for traction power supply system under distributed architecture
CN121906681A