A cloud-edge collaborative reinforcement learning control method and system for photovoltaic cluster access to a power distribution network

By constructing a hierarchical control architecture and a multi-agent deep reinforcement learning model, and combining centralized training in the cloud and decentralized execution at the edge, the problem of global coordination and local autonomy in photovoltaic cluster access to the power distribution network was solved, realizing a safe and economical control strategy and improving sample utilization and the convergence speed of the control strategy.

CN122437264APending Publication Date: 2026-07-21STATE GRID JIANGSU ELECTRIC POWER CO LTD RESEARCH INSTITUTE +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
STATE GRID JIANGSU ELECTRIC POWER CO LTD RESEARCH INSTITUTE
Filing Date
2026-03-27
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve a seamless integration of global coordination and local autonomy when photovoltaic clusters are connected to the power distribution network. They lack real-time security constraints and have low sample efficiency, making it difficult to adapt to rapid response and optimized control in large-scale photovoltaic cluster scenarios.

Method used

A hierarchical control architecture is constructed, employing a multi-agent deep reinforcement learning model. This model combines centralized training in the cloud with distributed execution at the edge. Through near-end policy optimization and security layer projection, secure and economical collaborative control decisions are achieved.

Benefits of technology

It has enabled the safe, economical, and intelligent operation of photovoltaic clusters connected to the power distribution network, improved sample utilization and the convergence speed of control strategies, and ensured the stability of the training process and the consistency of policy estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122437264A_ABST
    Figure CN122437264A_ABST
Patent Text Reader

Abstract

The application discloses a cloud-edge collaborative reinforcement learning control method and system for photovoltaic cluster access to a power distribution network, which comprises the following steps: constructing a hierarchical control architecture; embedding a multi-agent deep reinforcement learning model for centralized training and decentralized execution; establishing a distributed optimization control model combined with a multi-agent reinforcement learning mechanism to obtain an optimal control strategy through cloud-edge collaborative iteration; constructing a data-driven control model according to local observation data and a coordination signal issued by the cloud to complete the update of strategy network parameters; training and optimizing the strategy network through a cloud-edge collaborative iteration mechanism based on a proximal policy optimization to obtain an optimized control strategy; and constructing an experience replay pool to dynamically assign priorities to experience samples, introduce an importance sampling weight for bias correction, and obtain a stable sample optimization learning method. The method effectively promotes the safe, economic and intelligent operation of a high-proportion renewable energy power distribution network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a cloud-edge collaborative reinforcement learning control method and system for photovoltaic cluster access to distribution networks, belonging to the field of smart distribution network operation control and data-driven collaborative optimization technology. Background Technology

[0002] As the global energy structure shifts towards cleaner and lower-carbon energy, photovoltaic (PV) power generation, as an important form of renewable energy, is being integrated into power grids on a large scale in a distributed and clustered manner. PV clusters typically include PV power generation units, energy storage systems, and local loads, and possess the operational characteristics of "self-consumption with surplus power fed into the grid," which can improve energy efficiency and power supply reliability to a certain extent. However, photovoltaic power generation is characterized by strong intermittency, volatility, and randomness. Its large-scale clustered access also poses severe challenges to the safe, stable, and economical operation of the distribution network. These challenges are mainly manifested in the following aspects: photovoltaic clusters are widely distributed and numerous, making it difficult for traditional centralized optimization control methods to meet their decentralized, autonomous, and rapid response control requirements; rapid fluctuations in photovoltaic output and load can easily lead to safety issues such as voltage exceeding limits at distribution network nodes and overloaded branch power flows, and traditional model-based predictive control or optimization scheduling methods suffer from insufficient model accuracy; each photovoltaic cluster performs autonomous optimization with the goal of minimizing local costs, which may conflict with the overall economic objectives of the distribution network; existing methods often rely on accurate physical models for optimization, making them sensitive to model errors and uncertainties, while pure data-driven methods, although adaptable, lack explicit guarantees of safety constraints, making them difficult to reliably apply in actual operation.

[0003] To address the aforementioned challenges, existing technologies have proposed methods based on multi-agent systems, distributed optimization, and machine learning. However, these methods still have shortcomings in the following aspects: most methods still rely on iteratively solving distributed optimization problems, resulting in poor real-time performance and difficulty in adapting to fast time-scaled control; they lack mechanisms to organically integrate global collaborative signals with local autonomous decision-making; they lack real-time and reliable guarantees of operational safety constraints during the generation of control actions; and they have low sample efficiency and slow training convergence, making them unsuitable for large-scale photovoltaic cluster scenarios.

[0004] Therefore, there is an urgent need for a photovoltaic cluster collaborative control method that can organically combine cloud-based global collaboration with edge-based local autonomy, provide real-time security constraints, and efficiently utilize samples, so as to promote the safe, economical, and intelligent operation of high-proportion renewable energy distribution networks. Summary of the Invention

[0005] The purpose of this invention is to provide a cloud-edge collaborative reinforcement learning control method and system for photovoltaic cluster access to power distribution networks. It takes photovoltaic inverters and energy storage systems as the control objects, constructs a collaborative optimization control model based on a multi-agent reinforcement learning mechanism with centralized training and decentralized execution, learns policies based on local observation data and coordination signals sent from the cloud, and introduces a safety layer based on control obstacle functions to project actions into the feasible region. The policy is updated through a near-end policy optimization algorithm, enabling safe and economical collaborative control decisions to be made online based on the real-time operating status of the power distribution network.

[0006] To achieve the above objectives, the present invention is implemented using the following technical solution.

[0007] In a first aspect, the present invention provides a cloud-edge collaborative reinforcement learning control method for photovoltaic cluster access to the distribution network, comprising:

[0008] A hierarchical control architecture is constructed, which includes a cloud-based power grid dispatch layer and an edge photovoltaic cluster layer.

[0009] The hierarchical control architecture is embedded into a multi-agent deep reinforcement learning model for centralized training in the cloud and distributed execution at each edge.

[0010] A distributed optimization control model is established, and combined with a multi-agent reinforcement learning mechanism, the optimal control strategy of the distributed optimization control model is obtained through collaborative iterative solution between the cloud and the edge.

[0011] Each photovoltaic cluster at the edge constructs a data-driven control model based on local observation data and coordination signals sent from the cloud, and completes the update of strategy network parameters;

[0012] The policy network is trained and optimized through a cloud-edge collaborative iterative mechanism based on near-end policy optimization to obtain the optimized control policy.

[0013] An experience replay pool is constructed. Based on the optimized control strategy, the experience samples stored in the replay pool are dynamically assigned priorities, and importance sampling weights are introduced for bias correction, resulting in a stable sample optimization learning method.

[0014] Furthermore, the hierarchical control architecture includes a cloud-based distribution network scheduling layer and an edge-end photovoltaic cluster layer. The cloud-based distribution network scheduling layer, serving as the global coordination center of the distribution network, is deployed on the master station side. Its core function is to receive and aggregate boundary power statistics reported from each edge cluster based on a simplified distribution network equivalent model. Combined with ultra-short-term load and photovoltaic power generation forecast data, and aiming to minimize the total operating cost of the distribution network, it generates global collaborative scheduling instructions within a given scheduling cycle and issues them to each edge cluster to achieve cross-cluster power flow constraints and operating cost coordination. The distribution network equivalent model is constructed by equating each edge-end distributed cluster to an injection node and utilizing node voltage sensitivity parameters uploaded from the edge cluster layer. It achieves accurate estimation of global power flow and voltage state at the cloud level with low computational complexity. The core mathematical expression is as follows:

[0015] Consider a containing The distribution network system of each edge cluster, the cloud layer will connect each cluster Equivalent to a power injection point, its equivalent injected active power is .

[0016] Furthermore, the edge photovoltaic cluster control layer consists of multiple geographically dispersed distributed clusters containing photovoltaic power generation units, energy storage systems, and loads. Each cluster acts as an autonomous energy unit, aiming to minimize its internal operating costs. It utilizes local computing power to achieve real-time optimized scheduling of photovoltaics, energy storage, and loads, and after completing local optimization, it uploads its key boundary power information to the cloud-based distribution network scheduling layer.

[0017] Furthermore, the edge photovoltaic cluster control layer provides complete data to the distribution network layer through node voltage sensitivity estimation, and at the same time provides key model parameters for the coordinated voltage regulation of each photovoltaic cluster.

[0018] This invention employs a dual-layer rolling optimization mechanism of "slow timescale cloud and fast timescale edge," which further achieves a balance between global optimization accuracy and local response speed.

[0019] Furthermore, the hierarchical control architecture is embedded into a partially observable multi-agent deep reinforcement learning model, employing a centralized training and distributed execution mode.

[0020] Modeling multiple agents involves modeling the entire system as a collection of... A Markov game involving multiple agents, where agent 0 represents the cloud-based coordinating agent, corresponding to the cloud-based power grid scheduling layer; agents Representing the Each edge cluster agent, corresponding to the edge photovoltaic cluster layer, makes decisions based on its local observations.

[0021] Furthermore, observations were conducted on the cloud-coordinated intelligent agent and the edge-based cluster intelligent agent respectively:

[0022] The cloud-coordinated intelligent agent is observed. This includes boundary power statistics reported by each edge, network linearization power flow sensitivity parameters, and historical information on dual variables, as well as actions. The boundary power reference value vector to be distributed to each end. ;

[0023] The edge cluster agents are observed. This includes local photovoltaic data, real-time and forecast data of load, energy storage, local voltage, boundary power at the previous moment, and reference commands issued from the cloud. And coordination signals, actions This refers to the instruction vector for locally controllable devices.

[0024] Furthermore, centralized training in the cloud and distributed execution at various endpoints include:

[0025] The centralized training in the cloud refers to deploying a centralized commentator network in the cloud. During the training phase, the network can access the observations and actions of all agents to accurately estimate the global value of joint actions and guide the updates of each policy network. The global reward function is designed as a negative of the total operating cost of the distribution network and includes penalties for safety constraints such as voltage overruns.

[0026] The decentralized execution at each end refers to the deployment of the policy network separately after training, with the cloud-coordinated agent relying solely on its observations. Generate coordination signals; each edge cluster agent relies solely on its local observations. It can independently generate and execute local control actions based on received cloud signals, without the need for global communication, thus achieving scalable distributed control.

[0027] Furthermore, a distributed optimization control model is constructed. This model first performs physical and informational modeling of the photovoltaic clusters, then integrates the global economic objective of the distribution network with the local autonomous objectives of each cluster into a collaborative optimization framework. This framework does not rely on traditional distributed iterative solutions; instead, it transforms the optimization objective into long-term rewards for the agents through a multi-agent reinforcement learning mechanism. Finally, it achieves near-optimal attainment of the objective through cloud-based collaborative signals and edge-based security control. Specifically:

[0028] First, physical and information modeling is performed on the photovoltaic clusters. Each photovoltaic cluster i is regarded as an autonomous physical and information fusion unit that integrates energy production, storage, consumption and real-time control. The core components include: photovoltaic power generation unit, energy storage system and local load.

[0029] The maximum available output of the photovoltaic power generation unit is denoted as: The actual dispatch output is The possible amount of light wasted can be controlled as follows: The energy storage system has a state of charge denoted as... The charging and discharging power is The power requirement of the local load is: The cluster is connected to the distribution network through a common junction point, exchanging net active power, which is also known as boundary power. These are key variables that couple the internal operation of the cluster with the global state of the distribution network; each cluster is equipped with a local controller, which has the ability to sense, calculate and make decisions in real time and communicate with the cloud.

[0030] Secondly, a distributed optimization control model is defined, which includes a cloud-based global optimization model and an edge-based local optimization model. The model aims to minimize the total system operating cost within the scheduling cycle, and its mathematical expression is as follows:

[0031] ;

[0032] in, The total operating cost of the system within the scheduling period is represented by t; t is the time index. The total number of photovoltaic clusters; Let be the local operating cost of the i-th photovoltaic cluster at time t; The operating cost of the distribution network at time t mainly includes network loss costs; This is the global safety penalty term incurred at time t due to node voltage exceeding the limit;

[0033] Cluster local operating costs Further refined to:

[0034] ;

[0035] in, This is a function representing the operation and maintenance cost of photovoltaic power generation. Photovoltaic power generation; The cycle loss cost function for energy storage devices; For energy storage power, To address the issue of curtailed photovoltaic power The penalty function;

[0036] The model's operation must meet strict physical and security constraints. These constraints will be guaranteed in real time through a security layer at the edge control layer. Core constraints include power balance constraints within the cluster, expressed as follows:

[0037] ;

[0038] in, The active power exchanged at the boundary between photovoltaic cluster i and the distribution network during scheduling period t; Active power of node load;

[0039] Edge optimization also needs to meet the following local constraints, including node voltage safety constraints, power balance constraints, photovoltaic output constraints, energy storage operation constraints, photovoltaic reduction constraints, and cluster switching power constraints.

[0040] Although the framework proposed by the method of this invention includes an optimization model with a clear objective function and constraints, its purpose is not to perform traditional mathematical programming solutions, but to provide clear learning objectives and value orientation for subsequent data-driven algorithms.

[0041] Furthermore, for edge-distributed photovoltaic (PV) clusters, a data-driven control model is constructed, consisting of a policy network, a security layer, and a value network. Each edge PV cluster uses local observation data and coordination signals from the cloud as input, generates candidate control actions through the policy network, and the security layer performs runtime constraint verification and projection correction on these candidate actions. During the training phase, the value network and a near-end policy optimization algorithm are combined to update parameters. Specifically:

[0042] For each edge photovoltaic cluster i, construct a pair of policy networks based on deep neural networks. With value network The parameters are respectively and Its input state vector in each fast time-scaled control period t is defined as:

[0043] ;

[0044] in, , These are the photovoltaic output and load power measurement sequences of cluster i within the historical window, respectively; , In the prediction time domain Internal photovoltaic and load forecast data; The state of charge of the local energy storage system; The boundary active power of the previous scheduling cycle. This refers to the voltage amplitude at the local node. This is the boundary power reference signal issued by the cloud scheduling layer at the current slow time point, used to correspond to the actions of the cloud-coordinated intelligent agent; The component of the coordination / dual signal published in the cloud at cluster i is used to characterize the consistency requirements of global power flow and power balance constraints.

[0045] In a given state The policy network then outputs the candidate control actions for the cluster, expressed as:

[0046] ;

[0047] in, , These are the active and reactive power reference values ​​for the photovoltaic inverter. Power commands for energy storage charging and discharging; value network Used to assess the state Expected returns on long-term operating costs and coordination performance;

[0048] The method of this invention, while ensuring that the control actions generated by the policy network are physically feasible and meet operational safety requirements, connects a security layer in series after the policy network of each edge photovoltaic cluster for processing candidate control actions. Feasibility verification and correction are performed, and a set of obstacle functions characterizing operational safety constraints is constructed by introducing a control obstacle function framework, with the expression:

[0049] ;

[0050] Each of them This corresponds to a key operational constraint;

[0051] Define cluster The set of possible actions, expressed as:

[0052] ;

[0053] The set described in this invention is structurally consistent with the constraints of the local optimization model, thus achieving unified constraints on distribution network voltage security, branch power flow, energy storage operating range, and power change rate.

[0054] The security layer outputs candidate actions to the policy network. Perform a projection operation to obtain the final safety control action to be executed. The expression is:

[0055] ;

[0056] Among all possible actions that satisfy the control barrier function constraints, select the candidate action. The point with the minimum Euclidean distance is used to preserve the original decision intent of the policy network to the greatest extent possible while ensuring safety.

[0057] Finally, control commands The command is dispatched to the photovoltaic inverters and energy storage systems for execution, and the current boundary power is determined through the power balance relationship within the cluster. and will The relevant statistics are uploaded to the cloud.

[0058] Furthermore, based on the aforementioned hierarchical control architecture, multi-agent reinforcement learning mechanism, and distributed optimization control model, a cloud-edge collaborative iterative mechanism based on proximal policy optimization (PPO) is constructed to train and further optimize the policy network, specifically including:

[0059] 1) Slow timescale

[0060] The cloud-based coordinating agent acts as the global coordination center, aggregating boundary power statistics from each edge photovoltaic cluster. And related operating status, including node voltage exceeding limits and the degree to which branch power flow approaches thermal stability limits, the cloud-based calculation of the total operating cost and constraint violation degree of the distribution network within the current scheduling cycle is expressed as follows:

[0061] ;

[0062] in, For the cloud during the scheduling cycle The costs of electricity purchase and grid losses, etc. This represents the sum of the local operating costs of each cluster during this period. The constraint violation vector consists of voltage over-limit and power flow over-limit. Its norm;

[0063] Based on this, the cloud-based coordinating agent generates a new boundary active power reference value for each cluster i. and the corresponding penalty weights The boundary power reference value is determined based on the global power flow distribution and economic objectives, and the reference value of the previous cycle is corrected in one step along the direction of reducing network losses and voltage deviation; the penalty weight is adjusted according to the degree of constraint violation and reference deviation. When the voltage or power flow constraint violation caused by a certain cluster is more severe, its penalty weight is increased, as expressed in the following expression:

[0064] ;

[0065] in, For constraint violation metrics related to cluster i, As the allowable safety margin threshold, To adjust the step size; the result is These signals combine to form a coordination signal sent from the cloud to each distributed photovoltaic cluster, serving as constraints and guidance for fast-timescale edge strategy updates and control execution.

[0066] 2) Fast timescale

[0067] During the multiple fast-timescale control cycles between two updates in the cloud, each distributed photovoltaic cluster receives the coordination signal. Then, based on the state vector With policy network Generate candidate control actions The final action is obtained by projection through the security layer. It drives the operation of photovoltaic inverters and energy storage systems, and the corresponding boundary power is determined by the power balance relationship within the cluster. ;

[0068] The edge cluster constructs an instantaneous reward function in each control cycle based on local operating costs and boundary power tracking error, expressed as:

[0069] ;

[0070] in, The first term represents the local operating cost of the cluster at time t, including photovoltaic power generation costs, energy storage operating costs, and photovoltaic reduction penalties. The second term is a penalty for boundary power deviations from the cloud reference value, with penalty weights... The scheduling is determined by the cloud during the current scheduling period; each cluster collects multiple state-action-reward-next state trajectories within the current cloud scheduling period. Stored in the experience replay pool.

[0071] Furthermore, each distributed photovoltaic cluster uses the aforementioned empirical data to adjust the parameters of the strategy network. Implementing PPO strategy updates: based on value networks Estimating the advantage function Construct the probability ratio, expressed as:

[0072] ;

[0073] Introducing the clipping objective function, the expression is:

[0074] ;

[0075] in, This is the cutting factor;

[0076] By employing a trust region radius constraint, an upper bound is imposed on policy updates, expressed as:

[0077] ;

[0078] in, The radius of the trust region;

[0079] In the method of this invention, the use of trust region radius constraint can further limit the policy update step size, thereby improving training stability.

[0080] 3) Convergence criterion

[0081] After completing the edge policy update and control execution within a cloud scheduling cycle, each distributed photovoltaic cluster will display the boundary power trajectory for this cycle. The node voltage and branch power flow statistics, as well as the accumulated local operating costs, are transmitted back to the cloud-based coordinating agent. Based on the transmitted data, the cloud-based coordinating agent reassesses the overall operating costs of the distribution network and the degree of constraint violations, and adjusts the boundary power reference values ​​for the next scheduling cycle accordingly. With penalty weight This allows the boundary power to gradually converge toward a direction that satisfies the requirements of global economy and safety.

[0082] In the method of this invention, by constructing a PPO strategy, the strategy network of each distributed photovoltaic cluster is trained, enabling it to collaboratively achieve unified optimization of the global economic operation and local security constraints of the distribution network under the boundary power reference and penalty weight constraints issued by the cloud.

[0083] Furthermore, to address the issues of high interaction costs and limited training samples in multi-PV cluster scenarios, a sample efficiency optimization method combining priority experience replay and importance sampling is introduced. During the training phase, each distributed PV cluster generates a quadruple containing state-action rewards while executing local control. To improve sample utilization efficiency, in each cluster Side-building experience replay pool It is used to store the interaction trajectory over a recent period of time;

[0084] When updating the policy network, instead of using the traditional uniform random sampling method, a priority experience replay mechanism is introduced: specifically, for each experience sample in the replay pool... Assign a priority based on temporal-difference error (TD error). The TD error is defined as follows:

[0085] exist For each sample Assign a dynamically updated priority. The priority is determined by the centralized value network. The timing difference error is determined by the following expression:

[0086] ;

[0087] in, For the immediate reward of sample j, For the estimation of state value in a value network, The discount factor; the priority setting expression is:

[0088] ;

[0089] in, To prevent small constants with zero priority.

[0090] The method of this invention improves sample utilization and policy convergence speed by optimizing the empirical sampling method and loss weighting strategy, while ensuring the stability of the training process and the consistency of policy estimation.

[0091] Furthermore, the method of the present invention adopts a priority experience replay mechanism during sampling, and samples are sampled proportionally according to priority, so that samples with larger temporal difference errors are selected more frequently, thereby improving sample utilization efficiency and policy convergence speed. The larger the TD error, the larger the value estimation error of the region corresponding to the experience sample and the higher the learning potential, and the more attention should be paid to it in subsequent policy updates.

[0092] Priority This directly reflects the importance of each sample to the improvement of the current strategy. A proportional priority approach is used to allocate sampling probabilities based on priority, expressed as:

[0093] ;

[0094] in, The degree to which control priority affects the sampling distribution, when Degenerates into uniform sampling when Samples with larger TD errors are selected more frequently for training, thereby focusing on optimizing the state-action regions that are difficult to learn or have large prediction errors.

[0095] Furthermore, in the method of this invention, the non-uniform sampling with priority weighting will change the original data distribution of the experience. If such samples are directly used for gradient updates, estimation bias will be introduced. Based on this, this invention introduces importance sampling weights to correct each sampled experience, in order to effectively compensate for this bias:

[0096] For pressing from the experience replay pool The importance sampling weight of the sample j obtained from sampling is defined as follows:

[0097] ;

[0098] in, The total number of empirical samples in the replay pool. To control the coefficient of correction intensity, If no correction is performed, The distribution bias caused by priority sampling is fully compensated in time;

[0099] When updating the PPO strategy, the contribution of the PPO objective function to each sample is multiplied by the corresponding weight. Thus, the weighted objective is obtained, expressed as:

[0100] ;

[0101] in, Let be the policy probability ratio for sample j. The advantage function is based on value network estimation.

[0102] In this invention, the loss function is weighted by importance sampling. While prioritizing the use of high TD error samples, the distribution bias caused by non-uniform sampling is corrected. This ensures that the policy gradient estimation is consistent with the expectation under the original state-action distribution in the long-term statistical sense, thereby improving the stability of the training process and the consistency of the policy estimation results.

[0103] Within each cloud scheduling cycle, each distributed photovoltaic cluster, based on the aforementioned reward construction method and control execution process, uniformly stores the interaction experience collected in this cycle into a local experience replay pool. It also updates the TD error and priority of the corresponding samples in real time. When each cluster triggers a policy update, it no longer starts from... Instead of uniform sampling, it follows a priority probability. Extract batch experience and apply importance sampling weights to each sample when calculating the PPO objective function. .

[0104] Secondly, the present invention provides a cloud-edge collaborative reinforcement learning control system for photovoltaic cluster access to the power distribution network, comprising:

[0105] A control architecture construction module is used to build a hierarchical control architecture, which includes a cloud-based power grid dispatch layer and an edge photovoltaic cluster layer.

[0106] A centralized-distributed training module is used to embed the hierarchical control architecture into a multi-agent deep reinforcement learning model for centralized training in the cloud and distributed execution at each edge.

[0107] The distributed optimization solution module is used to establish a distributed optimization control model. By combining a multi-agent reinforcement learning mechanism, the optimal control strategy of the distributed optimization control model is obtained through collaborative iterative solution between the cloud and the edge.

[0108] The strategy network update module is used to build a data-driven control model for each edge photovoltaic cluster based on local observation data and coordination signals sent from the cloud, and to update the strategy network parameters.

[0109] The cloud-edge collaborative strategy optimization module is used to train and optimize the policy network through a cloud-edge collaborative iterative mechanism based on near-end policy optimization, so as to obtain the optimized control policy.

[0110] The sample efficiency optimization module is used to construct an experience replay pool. Based on the optimized control strategy, it dynamically assigns priorities to the experience samples stored in the replay pool and introduces importance sampling weights for bias correction, thereby obtaining a stable sample optimization learning method.

[0111] Thirdly, the present invention provides a computer-readable storage medium storing a computer program / instruction thereon, which, when executed by a processor, implements the steps of the cloud-edge collaborative reinforcement learning control method for photovoltaic cluster access to the distribution network as described in any of the first aspects.

[0112] Fourthly, the present invention provides a computer device, comprising:

[0113] Memory, used to store computer programs / instructions;

[0114] A processor is configured to execute the computer program / instructions to implement the steps of the cloud-edge collaborative reinforcement learning control method for photovoltaic cluster access to the distribution network as described in any one of the first aspects.

[0115] Compared with the prior art, the beneficial effects achieved by the present invention are as follows:

[0116] 1. The cloud-edge collaborative reinforcement learning control method for photovoltaic cluster access to distribution networks provided by this invention transforms the traditional centralized optimization problem into a multi-agent deep reinforcement learning mode of "centralized training and decentralized execution" by constructing a hierarchical control architecture including a cloud-based distribution network scheduling layer and an edge-based photovoltaic cluster layer. At the cloud-based distribution network scheduling layer, a centralized commentator network is used for global value estimation and collaborative signal generation. At each edge-based photovoltaic cluster layer, a policy update is performed using a PPO-based strategy, achieving local autonomous optimization and rapid response under security constraints. This invention improves the utilization rate of experience samples and the convergence speed of the policy by prioritizing and weighting experience samples in the replay pool, while ensuring the stability of the training process and the consistency of policy estimation. Furthermore, this invention enables safe and economical collaborative control decisions to be made online based on the real-time operating status of the distribution network, providing a stable and efficient optimization method for high-proportion renewable energy distribution networks.

[0117] 2. The computer-readable storage medium and computer device provided by the present invention can execute the steps of the cloud-edge collaborative reinforcement learning control method for photovoltaic cluster access to distribution network provided by the present invention. Attached Figure Description

[0118] Figure 1 This is an overall architecture diagram of a cloud-edge collaborative reinforcement learning control method for photovoltaic cluster access to power distribution networks, provided by an embodiment of the present invention. Detailed Implementation

[0119] It should be noted that:

[0120] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations thereof. In the absence of conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.

[0121] The term "and / or" simply describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0122] Example 1

[0123] like Figure 1 As shown in the figure, this embodiment introduces a cloud-edge collaborative reinforcement learning control method for photovoltaic cluster access to the distribution network, including:

[0124] A hierarchical control architecture is constructed, which includes a cloud-based power grid dispatch layer and an edge photovoltaic cluster layer.

[0125] The hierarchical control architecture is embedded into a multi-agent deep reinforcement learning model for centralized training in the cloud and distributed execution at each edge.

[0126] A distributed optimization control model is established, and combined with a multi-agent reinforcement learning mechanism, the optimal control strategy of the distributed optimization control model is obtained through collaborative iterative solution between the cloud and the edge.

[0127] Each photovoltaic cluster at the edge constructs a data-driven control model based on local observation data and coordination signals sent from the cloud, and completes the update of strategy network parameters;

[0128] The policy network is trained and optimized through a cloud-edge collaborative iterative mechanism based on near-end policy optimization to obtain the optimized control policy.

[0129] An experience replay pool is constructed. Based on the optimized control strategy, the experience samples stored in the replay pool are dynamically assigned priorities, and importance sampling weights are introduced for bias correction, resulting in a stable sample optimization learning method.

[0130] Furthermore, the hierarchical control architecture includes a cloud-based distribution network scheduling layer and an edge photovoltaic cluster layer. The cloud-based distribution network scheduling layer, as the global coordination center of the distribution network, is deployed on the master station side. Based on a simplified equivalent model of the distribution network, it receives and aggregates the boundary power statistics reported by each edge cluster, and combines ultra-short-term load and photovoltaic power generation forecast data. With the goal of minimizing the total operating cost of the distribution network, it generates global collaborative scheduling instructions within a given scheduling cycle and issues them to each edge cluster to achieve cross-cluster power flow constraints and operating cost coordination.

[0131] The equivalent model of the distribution network is constructed by equating each edge-end distributed cluster to an injection node and utilizing the node voltage sensitivity parameters uploaded from the edge-end cluster layer. This model achieves accurate estimation of global power flow and voltage state at the cloud level with low computational complexity. The core mathematical expression is as follows:

[0132] Consider a containing The distribution network system of each edge cluster, the cloud layer will connect each cluster Equivalent to a power injection point, its equivalent injected active power is .

[0133] Furthermore, the edge photovoltaic cluster control layer consists of multiple geographically dispersed distributed clusters containing photovoltaic power generation units, energy storage systems, and loads. Each cluster acts as an autonomous energy unit, aiming to minimize its internal operating costs. It utilizes local computing power to achieve real-time optimized scheduling of photovoltaics, energy storage, and loads, and after completing local optimization, it uploads its key boundary power information to the cloud-based distribution network scheduling layer.

[0134] Furthermore, the edge photovoltaic cluster control layer provides complete data to the distribution network layer through node voltage sensitivity estimation, and at the same time provides key model parameters for the coordinated voltage regulation of each photovoltaic cluster.

[0135] Furthermore, in terms of timing coordination, the cloud-based power distribution network scheduling layer performs a global optimization based on predictive data at relatively long time intervals to generate reference instructions for the next scheduling cycle.

[0136] The edge photovoltaic cluster control layer tracks and fine-tunes instructions issued from the cloud at short intervals based on the latest local real-time measurement data, and performs autonomous optimization under the premise of meeting local constraints, thereby effectively smoothing out rapid fluctuations in renewable energy and load.

[0137] Furthermore, in this embodiment, the hierarchical control architecture is embedded into a partially observable multi-agent deep reinforcement learning model, adopting a centralized training and distributed execution mode.

[0138] Modeling multiple agents involves modeling the entire system as a collection of... A Markov game involving multiple agents, where agent 0 represents the cloud-based coordinating agent, corresponding to the cloud-based power grid scheduling layer; agents Representing the Each edge cluster agent, corresponding to the edge photovoltaic cluster layer, makes decisions based on its local observations.

[0139] Furthermore, the cloud-coordinated intelligent agent is observed. This includes boundary power statistics reported by each edge, network linearization power flow sensitivity parameters, and historical information on dual variables, as well as actions. The boundary power reference value vector to be distributed to each end. ;

[0140] The edge cluster agents are observed. This includes local photovoltaic data, real-time and forecast data of load, energy storage, local voltage, boundary power at the previous moment, and reference commands issued from the cloud. And coordination signals, actions This refers to the instruction vector for locally controllable devices.

[0141] Furthermore, the centralized training and distributed execution model in this embodiment is as follows:

[0142] The cloud adopts a centralized training approach, deploying a centralized commentator network in the cloud. During the training phase, the network can access the observations and actions of all agents to accurately estimate the global value of joint actions and guide the updates of each policy network. The global reward function is designed as a negative of the total operating cost of the distribution network, and penalties for safety constraints such as voltage overruns are added.

[0143] Each endpoint adopts a distributed execution approach. After training, the policy network is deployed separately, and the cloud-based coordinating agent only relies on its observations. Generate coordination signals; each edge cluster agent relies solely on its local observations. It can independently generate and execute local control actions based on received cloud signals, without the need for global communication, thus achieving scalable distributed control.

[0144] Furthermore, a distributed optimization control model is constructed. In this embodiment, the photovoltaic cluster is first modeled physically and informationally. Each photovoltaic cluster i is regarded as an autonomous physical and information fusion unit that integrates energy production, storage, consumption and real-time control. The core components include: photovoltaic power generation unit, energy storage system and local load.

[0145] The maximum available output of the photovoltaic power generation unit is denoted as: The actual dispatch output is The amount of light discarded can be adjusted to ;

[0146] The energy storage system has a state of charge denoted as... The charging and discharging power is ;

[0147] The local load has the following power requirement: ;

[0148] Each cluster is connected to the distribution network via a common connection point, with a net exchanged active power of [missing information]. This power is a key variable that couples the internal operation of the cluster with the global state of the distribution network; in addition, each cluster in this embodiment is equipped with a local controller that has the ability to sense in real time, make calculation decisions and communicate with the cloud.

[0149] Furthermore, in this embodiment, a distributed optimization control model is defined, comprising a cloud-based global optimization model and an edge-based local optimization model. The model aims to minimize the total system operating cost within the scheduling cycle, and its mathematical expression is as follows:

[0150] ;

[0151] in, The total operating cost of the system within the scheduling period is represented by t; t is the time index. The total number of photovoltaic clusters; Let be the local operating cost of the i-th photovoltaic cluster at time t; The operating cost of the distribution network at time t includes network loss costs; This is the global safety penalty term incurred at time t due to node voltage exceeding the limit;

[0152] Cluster local operating costs Further refined to:

[0153] ;

[0154] in, This is a function representing the operation and maintenance cost of photovoltaic power generation. Photovoltaic power generation; The cycle loss cost function for energy storage devices; For energy storage power, To address the issue of curtailed photovoltaic power The penalty function;

[0155] The model's operation must meet strict physical and security constraints. These constraints will be guaranteed in real time through a security layer at the edge control layer. Core constraints include power balance constraints within the cluster, expressed as follows:

[0156] ;

[0157] in, The active power exchanged at the boundary between photovoltaic cluster i and the distribution network during scheduling period t; Active power of node load;

[0158] In this embodiment, edge optimization also needs to satisfy the following local constraints, including node voltage security constraints, power balance constraints, photovoltaic output constraints, energy storage operation constraints, photovoltaic reduction constraints, and cluster switching power constraints:

[0159] The expression for the node voltage safety constraint is:

[0160] ;

[0161] in, For node voltage, , These are the upper and lower limits of the voltage.

[0162] The power balance constraint expression is:

[0163] ;

[0164] in Photovoltaic power generation; This refers to the power of discarded light. For energy storage output power; For load power; Exchange of active power at the boundary.

[0165] The photovoltaic output constraint expression is:

[0166] ;

[0167] in, This is the upper limit of photovoltaic power output;

[0168] The energy storage operation constraint expression is as follows:

[0169] ;

[0170] ;

[0171] ;

[0172] ;

[0173] in, This represents the state of charge of the energy storage system during period t. , These are the charging and discharging efficiencies, respectively. This refers to the rated capacity of the energy storage. To schedule the time step, , The upper and lower limits of energy storage output;

[0174] The expression for the photovoltaic reduction constraint is:

[0175] ;

[0176] The expression for the cluster switching power constraint is:

[0177] ;

[0178] in, , These are the upper and lower limits for power exchange, respectively.

[0179] The framework proposed in this embodiment is not for traditional mathematical programming solutions. Instead, it employs a multi-agent reinforcement learning approach of "centralized training and distributed execution" to provide a structured and quantifiable learning guide; when the global objective function... The negative values ​​are mapped to the long-term cumulative reward of the multi-agent system, driving the agents to learn to maximize the reward;

[0180] The cloud-based centralized commentator network learns how to coordinate the behavior of each cluster to achieve the goal by evaluating the value of joint states and actions. At the same time, all physical and security constraints are formally embedded into the security layer of each edge agent, and control barrier functions ensure that all control actions generated by the policy network are within the feasible domain. Finally, guided by coordination signals generated by cloud agents and autonomous decision-making by edge agents under security constraints, the entire system approaches the optimal solution of the unified optimization model in a data-driven manner, achieving cloud-edge collaboration that is both economical and secure.

[0181] Furthermore, for edge-distributed photovoltaic clusters, a data-driven control model consisting of a policy network, a security layer, and a value network is constructed. In this embodiment, each edge-distributed photovoltaic cluster uses local observation data and coordination signals sent from the cloud as input. The policy network generates candidate control actions, the security layer verifies the operational constraints and performs projection corrections on the candidate actions, and during the training phase, the value network and the near-end policy optimization algorithm are combined to complete parameter updates. Specifically, this is manifested as follows:

[0182] For each edge photovoltaic cluster i, construct a pair of policy networks based on deep neural networks. With value network The parameters are respectively and Its input state vector in each fast time-scaled control period t is defined as:

[0183] ;

[0184] in, , These are the photovoltaic output and load power measurement sequences of cluster i within the historical window, respectively; , In the prediction time domain Internal photovoltaic and load forecast data; The state of charge of the local energy storage system; The boundary active power of the previous scheduling cycle. This refers to the voltage amplitude at the local node. This is the boundary power reference signal issued by the cloud scheduling layer at the current slow time point, used to correspond to the actions of the cloud-coordinated intelligent agent; The component of the coordination / dual signal published in the cloud at cluster i is used to characterize the consistency requirements of global power flow and power balance constraints.

[0185] In a given state The policy network then outputs the candidate control actions for the cluster, expressed as:

[0186] ;

[0187] in, , These are the active and reactive power reference values ​​for the photovoltaic inverter. Power commands for energy storage charging and discharging; value network Used to assess the state Expected returns on long-term operating costs and coordination performance;

[0188] In this embodiment, a security layer is connected in series after the policy network of each edge photovoltaic cluster to monitor candidate control actions. Feasibility verification and correction are performed, and a set of obstacle functions characterizing operational safety constraints is constructed by introducing a control obstacle function framework, with the expression:

[0189] ;

[0190] Each of them This corresponds to a key operational constraint;

[0191] Define cluster The set of possible actions, expressed as:

[0192] ;

[0193] Candidate actions output by the security layer to the policy network Perform a projection operation to obtain the final safety control action to be executed. The expression is:

[0194] ;

[0195] Among all possible actions that satisfy the control barrier function constraints, select the candidate action. The point with the minimum Euclidean distance is used to preserve the original decision intent of the policy network to the greatest extent possible while ensuring safety.

[0196] Finally, control commands The command is dispatched to the photovoltaic inverters and energy storage systems for execution, and the current boundary power is determined through the power balance relationship within the cluster. and will The relevant statistics are uploaded to the cloud.

[0197] Furthermore, based on the hierarchical control architecture, multi-agent reinforcement learning mechanism, and distributed optimization control model described above, this embodiment constructs a cloud-edge collaborative iterative mechanism for the PPO policy to train and further optimize the policy network.

[0198] The cloud-based coordinating agent acts as the global coordination center, aggregating boundary power statistics from each edge photovoltaic cluster. And related operating status, including node voltage exceeding limits and the degree to which branch power flow approaches thermal stability limits, the cloud-based calculation of the total operating cost and constraint violation degree of the distribution network within the current scheduling cycle is expressed as follows:

[0199] ;

[0200] in, For the cloud during the scheduling cycle The costs of electricity purchase and grid losses, etc. This represents the sum of the local operating costs of each cluster during this period. The constraint violation vector consists of voltage over-limit and power flow over-limit. Its norm;

[0201] Based on this, the cloud-based coordinating agent generates a new boundary active power reference value for each cluster i. and the corresponding penalty weights The boundary power reference value is determined based on the global power flow distribution and economic objectives, and the reference value of the previous cycle is corrected in one step along the direction of reducing network losses and voltage deviation; the penalty weight is adjusted according to the degree of constraint violation and reference deviation. When the voltage or power flow constraint violation caused by a certain cluster is more severe, its penalty weight is increased, as expressed in the following expression:

[0202] ;

[0203] in, For constraint violation metrics related to cluster i, As the allowable safety margin threshold, To adjust the step size; the result is These signals combine to form a coordination signal sent from the cloud to each distributed photovoltaic cluster, serving as constraints and guidance for fast-timescale edge strategy updates and control execution.

[0204] Furthermore, within the multiple fast-timescale control cycles between two updates in the cloud, each distributed photovoltaic cluster receives the coordination signal. Then, based on the state vector With policy network Generate candidate control actions The final action is obtained by projection through the security layer. It drives the operation of photovoltaic inverters and energy storage systems, and the corresponding boundary power is determined by the power balance relationship within the cluster. ;

[0205] The edge cluster constructs an instantaneous reward function in each control cycle based on local operating costs and boundary power tracking error, expressed as:

[0206] ;

[0207] in, The first term represents the local operating cost of the cluster at time t, including photovoltaic power generation costs, energy storage operating costs, and photovoltaic reduction penalties. The second term is a penalty for boundary power deviations from the cloud reference value, with penalty weights... The scheduling is determined by the cloud during the current scheduling period; each cluster collects multiple state-action-reward-next state trajectories within the current cloud scheduling period. Stored in the experience replay pool.

[0208] Furthermore, each distributed photovoltaic cluster uses the aforementioned empirical data to adjust the parameters of the strategy network. Implementing PPO strategy updates: based on value networks Estimating the advantage function Construct the probability ratio, expressed as:

[0209] ;

[0210] Introducing the clipping objective function, the expression is:

[0211] ;

[0212] in, This is the cutting factor;

[0213] By employing a trust region radius constraint, an upper bound is imposed on policy updates, expressed as:

[0214] ;

[0215] in, Let be the radius of the trust region.

[0216] Furthermore, after completing the edge policy update and control execution within a cloud scheduling cycle, each distributed photovoltaic cluster will display the boundary power trajectory for this cycle. The node voltage and branch power flow statistics, as well as the accumulated local operating costs, are transmitted back to the cloud-based coordinating agent. Based on the transmitted data, the cloud-based coordinating agent reassesses the overall operating costs of the distribution network and the degree of constraint violations, and adjusts the boundary power reference values ​​for the next scheduling cycle accordingly. With penalty weight This allows the boundary power to gradually converge toward a direction that satisfies the requirements of global economy and safety.

[0217] Furthermore, to address the issues of high interaction costs and limited training samples in multi-photovoltaic cluster scenarios, this embodiment introduces a sample efficiency optimization method, as follows:

[0218] During the training phase, each distributed photovoltaic cluster generates a quadruple containing state, action, and reward information while executing local control. To improve sample utilization efficiency, in each cluster Side-building experience replay pool It is used to store the interaction trajectory over a recent period of time;

[0219] When updating the policy network, a priority experience replay mechanism is introduced, whereby each experience sample in the replay pool is assigned a priority value. Assign a priority based on TD error. This allows samples with larger TD errors to be selected more frequently, thereby improving sample utilization efficiency and policy convergence speed. The TD error is defined as:

[0220] exist For each sample Assign a dynamically updated priority. The priority is determined by the centralized value network. The timing difference error is determined by the following expression:

[0221] ;

[0222] in, For the immediate reward of sample j, For the estimation of state value in a value network, Discount factor;

[0223] The priority setting expression is:

[0224] ;

[0225] in, To prevent small constants with zero priority.

[0226] Priority This directly reflects the importance of each sample to the improvement of the current strategy. In this embodiment, a proportional priority method is used to allocate sampling probabilities according to priority, as expressed by:

[0227] ;

[0228] in, The degree to which control priority affects the sampling distribution, when Degenerates into uniform sampling when Samples with larger TD errors are selected more frequently for training, thereby focusing on optimizing the state-action regions that are difficult to learn or have large prediction errors.

[0229] Furthermore, this embodiment introduces importance sampling weights to correct each sampled empirical sample, in order to effectively compensate for this bias:

[0230] For pressing from the experience replay pool The importance sampling weight of the sample j obtained from sampling is defined as follows:

[0231] ;

[0232] in, The total number of empirical samples in the replay pool. To control the coefficient of correction intensity, If no correction is performed, The distribution bias caused by priority sampling is fully compensated in time;

[0233] When updating the PPO strategy, the contribution of the PPO objective function to each sample is multiplied by the corresponding weight. Thus, the weighted objective is obtained, expressed as:

[0234] ;

[0235] in, Let be the policy probability ratio for sample j. This is the advantage function based on value network estimation.

[0236] In this embodiment, within each cloud scheduling cycle, each distributed photovoltaic cluster follows a predetermined reward mechanism and control process, stores the real-time collected system interaction experience in a local experience playback pool, and dynamically updates the TD error of each sample and its corresponding priority.

[0237] When the cluster triggers a policy update, it no longer uses uniform random sampling, but instead performs batch empirical sampling based on the probability determined by the sample priority.

[0238] When calculating the PPO objective function, each sampled empirical sample is given an importance sampling weight to effectively correct the estimation bias caused by non-uniform sampling, thereby ensuring the stability and convergence of policy learning.

[0239] Example 2

[0240] Based on the cloud-edge collaborative reinforcement learning control method for photovoltaic cluster access to distribution networks described in Example 1, this example introduces a cloud-edge collaborative reinforcement learning control method system for photovoltaic cluster access to distribution networks, including:

[0241] A control architecture construction module is used to build a hierarchical control architecture, which includes a cloud-based power grid dispatch layer and an edge photovoltaic cluster layer.

[0242] The centralized-distributed training module is used to embed the hierarchical control architecture into a multi-agent deep reinforcement learning model for centralized training in the cloud and distributed execution at each edge.

[0243] The distributed optimization solution module is used to establish a distributed optimization control model. By combining a multi-agent reinforcement learning mechanism, the optimal control strategy of the distributed optimization control model is obtained through collaborative iterative solution between the cloud and the edge.

[0244] The strategy network update module is used to build a data-driven control model for each edge photovoltaic cluster based on local observation data and coordination signals sent from the cloud, and to update the strategy network parameters.

[0245] The cloud-edge collaborative strategy optimization module is used to train and optimize the policy network through a cloud-edge collaborative iterative mechanism based on near-end policy optimization, so as to obtain the optimized control policy.

[0246] The sample efficiency optimization module is used to construct an experience replay pool. Based on the optimized control strategy, it dynamically assigns priorities to the experience samples stored in the replay pool and introduces importance sampling weights for bias correction, thereby obtaining a stable sample optimization learning method.

[0247] Example 3

[0248] Based on the cloud-edge collaborative reinforcement learning control method for photovoltaic cluster access to the distribution network described in Embodiment 1, this embodiment introduces a computer-readable storage medium storing a computer program / instruction. When the computer program / instruction is executed by a processor, it implements the steps of the cloud-edge collaborative reinforcement learning control method for photovoltaic cluster access to the distribution network as described in any of Embodiment 1.

[0249] Example 4

[0250] Based on the cloud-edge collaborative reinforcement learning control method for photovoltaic cluster access to the distribution network described in Embodiment 1, this embodiment provides a computer device, including:

[0251] Memory, used to store computer programs / instructions;

[0252] A processor is configured to execute the computer program / instructions to implement the steps of the cloud-edge collaborative reinforcement learning control method for photovoltaic cluster access to the distribution network as described in any one of Embodiments 1.

[0253] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0254] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0255] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1The function specified in one or more boxes.

[0256] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0257] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.

Claims

1. A cloud-edge collaborative reinforcement learning control method for photovoltaic cluster access to power distribution networks, characterized in that, include: A hierarchical control architecture is constructed, which includes a cloud-based power grid dispatch layer and an edge photovoltaic cluster layer. The hierarchical control architecture is embedded into a multi-agent deep reinforcement learning model for centralized training in the cloud and distributed execution at each edge. A distributed optimization control model is established, and combined with a multi-agent reinforcement learning mechanism, the optimal control strategy of the distributed optimization control model is obtained through collaborative iterative solution between the cloud and the edge. Each photovoltaic cluster at the edge constructs a data-driven control model based on local observation data and coordination signals sent from the cloud, and completes the update of strategy network parameters; The policy network is trained and optimized through a cloud-edge collaborative iterative mechanism based on near-end policy optimization to obtain the optimized control policy. An experience replay pool is constructed. Based on the optimized control strategy, the experience samples stored in the replay pool are dynamically assigned priorities, and importance sampling weights are introduced for bias correction, resulting in a stable sample optimization learning method.

2. The cloud-edge collaborative reinforcement learning control method for photovoltaic cluster access to distribution networks according to claim 1, characterized in that, The cloud-based distribution network scheduling layer, serving as the global coordination center of the distribution network, is deployed at the master station. Based on a simplified equivalent model of the distribution network, it receives and aggregates boundary power statistics reported from each edge cluster, and combines them with ultra-short-term load and photovoltaic power generation forecast data. With the goal of minimizing the total operating cost of the distribution network, it generates global collaborative scheduling instructions within a given scheduling cycle and distributes them to each edge cluster to achieve cross-cluster power flow constraints and operating cost coordination. The equivalent model of the distribution network is constructed by equating each edge distributed cluster to an injection node and utilizing node voltage sensitivity parameters uploaded from the edge cluster layer. This allows for accurate estimation of global power flow and voltage state at the cloud level with low computational complexity, as expressed below: Consider a containing The distribution network system of each edge cluster, the cloud layer will connect each cluster Equivalent to a power injection point, its equivalent injected active power is ; The edge photovoltaic cluster control layer consists of multiple geographically dispersed distributed clusters containing photovoltaic power generation units, energy storage systems, and loads. Each cluster acts as an autonomous energy unit, aiming to minimize internal operating costs. It utilizes local computing power to achieve real-time optimized scheduling of photovoltaics, energy storage, and loads. After completing local optimization, it uploads its key boundary power information to the cloud-based distribution network scheduling layer. Through node voltage sensitivity estimation, it provides complete data to the distribution network layer and provides key model parameters for the coordinated voltage regulation of each photovoltaic cluster.

3. The cloud-edge collaborative reinforcement learning control method for photovoltaic cluster access to distribution networks according to claim 1, characterized in that, The centralized training in the cloud refers to deploying a centralized commentator network in the cloud. During the training phase, this network can access the observations and actions of all agents to accurately estimate the global value of joint actions and guide the updates of each policy network. The global reward function is designed as a negative of the total operating cost of the distribution network and includes penalties for safety constraints such as voltage overruns. The decentralized execution at each edge refers to the deployment of the policy network separately after training, with the cloud-coordinated agent relying solely on observations. Generate coordination signals; each edge cluster agent relies solely on local observations. It can independently generate and execute local control actions based on received cloud signals, without the need for global communication, and is used to achieve scalable distributed control.

4. The cloud-edge collaborative reinforcement learning control method for photovoltaic cluster access to distribution networks according to claim 1, characterized in that, The distributed optimization control model includes a cloud-based global optimization model and an edge-based local optimization model. The cloud-based global optimization model provides clear learning objectives and value assessment criteria for the centralized training phase. The objective function is defined as minimizing the total operating cost of the distribution network, including the power exchange cost of each cluster, the total active power loss of the system, and the voltage deviation penalty at key nodes. This model is not directly used for numerical solutions but serves as the design basis for the global reward function in multi-agent reinforcement learning. A centralized commentator network is used to achieve global value assessment of joint actions. The objective function expression is: ; in This represents the total operating cost of the system during the scheduling period. t is the time index; The total number of photovoltaic clusters; Let be the local operating cost of the i-th photovoltaic cluster at time t. The operating cost of the distribution network at time t mainly includes network loss costs; This is the global safety penalty term incurred at time t due to node voltage exceeding the limit; The edge-local optimization model aims to minimize local operating costs. The objective function includes photovoltaic (PV) power generation costs, energy storage operating costs, PV reduction penalties, and costs associated with tracking cloud commands. It must also satisfy physical constraints such as power balance within the cluster, PV output, energy storage operation, and power exchange. The objective function expression is as follows: ; in It is a function of the operation and maintenance costs of photovoltaic power generation; Photovoltaic power generation; The cycle loss cost function for energy storage devices; For energy storage power, To address the issue of curtailed photovoltaic power The penalty function.

5. The cloud-edge collaborative reinforcement learning control method for photovoltaic cluster access to distribution networks according to claim 1, characterized in that, The data-driven control model is a local security reinforcement learning controller deployed on each distributed photovoltaic cluster side, consisting of a policy network, a security layer, and a value network. Define local observation states on each distributed photovoltaic cluster side, and configure a policy network based on the local observation states. The policy network generates candidate control actions such as photovoltaic inverter output and energy storage charging and discharging power according to the local observation states. A security layer is connected in series after the policy network. This security layer constructs a set of operational constraints based on a control barrier function to ensure the safe operation of the system, expressed as: ; Each of them This corresponds to a key operational constraint; Within each control cycle, the feasible domain of the control action is determined based on the set of operational constraints. The candidate control action is then projected onto the feasible domain according to the principle of minimum modification and output as the final control command to ensure that all operational constraints are always satisfied. The security layer outputs candidate actions to the policy network. The projection operation is performed, and the final safety control actions are executed. The expression is: ; in, For actionable sets, For each distributed photovoltaic cluster, a value network is configured with a policy network using a secure reinforcement learning framework, with the state vector as the input. The value network takes the local observation state as input and is used to evaluate the long-term operating costs and returns under the corresponding state.

6. The cloud-edge collaborative reinforcement learning control method for photovoltaic cluster access to distribution networks according to claim 1, characterized in that, The cloud-edge collaborative iterative mechanism uses the cloud-based coordinating agent as the global coordination center to aggregate boundary power statistics from each edge photovoltaic cluster. And related operating status, including node voltage exceeding limits and the degree to which branch power flow approaches the thermal stability limit, the cloud calculates the total operating cost and constraint violation degree of the distribution network within the current scheduling cycle, expressed as: ; in, For the cloud during the scheduling cycle The costs of electricity purchase and grid losses, etc. This represents the sum of the local operating costs of each cluster during this period. The constraint violation vector consists of voltage over-limit and power flow over-limit. It is a norm; During the multiple fast-timescale control cycles between two updates in the cloud, each distributed photovoltaic cluster receives the coordination signal. Then, based on the state vector With policy network Generate candidate control actions The final action is obtained by projection through the security layer. It drives the operation of photovoltaic inverters and energy storage systems, and the corresponding boundary power is determined by the power balance relationship within the cluster. .

7. The cloud-edge collaborative reinforcement learning control method for photovoltaic cluster access to distribution networks according to claim 1, characterized in that, The experience playback pool Used to store experience samples Utilizing value networks The temporal difference error for each empirical sample is calculated using the following expression: ; And define priorities: ; in, As a discount factor, To prevent small constants with zero priority, a priority experience replay mechanism is used during sampling. Samples are sampled proportionally according to their priority, so that samples with larger time difference errors are selected more frequently, thereby improving sample utilization efficiency and strategy convergence speed.

8. A cloud-edge collaborative reinforcement learning control system for photovoltaic cluster access to power distribution networks, characterized in that, include: A control architecture construction module is used to build a hierarchical control architecture, which includes a cloud-based power grid dispatch layer and an edge photovoltaic cluster layer. A centralized-distributed training module is used to embed the hierarchical control architecture into a multi-agent deep reinforcement learning model for centralized training in the cloud and distributed execution at each edge. The distributed optimization solution module is used to establish a distributed optimization control model. By combining a multi-agent reinforcement learning mechanism, the optimal control strategy of the distributed optimization control model is obtained through collaborative iterative solution between the cloud and the edge. The strategy network update module is used to build a data-driven control model for each edge photovoltaic cluster based on local observation data and coordination signals sent from the cloud, and to update the strategy network parameters. The cloud-edge collaborative strategy optimization module is used to train and optimize the policy network through a cloud-edge collaborative iterative mechanism based on near-end policy optimization, so as to obtain the optimized control policy. The sample efficiency optimization module is used to construct an experience replay pool. Based on the optimized control strategy, it dynamically assigns priorities to the experience samples stored in the replay pool and introduces importance sampling weights for bias correction, thereby obtaining a stable sample optimization learning method.

9. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instruction is executed by the processor, it implements the steps of the cloud-edge collaborative reinforcement learning control method for photovoltaic cluster access to the distribution network as described in any one of claims 1 to 7.

10. A computer device, characterized in that, include: Memory, used to store computer programs / instructions; A processor is configured to execute the computer program / instructions to implement the steps of the cloud-edge collaborative reinforcement learning control method for photovoltaic cluster access to the distribution network as described in any one of claims 1 to 7.