Photovoltaic cluster flexible grid-connected regulation and control method and system
By using data preprocessing and multi-objective optimization models, combined with reinforcement learning and decision tree algorithms, the control of photovoltaic inverters and energy storage units is coordinated, solving the environmental adaptability and comprehensive optimization problems in the grid-connected regulation of photovoltaic clusters. This achieves a balance between grid friendliness and power generation efficiency, and improves the stability and economy of the system.
Patent Information
- Application Number
- CN202511305125.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-12
- Publication Date
- 2025-12-23
AI Technical Summary
Existing photovoltaic cluster grid-connected control technologies are ill-suited to rapidly changing operating environments and lack the comprehensive optimization capabilities for grid friendliness and photovoltaic power generation benefits, leading to problems such as voltage over-limit, power backflow, and photovoltaic power curtailment.
By employing data cleaning and timestamp alignment preprocessing techniques, a multi-objective optimization model is constructed. Reinforcement learning and optimal decision tree algorithms are used to coordinate the control strategies of photovoltaic inverters and energy storage units to achieve flexible grid connection. The strategy is then verified through a hierarchical coordinated reactive power control system and a digital twin system.
It has improved the system's adaptability to environmental changes, achieved comprehensive optimization of grid friendliness and photovoltaic power generation benefits, solved the problems of voltage over-limit and power backflow, and ensured a balance between grid security and economy.
Smart Images

Figure CN121192811A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of new energy power generation, specifically to a method and system for flexible grid-connected control of photovoltaic clusters, and particularly to the use of reinforcement learning and optimal decision tree algorithms to achieve efficient and flexible grid-connected control of photovoltaic clusters. Background Technology
[0002] Distributed photovoltaic (PV) power generation, as an important component of clean energy, plays an increasingly vital role in the global energy transition. With the continuous development of PV technology and the sustained decline in costs, the penetration rate of distributed PV systems in power distribution networks has been increasing year by year, forming large-scale PV clusters.
[0003] Currently, common photovoltaic (PV) grid-connection technologies mainly include two modes: rigid grid connection and flexible grid connection. In rigid grid connection, PV systems typically employ a maximum power point tracking (MPPT) control strategy, injecting all possible power generation into the grid. Flexible grid connection, on the other hand, adjusts PV output to adapt to grid demand while considering grid constraints. Traditional PV grid-connection control usually employs simple droop control or centralized dispatching methods, lacking real-time perception and coordination of the overall distribution network status.
[0004] Existing advanced photovoltaic (PV) cluster grid-connected control technologies are typically based on deterministic programming models. These models establish the relationship between PV power generation and grid operating conditions, determining the operating parameters of the PV inverter by solving optimization problems. These technologies employ static optimization methods, calculating the optimal setpoints for the inverter based on a pre-defined grid model and predicted PV output.
[0005] However, this method has obvious technical defects: on the one hand, due to the intermittency and volatility of photovoltaic power generation and the dynamic changes in the distribution network load, static optimization is difficult to adapt to the rapidly changing operating environment; on the other hand, traditional methods often only focus on a single objective (such as voltage quality or photovoltaic absorption capacity), lacking the ability to comprehensively optimize grid friendliness and photovoltaic power generation benefits, which leads to problems such as voltage exceeding limits, power backfeeding, and photovoltaic power curtailment in the case of large-scale photovoltaic access. Summary of the Invention
[0006] The purpose of this invention is to provide a flexible grid-connected control method and system for photovoltaic clusters, aiming to solve the technical problems in the prior art where static optimization is difficult to adapt to rapidly changing operating environments and lacks comprehensive optimization capabilities for grid friendliness and photovoltaic power generation benefits.
[0007] To achieve the above objectives, the present invention provides a flexible grid-connected control method for photovoltaic clusters, comprising:
[0008] The system collects status data including voltage, current, and power parameters of distribution network nodes, power generation of photovoltaic clusters, and inverter operating parameters. The status data is then cleaned and preprocessed with timestamp alignment to obtain a comprehensive system status data stream.
[0009] Based on the system state integrated data stream, a multi-objective optimization model that includes grid friendliness and photovoltaic power generation benefits is constructed. The multi-objective optimization model is solved using a reinforcement learning algorithm to obtain the control strategies for the photovoltaic inverter and energy storage unit.
[0010] According to the control strategy, the active power output and reactive power output of the photovoltaic inverter and the charging and discharging power of the energy storage unit are coordinated and controlled to realize the flexible grid connection of the photovoltaic cluster and collect system response data after the control is executed.
[0011] Based on the system response data, the control strategy is optimized using the optimal decision tree algorithm, the reinforcement learning model is updated, and an optimization experience knowledge base is established.
[0012] Preferably, the step of performing data cleaning and timestamp alignment preprocessing on the state data to obtain a comprehensive system state data stream includes:
[0013] The state data is cleaned by detecting and processing missing and outlier values based on the 3σ criterion to obtain cleaned state data.
[0014] The cleaned state data is resampled according to a unified second-level time base to complete the timestamp alignment and obtain the time-aligned state data.
[0015] The time-aligned state data is structured using BWT-Runs spatial indexing technology, including discretizing the time-series data, applying BWT transformation to generate a sorted suffix array, identifying repetitive patterns and performing Run-length encoding, thereby achieving the fusion of multi-source heterogeneous data and obtaining the comprehensive system state data stream.
[0016] Preferably, after obtaining the system state integrated data stream, the method further includes:
[0017] Based on the system status integrated data stream, the key performance indicators of the power grid, such as node voltage deviation, line load rate, and power backflow, as well as photovoltaic power generation indicators such as photovoltaic absorption rate and power generation efficiency, are calculated to obtain a real-time system status assessment report.
[0018] Based on the real-time status assessment report of the system, a power flow equation for the distribution network is established, which includes voltage constraints (node voltage within ±7% of the rated value) and capacity constraints (line power not exceeding the rated capacity).
[0019] Preferably, the construction of a multi-objective optimization model that includes grid friendliness and photovoltaic power generation benefits includes:
[0020] A grid-friendly objective function is constructed, which includes sub-objectives of minimizing voltage deviation, minimizing reverse power transmission, and maximizing photovoltaic absorption.
[0021] A photovoltaic power generation benefit objective function is constructed, which includes sub-objectives of maximizing power generation revenue and minimizing equipment losses;
[0022] The grid-friendliness objective function and the photovoltaic power generation benefit objective function are integrated into a multi-objective optimization model using a weighted method. The weight coefficients of each objective function are dynamically adjusted according to the grid operation status of the integrated system state data stream. When a voltage over-limit risk is detected, the weight of the grid-friendliness objective function is increased.
[0023] Preferably, solving the multi-objective optimization model using a reinforcement learning algorithm includes:
[0024] The multi-objective optimization problem is transformed into a Markov decision process, and the integrated data flow of the system state is mapped to the state space, and the control parameters of the photovoltaic inverter and energy storage unit are mapped to the action space.
[0025] A reward function is constructed based on the objective function of the multi-objective optimization model. A deep reinforcement learning network containing a policy network and a value network is used for policy training to obtain the control policy. The policy network and the value network both adopt a multilayer perceptron structure.
[0026] The control strategy was simulated and verified to evaluate its voltage control accuracy, power fluctuation suppression effect, and economic performance.
[0027] Preferably, the step of coordinating and controlling the active power output and reactive power output of the photovoltaic inverter according to the control strategy includes:
[0028] The active power setpoint of each photovoltaic inverter is generated based on the control strategy.
[0029] The active power setpoint is sent to the control system of each photovoltaic inverter via a communication network;
[0030] Adjust the maximum power point tracking control strategy of each photovoltaic inverter to achieve active power control;
[0031] The system monitors the execution of active power control commands in real time. When a communication interruption or equipment failure is detected, it initiates a local control mode based on local measurement data to adjust the active power to a preset safe operating point.
[0032] Preferably, the step of coordinating and controlling the active power output and reactive power output of the photovoltaic inverter according to the control strategy includes:
[0033] Determine the global reactive power optimization target based on the control strategy described above;
[0034] Based on the local voltage measurement results of each photovoltaic inverter, a reactive power control command or power factor setting value is generated.
[0035] Establish a hierarchical coordinated reactive power control system from a single inverter to a photovoltaic cluster, including first-level single-unit reactive power regulation, second-level substation reactive power coordination, and third-level cluster reactive power optimization, to achieve dynamic support for grid voltage, wherein the voltage support target is to maintain the voltage of each node within ±5% of the rated value.
[0036] Real-time monitoring of the execution effect of reactive power control commands and dynamic adjustment of reactive power allocation strategies.
[0037] Preferably, the step of coordinating and controlling the active power output and reactive power output of the photovoltaic inverter and the charging and discharging power of the energy storage unit according to the control strategy includes:
[0038] Based on grid load conditions, photovoltaic power generation forecasts, and electricity price signals, the timing of charging and discharging of energy storage units is determined;
[0039] The energy storage unit's charging and discharging power commands are generated according to the control strategy.
[0040] Coordinate and control the charging and discharging behavior of each energy storage unit to smooth out photovoltaic power output fluctuations. The goal of power output fluctuation smoothing is to control the power change rate within 10% / minute of the rated capacity.
[0041] Monitor the state of charge of the energy storage unit to ensure it is within the safe operating range of 20% to 80%.
[0042] Preferably, optimizing the control strategy using the optimal decision tree algorithm includes:
[0043] Collect system response data after control execution, including changes in grid operating parameters, photovoltaic output adjustment, and energy storage unit status changes;
[0044] Calculate quantitative indicators that include voltage deviation improvement rate, photovoltaic absorption rate improvement, power backflow reduction, and economic benefit increase.
[0045] The optimal decision tree algorithm is used to evaluate each decision point of the control strategy. An evaluation threshold is set based on the quantitative index to identify decision branches with poor performance that are below the threshold.
[0046] For decision branches with poor performance, pruning is performed by minimizing the depth of the decision tree, and branches are merged by merging similar decision rules to optimize the decision process while maintaining control accuracy.
[0047] The present invention also provides a flexible grid-connected control system for photovoltaic clusters, comprising:
[0048] The data acquisition module is used to collect status data including voltage, current, and power parameters of distribution network nodes, power generation of photovoltaic clusters, and inverter operating parameters. The status data is cleaned and preprocessed with timestamp alignment to obtain a comprehensive system status data stream.
[0049] The optimization control module is used to construct a multi-objective optimization model that includes grid friendliness and photovoltaic power generation benefits based on the comprehensive data flow of the system state, and to solve the multi-objective optimization model using a reinforcement learning algorithm to obtain the control strategy of the photovoltaic inverter and energy storage unit.
[0050] The execution control module is used to coordinate and control the active power output and reactive power output of the photovoltaic inverter and the charging and discharging power of the energy storage unit according to the control strategy, so as to realize the flexible grid connection of the photovoltaic cluster and collect system response data after the control is executed.
[0051] The evaluation and optimization module is used to optimize the control strategy based on the system response data using the optimal decision tree algorithm, update the reinforcement learning model, and establish an optimization experience knowledge base.
[0052] The beneficial effects of this invention include:
[0053] 1. Efficient historical data retrieval and pattern recognition based on BWT-Runs spatial indexing technology: Upgrade the historical data management in the traditional photovoltaic grid-connected control system to an efficient indexing structure based on BWT-Runs (Burrows-WheelerTransform with Runs), realize the rapid retrieval and similar pattern recognition of massive heterogeneous time series data, greatly improve the query efficiency of similar scenarios, and provide accurate historical experience reference for multi-objective optimization models;
[0054] 2. Adaptive control strategy optimization combined with the simple approximation algorithm of the optimal decision tree: The photovoltaic cluster control strategy is dynamically evaluated and optimized by the simple approximation algorithm of the optimal decision tree. Inefficient decision branches are automatically identified and improved. While keeping the computational complexity under control, the system approaches the global optimal control effect, thereby improving the system's adaptability to environmental changes and control stability.
[0055] 3. Flexible grid-connected control model with multi-objective collaborative optimization: An innovative comprehensive optimization model is constructed that simultaneously considers grid friendliness (voltage stability, reduction of backfeed power, and maximization of absorption capacity) and photovoltaic power generation benefits. The model is dynamically solved through reinforcement learning algorithm to achieve coordinated control of the active / reactive output of photovoltaic inverter and the charging and discharging behavior of energy storage, thus solving the system imbalance problem caused by single-objective optimization in traditional methods.
[0056] 4. Hierarchical and coordinated reactive power control system: Based on local voltage measurement and global reactive power optimization objectives, a hierarchical and coordinated reactive power control system from a single inverter to a photovoltaic cluster is designed to achieve an organic combination of voltage support and reactive power optimization, effectively solving the voltage limit exceeding problem caused by large-scale photovoltaic access.
[0057] 5. Control strategy verification mechanism based on digital twin: Before the control strategy is executed, simulation verification is carried out through the distribution network digital twin system to evaluate the effectiveness and stability of the strategy, reduce actual operation risks, and ensure the balance between grid security and photovoltaic system economy. Attached Figure Description
[0058] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0059] Figure 1 A flowchart of a flexible grid-connected control method for photovoltaic clusters provided in an embodiment of the present invention;
[0060] Figure 2 The diagram shows the structure of a flexible grid-connected control system for a photovoltaic cluster, as provided in an embodiment of the present invention. Detailed Implementation
[0061] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments.
[0062] like Figure 1 As shown, this embodiment of the invention provides a flexible grid-connected control method for photovoltaic clusters, comprising the following steps:
[0063] Step S1: Collect status data including voltage, current, and power parameters of distribution network nodes and power generation and inverter operating parameters of photovoltaic clusters; perform data cleaning and timestamp alignment preprocessing on the status data to obtain a comprehensive system status data stream.
[0064] In this embodiment, intelligent sensing devices, including voltage / current transformers and power quality analyzers, need to be deployed at key nodes of the distribution network to collect real-time data such as voltage, current, active / reactive power, power factor, and harmonics, forming a power grid operation status dataset. These sensing devices are typically installed at transformers, critical lines, and connection points to comprehensively monitor the operation status of the distribution network. Simultaneously, a photovoltaic cluster monitoring system collects real-time power generation, inverter operating parameters (including active / reactive output and power factor), environmental parameters (including illuminance and temperature), and the charging / discharging status and state of charge (SOC) of energy storage units from each photovoltaic power station, forming a photovoltaic cluster operation status dataset. These data acquisition devices typically sample at a frequency ranging from 1 second to 1 minute and transmit the data to the central control system via a dedicated communication network, providing a data foundation for subsequent analysis and decision-making.
[0065] After collecting the raw data, a comprehensive data cleaning process is performed. This process first detects and processes missing values in the data, using time series interpolation algorithms to complete the data. For example, if a voltage sensor has missing data for a short period, it will be reasonably filled in based on the historical data trend of that node to ensure data continuity. At the same time, statistical methods such as the 3σ criterion are used to identify outliers, and they are repaired or marked according to the detection results. For unreasonable jumps in photovoltaic power data (such as a sudden drop in power from 80% to 0% and then a rapid recovery within a short period), such anomalies are automatically identified, and smooth corrections are made based on the trend of the preceding and following data to avoid interference from abnormal data in subsequent analysis. This data cleaning process not only improves data quality but also identifies potential equipment failures or communication problems, providing a reference for maintenance.
[0066] Next, the issue of inconsistent timestamps from different data sources is addressed. Since the sampling frequencies and times of power grid sensing devices and photovoltaic monitoring systems may differ, time alignment is necessary. The system employs a unified time base and resampling technology to adjust all data to the same time scale (typically at the second level), ensuring data consistency across the time dimension. For example, when the power grid monitoring device samples at 5-second intervals while the photovoltaic monitoring device samples at 1-minute intervals, both are unified to the same time scale, and the difference in time resolution is addressed through interpolation or averaging. This time alignment ensures the temporal correlation between data from different sources, laying the foundation for subsequent multi-source data fusion and pattern recognition.
[0067] After data cleaning and time alignment, BWT-Runs spatial indexing technology is introduced to organize the data in a structured manner. This technology first discretizes various types of time-series data appropriately, mapping continuous values to a finite character set; then, it applies BWT transformation to the discretized data sequence to generate a sorted suffix array; next, it identifies repeating patterns in the transformed sequence and applies Run-length encoding for compressed storage; finally, it constructs a multi-level index structure to support rapid location and retrieval of similar data patterns. The advantage of this index structure lies in its ability to quickly retrieve historical scenarios similar to the current state, providing efficient support for subsequent pattern recognition and decision optimization. For example, when the system needs to find historical operating data under combined conditions such as "sunny day, high load, voltage fluctuation," traditional methods require traversing the entire database, while the BWT-Runs index can locate the target data segment in sublinear time, significantly improving retrieval efficiency, saving computational resources, and accelerating response speed.
[0068] Finally, different types of data are fused and integrated to form a unified system state representation. The fusion process is based on the physical meaning and spatiotemporal correlation of the data, employing a combination of feature-level fusion and decision-level fusion. For example, data such as voltage, power, and photovoltaic output on the same distribution feeder are correlated to construct a feeder state vector; or data from multiple geographically proximate photovoltaic sites are fused to form regional photovoltaic cluster characteristics. Through this multi-source heterogeneous data fusion, a more comprehensive and accurate understanding of the system state can be obtained, providing a reliable foundation for subsequent optimization decisions. After data preprocessing, various key performance indicators, such as grid indicators like node voltage deviation, line load rate, and power backflow, and photovoltaic power generation indicators like photovoltaic absorption rate and power generation efficiency, are calculated based on the comprehensive system state data stream, generating a real-time system state assessment report to provide a basis for subsequent optimization model construction.
[0069] Step S2: Based on the system state integrated data stream, construct a multi-objective optimization model that includes grid friendliness and photovoltaic power generation benefits, and use reinforcement learning algorithm to solve the multi-objective optimization model to obtain the control strategy of photovoltaic inverter and energy storage unit;
[0070] In this embodiment, a grid-friendly objective function is first constructed. This objective function includes three key sub-objectives: minimizing voltage deviation, minimizing reverse power flow, and maximizing photovoltaic (PV) absorption. Minimizing voltage deviation aims to bring the voltage of all nodes as close as possible to the ideal value (usually the rated value), reducing the impact of voltage fluctuations on electrical equipment. In actual operation, excessively high or low voltages can adversely affect electrical equipment, such as shortening equipment lifespan, increasing losses, or causing protection system activation. Minimizing reverse power flow aims to reduce power backflow from the low-voltage side to the high-voltage side, avoiding increased network losses and protection coordination issues. When PV power generation exceeds local load demand, excess power flows from the low-voltage side to the high-voltage side; this reverse flow can lead to increased line losses and malfunctions of protection devices. Maximizing PV absorption, under the premise of ensuring grid security, maximizes the utilization of renewable energy, improves the utilization efficiency of PV power generation, reduces curtailment, and achieves both environmental and economic benefits. These three sub-objectives are combined in a weighted manner to form a comprehensive grid-friendly objective function.
[0071] Next, a photovoltaic (PV) power generation benefit objective function is constructed. This objective function mainly considers two aspects: maximizing power generation revenue and minimizing equipment losses. Regarding power generation revenue, the system calculates the economic benefits of PV power generation based on local electricity pricing policies (such as fixed tariffs, time-of-use tariffs, or market-based tariffs) and possible subsidy mechanisms. Under China's current electricity pricing mechanism, PV power generation revenue mainly comes from the sale of grid-connected electricity and renewable energy subsidies (if applicable). Regarding equipment losses, factors such as inverter efficiency variations at different load rates, the impact of frequent operating point adjustments on equipment lifespan, and charging and discharging losses of the energy storage system are considered. For example, the inverter efficiency curve exhibits an inverted U-shape at different load rates, typically reaching its highest efficiency at 70%-90% load rates, while efficiency decreases at extremely low or high load rates. Frequent adjustments to PV output or energy storage charging and discharging states not only increase control costs but may also accelerate equipment aging. These factors are comprehensively considered to form a quantitative expression of PV power generation benefits.
[0072] To balance the potentially conflicting objectives of grid friendliness and photovoltaic (PV) power generation benefits, the system employs a weighting method to integrate the two objective functions into a multi-objective optimization model. The weighting coefficients are dynamically adjusted based on the grid's operating status. For example, when a voltage risk of exceeding limits is detected, the weight of the grid friendliness objective function is automatically increased to prioritize grid safety; conversely, when the grid is in good condition, the weight of the power generation benefit objective is appropriately increased to maximize the economic returns of PV power generation. This dynamic weighting mechanism allows the system to flexibly adjust its optimization strategy based on real-time conditions, achieving the optimal balance between grid safety and economic benefits.
[0073] In addition to the objective function, a series of constraints are set, including physical constraints (such as node voltage constraints and line capacity constraints), equipment constraints (such as inverter capacity constraints and energy storage SOC constraints), operational constraints (such as photovoltaic power output change rate constraints), and safety margin constraints. These constraints ensure that the optimization results are feasible in the actual system and avoid unrealistic control strategies. For example, the node voltage constraint requires that the voltage amplitude of all nodes be kept within ±7% of the rated value, the line capacity constraint requires that the power of all lines not exceed their rated capacity, and the energy storage system SOC constraint requires that the battery state of charge be kept within a safe range of 20% to 80%. These constraints work together to ensure the safe and stable operation of the system.
[0074] After constructing the multi-objective optimization model, a reinforcement learning algorithm is used for solving it. First, the multi-objective optimization problem is transformed into a Markov decision process, mapping the system state integrated data flow to a state space, and mapping the control parameters of the photovoltaic inverter and energy storage unit to an action space. The state space includes grid state parameters (voltages at each node, line power flow, etc.), photovoltaic system parameters (power generation at each site, inverter status, etc.), energy storage system parameters (SOC level, charge / discharge power, etc.), environmental parameters (illuminance, temperature, etc.), and load parameters. The action space includes active power adjustment commands for the photovoltaic inverter, reactive power or power factor setpoints, and charge / discharge power commands for the energy storage system.
[0075] Based on the objective function of a multi-objective optimization model, a reward function for reinforcement learning is constructed, comprehensively considering rewards for grid-friendliness (such as improved voltage compliance rate and reduced backfeed power), rewards for power generation efficiency, and penalties for constraint violations. The system employs a deep reinforcement learning network for policy training, including a policy network (responsible for selecting the optimal action based on the current state) and a value network (evaluating the long-term value of state-action pairs). Both networks utilize a multilayer perceptron structure, continuously optimizing the control strategy through extensive interactive learning. The training process includes offline pre-training (constructing an initial policy using historical data), simulation environment training (trial and error learning in a digital twin system), exploration-exploitation balancing (large-scale exploration in the initial stage, followed by fine-tuning in the later stage), hierarchical training (training single-site control first, then coordinated control), and safety constraints (ensuring that the exploration process does not violate key safety constraints). Through the training mechanism designed above, the reinforcement learning algorithm can learn near-optimal control strategies in complex and ever-changing photovoltaic grid-connected environments.
[0076] After training, the learned control strategy is simulated and verified to evaluate its voltage control accuracy (e.g., the degree of voltage deviation reduction), power fluctuation suppression effect (e.g., power change rate control), and economic indicators (e.g., the degree of improvement in power generation revenue). The verification process covers various scenarios, including normal operation, extreme operation (e.g., the combination of maximum solar radiation and minimum load), fault scenarios (e.g., line short circuits, equipment tripping), communication interruption scenarios, and parameter uncertainty scenarios, ensuring that the control strategy can work stably and effectively under various conditions. If the verification finds that the strategy performs poorly in certain scenarios, the system will apply a simple approximation algorithm of the optimal decision tree to adjust the strategy, transforming the complex neural network strategy into interpretable and adjustable decision rules, thus maintaining the high performance of the strategy while improving its reliability and interpretability.
[0077] Step S3: According to the control strategy, coordinate the active power output and reactive power output of the photovoltaic inverter, as well as the charging and discharging power of the energy storage unit, to realize the flexible grid connection of the photovoltaic cluster, and collect the system response data after the control is executed;
[0078] In this embodiment, the photovoltaic cluster coordinated control adopts a hierarchical distributed architecture, including three levels: the cluster layer, the site area layer, and the equipment layer. The cluster layer is responsible for the formulation and coordination of global control strategies; the site area layer is responsible for the coordinated control of multiple photovoltaic sites within the same area; and the equipment layer is responsible for the local control of individual inverters or energy storage units. This hierarchical architecture ensures both global optimality and good robustness and real-time performance.
[0079] In terms of active power control, the total active power adjustment for each photovoltaic site is first determined based on the grid status and optimization objectives. Then, based on the principles of fairness and efficiency, the total adjustment is allocated to each inverter. The allocation algorithm considers the rated capacity, current output level, and regulation capability of each inverter, employing a weighted proportional allocation method, with weights related to inverter capacity and efficiency. For example, inverters with larger capacities or currently operating at high efficiency receive smaller reduction ratios to maximize overall power generation efficiency.
[0080] For reactive power control, an innovative hierarchical coordinated control system has been implemented. The first level is individual inverter reactive power regulation, where each inverter autonomously adjusts its reactive power output based on local voltage measurements, following a characteristic curve similar to droop control. The second level is substation reactive power coordination, where multiple inverters under the same transformer collaboratively adjust reactive power distribution to avoid competition or cancellation. The third level is cluster reactive power optimization, which optimizes the reactive power regulation targets of each substation based on global voltage distribution and network topology. These three levels of control operate on different time scales: individual inverter control response time is in the millisecond range, substation coordination is in the second range, and cluster optimization is in the minute range.
[0081] The reactive power allocation algorithm employs a sensitivity-based approach. It first calculates the influence coefficient matrix of reactive power injection at each node on the critical voltage point, and then, based on this matrix and the target voltage deviation, solves for the optimal reactive power allocation scheme. The algorithm design considers the priority order of reactive power compensation devices: first, dedicated reactive power compensation devices such as SVG are utilized; second, the remaining capacity of photovoltaic inverters is utilized; and finally, the reactive power capacity of energy storage converters is invoked when necessary.
[0082] Energy storage unit control strategies encompass two key dimensions: timing of charging and discharging, and power allocation, both spatially and temporally. Timing-based decisions are based on electricity price signals, load forecasting, and photovoltaic (PV) forecasting, employing dynamic programming algorithms to optimize charging and discharging sequences and maximize economic benefits. Spatial-based decisions consider the state of charge (SOC), cycle life characteristics, and efficiency curves of each energy storage unit, balancing the utilization rate of each unit while ensuring system safety.
[0083] In actual control, the execution of control commands is monitored in real time. For devices with abnormal responses, such as inverters with communication interruptions or excessive control deviations, the control strategy is automatically adjusted, and their responsibilities are redistributed to other available devices. At the same time, a certain percentage (usually 10%) of control margin is maintained to cope with emergencies and prediction errors.
[0084] Step S4: Based on the system response data, the control strategy is optimized using the optimal decision tree algorithm, the reinforcement learning model is updated, and an optimization experience knowledge base is established.
[0085] In this embodiment, the optimal decision tree algorithm is an improved version of the traditional CART (Classification and Regression Tree) algorithm, specifically optimized for photovoltaic grid-connected control scenarios. The core of the algorithm is to simultaneously consider the complexity of the tree and the prediction accuracy during the decision tree construction process, seeking the optimal balance point.
[0086] In practice, (state, action) pairs are first extracted from the reinforcement learning policy as training samples, typically numbering 5000-10000 pairs. Then, the importance of each state feature is calculated using SHAP (SHapley Additive exPlanations) values, and the top 15-20 most important features are selected as the candidate set for splitting features in the decision tree. Feature importance calculation employs a permutation-based method, evaluating the influence of a feature by randomly shuffling individual features and observing changes in the model output.
[0087] The decision tree construction employs a top-down recursive splitting process, but unlike the traditional CART, the splitting criteria use a weighted Gini index or mean squared error, with weights related to the number of node samples and the tree depth. This weighting mechanism encourages the algorithm to perform more refined splits in shallower nodes with a larger number of samples, while maintaining a coarser decision boundary in deeper nodes, effectively controlling the risk of overfitting.
[0088] The tree growth process is subject to multiple constraints: the maximum depth is limited to 8 layers to ensure that the decision-making process is not too complex; the minimum number of leaf node samples is set to 1% of the total samples to avoid generating rules that only apply to a very small number of cases; and the information gain threshold is set to 0.01 to prevent invalid splits that have little to no improvement on the results.
[0089] After the tree is constructed, post-pruning is performed. Pruning is based on cross-validation error, using a cost complexity parameter α to balance the complexity and accuracy of the tree. The pruning process starts with a fully grown tree, progressively calculating the α value of each non-leaf node (the ratio of the increase in error after deleting the node's subtree to the subtree's complexity), selecting the node with the smallest α value for pruning, until the optimal balance point is reached.
[0090] In addition, branch merging techniques are employed to identify adjacent branches with similar decision results, simplifying the tree structure by adjusting the splitting threshold or directly merging nodes. The merging criteria are based on the similarity of node predictions and sample distribution characteristics, typically requiring that the difference in predictions between two candidate merged nodes be less than a preset threshold (e.g., 5% of the control amount) and that there be no significant difference in sample distribution.
[0091] The resulting decision tree is converted into an IF-THEN rule set for easy implementation and adjustment. For example, a typical rule might be 'IF bus voltage > 1.05 pu AND PV output > 75% THEN PV active power reduced to 80% of rated'. These rules are organized into a hierarchical structure, first evaluating rules related to grid security, then considering economic optimization rules to ensure that economic benefits are pursued while ensuring security.
[0092] In this embodiment, a complete closed loop for control effect evaluation and strategy optimization is also constructed to achieve continuous self-optimization of the system. The evaluation process is based on a multi-dimensional indicator system, including grid-friendliness indicators (such as voltage qualification rate, power backflow, grid loss rate, etc.), photovoltaic utilization indicators (such as curtailment rate, power generation efficiency, etc.), and economic indicators (such as power generation revenue, equipment losses, etc.). These indicators are weighted and summed to form a comprehensive score, with the weights dynamically adjusted according to the current operating objectives.
[0093] The control strategy evaluation adopts a comparative analysis method, comparing the actual control effect with the following three benchmark schemes: first, the benchmark scheme without any control measures, reflecting the absolute improvement brought about by control; second, the control scheme based on traditional deterministic rules, demonstrating the advantages of intelligent algorithms; and third, the control strategy of the previous version, showing the incremental benefits of the optimization process.
[0094] The evaluation results are directly fed back to the reinforcement learning model, updating its value function and policy function. The update process employs importance sampling, weighting samples based on their control effectiveness, with higher weights assigned to samples performing better, guiding the model to learn from successful experiences. Simultaneously, a priority experience replay buffer is maintained to preferentially store and replay samples that perform exceptionally well or poorly, accelerating the model's learning of key scenarios.
[0095] While updating the reinforcement learning model, the optimal decision tree is also adjusted. Adjustment methods include: recalculating feature importance, which may introduce new key features or eliminate features that are no longer important; relearning the tree structure based on new data, paying special attention to decision branches that previously performed poorly; and fine-tuning the decision threshold to better adapt it to the current system state and control objectives.
[0096] To avoid overfitting to recent data and forgetting previously learned useful experiences, a continuous learning mechanism based on knowledge distillation is implemented. Specifically, when updating the model, not only is the loss function for new data considered, but also a KL divergence term with the predictions of the old model is added to balance the learning of new knowledge and the retention of old knowledge.
[0097] Finally, successful control experiences are extracted into knowledge rules and stored in an optimization experience knowledge base. The knowledge base uses a graph database structure to represent the complex relationships between scenarios, control strategies, and effects. Each knowledge rule includes: a description of the applicable scenario, control parameter configuration, expected effect, and confidence level. When formulating a new control strategy, the system first queries relevant experiences in the knowledge base as a decision-making reference.
[0098] The retrieval of the experience knowledge base is also based on the BWT-Runs indexing technology, but it extends the semantic similarity measurement, considering not only the similarity of numerical features, but also the semantic equivalence of the scene. For example, although the specific parameters of 'high light intensity at midday on summer weekdays' and 'strong radiation at midday on weekends and holidays' are different, they may require similar processing methods from the perspective of control strategies.
[0099] Through this closed-loop mechanism of continuous evaluation, optimization, and accumulation, the control effect can be continuously improved, adapting to the dynamic changes of the distribution network and photovoltaic clusters, and achieving long-term, stable, efficient, and flexible grid-connected regulation.
[0100] Specifically, in step S1, status data including voltage, current, and power parameters of distribution network nodes and power generation and inverter operating parameters of photovoltaic clusters are collected. The status data undergoes data cleaning and timestamp alignment preprocessing to obtain a comprehensive system status data stream, including:
[0101] The state data is cleaned by detecting and processing missing and outlier values based on the 3σ criterion to obtain cleaned state data.
[0102] The cleaned state data is resampled according to a unified second-level time base to complete the timestamp alignment and obtain the time-aligned state data.
[0103] The time-aligned state data is structured using BWT-Runs spatial indexing technology, including discretizing the time-series data, applying BWT transformation to generate a sorted suffix array, identifying repetitive patterns and performing Run-length encoding, thereby achieving the fusion of multi-source heterogeneous data and obtaining the comprehensive system state data stream.
[0104] In practice, the detailed steps of the BWT-Runs spatial indexing technology are as follows: First, the time-series data is discretized, mapping continuous values to a finite character set. This embodiment uses the Adaptive Piecewise Constant Approximation (APCA) method to automatically determine the number of segments and segment boundaries based on the data distribution characteristics. Typically, continuous data such as photovoltaic power and voltage are divided into 8-12 discrete levels, with each level represented by a single character.
[0105] Then, the BWT transform is applied to the discretized sequence. The core of the BWT transform is to construct a cyclic shift matrix, sort the matrix lexicographically, and then extract the last column. In the specific implementation, it is not necessary to explicitly construct the complete matrix; instead, the BWT result is directly calculated using suffix array technology, which significantly reduces computational complexity. This system uses the SA-IS algorithm with induced sorting to construct the suffix array, which has near-linear time complexity when processing long sequences.
[0106] Next, repetitive patterns in the BWT transform results are identified, and run-length encoding is applied for compression. The system employs a two-level encoding strategy: first, identical consecutive characters are identified (first-level run), and then sequence segments with similar patterns are identified (second-level run). The encoding process uses adaptive dictionary technology, constructing dedicated encoding dictionaries for different types of data (such as voltage, power, etc.) to further improve compression efficiency.
[0107] Finally, a multi-level index structure is constructed based on the encoding results. The system uses FM-Index (Full-text index in Minute space) as the basic index structure and extends it to support approximate matching queries. The index structure contains three key components: BWT conversion results, a cumulative frequency table, and a checkpoint location table. To support efficient queries, the system sets a checkpoint every 256 characters, storing complete index information for that position.
[0108] This index structure is particularly suitable for pattern retrieval of time-series data in power systems. For example, when searching for historical scenarios such as 'voltage exceeding 1.05 pu and photovoltaic output greater than 80% for three consecutive days', the system converts the query pattern into a discrete sequence and uses the FM-Index to locate all matching positions within logarithmic time complexity. For approximate queries (allowing for a small number of mismatches), the system employs a backtracking search algorithm to return the best matching result within an acceptable time.
[0109] In flexible grid-connected control systems for photovoltaic clusters, multi-source heterogeneous data preprocessing and fusion are fundamental to achieving precise control. This step receives raw datasets from sensing devices at key grid nodes and the photovoltaic cluster monitoring system, and transforms them into a high-quality integrated system status data stream through a series of technical means.
[0110] The raw data is first cleaned. This includes detecting and handling missing values, and using time series interpolation algorithms (such as linear interpolation, spline interpolation, or imputation based on historical similarity patterns) to complete the data. Simultaneously, statistical methods (such as the 3σ criterion, box plots, or the density-based DBSCAN algorithm) are used to identify outliers, and these outliers are repaired or marked based on the detection results. For example, for unreasonable jumps in photovoltaic power data (such as a sudden drop in power from 80% to 0% and then recovering within a short period), the system will automatically identify and smooth the data based on the trend before and after the jumps.
[0111] Next, the issue of inconsistent timestamps from different data sources is addressed. Since the sampling frequencies and times of grid sensing devices and photovoltaic monitoring systems may differ, time alignment is necessary. Here, a unified time base and resampling technique are used to adjust all data to the same time scale (e.g., seconds or minutes), ensuring data consistency across the time dimension and providing a foundation for subsequent analysis.
[0112] The core innovation of this step lies in introducing BWT-Runs spatial indexing technology to organize data in a structured manner. Burrows-WheelerTransform (BWT) is a string transformation algorithm; combined with run-length encoding, the resulting BWT-Runs structure can efficiently compress and index sequential data. In this system, time-series data is treated as a special string sequence, and an index is built using BWT-Runs technology to achieve compact data storage and fast retrieval. The specific implementation is as follows:
[0113] Appropriate discretization processing is performed on various types of time-series data to map continuous values to a finite character set;
[0114] Apply the BWT transform to the discretized data sequence to generate a sorted suffix array;
[0115] Identify repetition patterns in the transformed sequence and apply Run-length encoding for compressed storage;
[0116] Build a multi-level index structure to support quick location and retrieval of similar data patterns.
[0117] The advantage of this index structure lies in its ability to quickly retrieve historical scenarios similar to the current state and to support long-pattern exact match queries (LEM queries), providing efficient support for subsequent pattern recognition and decision optimization. For example, when the system needs to find historical operating data under combined conditions such as "sunny day, high load, voltage fluctuation," traditional methods require traversing the entire database, while the BWT-Runs index can locate the target data segment in sublinear time.
[0118] Finally, different types of data are merged and integrated to form a unified system state representation. The fusion process is based on the physical meaning and spatiotemporal correlation of the data, and adopts a combination of feature-level fusion and decision-level fusion. For example, data such as voltage, power, and photovoltaic output on the same distribution feeder are correlated to construct a feeder state vector; or data from multiple photovoltaic sites in geographically close proximity are merged to form regional photovoltaic cluster characteristics.
[0119] Through the above processing, the original multi-source heterogeneous data is transformed into a high-quality integrated system state data stream, providing a reliable data foundation for subsequent multi-objective optimization models. This data stream not only includes the current state of the power grid and photovoltaic system, but also contains hidden spatiotemporal patterns and dynamic characteristics, enabling more accurate system state perception and prediction.
[0120] In step S1, after obtaining the system state integrated data stream, the following steps are also included:
[0121] Based on the system status integrated data stream, the key performance indicators of the power grid, such as node voltage deviation, line load rate, and power backflow, as well as photovoltaic power generation indicators such as photovoltaic absorption rate and power generation efficiency, are calculated to obtain a real-time system status assessment report.
[0122] Based on the real-time status assessment report of the system, a power flow equation for the distribution network is established, which includes voltage constraints (node voltage within ±7% of the rated value) and capacity constraints (line power not exceeding the rated capacity).
[0123] In step S2, a multi-objective optimization model incorporating grid friendliness and photovoltaic power generation benefits is constructed, including:
[0124] A grid-friendly objective function is constructed, which includes sub-objectives of minimizing voltage deviation, minimizing reverse power transmission, and maximizing photovoltaic absorption.
[0125] A photovoltaic power generation benefit objective function is constructed, which includes sub-objectives of maximizing power generation revenue and minimizing equipment losses;
[0126] The grid-friendliness objective function and the photovoltaic power generation benefit objective function are integrated into a multi-objective optimization model using a weighted method. The weight coefficients of each objective function are dynamically adjusted according to the grid operation status of the integrated system state data stream. When a voltage over-limit risk is detected, the weight of the grid-friendliness objective function is increased.
[0127] This step is fundamental to the construction of the multi-objective optimization model, aiming to establish a mathematical model of the distribution network and construct a grid-friendly objective function. Taking the system's real-time state assessment report as input, this step outputs a grid-friendly optimization sub-model through grid constraint modeling and grid-friendly objective function construction.
[0128] First, a network model is established based on the distribution network topology. The node admittance matrix method is used to represent the distribution network as a network diagram composed of nodes and branches. Nodes correspond to key locations such as transformers and tap points, while branches correspond to lines. For a distribution network with n nodes, its node admittance matrix Y is an n×n complex matrix, where the off-diagonal element Y[i,j] represents the admittance between nodes i and j, and the diagonal element Y[i,i] represents the sum of the admittances of node i and all connected nodes. This matrix can be used to establish the relationship between node voltage and injected power.
[0129] Secondly, the power flow equations of the distribution network are constructed based on the admittance matrix. The power flow equations in the distribution network describe the mathematical relationship between node power and voltage. For node i, its complex power injection Si can be expressed as:
[0130] S i =P i +jQ i=V i ×∑(V j ×Y [i,j] )*
[0131] Where P i and Q i These represent active and reactive power injections, respectively, V i and V j Y is the complex representation of the node voltage. [i,j] For each element of the admittance matrix, * denotes complex conjugate. Due to the radial structure and high R / X ratio of distribution networks, the traditional Newton-Raphson method may have convergence problems. Therefore, this system uses the forward / backward scanning method or an improved Newton's method to solve the power flow equations, which is more suitable for the characteristics of distribution networks containing distributed photovoltaic power.
[0132] After establishing the basic power flow model, the system sets various operational constraints, mainly including:
[0133] Node voltage constraints: All node voltage amplitudes must be kept within the allowable range, i.e., Vmin≤|Vi|≤Vmax, typically ±7% or ±5% of the rated voltage;
[0134] Line capacity constraint: The apparent power of all lines shall not exceed their rated capacity, i.e., |Sij|≤Sijmax;
[0135] Transformer capacity constraint: The power passing through the transformer shall not exceed its rated capacity;
[0136] Power balance constraint: The balance between the total power generation of the system and the total load and losses.
[0137] Then, a grid-friendly objective function is constructed, which includes three key sub-objectives:
[0138] The first sub-objective is to minimize voltage deviation. This objective aims to make all node voltages as close as possible to their ideal values (typically nominal values). It is defined as:
[0139] f V =min(∑ω i (|V i |-V ref ) 2 )
[0140] Where ω i This is the node weight coefficient, which can be set according to the importance of the node, |V i | represents the voltage magnitude at node i, V ref This is the reference voltage.
[0141] The second sub-objective is to minimize reverse power transmission. Large-scale photovoltaic (PV) grid connection can lead to power being transmitted from the low-voltage side to the high-voltage side, increasing network losses and causing protection coordination issues. This objective is defined as:
[0142] f R =min(∑max(0,-P) ij ))
[0143] Where P ij Let represent the active power from node i to j, with negative values indicating reverse power flow. This function only penalizes reverse power, while forward power remains unaffected.
[0144] The third sub-objective is to maximize photovoltaic (PV) power consumption. Under the premise of ensuring grid security, the amount of PV power generation consumed should be maximized. This is defined as:
[0145]
[0146] in To contribute to actual photovoltaic power, To maximize possible output (based on current lighting conditions).
[0147] Finally, these three sub-objectives are integrated into a comprehensive objective function for grid friendliness:
[0148] f G =α·f V +β·f R -γ·f C
[0149] α, β, and γ are weighting coefficients used to balance the importance of different sub-objectives and can be dynamically adjusted according to the grid operation status. For example, the value of α is increased when the voltage is close to exceeding the limit, and the value of β is increased during peak photovoltaic output.
[0150] This step takes the system's real-time status assessment report as input, and outputs a photovoltaic power generation benefit optimization sub-model through photovoltaic power generation benefit modeling and objective function construction. The core of this step is to construct a quantitative objective function for economic benefits and equipment efficiency from the perspective of photovoltaic operators.
[0151] First, we construct a photovoltaic (PV) power generation revenue model. Under the current electricity pricing mechanism, PV power generation revenue mainly comes from the sale of grid-connected electricity and renewable energy subsidies (if applicable). The power generation revenue function can be expressed as:
[0152] R = ∑ t [P t ·E t ·(π t +s t )]
[0153] Where R is the total revenue, P tE represents the actual photovoltaic output power at time t. t π represents the duration of that time period (e.g., 15 minutes). t Let s be the electricity price at time t. t To subsidize the unit price.
[0154] Electricity price π t The pricing varies depending on the region and time of day, and may include fixed tariffs, time-of-use tariffs, or market-based tariffs. For large-scale photovoltaic power plants participating in the electricity market, the price may also be affected by day-ahead reporting and real-time deviations. The system will automatically select the applicable tariff model based on regional policies and market rules. For example, in regions implementing time-of-use tariffs, π t It will automatically adjust according to peak, flat, and off-peak periods; when participating in the spot market, π t It will change dynamically based on the market clearing results.
[0155] Secondly, a model is established to influence equipment losses and lifespan. Frequent adjustments to the operating point of a photovoltaic inverter may increase losses and affect equipment lifespan. The inverter's loss model can be simplified as follows:
[0156] Loss inv =∑ t [P t ·(1-η t )+S inv ·δ t ]
[0157] Where, η t Let S be the inverter's operating efficiency at time t (related to load rate). inv δ is the rated capacity of the inverter. t This is the additional loss coefficient caused by adjusting the working point per unit time.
[0158] For energy storage systems, the effects of losses and lifetime are more complex, involving charge / discharge efficiency, cycle count, and depth. The energy storage loss model can be expressed as:
[0159] Loss bat =∑ t [|P bat,t |·(1-η bat )+λ·|DOD t |]
[0160] Among them, P bat,t Let η be the energy storage charging and discharging power at time t. bat For charge and discharge efficiency, DOD t λ represents the depth of discharge, and λ is the cycle life influence coefficient.
[0161] Taking all the above factors into account, the objective function for photovoltaic power generation benefits is constructed as follows:
[0162] f E =max(RC ossloss )
[0163] Among them, C ossloss The equivalent cost of equipment wear and tear and its lifespan can be calculated by converting physical wear and tear into economic value.
[0164] This objective function also needs to consider various constraints, including:
[0165] Photovoltaic inverter capacity constraint: 0≤P t ≤P max,t , where P max,t The maximum possible output at time t (related to lighting conditions);
[0166] Power factor or reactive power range constraints: Typically, a power factor of 0.95 or higher, or a Q / P ratio within a specific range is required.
[0167] Energy storage system SOC constraint: SOC min ≤SOC t ≤SOC max This ensures that the energy storage system operates within a safe range.
[0168] Energy storage system charge / discharge power constraint: P bat,min ≤P bat,t ≤P bat,max .
[0169] The key feature of this objective function is the introduction of a dynamic efficiency model for the equipment, rather than simply assuming a fixed efficiency, which better reflects actual operating conditions. For example, the efficiency curve of an inverter at different load rates exhibits an inverted U-shape, typically reaching its highest efficiency at 70%-90% load rates, while efficiency decreases at extremely low or high load rates.
[0170] Furthermore, the objective function also considers the impact of the control frequency on the equipment. Frequent adjustments to photovoltaic output or energy storage charge / discharge states not only increase control costs but may also accelerate equipment aging. This is addressed by introducing δ... t The λ coefficient is used to quantify and optimize the "intensity" of control behavior.
[0171] Ultimately, the output photovoltaic power generation benefit optimization sub-model is a mathematical model that comprehensively considers maximizing revenue and minimizing equipment losses. It can accurately represent the economic demands of photovoltaic operators and provide a necessary component for subsequent multi-objective optimization.
[0172] This step integrates the grid-friendly optimization sub-model and the photovoltaic power generation benefit optimization sub-model into a unified multi-objective optimization overall model. The step receives the grid-friendly optimization sub-model and the photovoltaic power generation benefit optimization sub-model as input, and through the application of multi-objective optimization methods and the comprehensive setting of constraints, outputs a complete multi-objective optimization overall model.
[0173] First, it's necessary to resolve potential conflicting objectives between the two sub-models. For example, grid friendliness might require reducing photovoltaic output to avoid voltage overshoot, while power generation efficiency tends to maximize output. To resolve this conflict, the system employs one of the three main methods in multi-objective optimization theory:
[0174] Weighted Sum Method: This method linearly combines multiple objective functions into a single objective function.
[0175] F = w1·f G +w2·f E
[0176] Where f G Let f be the objective function for grid friendliness. E Let w1 and w2 be the objective function for photovoltaic power generation benefits, and w1 + w2 = 1. The weights can be dynamically adjusted based on grid operating conditions and policy requirements, such as increasing w1 when the grid is under high load and increasing w2 when photovoltaic revenue is prioritized.
[0177] ε-constraint method: Select one objective function as the primary optimization objective, and transform other objective functions into constraints.
[0178] Maximize f E ;
[0179] Constraints: f G ≥ε;
[0180] Where ε is the minimum threshold required for grid friendliness. This method is particularly suitable for scenarios that maximize photovoltaic benefits while ensuring grid security.
[0181] Pareto optimality method: Directly calculates the Pareto front, i.e., there is no solution set that can improve one objective without harming others, and then selects the most suitable solution from it. The system uses algorithms such as NSGA-II (Non-dominated sorting genetic algorithm II) or MOEA / D (decomposition-based multi-objective evolutionary algorithm) to calculate the Pareto front, and then selects the final solution according to the decision-maker's preferences or specific rules.
[0182] The above methods should be flexibly selected based on actual operational needs. In emergency situations, the system tends to use the more computationally efficient weighting method; in regular operation and offline planning, the Pareto optimal method is more often used to obtain a more comprehensive solution set.
[0183] Next, we will set comprehensive operational constraints, which include, but are not limited to:
[0184] Physical constraints include grid constraints (such as node voltage constraints and line capacity constraints) and equipment constraints (such as inverter capacity constraints and energy storage SOC constraints).
[0185] Operational constraints: such as photovoltaic output change rate constraints (usually limited to 10% of rated capacity / minute), energy storage charge and discharge cycle constraints, etc.
[0186] Control constraints: Practical control constraints that take into account communication delays and control accuracy, such as minimum control step size and control signal update period;
[0187] Safety constraints: To ensure system robustness, safety margin constraints are introduced, such as reserving a certain adjustment capacity to cope with emergencies;
[0188] These constraints are mathematically represented and incorporated into the optimization model to ensure that the solution is feasible in the actual system.
[0189] Finally, the set of decision variables is constructed, mainly including:
[0190] The active power output setpoint P of the photovoltaic inverter pv,i (i = 1, 2, ..., n);
[0191] The reactive power setpoint Q of the photovoltaic inverter pv,i or power factor setpoint PF i ;
[0192] The charging and discharging power P of the energy storage system batt,j (j = 1, 2, ..., m).
[0193] Through the above steps, the two sub-models of grid friendliness and photovoltaic power generation benefits were successfully integrated, and a complete multi-objective optimization overall model was constructed:
[0194] Optimization objective: The comprehensive objective function determined based on the selected multi-objective optimization method;
[0195] Constraints: Includes all physical, operational, control, and safety constraints;
[0196] Decision variables: Control parameters of all controllable resources.
[0197] This multi-objective optimization model can balance the safe and stable operation of the power grid and the economic benefits of photovoltaic power generation. It is the mathematical basis for the flexible grid-connected regulation of photovoltaic clusters and provides a clear problem definition for subsequent optimization algorithms.
[0198] In step S2, the multi-objective optimization model is solved using a reinforcement learning algorithm, including:
[0199] The multi-objective optimization problem is transformed into a Markov decision process, and the integrated data flow of the system state is mapped to the state space, and the control parameters of the photovoltaic inverter and energy storage unit are mapped to the action space.
[0200] A reward function is constructed based on the objective function of the multi-objective optimization model. A deep reinforcement learning network containing a policy network and a value network is used for policy training to obtain the control policy. The policy network and the value network both adopt a multilayer perceptron structure.
[0201] The control strategy was simulated and verified to evaluate its voltage control accuracy, power fluctuation suppression effect, and economic performance.
[0202] In this embodiment, both the policy network and the value network employ a multilayer perceptron structure. Specifically, the policy network contains three hidden layers with 256, 128, and 64 neurons in each layer, respectively, and uses the ReLU activation function. The output layer uses the tanh activation function to generate normalized action values in the range [-1, 1]. The value network also contains three hidden layers, with the same structure as the policy network, but its output layer is a single neuron used to estimate state values.
[0203] The input layers of both networks receive a system state vector containing normalized grid voltage data (10-20 key nodes), photovoltaic output data (real-time power and adjustable range for each site), energy storage status data (SOC level and available capacity), and environmental data (illuminance, temperature, etc.). To improve training efficiency, principal component analysis is used to compress the original high-dimensional state vector (potentially containing hundreds of measurement points) to within 50 dimensions.
[0204] Model training employs a heterogeneous policy approach, using an experience replay buffer to store interaction samples with a capacity of 100,000 records. During training, a batch of 64 samples is randomly selected for each iteration to update the network. The policy network is updated using the policy gradient method with a learning rate of 0.0003; the value network uses the mean squared error loss function with a learning rate of 0.001. A target network mechanism is used during training, updating the target network every 100 steps to improve training stability.
[0205] The training data comes from two sources: first, an offline dataset built based on historical operational data, containing records of system status, control actions, and system responses under different scenarios; and second, interactive data generated in real time within the digital twin environment of the power distribution network. Initially, offline data is the primary source, with the proportion of real-time interactive data gradually increasing as training progresses. The complete training process involves approximately 500,000 interactions, requiring 8-12 hours on a typical computing platform.
[0206] To handle environmental uncertainties, domain randomization is introduced during model training to enhance the model's generalization ability by randomly adjusting environmental parameters (such as load fluctuations and lighting changes). Simultaneously, to ensure control safety, a safety constraint layer is implemented during training to automatically filter out control actions that might cause the system to exceed its limits.
[0207] This step designs a dynamic solution algorithm based on reinforcement learning, transforming the multi-objective optimization problem into a Markov Decision Process (MDP), and learning the optimal control strategy through interaction with the environment. The step receives the overall multi-objective optimization model and a set of historical experience features as input, and outputs the optimal control strategy for the photovoltaic cluster through MDP modeling, reinforcement learning algorithm design and training.
[0208] First, the multi-objective optimization problem is transformed into a Markov decision process (MDP). An MDP is a quintuple (S, A, P, R, γ), where S is the state space, A is the action space, P is the state transition probability, R is the reward function, and γ is the discount factor. In the photovoltaic cluster regulation problem, these elements are defined as follows:
[0209] State space S: Contains all relevant parameters describing the current state of the system, specifically including:
[0210] Power grid status parameters: voltage at each node, line power flow, transformer load, etc.
[0211] Photovoltaic system parameters: real-time power generation, controllable range, inverter status, etc. at each site;
[0212] Energy storage system parameters: SOC level, chargeable / dischargeable power, etc.;
[0213] Environmental parameters: light intensity, temperature, weather forecast, etc.;
[0214] Load parameters: power demand and predicted trends for various types of loads;
[0215] The state space has a high dimensionality. To improve learning efficiency, the system uses dimensionality reduction techniques such as principal component analysis (PCA) or autoencoders to compress the state representation.
[0216] Action space A: Contains control parameters for all controllable resources, specifically including:
[0217] Photovoltaic inverter active power regulation command: can be expressed as a percentage relative to the maximum possible output;
[0218] Reactive power or power factor setting value of photovoltaic inverter;
[0219] Energy storage system charging and discharging power commands;
[0220] To avoid the curse of dimensionality, the system discretizes the action space appropriately, for example, discretizing the continuous power adjustment range into 10 levels.
[0221] State transition probability P: describes the probability of transitioning to state s' after taking action a in the current state s, P(s'|s,a). Due to the complexity of distribution network dynamics, which is difficult to model accurately, the system adopts a sample-based learning method instead of explicitly defining the state transition function.
[0222] Reward function R: Constructed based on the objective function of the multi-objective optimization overall model, it needs to balance grid friendliness and photovoltaic power generation benefits.
[0223] R(s,a,s')=w1·R G (s,a,s')+w2·R E (s,a,s')-w3·R penalty (s')
[0224] Where R G RE is a grid-friendly incentive (such as improved voltage compliance rate, reduced backfeed power, etc.), and R is a power generation benefit incentive. penalty For the penalty items for violating constraints (such as voltage over-limit penalty), w1, w2, and w3 are weighting coefficients.
[0225] Discount factor γ: This balances the importance of immediate rewards versus future rewards and is typically set between 0.9 and 0.99. A higher γ value makes the algorithm focus more on long-term gains, but may increase the difficulty of training.
[0226] Based on the aforementioned MDP model, a reinforcement learning algorithm suitable for photovoltaic cluster regulation is designed. Considering the high dimensionality and continuity of the state space, as well as the mixed characteristics (continuous and discrete) of the action space, a deep reinforcement learning algorithm is selected, specifically including:
[0227] Value function-based methods, such as Deep Q-Network (DQN) and its variants (Double DQN, Dueling DQN, etc.), are suitable for discretized action spaces. DQN uses a neural network to approximate the Q-value function Q(s,a) and improves learning stability through temporal difference learning and experience replay mechanisms.
[0228] Policy gradient-based methods, such as Deep Deterministic Policy Gradient (DDPG) and Proximal Policy Optimization (PPO), are suitable for continuous action spaces. These algorithms directly learn the policy function π(a|s) and optimize the expected reward through gradient ascent.
[0229] Hybrid architecture: For photovoltaic clusters, which involve multiple types of control parameters, the system adopts a hybrid architecture, such as using DDPG for continuous parameters (power regulation) and DQN for discrete parameters (operation mode selection).
[0230] The reinforcement learning training process includes:
[0231] Offline pre-training: Utilizes historical data to construct the initial policy network and value network, avoiding inefficient exploration caused by random initialization;
[0232] Simulation Environment Training: A high-fidelity simulation environment is constructed based on the digital twin system of the power distribution network, and the agent improves its strategy through trial and error in this environment;
[0233] Exploration-Exploitation Balance: Employ methods such as ε-greedy strategy or noise injection to balance exploration and exploitation. Initially, a larger exploration rate is used to discover potential good strategies, and later the exploration rate is reduced for fine optimization.
[0234] Layered training: First train the control strategy of a single photovoltaic site, then train the coordination strategy between sites to reduce the learning difficulty;
[0235] Transfer learning: Using knowledge learned in a specific scenario to accelerate learning in other scenarios and improve training efficiency;
[0236] Safety constraints: A safety layer is introduced during training to ensure that critical constraints are not violated during exploration.
[0237] In the algorithm implementation, deep learning frameworks such as TensorFlow or PyTorch are used to build neural network models, and distributed training frameworks such as Ray RLlib are used to accelerate the training process. The network architecture is based on a fully connected network, and depending on the data characteristics, LSTM layers may be introduced to capture temporal dependencies, or graph neural network layers may be used to model the power grid topology.
[0238] Finally, the fully trained reinforcement learning model outputs a policy function π(a|s), which represents the mapping relationship between the current system state s and the optimal control action a. This policy function is stored in the form of a neural network, enabling efficient inference calculations and achieving millisecond-level control decisions to meet real-time control requirements.
[0239] Compared with traditional optimization algorithms, the advantages of reinforcement learning-based dynamic solution methods are: (1) they do not rely on precise system models and adapt to complex environments through interactive learning; (2) they can handle high-dimensional nonlinear decision problems; (3) they have fast computation and inference speeds, making them suitable for real-time control; and (4) they possess adaptive capabilities and can continuously improve strategies. These characteristics make them particularly suitable for complex, dynamic, and uncertain control scenarios such as photovoltaic clusters.
[0240] This step uses a digital twin system of the distribution network to simulate and verify the optimal control strategy output by the reinforcement learning algorithm, evaluating its effectiveness and stability. Necessary adjustments are made using a simple approximation algorithm based on the optimal decision tree, ultimately outputting a verified and reliable control strategy scheme. The step receives the optimal control strategy output by the reinforcement learning algorithm as input, and through simulation verification, strategy evaluation, and adjustment processes, outputs a verified and reliable control strategy scheme.
[0241] First, a high-fidelity digital twin simulation platform for the distribution network is constructed. This platform is a virtual mapping of the physical distribution network and can accurately simulate the dynamic characteristics of the real network. The platform mainly includes the following modules:
[0242] Network Model Module: An electrical model built based on the actual distribution network topology and parameters, supporting power flow calculation and transient simulation, and accurately reflecting the steady-state and dynamic response characteristics of the system;
[0243] Photovoltaic Model Module: Contains detailed models of photovoltaic panels, MPPT controllers, and inverters, capable of simulating the power generation characteristics and dynamic response of photovoltaic systems under different light and temperature conditions;
[0244] Energy storage model module: includes battery characteristic model (state of charge, cycle life, etc.) and energy storage converter model, which can simulate the charging and discharging process and electrical characteristics of energy storage system;
[0245] Load Model Module: Capable of simulating the characteristics and dynamic behavior of different types of loads (such as constant power, constant impedance, constant current and motor loads);
[0246] Control and Communication Module: Simulates the actual communication characteristics such as transmission delay and data packet loss of control commands, as well as the response characteristics of the controller.
[0247] The digital twin platform employs real-time simulation technologies, such as RTDS (Real Time Digital Simulator) or OPAL-RT, to ensure a high degree of consistency between the simulation process and the actual system behavior. The platform supports simulations at multiple time scales, from millisecond-level electrical transients to hourly-level power balance, comprehensively evaluating the performance of control strategies.
[0248] Next, a comprehensive set of verification scenarios will be designed to ensure the effectiveness of the control strategy under various operating conditions. Verification scenarios include:
[0249] Normal operating scenarios: such as system response under typical daily load curves and typical solar illumination curves;
[0250] Extreme operating scenarios: such as combinations of maximum sunlight and minimum load, rapidly changing weather conditions, etc.
[0251] Fault scenarios: such as emergency situations like line short circuits or equipment tripping;
[0252] Communication interruption scenario: Simulate system behavior when partial or complete communication is interrupted;
[0253] Parameter uncertainty scenario: Consider the impact of model parameter error and measurement noise.
[0254] The system executes control strategies in various verification scenarios and collects key performance indicators, including:
[0255] Power grid safety indicators include: node voltage qualification rate, line load rate, and power backflow level.
[0256] Photovoltaic grid integration indicators: photovoltaic curtailment rate, grid integration capacity, etc.;
[0257] Economic benefit indicators: power generation revenue, equipment wear and tear, etc.
[0258] Control performance indicators: response time, settling time, overshoot, etc.
[0259] Robustness metrics: the ability to adapt to disturbances and uncertainties.
[0260] Based on the verification results, the advantages and disadvantages of the control strategy are evaluated. If the strategy is found to perform poorly in certain scenarios (such as overly aggressive response or sensitivity to certain types of disturbances), the strategy adjustment phase begins.
[0261] The core innovation in the strategy adjustment phase lies in applying a simple approximation algorithm for the optimal decision tree. A decision tree is an intuitive decision model that represents the decision-making process in a tree structure, where internal nodes represent feature tests and leaf nodes represent decision outcomes. The optimal decision tree is the tree structure that minimizes a predefined loss function among all possible decision trees.
[0262] In photovoltaic cluster control, the policy generated by reinforcement learning can be viewed as a complex black-box function, difficult to interpret and fine-tune. A simple approximation algorithm using optimal decision trees can approximate this complex policy into a more concise and interpretable decision rule. The specific steps include:
[0263] Data generation: Run reinforcement learning policies in a simulation environment to collect (state, action) pairs of samples;
[0264] Feature importance analysis: Using techniques such as SHAP (SHapley Additive exPlanations) values to analyze which state features have the greatest impact on decision-making;
[0265] Decision tree training: Based on important features and action samples, train a decision tree model to approximate the original policy function;
[0266] Tree structure optimization: The decision tree structure is optimized through operations such as pruning and merging to simplify the model while maintaining performance;
[0267] Rule extraction and adjustment: Extract intuitive IF-THEN rules from the decision tree and manually adjust the issue rules based on the verification results.
[0268] For example, a simple rule might be extracted from a complex strategy: "IF bus voltage > 1.05 pu AND PV output > 80% THEN PV active power reduced by 20%". Experts can evaluate and fine-tune these rules, such as adjusting the reduction ratio from 20% to 15% to balance control strength and system response.
[0269] The adjusted policy rules are then validated again through the digital twin platform to ensure that the modified performance meets expectations. This validation-adjustment process may require multiple iterations until the policy performance meets all key performance indicators.
[0270] Finally, the system outputs a verified and reliable control strategy, including:
[0271] Strategy execution logic: can be represented as a decision tree or a set of rules;
[0272] Parameter settings: such as gain, threshold, time constant, etc. for each control element;
[0273] Execution conditions: The system state range and preconditions to which the strategy applies;
[0274] Expected results: Prediction of control effects in typical scenarios;
[0275] Security boundary: The boundary conditions for the effectiveness of the strategy; if this range is exceeded, an emergency plan must be activated.
[0276] This hybrid approach of "reinforcement learning + decision tree approximation" combines the advantages of both: reinforcement learning provides an initial policy close to the global optimum, while decision tree approximation provides interpretability and adjustability, together ensuring the high performance and reliability of the control policy.
[0277] In step S3, according to the control strategy, the active power output and reactive power output of the photovoltaic inverter are coordinated and controlled, including:
[0278] The active power setpoint of each photovoltaic inverter is generated based on the control strategy.
[0279] The active power setpoint is sent to the control system of each photovoltaic inverter via a communication network;
[0280] Adjust the maximum power point tracking control strategy of each photovoltaic inverter to achieve active power control;
[0281] The system monitors the execution of active power control commands in real time. When a communication interruption or equipment failure is detected, it initiates a local control mode based on local measurement data to adjust the active power to a preset safe operating point.
[0282] In step S3, according to the control strategy, the active power output and reactive power output of the photovoltaic inverter are coordinated and controlled, including:
[0283] Determine the global reactive power optimization target based on the control strategy described above;
[0284] Based on the local voltage measurement results of each photovoltaic inverter, a reactive power control command or power factor setting value is generated.
[0285] Establish a hierarchical coordinated reactive power control system from a single inverter to a photovoltaic cluster, including first-level single-unit reactive power regulation, second-level substation reactive power coordination, and third-level cluster reactive power optimization, to achieve dynamic support for grid voltage, wherein the voltage support target is to maintain the voltage of each node within ±5% of the rated value.
[0286] Real-time monitoring of the execution effect of reactive power control commands and dynamic adjustment of reactive power allocation strategies.
[0287] In step S3, according to the control strategy, the active power output and reactive power output of the photovoltaic inverter, as well as the charging and discharging power of the energy storage unit, are coordinated and controlled, including:
[0288] Based on grid load conditions, photovoltaic power generation forecasts, and electricity price signals, the timing of charging and discharging of energy storage units is determined;
[0289] The energy storage unit's charging and discharging power commands are generated according to the control strategy.
[0290] Coordinate and control the charging and discharging behavior of each energy storage unit to smooth out photovoltaic power output fluctuations. The goal of power output fluctuation smoothing is to control the power change rate within 10% / minute of the rated capacity.
[0291] Monitor the state of charge of the energy storage unit to ensure it is within the safe operating range of 20% to 80%.
[0292] In step S4, the control strategy is optimized using the optimal decision tree algorithm, including:
[0293] Collect system response data after control execution, including changes in grid operating parameters, photovoltaic output adjustment, and energy storage unit status changes;
[0294] Calculate quantitative indicators that include voltage deviation improvement rate, photovoltaic absorption rate improvement, power backflow reduction, and economic benefit increase.
[0295] The optimal decision tree algorithm is used to evaluate each decision point of the control strategy. An evaluation threshold is set based on the quantitative index to identify decision branches with poor performance that are below the threshold.
[0296] For decision branches with poor performance, pruning is performed by minimizing the depth of the decision tree, and branches are merged by merging similar decision rules to optimize the decision process while maintaining control accuracy.
[0297] like Figure 2 As shown, this embodiment of the invention also provides a flexible grid-connected control system for photovoltaic clusters, comprising:
[0298] Data acquisition module 201 is used to acquire status data including voltage, current, and power parameters of distribution network nodes and power generation and inverter operating parameters of photovoltaic clusters, and to perform data cleaning and timestamp alignment preprocessing on the status data to obtain a comprehensive system status data stream;
[0299] The optimization control module 202 is used to construct a multi-objective optimization model that includes grid friendliness and photovoltaic power generation benefits based on the system state integrated data stream, and to solve the multi-objective optimization model using a reinforcement learning algorithm to obtain the control strategy for the photovoltaic inverter and energy storage unit.
[0300] The execution control module 203 is used to coordinate and control the active power output and reactive power output of the photovoltaic inverter and the charging and discharging power of the energy storage unit according to the control strategy, so as to realize the flexible grid connection of the photovoltaic cluster and collect the system response data after the control is executed.
[0301] The evaluation and optimization module 204 is used to optimize the control strategy based on the system response data using the optimal decision tree algorithm, update the reinforcement learning model, and establish an optimization experience knowledge base.
[0302] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A flexible grid-connected control method for photovoltaic clusters, characterized in that, include: The system collects status data including voltage, current, and power parameters of distribution network nodes, power generation of photovoltaic clusters, and inverter operating parameters. The status data is then cleaned and preprocessed with timestamp alignment to obtain a comprehensive system status data stream. Based on the system state integrated data stream, a multi-objective optimization model that includes grid friendliness and photovoltaic power generation benefits is constructed. The multi-objective optimization model is solved using a reinforcement learning algorithm to obtain the control strategies for the photovoltaic inverter and energy storage unit. According to the control strategy, the active power output and reactive power output of the photovoltaic inverter and the charging and discharging power of the energy storage unit are coordinated and controlled to realize the flexible grid connection of the photovoltaic cluster and collect system response data after the control is executed. Based on the system response data, the control strategy is optimized using the optimal decision tree algorithm, the reinforcement learning model is updated, and an optimization experience knowledge base is established.
2. The method according to claim 1, characterized in that, The process of cleaning and timestamp-aligning the state data to obtain a comprehensive system state data stream includes: The state data is cleaned by detecting and processing missing and outlier values based on the 3σ criterion to obtain cleaned state data. The cleaned state data is resampled according to a unified second-level time base to complete the timestamp alignment and obtain the time-aligned state data. The time-aligned state data is structured using BWT-Runs spatial indexing technology, including discretizing the time-series data, applying BWT transformation to generate a sorted suffix array, identifying repetitive patterns and performing Run-length encoding, thereby achieving the fusion of multi-source heterogeneous data and obtaining the comprehensive system state data stream.
3. The method according to claim 1, characterized in that, After obtaining the system state integrated data stream, the process also includes: Based on the system status integrated data stream, the key performance indicators of the power grid, such as node voltage deviation, line load rate, and power backflow, as well as photovoltaic power generation indicators such as photovoltaic absorption rate and power generation efficiency, are calculated to obtain a real-time system status assessment report. Based on the real-time status assessment report of the system, a power flow equation for the distribution network is established, which includes voltage constraints (node voltage within ±7% of the rated value) and capacity constraints (line power not exceeding the rated capacity).
4. The method according to claim 1, characterized in that, The construction of a multi-objective optimization model that incorporates grid friendliness and photovoltaic power generation benefits includes: A grid-friendly objective function is constructed, which includes sub-objectives of minimizing voltage deviation, minimizing reverse power transmission, and maximizing photovoltaic absorption. A photovoltaic power generation benefit objective function is constructed, which includes sub-objectives of maximizing power generation revenue and minimizing equipment losses; The grid-friendliness objective function and the photovoltaic power generation benefit objective function are integrated into a multi-objective optimization model using a weighted method. The weight coefficients of each objective function are dynamically adjusted according to the grid operation status of the integrated system state data stream. When a voltage over-limit risk is detected, the weight of the grid-friendliness objective function is increased.
5. The method according to claim 1, characterized in that, Solving the multi-objective optimization model using reinforcement learning algorithms includes: The multi-objective optimization problem is transformed into a Markov decision process, and the integrated data flow of the system state is mapped to the state space, and the control parameters of the photovoltaic inverter and energy storage unit are mapped to the action space. A reward function is constructed based on the objective function of the multi-objective optimization model. A deep reinforcement learning network containing a policy network and a value network is used for policy training to obtain the control policy. The policy network and the value network both adopt a multilayer perceptron structure. The control strategy was simulated and verified to evaluate its voltage control accuracy, power fluctuation suppression effect, and economic performance.
6. The method according to claim 1, characterized in that, The step of coordinating and controlling the active power output and reactive power output of the photovoltaic inverter according to the control strategy includes: The active power setpoint of each photovoltaic inverter is generated based on the control strategy. The active power setpoint is sent to the control system of each photovoltaic inverter via a communication network; Adjust the maximum power point tracking control strategy of each photovoltaic inverter to achieve active power control; The system monitors the execution of active power control commands in real time. When a communication interruption or equipment failure is detected, it initiates a local control mode based on local measurement data to adjust the active power to a preset safe operating point.
7. The method according to claim 1, characterized in that, The step of coordinating and controlling the active power output and reactive power output of the photovoltaic inverter according to the control strategy includes: Determine the global reactive power optimization target based on the control strategy described above; Based on the local voltage measurement results of each photovoltaic inverter, a reactive power control command or power factor setting value is generated. Establish a hierarchical coordinated reactive power control system from a single inverter to a photovoltaic cluster, including first-level single-unit reactive power regulation, second-level substation reactive power coordination, and third-level cluster reactive power optimization, to achieve dynamic support for grid voltage, wherein the voltage support target is to maintain the voltage of each node within ±5% of the rated value. Real-time monitoring of the execution effect of reactive power control commands and dynamic adjustment of reactive power allocation strategies.
8. The method according to claim 1, characterized in that, The step of coordinating and controlling the active power output and reactive power output of the photovoltaic inverter, as well as the charging and discharging power of the energy storage unit, according to the control strategy, includes: Based on grid load conditions, photovoltaic power generation forecasts, and electricity price signals, the timing of charging and discharging of energy storage units is determined; The energy storage unit's charging and discharging power commands are generated according to the control strategy. Coordinate and control the charging and discharging behavior of each energy storage unit to smooth out photovoltaic power output fluctuations. The goal of power output fluctuation smoothing is to control the power change rate within 10% / minute of the rated capacity. Monitor the state of charge of the energy storage unit to ensure it is within the safe operating range of 20% to 80%.
9. The method according to claim 1, characterized in that, The optimization of the control strategy using the optimal decision tree algorithm includes: Collect system response data after control execution, including changes in grid operating parameters, photovoltaic output adjustment, and energy storage unit status changes; Calculate quantitative indicators that include voltage deviation improvement rate, photovoltaic absorption rate improvement, power backflow reduction, and economic benefit increase. The optimal decision tree algorithm is used to evaluate each decision point of the control strategy. An evaluation threshold is set based on the quantitative index to identify decision branches with poor performance that are below the threshold. For decision branches with poor performance, pruning is performed by minimizing the depth of the decision tree, and branches are merged by merging similar decision rules to optimize the decision process while maintaining control accuracy.
10. A flexible grid-connected control system for photovoltaic clusters, characterized in that, include: The data acquisition module is used to collect status data including voltage, current, and power parameters of distribution network nodes, power generation of photovoltaic clusters, and inverter operating parameters. The status data is cleaned and preprocessed with timestamp alignment to obtain a comprehensive system status data stream. The optimization control module is used to construct a multi-objective optimization model that includes grid friendliness and photovoltaic power generation benefits based on the comprehensive data flow of the system state, and to solve the multi-objective optimization model using a reinforcement learning algorithm to obtain the control strategy of the photovoltaic inverter and energy storage unit. The execution control module is used to coordinate and control the active power output and reactive power output of the photovoltaic inverter and the charging and discharging power of the energy storage unit according to the control strategy, so as to realize the flexible grid connection of the photovoltaic cluster and collect system response data after the control is executed. The evaluation and optimization module is used to optimize the control strategy based on the system response data using the optimal decision tree algorithm, update the reinforcement learning model, and establish an optimization experience knowledge base.
Citation Information
Cited By
Cooperative distributed photovoltaic and energy storage regulation and control strategy generation method and device
CN121546810A
Network construction type active support control method and system suitable for flexible direct current
CN121769881A
Cooperative control method, system and equipment for grid connection of multiple micro-grids and storage medium
CN122026388A