Optical storage charging and discharging station aggregation control and optimization method based on virtual power plant

Through adaptive clustering and dynamic resource aggregation of Markov decision-making processes, combined with multi-objective optimization and hierarchical control, the resource scheduling instability and computing complexity of optical storage charging stations in virtual power plants is solved, and efficient and stable power system scheduling is achieved.

CN120498043APending Publication Date: 2025-08-15NANJING INST OF MECHATRONIC TECH

Patent Information

Application Number
CN202510624437.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing virtual power plant control methods have challenges in resource aggregation, dynamic response and multi-objective optimization of optical storage charging stations. It is difficult to effectively coordinate the multi-party goals of photovoltaic power generation, energy storage systems and electric vehicle charging piles, resulting in unstable system scheduling and insufficient economicality, and high computational complexity, making it difficult to meet real-time control needs.

Method used

Adaptive clustering algorithm and Markov decision-making process are used to dynamic resource aggregation, combined with multi-objective optimization model and hierarchical control architecture, rolling time domain control and deep reinforcement learning are introduced, and scheduling strategies are optimized through edge-cloud collaborative computing to reduce communication load.

Benefits of technology

It realizes efficient coordinated scheduling of photovoltaic, energy storage and charging pile resources, improves the stability and economy of the system, enhances the adaptability to grid load and market changes, and reduces operating costs and calculation complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120498043A_ABST
    Figure CN120498043A_ABST
Patent Text Reader

Abstract

The invention provides an optical storage charging and discharging station aggregation control and optimization method based on a virtual power plant, and aims to solve the problems of multi-target collaborative optimization, dynamic resource response and uncertainty robustness. By introducing a Markov decision process and an adaptive clustering algorithm, the system can dynamically aggregate photovoltaic, energy storage and charging pile resources according to equipment characteristics, and power dispatching is optimized. A multi-objective optimization model is adopted, economical, technical and environmental objectives are combined, a dynamic weight factor is introduced, and optimal scheduling is generated in combination with a fuzzy decision theory. And real-time compensation is carried out by adopting a rolling time domain control framework and deep reinforcement learning, so that the scheduling precision and the response speed are improved. The edge computing and cloud collaboration mechanism reduces the communication load through a lightweight federated learning model, and improves the scheduling response efficiency. According to the invention, the scheduling efficiency of the optical storage charging station can be obviously improved, the operation cost is reduced, the system stability is improved, and the system has good adaptability and expandability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of power system automation, and in particular to a photovoltaic storage charging and discharging station aggregation control and optimization method based on a virtual power plant. Background Art

[0002] With the acceleration of the global energy transition and the widespread adoption of distributed energy resources such as photovoltaics, energy storage, and electric vehicle charging stations, traditional power systems are increasingly facing increasingly complex management and scheduling challenges. With the increasing penetration of renewable energy, the coordinated control of distributed energy resources has become a major challenge for energy management systems. Traditional independent scheduling strategies struggle to fully leverage the flexibility of distributed resources. In the context of photovoltaic-storage-charging systems, virtual power plants (VPPs), as an emerging power system management technology, integrate distributed energy resources (such as photovoltaics, energy storage, and electric vehicle charging stations) to achieve resource aggregation, control, and optimization. However, existing VPP control methods still face numerous challenges in resource aggregation, dynamic response, and multi-objective optimization of photovoltaic-storage-charging stations. Existing methods typically focus on a single objective (such as economic optimization or stability optimization) and fail to consider the synergistic effects of multiple devices, such as photovoltaic power generation, energy storage systems, and electric vehicle charging stations, under different operating modes. Photovoltaic resources are significantly affected by weather, while energy storage systems are constrained by factors such as charging efficiency, discharge depth, and lifespan during charging and discharging. Charging station loads are affected by fluctuations in user demand. How to effectively coordinate these multiple objectives and achieve a balance between economic, environmental, and technical considerations during the optimization process is a core issue in current PV-storage charging station control. Multi-timescale collaborative optimization control is also a challenge. Due to the differences in the operating characteristics of photovoltaics, energy storage, and electric vehicles, as well as the influence of external environments, achieving effective collaborative optimization across day-ahead, intraday, and real-time scenarios presents a technical bottleneck. Regarding dynamic resource response, the operating state of the PV-storage charging system varies with changes in sunlight, the charge and discharge status of energy storage devices, and fluctuations in user demand. How to collect real-time operational data from various resources and dynamically adjust resource aggregation and scheduling strategies based on this data is a pressing issue facing existing virtual power plant control methods. In particular, the uncertainty of photovoltaic power generation and the fluctuations in charging pile load often lead to unstable system scheduling, impacting grid balance and achieving optimization goals. The uncertainty introduced by dynamic electricity market changes and user-side flexibility demands makes existing control strategies difficult to respond quickly to market fluctuations and underutilizes user-side flexible resources, impacting economic efficiency. Uncertainty robustness is also crucial in the optimization of PV-storage charging stations. Traditional optimization methods fail to effectively address external uncertainties such as short-term fluctuations in photovoltaic power generation, sudden changes in electric vehicle charging demand, fluctuations in grid load, and equipment failures, resulting in instability and inefficiency in optimization results. The complexity and variability of photovoltaic and energy-storage charging stations, in particular, require robust optimization control methods to ensure stable system operation and maximize economic benefits in a volatile external environment.Finally, there is the problem of computational complexity of optimization algorithms. As the scale of virtual power plants expands, traditional algorithms are difficult to meet real-time control requirements. There is an urgent need to develop efficient optimization algorithms, such as model predictive control or distributed optimization algorithms, to improve the real-time control performance of large-scale virtual power plants.

[0003] After searching, the invention patent with Chinese patent publication number CN104603455B discloses a wind power station control system, a wind power station including a wind power station control system, and a method for controlling a wind power station. This patent mainly focuses on estimating electrical output parameters and reference signal scheduling of wind turbine generators in wind power stations.

[0004] The technical comparison between the above-mentioned reference documents and this application is as follows:

[0005] 1. The photovoltaic storage charging and discharging station aggregation control and optimization method based on the virtual power plant proposed in this application adopts the Markov decision process and adaptive clustering algorithm. It can dynamically aggregate resources according to the different characteristics of various types of equipment (photovoltaic, energy storage and charging piles). It is not limited to a single energy type and has higher resource utilization and system flexibility. Patent CN104603455B mainly estimates the electrical output parameters and schedules reference signals of wind turbine generators in wind power stations. Its control strategy focuses on the operating status monitoring and scheduling of a single resource in the wind farm - wind turbine generators.

[0006] Second, this invention uses a multi-objective optimization model that simultaneously considers economic, technical, and environmental objectives. It also introduces dynamic weighting factors and fuzzy decision theory to generate an optimal scheduling solution. This solution, combined with a rolling horizon control framework and deep reinforcement learning for real-time compensation, significantly improves scheduling accuracy and response speed, achieving multi-objective collaborative optimization and robustness improvements under uncertainty. The scheduling strategy in patent CN104603455B is relatively fixed, focusing on modeling wind power output and generating reference signals.

[0007] After searching, the invention patent of China Patent Publication No. CN115395566A discloses a photovoltaic power station control system and method, which issues adjustment instructions to the voltage outer loop controller through the power station control layer, and distributes current instructions for the grid-connected scheduling of the grid-type photovoltaic power station.

[0008] The technical comparison between the above-mentioned reference documents and this application is as follows:

[0009] 1. In addition to photovoltaic resources, this application integrates energy storage and charging pile resources, and uses the virtual power plant architecture to achieve resource aggregation and regulation, which can respond more flexibly to electricity market demand and grid fluctuations; while patent CN115395566A mainly issues adjustment instructions to the voltage outer loop controller through the power station control layer, distributes current instructions for the grid-connected scheduling of grid-type photovoltaic power stations, and focuses on ensuring the smooth output of the grid-connected current of the photovoltaic inverter.

[0010] Second, the present invention constructs an optimization model that includes economic, technical, and environmental multi-dimensional objectives, and introduces deep reinforcement learning and rolling time domain control to achieve real-time compensation and dynamic optimization scheduling, thereby significantly reducing overall operating costs while improving system stability and operating efficiency; while the solution of patent CN115395566A mainly focuses on the regulation between photovoltaic inverters and grid connection points. Its technical implementation has a positive contribution to system stability, but the scheduling optimization dimension is single.

[0011] After searching, the invention patent with Chinese patent publication number CN115395566A discloses a substation control cabinet humidity monitoring and risk warning system based on digital twin, which uses digital twin technology to realize real-time monitoring and risk warning of the humidity inside and outside the substation control cabinet.

[0012] The technical comparison between the above-mentioned reference documents and this application is as follows:

[0013] The research object and technical application scenario of this application are significantly different. It is to aggregate control and optimize the scheduling of photovoltaic storage charging stations through the concept of virtual power plants. By introducing Markov decision process, dynamic resource aggregation, multi-objective optimization, rolling time domain control and deep reinforcement learning, real-time optimization of the coordinated scheduling of photovoltaic, energy storage and charging piles is achieved. At the same time, the lightweight federated learning model in the edge computing and cloud collaboration mechanism is adopted to effectively reduce the communication load and improve the scheduling response efficiency. Compared with patent CN115601015A, the focus has shifted from environmental monitoring and early warning to resource scheduling optimization and system robustness improvement, with greater adaptability and scalability. Summary of the Invention

[0014] Purpose of the invention: The purpose of the present invention is to provide a method for controlling and optimizing photovoltaic storage charging and discharging stations based on a virtual power plant to solve the problems existing in the above-mentioned prior art. By solving the problems of multi-objective collaborative optimization, dynamic response of resources and robustness to uncertainty in the prior art, the stability and economy of the power system can be improved. This method optimizes the resource scheduling and control strategy of photovoltaic storage charging and discharging stations by introducing technical means such as adaptive clustering algorithm, Markov decision process, NSGA-III multi-objective optimization algorithm, and hierarchical control architecture, thereby improving the economy, stability and flexibility of the virtual power plant system.

[0015] In order to achieve the above object, the solution of the present invention includes the following steps:

[0016] S1. Dynamic Resource Aggregation Modeling: Using an adaptive clustering algorithm, we dynamically divide resource clusters based on device characteristics (charge and discharge efficiency, SOC constraints, response delay, etc.). We also construct a cluster state transition model based on a Markov decision process to characterize the probability of device operating mode switching.

[0017] S2. Construction of multi-objective optimization model: A multi-objective optimization model with economic (operating income), technical (minimizing network losses), environmental (carbon emission reduction) and other objectives is constructed through the weighted summation method. Dynamic weight factors are used to adjust the priority of different objectives to adapt to the grid scheduling needs.

[0018] S3. Hierarchical Collaborative Control Architecture: This architecture employs a hierarchical control structure. The global optimization layer uses an improved NSGA-III algorithm to solve the Pareto frontier and combines it with fuzzy decision theory to generate optimal scheduling instructions. The middle-level control layer performs distributed robust optimization through edge computing nodes and utilizes scenario reduction technology to address photovoltaic power generation forecast errors. The lower-level control layer uses a local controller to implement second-level power adjustment based on a consensus protocol to ensure power balance within the cluster.

[0019] S4. Rolling Horizon Optimization and Dynamic Correction: Through the Rolling Horizon Control (RHC) framework, the global scheduling plan is updated every 15 minutes, and prediction deviations are compensated online based on deep reinforcement learning (DRL).

[0020] S5. Edge-cloud collaborative computing: The cloud is responsible for long-term strategy generation and historical data analysis, while edge nodes perform short-term and rapid optimization, reducing communication load through lightweight federated learning models.

[0021] In modern virtual power plants, the integration and scheduling of distributed energy resources, such as photovoltaics, energy storage, and charging stations, is key to optimizing grid operation. Because these devices have varying characteristics (such as charge and discharge efficiency, SOC constraints, and response latency), efficient resource utilization requires dynamic aggregation and optimized scheduling. This paper proposes a resource aggregation modeling method based on an adaptive clustering algorithm and a Markov decision process. This method can adjust resource clusters in real time based on device characteristics, optimizing the operation of photovoltaic and energy storage charging and discharging stations.

[0022] Constructing a hybrid resource model: This model comprehensively covers distributed energy components such as photovoltaics, energy storage, and charging piles. Key characteristics of each device are analyzed in depth, including parameters such as charge and discharge efficiency, state of charge constraints, and response delays. Based on this, an adaptive clustering algorithm is introduced. This algorithm can dynamically divide resource clusters based on the real-time monitoring of device operating status, grouping devices with similar characteristics to achieve preliminary resource integration and refined management. For example, when it is detected that some photovoltaic panels have similar power generation characteristics due to similar orientations and shading conditions, and the matching energy storage devices have matching charge and discharge performance, they are aggregated into a cluster to facilitate subsequent unified control.

[0023] Establishing a cluster state transition model: Based on a Markov decision process, this model draws on extensive historical and real-time operational data to accurately characterize the probability of device operating mode transitions. For example, under varying light intensity, electricity price periods, and load demand scenarios, it predicts the likelihood of photovoltaic power generation switching to standby mode, and energy storage switching from charging to discharging mode. This provides a forward-looking basis for upper-level optimization decisions, allows for pre-planning resource allocation, and enhances the system's dynamic adaptability.

[0024] Devices are dynamically grouped based on their characteristics (e.g., charge / discharge efficiency, SOC constraint, response delay). The present invention uses an adaptive clustering algorithm to flexibly group devices, thereby adjusting the aggregation strategy based on the current system state.

[0025] Furthermore, the dynamic aggregation of resources includes the following steps:

[0026] S111, device characteristics description: set the device's characteristic vector X i =[η i ,SOC i ,Δt i ],in:

[0027] η i represents the charge and discharge efficiency of device i;

[0028] SOC i Indicates the current battery status of device i;

[0029] Δt i It represents the response delay of device i, that is, the time required from receiving the scheduling signal to completing the action.

[0030] S112, Adaptive Clustering Algorithm: Use adaptive clustering methods (such as K-means or Fuzzy C-means) to divide resources into multiple clusters. The purpose of clustering is to divide devices into subsets with similar characteristics. Devices within each cluster have similar response behaviors, which facilitates optimized scheduling.

[0031] The clustering goal is to minimize the differences between device characteristics:

[0032]

[0033] Among them, C k represents the kth cluster, μ k It is cluster C k The center point, X i is the feature vector of device i, and K is the number of clusters.

[0034] S113. Clustering Update: Clustering results are dynamically updated based on the device's operating status and grid demand. Clustering is adjusted in real time as device characteristics (such as charging status, load demand, etc.) change.

[0035] In order to improve the scheduling flexibility of the device resource cluster, a Markov decision process is used to model the probability of switching the device's operating mode. The state transition model of the device describes the transition probability from one state to another, taking into account the physical constraints of the device, scheduling requirements, and external disturbances. In this invention, the Markov state transition mechanism combines the dynamic aggregation of photovoltaic, energy storage, and charging pile resources, and uses the Markov decision process to describe the transition between each device state and optimize resource scheduling. The following are the detailed steps:

[0036] S121. Define the system state space S, which contains all possible system states. In the present invention, each device (such as photovoltaic, energy storage, charging pile, etc.) has its own independent state, which can include the device's operating mode, power output, energy storage capacity, and load demand.

[0037] The definition of the state space can be expressed as a vector:

[0038] S={s1,s2,...,s n}

[0039] Among them, s i Represents the i-th state of the system.

[0040] S122. Define state transition probability

[0041] State transition probability P(s t ,s t+1 ) describes the system from the current state s t Transfer to the next state s t+1 In a solar-storage-charging system, the probability of state transitions is determined by the physical characteristics and operating rules of each device. For example, the charging and discharging behavior of the energy storage system is limited by the battery's SOC range and charge / discharge efficiency, the power output of photovoltaics is affected by ambient light conditions, and the power demand of charging piles is affected by load demand and charging strategy.

[0042] P(s t,s t+1 )=P(s t+1 |s t )

[0043] Furthermore, the device state transition probability P(s i (t+1)|s i (t),a i (t)) can be estimated through historical data or device models. Assume that the device’s action at each time t is a i (t) is determined by the following formula:

[0044]

[0045] Among them, V(s′) represents the value of state s′, R(s i (t),a) is the immediate reward, and γ is the discount factor.

[0046] The cluster state transition model is based on the MDP model of each device and is extended to describe the dynamic behavior of the resource cluster by aggregating the state transition probabilities of the devices:

[0047]

[0048] S123. Define state transfer function

[0049] State transition can be represented by a state transition function T(s t ,s t+1 ) to describe, the function represents the transition from state s t Transfer to s t+1 The expected reward or utility of . For different device types, the state transition function will be different, depending on the physical characteristics of the device and the scheduling policy.

[0050] For example, the transfer function of the energy storage system can be expressed as:

[0051] α·(SOC t+1 -SOC t )+β·P charge -γ·P discharge

[0052] Among them, α, β, and γ are the weight coefficients of the energy storage equipment, which represent the impact of charging, discharging, and SOC changes on the system benefits.

[0053] The photovoltaic transfer function takes into account the effects of lighting conditions and prediction errors:

[0054]

[0055] Where δ is the weight coefficient of photovoltaic power change, ∈ pv is the prediction error.

[0056] The transfer function of the charging pile can consider the matching between load demand and charging power:

[0057]

[0058] Among them, ζ is the weight coefficient of load demand change, which represents the impact of load change on charging pile operation.

[0059] S124. Define reward function

[0060] Reward function R(s t ,s t+1 ) represents the utility or reward the system receives after transitioning from the current state to the next state. In a solar-storage-charging system, the reward function takes into account the system’s economic efficiency, stability, and environmental benefits.

[0061] R(s t ,s t+1 )=w1·C Profit -w2·C Cost +w3·C Benefit

[0062] Among them, w1, w2, w3 are the weight coefficients of economic, cost and environmental benefits respectively, C Profit is operating income, C Cost is the operating cost, C Benefit It is the carbon emission reduction benefit.

[0063] S125. Iteration and update of state transfer matrix

[0064] Based on a Markov decision process, the system iteratively updates the state transition matrix P to improve its decision-making strategy. Through multiple experiments and feedback mechanisms, the system dynamically adjusts the state transition matrix based on factors such as device performance and environmental changes, optimizing the system's scheduling strategy.

[0065] The optimal policy for the system is obtained by solving the value iteration or policy iteration problem of the Markov decision process. Given a state space and state transition probability matrix, the optimal resource scheduling policy, i.e., the best action to take in each state, is obtained by optimizing the cumulative reward function. The basic formula for value iteration is as follows:

[0066]

[0067] Among them, V(s t ) is the state s t value.

[0068] In the aggregated control and optimization of solar-to-storage charging and discharging stations, a multi-objective optimization model is employed to simultaneously optimize multiple objectives, including economic, technical, and environmental performance. These objectives are comprehensively considered through a weighted summation approach to ensure overall system performance. To adapt to the dynamic changes in grid dispatch requirements, dynamic weighting factors are introduced to adjust the importance of each objective. The optimization model not only needs to balance multiple objectives but also must satisfy a series of physical and grid security constraints.

[0069] Furthermore, the objective function of the multi-objective optimization model takes into account the economic, technical and environmental aspects and integrates them using a weighted summation method. The objective function is as follows:

[0070] maxJ(u(t))=α·J eco (u(t))+β·J tech (u(t))+γ·J env (u(t))

[0071] st

[0072] SOC min ≤SOC i (t)≤SOC max

[0073]

[0074] V min ≤V k (t)≤V ma

[0075] Among them: α, β, γ are dynamic weight factors, and these weights are adjusted according to the grid dispatch requirements.

[0076] Constraints include energy storage SOC constraints, photovoltaic power output constraints, charging pile power constraints, grid line flow constraints and voltage constraints.

[0077] The economic objective function aims to maximize the system's operating revenue. Assume that at time t, the equipment's operating revenue primarily comes from photovoltaic power generation, energy storage charging / discharging, and electricity sales from charging stations. The economic objective function can be expressed as:

[0078]

[0079] in:

[0080] P pv,i (t) represents the photovoltaic power generation of device i at time t;

[0081] P storage,i (t) represents the energy storage charging power of device i at time t;

[0082] P charge,i (t) represents the charging power of the charging pile of device i at time t;

[0083] λ pv ,λ storage ,λ charge The electricity prices are for photovoltaic power generation, energy storage charging / discharging, and charging pile charging respectively.

[0084] The technical objective function aims to minimize grid losses. Grid losses are primarily related to power transmission, and they increase with increasing current. Assume that grid power losses can be expressed using the following formula:

[0085]

[0086] in:

[0087] R k is the resistance of circuit k;

[0088] I k is the current on line k, current and power P k The relationship between (t) is V k is the voltage on line k.

[0089] By minimizing network losses, energy waste in power transmission is reduced and the technical performance of the system is improved.

[0090] The environmental objective function aims to maximize carbon emission reduction. Photovoltaic power generation and energy storage charging and discharging have significant carbon emission reduction effects. Assuming that the carbon emission reduction C reduce (t) Determined by the amount of electricity generated by photovoltaic power generation and energy storage scheduling:

[0091]

[0092] in:

[0093] η pv and η storage are the carbon emission reduction coefficients of photovoltaic power generation and energy storage charging and discharging, respectively.

[0094] By optimizing carbon emission reduction targets, we can reduce the environmental impact of the system and promote the development of green energy.

[0095] Due to the ever-changing demands of grid dispatch, the weighting factors α, β, and γ in the objective function need to be dynamically adjusted. The introduction of dynamic weighting factors allows for adjustments in the priorities among economic, technical, and environmental objectives based on real-time grid demand and electricity market fluctuations. For example, when the grid experiences load fluctuations or power shortages, the weighting of economic and technical objectives may need to be increased; whereas, when carbon emission pressures are high, environmental objectives may receive a higher weighting.

[0096] The dynamic weight factor can be adjusted in the following ways:

[0097] γ(t)=1-α(t)-β(t)

[0098] in:

[0099] P load (t) is the load demand of the power grid;

[0100] P loss (t) is the power loss of the power grid;

[0101] P total (t) is the total power of the grid.

[0102] This dynamic adjustment method automatically adjusts the weights between various objectives according to the real-time status of the power grid and scheduling requirements to ensure the optimal performance of the system.

[0103] To achieve efficient, aggregated control and optimization of PV-storage charging and discharging stations, collaboration is required across multiple layers, from global optimization to local control. A hierarchical control architecture fully leverages the computing power and responsiveness of different layers, ensuring efficient system scheduling across a large area while also enabling rapid local power adjustments to achieve overall optimization goals.

[0104] Furthermore, the hierarchical control architecture is divided into three main layers: global optimization layer, edge computing layer, and local control layer. Each layer has different tasks, and through precise interfaces and coordination mechanisms, it ensures the efficient operation of the entire system.

[0105] The overall process of the hierarchical collaborative control architecture is as follows:

[0106] S31, the global optimization layer generates multi-objective optimization scheduling instructions through the improved NSGA-III algorithm and optimizes the solution through fuzzy decision making;

[0107] S32, the edge computing layer performs distributed robust optimization, uses scenario reduction technology to handle photovoltaic forecast errors, and optimizes the scheduling strategy;

[0108] S33, the local control layer implements second-level power adjustment through the consistency protocol to ensure power balance and fast response between devices.

[0109] The upper global optimization layer uses an improved NSGA-III algorithm to solve the Pareto frontier of multi-objective optimization problems. This algorithm, based on the traditional NSGA-III algorithm, optimizes selection, crossover, and mutation operations, enabling a more efficient search for evenly distributed, near-optimal solutions. Incorporating fuzzy decision theory, it generates optimal dispatch instructions from the Pareto frontier based on factors such as decision-maker preferences and the real-time state of the power grid. This provides macro-control strategies for the middle and lower layers, achieving global resource optimization.

[0110] Mid-layer edge computing nodes perform distributed robust optimization, leveraging scenario reduction techniques to address the uncertainty introduced by photovoltaic forecast errors. Cluster analysis of extensive historical photovoltaic output and meteorological data identifies a small number of representative scenarios. Based on these scenarios, robust optimization models are constructed to ensure stable system operation under varying light fluctuations, minimizing scheduling biases caused by inaccurate forecasts. Furthermore, a lightweight federated learning model interacts with the cloud, sharing computational burdens and reducing communication overhead.

[0111] The lower-level local controller implements power adjustment within seconds based on a consistency protocol, monitors the power status of each device in the cluster in real time, and ensures power balance within the cluster through fast communication and collaborative control. For example, when the load on the charging pile suddenly changes, it quickly coordinates the charging and discharging of the energy storage device or adjusts the photovoltaic output to ensure the stability of the local power supply and avoid impact on the upper power grid.

[0112] Through the coordinated work of these three levels, the control system of the entire photovoltaic storage charging and discharging station can improve response speed and adaptability while ensuring system performance, ensuring system stability and reliability.

[0113] The global optimization layer is responsible for high-level scheduling decisions for the entire system. It uses a multi-objective optimization algorithm to solve the optimal scheduling strategy and combines it with fuzzy decision theory to optimize the results. The goal of global optimization is to calculate the charging and discharging instructions for each device based on grid demand, device status, and external market conditions.

[0114] In this invention, NSGA-III is improved to adapt to the dynamic and diverse nature of resources in the power system. The improved NSGA-III algorithm mainly consists of the following steps:

[0115] S3111. Initial population generation: Generate an initial set of candidate solutions based on the current state of the equipment (such as SOC, power output, load demand, etc.).

[0116] S3112, Non-dominated sorting: Perform non-dominated sorting on each solution in the population and assign different levels and congestion levels.

[0117] S3113. Crossover and mutation operations: Utilize crossover and mutation operations to generate new solutions, and select the optimal solution based on the fitness function. Preferably, the crossover operation utilizes simulated binary crossover (SBX), which combines information from parent individuals by generating new solutions. The mutation operation utilizes a non-uniform mutation method, which randomly selects and adjusts certain genes in the solution.

[0118] S3114, selection operation: Use congestion sorting and non-dominated sorting methods to make selections to ensure that the selected solution is on the Pareto front.

[0119] To maintain solution diversity, the crowding distance is used to measure the density of solutions in the target space. The crowding distance is defined by calculating the relative distance between solutions in the target space. Solutions with higher crowding are typically preferred to maintain population diversity.

[0120] The calculation formula for the congestion distance is:

[0121]

[0122] in:

[0123] f is the function value of solution x on the kth target;

[0124] M is the dimension of the target;

[0125] x i+1 and x i-1 are the front and back neighbors of solution x on the k-th target respectively.

[0126] After multiple iterations, NSGA-III generates a set of optimal solutions, which form the Pareto front. To determine the optimal solution, fuzzy decision theory is used to perform a weighted evaluation of multiple objectives and generate the final scheduling instructions.

[0127] The basic steps of fuzzy decision making are as follows:

[0128] S3121. Input fuzzification: Map the Pareto frontier results output by the optimization algorithm into fuzzy values through fuzzy membership functions.

[0129] S3122, Rule-based reasoning: Reasoning the most appropriate decision-making strategy based on a set of predefined rule bases.

[0130] S3123, output defuzzification: Defuzzify the inference results to obtain specific scheduling instructions, such as the charging and discharging power of each device, target weight, etc.

[0131] The fuzzy inference results are converted into specific scheduling instructions through defuzzification. The goal of the defuzzification process is to convert fuzzy sets into specific values or decisions. The present invention uses the centroid method to defuzzify:

[0132]

[0133] Among them, μ i is the fuzzy membership, x i is the corresponding specific scheduling value. The center of gravity method obtains the final scheduling decision through weighted average.

[0134] Finally, the scheduling instructions generated by the global optimization layer are passed to the edge computing layer for further processing.

[0135] In complex power systems, real-time updates and adaptability of dispatch plans are crucial. To this end, this paper proposes a dynamic optimization strategy that combines a rolling horizon control framework with deep reinforcement learning. This rolling horizon optimization enables continuous adjustment of dispatch plans to respond to changes in grid load, power generation resources, and market demand. Furthermore, by introducing deep reinforcement learning technology, online compensation and corrections based on real-time electricity prices and load forecasts can be performed, further improving dispatch accuracy and system robustness.

[0136] Based on real-time data and system status, the scheduling plan for a certain period of time in the future is continuously optimized. Whenever a new time window opens, the scheduling plan is recalculated and revised based on the new information.

[0137] Furthermore, the steps of rolling time domain optimization in the present invention are as follows:

[0138] S411. Time Window Definition: Define a fixed time window, such as 15 minutes, for optimization calculation. Within each time window, perform optimization scheduling based on the current system status (such as device SOC, electricity price, load demand, etc.).

[0139] S412, scheduling update: During each update, the scheduling plan within the future time window is recalculated to determine the scheduling target at each moment, such as power output, charging and discharging strategy, etc.

[0140] S413. Optimization model: The optimization model is usually a multi-objective optimization problem, considering multiple objectives such as economic (minimizing operating costs), technical (minimizing network losses, load balancing), and environmental (maximizing carbon emission reduction).

[0141] A rolling horizon control framework is applied; the global scheduling plan is updated every 15 minutes, breaking down long-term optimization problems into multiple shorter sub-problems that are solved sequentially. Within each sub-period, an optimization strategy is formulated based on the latest system status and forecast information. After the period ends, subsequent period plans are revised based on actual operational feedback, effectively responding to system dynamics and improving optimization timeliness.

[0142] The steps for deep reinforcement learning-assisted dynamic correction are as follows:

[0143] S421. Use the DRL algorithm to compensate for the prediction deviation online, and adjust the control strategy in real time through repeated interactive learning between the intelligent agent and the environment.

[0144] S422. At the same time, a real-time electricity price signal and load forecast feedback mechanism is embedded to dynamically adjust the optimization target weight according to market price fluctuations and load change trends, ensuring that the optimization plan always meets actual operating needs and enhances system adaptability and economy.

[0145] Assume T w is the time window of the rolling time domain control, T current If it is the current time, the rolling time domain optimization is updated every 15 minutes, and the new scheduling plan is:

[0146]

[0147] in:

[0148] C oper (T) is the operating cost,

[0149] L loss (T) is the grid loss,

[0150] δ CO2 (T) is carbon emissions,

[0151] P is the scheduling decision vector,

[0152] α1, α2, α3 are the weight factors of the target.

[0153] Through rolling optimization, the optimization model dynamically updates the scheduling plan and adjusts system behavior based on real-time data.

[0154] To further improve dispatch accuracy, deep reinforcement learning (DRL) is used to perform online compensation and corrections, combining real-time electricity price signals and load forecasts. DRL can adaptively adjust based on real-time feedback to better cope with dynamically changing environments.

[0155] Deep reinforcement learning uses proximal policy optimization to update the policy to optimize system scheduling. According to the reinforcement learning framework, the policy is continuously adjusted based on environmental feedback.

[0156]

[0157] Among them, α is the learning rate, is the gradient of parameter update, Expectation of reward.

[0158] Through deep reinforcement learning, the system adjusts the weights of optimization objectives and scheduling strategies based on real-time electricity prices and load forecasts. After each decision, the system updates its strategy based on the latest environmental conditions, enabling dynamic corrections.

[0159] By combining rolling horizon control (RHC) with deep reinforcement learning (DRL), the system continuously updates and refines the dispatch plan based on multi-objective optimization. Every 15 minutes, the system updates the optimization target weights and dispatch plan based on the latest system status (such as grid load, SOC of PV and storage equipment, and real-time electricity prices). After each update, DRL dynamically adjusts the dispatch plan by receiving real-time feedback signals. In the event of market price fluctuations or load forecast errors, DRL can compensate online and dynamically refine the target weights to further optimize system dispatch. Combining rolling horizon control (RHC) with DRL enables adaptive optimization of the power system under dynamic changes. By updating the dispatch plan every 15 minutes and leveraging reinforcement learning to adjust for dynamic factors such as price fluctuations and load changes, the dispatch system maintains greater flexibility and adaptability. This dynamic optimization mechanism improves dispatch accuracy, reduces operating costs, and enhances the economy and stability of the system.

[0160] This invention utilizes edge-cloud collaborative computing. The synergy between edge and cloud computing is crucial for improving the efficiency, responsiveness, and accuracy of the scheduling system. Cloud computing primarily handles long-term strategy generation and historical data analysis, while edge computing handles short-term, rapid optimization tasks. A lightweight federated learning model reduces communication overhead and improves system responsiveness.

[0161] Cloud computing is responsible for the generation of long-term strategies and the analysis of historical data. The cloud system can access large-scale data and computing resources, use global optimization algorithms to generate long-term scheduling plans, and optimize the operation of the entire virtual power plant. Specific functions include: based on the system's historical data (such as power demand, grid load, electricity price fluctuations, etc.), the cloud system uses global optimization algorithms (such as genetic algorithms, particle swarm algorithms, NSGA-III, etc.) to generate long-term scheduling strategies. By analyzing historical data, the cloud can extract some potential patterns and trends to help improve future scheduling plans. Cloud computing generates global scheduling strategies based on historical data and forecast information. The goal of long-term strategies is usually to reduce the overall operating costs of the system, improve efficiency, and enhance sustainability.

[0162] Edge computing is primarily responsible for short-term, rapid optimization, executing localized tasks, reducing communication latency, and improving the system's real-time responsiveness. Since cloud computing is often limited by communication latency, edge computing takes on the task of rapidly processing real-time data. Specific functions include: Edge nodes use local optimization algorithms to rapidly process real-time data and optimize local system scheduling tasks, such as power scheduling for photovoltaics, energy storage, and charging stations. To reduce the communication burden, edge computing collaborates with the cloud, employing a lightweight federated learning model for local data training and global model aggregation. Through a distributed learning approach, each edge node trains a model based on local data and uploads the parameters to the cloud for global model updates. This allows scheduling models to be updated without directly transmitting large amounts of data. Edge computing can quickly respond to real-time changes in grid load, electricity price fluctuations, and changes in device status, improving system adaptability.

[0163] The cloud is responsible for long-term strategy generation. Based on massive amounts of historical data, it analyzes energy market patterns, long-term equipment operating characteristics, and regional energy demand trends to provide macro-strategic guidance for the system, such as developing seasonal and annual energy reserve and scheduling plans. Edge nodes focus on short-term, rapid optimization. Leveraging their low-latency proximity to the device, they respond in real time to changes in local device status and real-time task requirements, performing operations such as real-time power balancing adjustments and emergency fault handling. By sharing key model parameters with the cloud through a lightweight federated learning model, they reduce full data transmission, alleviate communication pressure, accelerate overall system decision-making, and achieve efficient collaboration between the edge and cloud.

[0164] In each round of training on the edge node, the local model update formula is as follows:

[0165]

[0166] Among them, D k is the local dataset of the kth edge node, is the local model parameter of the loss function L, and the local learning rate is η k ,θ k,t It is the local dataset D k The gradient on , the loss function L is defined according to the specific task, the regression task uses the mean square error, and the classification task uses the cross entropy loss;

[0167] After each round of training, each edge node uploads the updated local model parameters to the cloud for aggregation. The global model aggregation formula is:

[0168]

[0169] Among them, θ t+1 is the aggregated global model parameter, |Dk | is the size of the local dataset of edge node k, |D| is the total size of the datasets of all edge nodes;

[0170] In order to further reduce the communication load, the model parameters are compressed, the precision of the model parameters is reduced by quantization, and the unimportant connections or parameters in the model are removed by pruning technology. Since quantization and pruning operations will introduce certain errors, error compensation is required to ensure model performance. The error introduced by quantization or pruning is set to ε , compensation is performed when updating the model parameters, the formula is as follows:

[0171]

[0172] in, It is the local model parameter after error compensation. By reasonably estimating and compensating the error, the impact of quantization and pruning on model performance can be minimized.

[0173] In the federated learning framework, edge nodes upload local model parameters to the cloud, which then calculates global model parameters and returns them to each edge node. This allows edge and cloud computing to collaborate efficiently, reducing communication overhead while improving system response speed and scheduling accuracy.

[0174] Beneficial effects: The present invention has the following advantages:

[0175] This invention, by introducing a Markov decision process and a dynamic clustering algorithm, achieves dynamic aggregation and optimized scheduling of photovoltaic, energy storage, and charging pile resources, significantly improving the flexibility and optimization of resource scheduling. By employing a multi-objective optimization approach, combining economic, technical, and environmental objectives, and introducing dynamic weighting factors, the scheduling strategy can achieve optimal balance based on changes in grid demand. Furthermore, through rolling horizon control and deep reinforcement learning techniques, the invention enhances the system's real-time responsiveness and scheduling accuracy, updating the global scheduling plan every 15 minutes and enhancing adaptability to load fluctuations and electricity price changes. The edge-cloud collaborative computing architecture reduces communication latency, improving the system's response speed and computational efficiency. Furthermore, fuzzy decision theory and an improved NSGA-III algorithm optimize multi-objective scheduling, reducing operating costs and technical losses. A lightweight federated learning model reduces communication load while ensuring data privacy and security. Overall, the invention not only improves the efficiency and stability of grid scheduling, but also effectively reduces system operating costs and environmental impact, demonstrating excellent scalability and adaptability. BRIEF DESCRIPTION OF THE DRAWINGS

[0176] Figure 1 is a flow chart of the steps of the present invention;

[0177] Figure 2 This is a diagram of the layered cloud-edge collaborative control architecture;

[0178] Figure 3 It is a diagram of the multi-objective optimization model architecture;

[0179] Figure 4 It is a timing diagram of rolling time domain optimization and dynamic correction. DETAILED DESCRIPTION

[0180] The technical solutions in the embodiments of the present invention will be described clearly and completely below with reference to the accompanying drawings.

[0181] This embodiment provides a method for controlling and optimizing photovoltaic storage charging and discharging stations based on a virtual power plant. Figure 1 As shown, the following steps are included:

[0182] S1. Dynamic resource aggregation modeling: Using an adaptive clustering algorithm, resource clusters are dynamically divided according to device characteristics. A cluster state transition model is constructed based on a Markov decision process to characterize the probability of device operation mode switching. Device characteristics include charge and discharge efficiency, SOC constraint, and response delay. In step S1, resource clusters are dynamically divided according to device characteristics using an adaptive clustering algorithm. The specific method is as follows:

[0183] S111. Device Characteristics Description

[0184] Set the device's characteristic vector X i =[η i ,SOC i ,Δt i ],in:

[0185] η i represents the charge and discharge efficiency of device i;

[0186] SOC i Indicates the current battery status of device i;

[0187] Δt i represents the response delay of device i, that is, the time required from receiving the scheduling signal to completing the action;

[0188] S112, Adaptive Clustering Algorithm

[0189] An adaptive clustering method is used to divide resources into multiple clusters. The clustering goal is to minimize the differences between device characteristics:

[0190]

[0191] Among them, C k represents the kth cluster, μ k It is cluster C k The center point, Xi is the feature vector of device i, K is the number of clusters;

[0192] S113, cluster division update

[0193] Clustering results are dynamically updated based on the device's operating status and grid demand. Clustering is adjusted in real time as device characteristics change.

[0194] At the same time, a cluster state transition model is constructed based on the Markov decision process to characterize the probability of device operation mode switching. The specific method is as follows:

[0195] S121. Define the state space S of the system

[0196] The state space S contains all possible system states, represented by vectors:

[0197] S={s1,s2,...,s n}

[0198] Among them, s i represents the i-th state of the system;

[0199] S122. Define state transition probability

[0200] The device state transition probability P(s) is calculated by historical data or device model. i (t+1)|s i (t),a i (t)) make an estimate;

[0201] Assume that the device takes action a at each time t i (t) is determined by the following formula:

[0202]

[0203] Among them, V(s′) represents the value of state s′, R(s i (t),a) is the immediate reward, γ is the discount factor;

[0204] The cluster state transition model is based on the MDP model of each device and is extended to describe the dynamic behavior of the resource cluster by aggregating the state transition probabilities of the devices:

[0205]

[0206] S123. Define state transfer function

[0207] Define state transition functions based on the physical characteristics of the device and the scheduling strategy;

[0208] The transfer function of the energy storage system is as follows:

[0209] α·(SOC t+1 -SOC t )+β·P charge -γ·P discharge

[0210] Among them, α, β, and γ are the weight coefficients of the energy storage device, which represent the impact of charging, discharging, and SOC changes on the system benefits;

[0211] The transfer function of photovoltaic can be given as follows:

[0212]

[0213] Where δ is the weight coefficient of photovoltaic power change, ∈ pv is the prediction error;

[0214] The transfer function of the charging pile is as follows:

[0215]

[0216] Where ζ is the weight coefficient of load demand change, which represents the impact of load change on charging pile operation;

[0217] S124. Define reward function

[0218] The reward function is defined based on the system's economy, stability, and environmental benefits:

[0219] R(s t ,s t+1 )=w1·C Profit -w2·C Cost +w3·C Benefit

[0220] Among them, w1, w2, w3 are the weight coefficients of economic, cost and environmental benefits respectively, C Profit is operating income, C Cost is the operating cost, C Benefit It is the carbon emission reduction benefit;

[0221] S125. Iteration and update of state transfer matrix

[0222] Based on the Markov decision process, the system iteratively updates the state transition matrix P to improve the decision-making strategy. Through multiple experiments and feedback mechanisms, the system dynamically adjusts the state transition matrix based on device performance and environmental changes to optimize the system's scheduling strategy.

[0223] The optimal strategy of the system is obtained by solving the value iteration or policy iteration problem of the Markov decision process. Under a given state space and state transition probability matrix, the optimal resource scheduling strategy is obtained by optimizing the cumulative reward function, that is, the best action to take in each state. The basic formula of value iteration is as follows:

[0224]

[0225] Among them, V(s t ) is the state s t value.

[0226] S2. Construction of multi-objective optimization model: A multi-objective optimization model with economic, technical and environmental objectives is constructed through the weighted summation method, and dynamic weight factors are used to adjust the priorities of different objectives to meet the needs of power grid scheduling.

[0227] The objective function of the multi-objective optimization model is as follows:

[0228] maxJ(u(t))=α·J eco (u(t))+β·J tech (u(t))+γ·J env (u(t))

[0229] st

[0230] SOC min ≤SOC i (t)≤SOC max

[0231]

[0232] V min ≤V k (t)≤V ma

[0233] Among them: α, β, γ are dynamic weight factors, and the weight factors α, β, γ are adjusted in the following way:

[0234] γ(t)=1-α(t)-β(t)

[0235] in:

[0236] P load (t) is the load demand of the power grid;

[0237] P loss (t) is the power loss of the power grid;

[0238] P total (t) is the total power of the grid.

[0239] Constraints include energy storage SOC constraints, photovoltaic power output constraints, charging pile power constraints, grid line flow constraints and voltage constraints;

[0240] The economic objective function aims to maximize the operating benefits of the system. The economic objective function is as follows:

[0241]

[0242] in:

[0243] P pv,i (t) represents the photovoltaic power generation of device i at time t;

[0244] P storage,i (t) represents the energy storage charging power of device i at time t;

[0245] P charge,i (t) represents the charging power of the charging pile of device i at time t;

[0246] λ pv ,λ storage ,λ charge The electricity prices for photovoltaic power generation, energy storage charging / discharging, and charging pile charging are respectively;

[0247] The technical objective function aims to minimize the network loss of the power grid. The technical objective function is as follows:

[0248]

[0249] in:

[0250] R k is the resistance of circuit k;

[0251] I k is the current on line k, current and power P k The relationship between (t) is V k is the voltage on line k;

[0252] The environmental objective function aims to maximize carbon emission reduction. The environmental objective function is as follows:

[0253]

[0254] in:

[0255] η pv and η storage are the carbon emission reduction coefficients of photovoltaic power generation and energy storage charging and discharging, respectively.

[0256] like Figure 3The diagram below shows the architecture of the multi-objective optimization model used in this step, presenting the optimization model system constructed to maximize the comprehensive benefits of the photovoltaic storage charging and discharging station. First, the model defines three main objectives: economic efficiency, which focuses on operational benefits, including the grid-connected revenue of photovoltaic power generation, the price difference between charging and discharging of energy storage equipment, and the charging service revenue of charging piles; technical efficiency, which aims to minimize network losses and reduces power loss during transmission by rationally planning the power distribution and operation mode of equipment within the photovoltaic storage charging and discharging station; and environmental efficiency, which takes carbon emission reduction as the goal, evaluates the environmental impact of system operation, and calculates the carbon emissions reduced by replacing traditional fossil fuel power generation with photovoltaic power generation, as well as the carbon emissions of energy storage equipment and charging piles. To adapt to different grid scheduling needs and operating scenarios, a dynamic weighting factor is introduced, which can flexibly adjust the weight of each objective based on factors such as real-time grid status, electricity price information, and environmental policies. In terms of constraints, the system comprehensively considers the physical limitations of the equipment, such as the maximum power output limit of photovoltaic equipment, the charging and discharging power and capacity limits of energy storage equipment, and the charging power limit of charging piles. Grid security constraints, including voltage amplitude limits and line flow limits, ensure that the access of photovoltaic and energy storage charging and discharging stations does not threaten the safe and stable operation of the grid. Furthermore, demand-side response agreements require that if photovoltaic and energy storage charging and discharging stations participate in demand-side response, they must comply with relevant contractual provisions and adjust power output as required during specific time periods. By combining these objective functions with the constraints, a complete multi-objective optimization model is constructed, providing a decision-making basis for the optimal scheduling of photovoltaic and energy storage charging and discharging stations, and achieving comprehensive optimization of the system in terms of economy, technology, and environment.

[0257] S3. Hierarchical collaborative control architecture: This architecture comprises a global optimization layer, an edge computing layer, and a local control layer. The global optimization layer uses an improved NSGA-III algorithm to solve the Pareto frontier and generates optimal scheduling instructions in combination with fuzzy decision theory. The middle-layer control layer performs distributed robust optimization through edge computing nodes and uses scenario reduction technology to address photovoltaic power generation prediction errors. The lower-layer control layer, based on the local controller, implements second-level power adjustment based on a consistency protocol to ensure power balance within the cluster.

[0258] S31, the global optimization layer generates multi-objective optimization scheduling instructions through the improved NSGA-III algorithm and optimizes the solution through fuzzy decision making;

[0259] The improved NSGA-III algorithm includes the following steps:

[0260] S3111. Initial population generation: Generate an initial set of candidate solutions based on the current state of the device.

[0261] S3112, Non-dominated sorting: Perform non-dominated sorting on each solution in the population and assign different levels and congestion levels.

[0262] S3113, Crossover and Mutation Operation: Generate new solutions using simulated binary crossover and non-uniform mutation methods, and select the optimal solution based on the fitness function;

[0263] S3114, selection operation: Use congestion sorting and non-dominated sorting methods to select, ensuring that the selected solution is on the Pareto front;

[0264] The crowding distance is used to measure the distribution density of solutions in the target space. The crowding distance is defined by calculating the relative distance between solutions in the target space, and the solution with larger crowding is given priority.

[0265] The calculation formula for the congestion distance is:

[0266]

[0267] in:

[0268] f is the function value of solution x on the kth target;

[0269] M is the dimension of the target;

[0270] x i+1 and x i-1 are the front and back neighbors of solution x on the k-th target respectively.

[0271] The steps of fuzzy decision making are as follows:

[0272] S3121. Input fuzzification: Map the Pareto frontier results output by the optimization algorithm into fuzzy values through fuzzy membership functions.

[0273] S3122, Rule-Based Reasoning: Reasoning the most appropriate decision strategy based on a set of predefined rule bases;

[0274] S3123, output defuzzification: Defuzzify the inference results using the center of gravity method to obtain specific scheduling decision instructions;

[0275] The centroid method defuzzification is performed as follows:

[0276]

[0277] Among them, μ i is the fuzzy membership, x i is the corresponding specific scheduling value.

[0278] S32, the edge computing layer performs distributed robust optimization, uses scenario reduction technology to handle photovoltaic forecast errors, and optimizes the scheduling strategy;

[0279] S33, the local control layer implements second-level power adjustment through the consistency protocol to ensure power balance and fast response between devices.

[0280] S4, Rolling Horizon Optimization and Dynamic Correction: Through the rolling horizon control framework, the global scheduling plan is updated every 15 minutes, and prediction deviations are compensated online based on deep reinforcement learning.

[0281] S411. Time window definition: Define a fixed time window of 15 minutes. Within each time window, perform optimized scheduling based on the current system status.

[0282] S412, scheduling update: During each update, the scheduling plan within the future time window is recalculated to determine the scheduling target at each moment;

[0283] S413, Optimization model: Perform multi-objective optimization of the model based on economic, technical and environmental objectives, and perform optimization according to the following formula;

[0284]

[0285] Where: T w is the time window of the rolling time domain control, T current is the current moment, C oper (T) is the operating cost, L loss (T) is the grid loss, δ CO2 (T) is the carbon emission, P is the scheduling decision vector, and α1, α2, and α3 are the weight factors of the target;

[0286] like Figure 4 The following diagram illustrates the dynamic process of scheduling optimization in a virtual power plant (VPP). This process, based on a rolling horizon control framework, regularly updates the scheduling plan to respond to changing grid demand and resource status. Each optimization cycle typically lasts 15 minutes. At the beginning of each cycle, the system recalculates the global scheduling plan based on current grid load, device status, and forecast data. First, the system collects information from each device and the grid, including real-time electricity prices, load demand, and device status. Next, the cloud computing layer generates a new scheduling plan based on this collected information. The edge computing layer locally optimizes this information and refines it based on real-time data. The optimized scheduling plan is then transmitted to the local control end for specific resource scheduling. A new scheduling update is performed every 15 minutes to ensure that the system can dynamically adjust to the latest grid demand and respond quickly to unexpected events. This rolling horizon optimization approach enables the VPP to adapt to load fluctuations and resource status changes, ensuring efficient and flexible scheduling.

[0287] S421. Use the DRL algorithm to compensate for prediction deviations online, and adjust the control strategy in real time through repeated interactive learning between the agent and the environment.

[0288] S422. Embed real-time electricity price signals and load forecast feedback mechanisms to dynamically adjust optimization target weights based on market price fluctuations and load change trends, ensuring that the optimization plan always meets actual operating needs and enhancing system adaptability and economy.

[0289] S5. Edge-Cloud Collaborative Computing: Cloud computing is responsible for generating long-term strategies and analyzing historical data. The cloud system has access to large-scale data and computing resources, uses global optimization algorithms to generate long-term scheduling plans, and optimizes the operation of the entire virtual power plant. Specific functions include: Based on the system's historical data, the cloud system uses global optimization algorithms to generate long-term scheduling strategies. By analyzing historical data, it extracts potential patterns and trends and generates a global scheduling strategy.

[0290] Edge nodes use local optimization algorithms to quickly process real-time data and optimize the scheduling tasks of local systems. They adopt lightweight federated learning models for local data training and global model aggregation. Through distributed learning, each edge node trains a model based on local data and uploads parameters to the cloud for global model updates.

[0291] like Figure 2The following diagram shows the hierarchical cloud-edge collaborative control architecture involved in step S5. This diagram illustrates the collaboration between edge computing nodes and cloud servers in a PV-storage charging and discharging station to achieve efficient computing and data processing. In this architecture, the cloud server undertakes important macro-management and data analysis tasks. It is responsible for generating long-term strategies, such as long-term charging and discharging strategies for energy storage devices and charging pile layout planning, based on factors such as the long-term development goals of the PV-storage charging and discharging station, grid planning, and energy market changes. Simultaneously, the cloud server conducts in-depth analysis of the extensive historical data accumulated during the operation of the PV-storage charging and discharging station, mining valuable information such as device performance, user behavior patterns, and energy consumption patterns, providing strong support for system optimization and decision-making. Edge computing nodes, on the other hand, focus on real-time, localized computing tasks. They connect to photovoltaic (PV), energy storage, charging piles, and other equipment, and can obtain real-time operating status data, such as the real-time output power of PV, the SOC value of energy storage devices, and the charging current of charging piles. Edge nodes leverage this real-time data to perform short-term, rapid optimization. When a device's operating status changes, such as a sudden fluctuation in PV power or an increase in charging pile load, the edge nodes react quickly and adjust the device's operating parameters to maintain stable system operation. To achieve effective collaboration between edge nodes and the cloud, a lightweight federated learning model is employed. In this model, edge nodes train models locally using collected data and then upload the resulting model parameters to the cloud. The cloud aggregates these model parameters from different edge nodes to generate a global model, which is then sent back to the edge nodes. The edge nodes then update their local models based on the global model and continue the next round of training. This approach reduces the communication overhead associated with the large amount of raw data transmitted between edge nodes and the cloud, while enabling continuous model optimization and updates, thereby improving the overall performance and efficiency of the PV-storage charging and discharging station system.

[0292] The basic principles, main features, and advantages of the present invention are shown and described above. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely illustrative of the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for controlling and optimizing photovoltaic storage charging and discharging stations based on a virtual power plant, characterized in that: The following steps are involved: S1. Dynamic resource aggregation modeling: Using an adaptive clustering algorithm, resources are dynamically divided into clusters based on device characteristics. Furthermore, a cluster state transition model is constructed based on a Markov decision process to characterize the probability of device operation mode switching. S2. Multi-objective optimization model construction: A multi-objective optimization model with economic, technical, and environmental objectives is constructed using a weighted summation method. Dynamic weight factors are used to adjust the priorities of different objectives to meet the needs of power grid dispatching. S3. Hierarchical collaborative control architecture: This architecture comprises a global optimization layer, an edge computing layer, and a local control layer. The global optimization layer uses an improved NSGA-III algorithm to solve the Pareto frontier and generates optimal scheduling instructions in combination with fuzzy decision theory. The middle-layer control layer performs distributed robust optimization through edge computing nodes and uses scenario reduction technology to address photovoltaic power generation prediction errors. The lower-layer control layer, based on the local controller, implements second-level power adjustment based on a consistency protocol to ensure power balance within the cluster. S4, Rolling Horizon Optimization and Dynamic Correction: Through the rolling horizon control framework, the global scheduling plan is updated every 15 minutes, and prediction deviations are compensated online based on deep reinforcement learning. S5. Edge-cloud collaborative computing: The cloud is responsible for long-term strategy generation and historical data analysis, while edge nodes perform short-term and rapid optimization, reducing communication load through lightweight federated learning models.

2. The method for controlling and optimizing photovoltaic storage charging and discharging stations based on a virtual power plant according to claim 1, characterized in that: The device characteristics include charge and discharge efficiency, SOC constraint, and response delay. In step S1, an adaptive clustering algorithm is used to dynamically divide resource clusters according to the device characteristics. The specific method is as follows: S111. Device Characteristics Description Set the device's characteristic vector X i =[η i ,SOC i ,Δt i ],in: η i represents the charge and discharge efficiency of device i; SOC i Indicates the current battery status of device i; Δt i represents the response delay of device i, that is, the time required from receiving the scheduling signal to completing the action; S112, Adaptive Clustering Algorithm An adaptive clustering method is used to divide resources into multiple clusters. The clustering goal is to minimize the differences between device characteristics: Among them, C k represents the kth cluster, μ k It is cluster C k The center point, X i is the feature vector of device i, K is the number of clusters; S113, cluster division update The clustering results are dynamically updated based on the operating status of the equipment and the grid demand. As the characteristics of the equipment change, the cluster division is adjusted in real time. At the same time, a cluster state transition model is constructed based on the Markov decision process to characterize the probability of device operation mode switching. The specific method is as follows: S121. Define the state space S of the system The state space S contains all possible system states, represented by vectors: S={s1,s2,...,s n } Among them, s i represents the i-th state of the system; S122. Define state transition probability The device state transition probability P(s) is calculated by historical data or device model. i (t+1)|s i (t),a i (t)) make an estimate; Assume that the device takes action a at each time t i (t) is determined by the following formula: Among them, V(s′) represents the value of state s′, R(s i (t),a) is the immediate reward, γ is the discount factor; The cluster state transition model is based on the MDP model of each device and is extended to describe the dynamic behavior of the resource cluster by aggregating the state transition probabilities of the devices: S123. Define state transfer function Define state transition functions based on the physical characteristics of the device and the scheduling strategy; The transfer function of the energy storage system is as follows: α·(SOC t+1 -SOC t )+β·P charge -γ·P discharge Among them, α, β, and γ are the weight coefficients of the energy storage device, which represent the impact of charging, discharging, and SOC changes on the system benefits; The transfer function of photovoltaic can be given as follows: Where δ is the weight coefficient of photovoltaic power change, ∈ pv is the prediction error; The transfer function of the charging pile is as follows: Where ζ is the weight coefficient of load demand change, which represents the impact of load change on charging pile operation; S124. Define reward function The reward function is defined based on the system's economy, stability, and environmental benefits: R(s t ,s t+1 )=w1·C Profit -w2·C Cost +w3·C Benefit Among them, w1, w2, w3 are the weight coefficients of economic, cost and environmental benefits respectively, C Profit is operating income, C Cost is the operating cost, C Benefit It is the carbon emission reduction benefit; S125. Iteration and update of state transfer matrix Based on the Markov decision process, the system iteratively updates the state transition matrix P to improve the decision-making strategy. Through multiple experiments and feedback mechanisms, the system dynamically adjusts the state transition matrix based on device performance and environmental changes to optimize the system's scheduling strategy. The optimal strategy of the system is obtained by solving the value iteration or policy iteration problem of the Markov decision process. Under a given state space and state transition probability matrix, the optimal resource scheduling strategy is obtained by optimizing the cumulative reward function, that is, the best action to take in each state. The basic formula of value iteration is as follows: Among them, V(s t ) is the state s t value.

3. The method for controlling and optimizing photovoltaic storage charging and discharging stations based on a virtual power plant according to claim 1, characterized in that: The objective function of the multi-objective optimization model in step S2 is as follows: maxJ(u(t))=α·J eco (u(t))+β·J tech (u(t))+γ·J env (u(t)) st SOC min ≤SOC i (t)≤SOC max V min ≤V k (t)≤V ma Among them: α, β, γ are dynamic weight factors, which are adjusted according to the grid dispatch requirements; Constraints include energy storage SOC constraints, photovoltaic power output constraints, charging pile power constraints, grid line flow constraints and voltage constraints; The economic objective function aims to maximize the operating benefits of the system. The economic objective function is as follows: in: P pv,i (t) represents the photovoltaic power generation of device i at time t; P storage,i (t) represents the energy storage charging power of device i at time t; P charge,i (t) represents the charging power of the charging pile of device i at time t; λ pv ,λ storage ,λ charge The electricity prices for photovoltaic power generation, energy storage charging / discharging, and charging pile charging are respectively; The technical objective function aims to minimize the network loss of the power grid. The technical objective function is as follows: in: R k is the resistance of circuit k; I k is the current on line k, current and power P k The relationship between (t) is V k is the voltage on line k; The environmental objective function aims to maximize carbon emission reduction. The environmental objective function is as follows: in: η pv and η storage are the carbon emission reduction coefficients of photovoltaic power generation and energy storage charging and discharging, respectively.

4. The method for controlling and optimizing photovoltaic storage charging and discharging stations based on a virtual power plant according to claim 3 is characterized by: The weight factors α, β, and γ in the objective function of the multi-objective optimization model in step S2 are adjusted in the following manner: γ(t)=1-α(t)-β(t) in: P load (t) is the load demand of the power grid; P loss (t) is the power loss of the power grid; P total (t) is the total power of the grid.

5. The method for controlling and optimizing photovoltaic storage charging and discharging stations based on a virtual power plant according to claim 1, characterized in that: The overall process of the hierarchical collaborative control architecture in step S3 is as follows: S31, the global optimization layer generates multi-objective optimization scheduling instructions through the improved NSGA-III algorithm and optimizes the solution through fuzzy decision making; S32, the edge computing layer performs distributed robust optimization, uses scenario reduction technology to handle photovoltaic forecast errors, and optimizes the scheduling strategy; S33, the local control layer implements second-level power adjustment through the consistency protocol to ensure power balance and fast response between devices.

6. The method for controlling and optimizing photovoltaic storage charging and discharging stations based on a virtual power plant according to claim 5, characterized in that: The improved NSGA-III algorithm in step S31 includes the following steps: S3111, initial population generation: Generate an initial candidate solution set based on the current state of the device; S3112, Non-dominated sorting: Perform non-dominated sorting on each solution in the population and assign different ranks and congestion degrees; S3113, crossover and mutation operations: Generate new solutions using crossover and mutation operations, and select the optimal solution based on the fitness function; S3114, selection operation: Use congestion sorting and non-dominated sorting methods to select, ensuring that the selected solution is on the Pareto front; The steps of fuzzy decision making in step S31 are as follows: S3121, Input Fuzzification: Map the Pareto frontier results output by the optimization algorithm into fuzzy values through fuzzy membership functions; S3122, Rule-Based Reasoning: Reasoning the most appropriate decision strategy based on a set of predefined rule bases; S3123, output defuzzification: Defuzzify the inference results using the center of gravity method to obtain specific scheduling decision instructions; The centroid method defuzzification is performed as follows: Among them, μ i is the fuzzy membership, x i is the corresponding specific scheduling value.

7. The method for controlling and optimizing photovoltaic storage charging and discharging stations based on a virtual power plant according to claim 6, characterized in that: The crossover operation in step S3113 adopts simulated binary crossover, and the mutation operation adopts a non-uniform mutation method; In step S3114, the congestion distance is used to measure the distribution density of the solutions in the target space. The congestion distance is defined by calculating the relative distance between the solutions in the target space, and the solution with the larger congestion is preferentially selected. The calculation formula for the congestion distance is: in: f is the function value of solution x on the kth target; M is the dimension of the target; x i+1 and x i-1 are the front and back neighbors of solution x on the k-th target respectively.

8. The method for controlling and optimizing photovoltaic storage charging and discharging stations based on a virtual power plant according to claim 1, characterized in that: The specific steps of step S4 are as follows: S411. Time window definition: Define a fixed time window of 15 minutes. Within each time window, perform optimized scheduling based on the current system status. S412, scheduling update: During each update, the scheduling plan within the future time window is recalculated to determine the scheduling target at each moment; S413, Optimization model: Perform multi-objective optimization of the model based on economic, technical and environmental objectives, and perform optimization according to the following formula; Where: T w is the time window of the rolling time domain control, T current is the current moment, C oper (T) is the operating cost, L loss (T) is the grid loss, δ CO2 (T) is the carbon emission, P is the scheduling decision vector, and α1, α2, and α3 are the weight factors of the target; S421. Use the DRL algorithm to compensate for prediction deviations online. Through repeated interactive learning between the agent and the environment, the control strategy is adjusted in real time. Deep reinforcement learning uses proximal policy optimization to update the strategy to optimize system scheduling. According to the reinforcement learning framework, the strategy will be continuously adjusted according to environmental feedback: Among them, α is the learning rate, is the gradient of parameter update, for the expectation of reward; S422. Embed real-time electricity price signals and load forecast feedback mechanisms to dynamically adjust optimization target weights based on market price fluctuations and load change trends, ensuring that the optimization plan always meets actual operating needs and enhancing system adaptability and economy.

9. The method for controlling and optimizing photovoltaic storage charging and discharging stations based on a virtual power plant according to claim 1, characterized in that: The specific method of step S5 is as follows: Cloud computing is responsible for generating long-term strategies and analyzing historical data. The cloud system can access large-scale data and computing resources, use global optimization algorithms to generate long-term scheduling plans, and optimize the operation of the entire virtual power plant to generate a global scheduling strategy. Edge nodes use local optimization algorithms to quickly process real-time data and optimize the scheduling tasks of local systems. They adopt lightweight federated learning models for local data training and global model aggregation. Through distributed learning, each edge node trains a model based on local data and uploads parameters to the cloud for global model updates.

10. According to the method for controlling and optimizing photovoltaic and energy storage charging and discharging stations based on a virtual power plant in claim 1, the federated learning framework is characterized by: In each round of training on the edge node, the local model update formula is as follows: in, D k is the local dataset of the kth edge node, is the local model parameter of the loss function L, and the local learning rate is η k ,θ k,t It is the local dataset D k The gradient on , the loss function L is defined according to the specific task, the regression task uses the mean square error, and the classification task uses the cross entropy loss; After each round of training, each edge node uploads the updated local model parameters to the cloud for aggregation. The global model aggregation formula is: Among them, θ t+1 is the aggregated global model parameter, |D k | is the size of the local dataset of edge node k, |D| is the total size of the datasets of all edge nodes; In order to further reduce the communication load, the model parameters are compressed, the precision of the model parameters is reduced by quantization, and the unimportant connections or parameters in the model are removed by pruning technology. Since operations such as quantization and pruning will introduce certain errors, error compensation is required to ensure model performance. The error introduced by quantization or pruning is set to ε, and compensation is performed when updating the model parameters. The formula is as follows: in, It is the local model parameter after error compensation. By reasonably estimating and compensating the error, the impact of quantization and pruning on model performance can be minimized. In the federated learning framework, edge nodes upload local model parameters to the cloud, which calculates global model parameters and returns them to each edge node.

Citation Information

Patent Citations

  • Wind power plant control system, wind power plant including wind power plant control system, and method for controlling wind power plant.

    CN104603455B

  • Photovoltaic power station control system and method

    CN115395566A

Cited By

  • Layered optimization regulation and control method for electricity-hydrogen coupling in multi-energy complementary system

    CN120745439A

  • Low-carbon economic operation optimization method for electricity-hydrogen-heat comprehensive energy system

    CN120850811A

  • Flexible resource scheduling method and system considering state time sequence coupling and medium

    CN120851562A

  • Distributed photovoltaic hierarchical coordination control method based on micro-grid architecture

    CN120914859A

  • Distributed scheduling method for whole vehicle manufacturing stamping resources under cloud side end cooperation

    CN120952497A