Optical storage direct flexible system optimization method based on distributed space-time cooperative control
By employing a distributed spatiotemporal collaborative control optimization method for photovoltaic-storage-DC-flexible systems, utilizing graph neural networks and federated Kalman filtering for load forecasting, and combining multi-agent deep reinforcement learning for scheduling optimization, this approach addresses the grid stability and energy efficiency issues caused by the volatility of renewable energy in photovoltaic-storage-DC-flexible systems, achieving efficient load management and system coordination.
Patent Information
- Application Number
- CN202511919939.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-18
- Publication Date
- 2026-02-10
AI Technical Summary
Existing photovoltaic-storage-DC-flexible systems face the intermittency and volatility of renewable energy, making it difficult to improve grid stability and energy efficiency in a coordinated manner. Traditional models are sensitive to data quality and lack generalization ability, making it difficult to achieve coordinated scheduling of large-scale distributed nodes.
A distributed spatiotemporal cooperative control method is adopted, which constructs a load prediction model through graph neural network, combines federated Kalman filtering and multi-agent deep reinforcement learning to perform load prediction and scheduling optimization, and achieves system stability and energy efficiency improvement through edge-cloud collaborative verification.
It achieves high-precision load forecasting and distributed collaborative scheduling of photovoltaic-storage-DC-flexible systems, reduces load variance, improves cross-regional collaborative efficiency and photovoltaic absorption rate, and enhances system stability and privacy protection.
Smart Images

Figure CN121507985A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of system optimization and new energy technology, specifically focusing on an optimization method for a photovoltaic-storage-direct-flexible system based on distributed spatiotemporal collaborative control. Background Technology
[0002] Against the backdrop of global energy transition, the development and utilization of new energy sources have become a focal point. Photovoltaic-storage-DC-flexible systems, with their unique advantages, have demonstrated enormous potential in improving energy efficiency. However, these systems currently face numerous severe challenges in actual operation. On the one hand, with the rapid advancement of industrialization, the proportion of building energy consumption in total social energy consumption continues to rise, with industrial buildings, due to their large scale and highly dense equipment, becoming the main energy consumers. On the other hand, the inherent intermittency and volatility of renewable energy sources greatly disrupt the stable operation of the power grid, leading to frequent frequency deviations and voltage fluctuations. Existing technologies have made some progress in addressing the issues of renewable energy consumption and energy reduction. For example, the application of deep learning in smart microgrids has demonstrated the potential of deep learning models in predicting renewable energy generation and power load; however, it also reveals the high dependence of deep learning models on data quality and quantity, and the problems of model training and computational costs, which limit the feasibility of their widespread application. A DC grid protection method based on dynamic state estimation has been proposed, reducing fault detection time to within 3ms, but the performance of this method is constrained by the accuracy of line parameters. Furthermore, opportunity-constrained programming models can reduce operating costs by 12.3% through demand response, but their robustness is significantly affected by renewable energy prediction errors. These existing technologies have three limitations: high sensitivity to data quality, insufficient model generalization ability, and a lack of mechanisms for handling efficiency uncertainties. These limitations make it difficult to comprehensively achieve synergistic improvements in energy efficiency and stability of photovoltaic-storage-DC-flexible systems. Summary of the Invention
[0003] This invention addresses the aforementioned technical problems by proposing an optimization method for a photovoltaic-storage-direct-current-flexible system based on distributed spatiotemporal coordinated control. The method is based on a distributed sensing-hierarchical decision-making-spatiotemporal coordinated optimization approach for photovoltaic-storage-direct-current-flexible systems. It constructs a load forecasting model using a graph neural network and combines it with federated Kalman filtering for long-term forecasting. After data preprocessing and stage division, a multi-agent deep reinforcement learning method based on a multi-agent deep deterministic policy gradient algorithm is employed for scheduling optimization. Finally, edge-cloud collaborative verification further improves the system's energy efficiency and stability.
[0004] This invention provides an optimization method for a photovoltaic-storage-direct-flexible system based on distributed spatiotemporal cooperative control, comprising: Step 110: Data collection is performed, including obtaining historical load data; a load prediction model is constructed using a graph neural network; and the model parameters are optimized using an adaptive moment estimation algorithm to obtain the initial load prediction model. Step 120: Based on the initial load forecasting model, use federated Kalman filtering to fuse local information from edge nodes for long-term load forecasting and generate global load forecasting data. Step 130: Based on global load forecast data, construct a continuous load function using adaptive cubic spline interpolation; and use density clustering algorithm to divide the continuous load function into stages. Step 140: Multi-agent deep deterministic policy gradient algorithm is used to optimize multi-agent cooperative scheduling, including: defining the heterogeneous state space and hybrid action space of multi-agents based on the continuous load function of stage division; setting up multi-agent experience replay pool and cooperative evaluator; and designing a dynamic adjustment mechanism that includes local rewards and global cooperative rewards.
[0005] Step 150: Evaluate the optimization effect by combining multiple indicators; the indicators include at least: load variance reduction rate, cross-regional collaborative efficiency, photovoltaic absorption rate, energy storage health degradation rate, and privacy protection level. Step 160: Perform edge-cloud collaborative verification: upload the action execution records of edge nodes in real time; verify global consistency in the cloud through a digital twin model and send correction instructions for abnormal actions; dynamically adjust node trust and reward allocation coefficients.
[0006] Compared with the prior art, the beneficial effects obtained by the present invention include, but are not limited to: (1) A multi-agent collaborative decision-making framework for photovoltaic-energy storage-load is adopted, and a sub-regional interconnected topology is constructed by combining graph neural networks (GNNs) (edge features are embedded with building physical parameters and grid impedance) to realize the technological leap from centralized single decision-making to distributed heterogeneous collaboration of photovoltaic-energy storage DC-flexible systems. Through multi-agent experience sharing and global evaluation mechanism, the problem of traditional single agents being unable to adapt to the collaborative scheduling of large-scale distributed nodes is solved, supporting the dynamic access and autonomous collaboration of different types of photovoltaic-energy storage units.
[0007] (2) An innovative approach integrates federated Kalman filtering and adaptive feature methods, replacing the information redundancy processing of traditional centralized filtering through iterative local filtering at edge nodes and weight aggregation in the cloud. Combined with node feature extension of graph neural networks (embedding equipment cycle and environmental sensitivity parameters), it addresses the insufficient adaptability of traditional models to extreme weather and sudden load changes. Technically, it achieves real-time matching between the prediction model and the dynamic characteristics of the system, overcoming the robustness limitations of fixed algorithms in complex scenarios.
[0008] (3) Construct a hierarchical execution system that combines rapid local response at the edge with global verification in the cloud. Deploy local (local) decision modules at the edge nodes to achieve millisecond-level load fluctuation response. The cloud performs global consistency verification through a digital twin model and combines dynamic adjustment of node trust and reward allocation coefficients to influence system decision-making. Attached Figure Description
[0009] To more clearly illustrate the optimization method for the optical-storage-direct-flexible system based on distributed spatiotemporal coordinated control, the following diagram illustrates the optimization method: Figure 1 This is a flowchart illustrating the steps of an optimization method for a photovoltaic-storage-direct-flexible system based on distributed spatiotemporal cooperative control in one embodiment of the present invention. Figure 2 This is a schematic diagram showing the load standard deviation before and after overall data scheduling optimization in the experiment of this invention; Figure 3 This is a schematic diagram showing the load standard deviation before and after data scheduling optimization during the high-load phase of the experiment in this invention. Detailed Implementation
[0010] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0011] In one embodiment, the present invention provides an optimization method for a photovoltaic-storage-direct-flexible system based on distributed spatiotemporal cooperative control, such as... Figure 1 As shown, the method includes: Step 110: Data collection is performed, including obtaining historical load data; a load prediction model is constructed using a graph neural network; and the model parameters are optimized using an adaptive moment estimation algorithm to obtain the initial load prediction model. Step 120: Based on the initial load forecasting model, use federated Kalman filtering to fuse local information from edge nodes for long-term load forecasting and generate global load forecasting data. Step 130: Based on global load forecast data, construct a continuous load function using adaptive cubic spline interpolation; and use density clustering algorithm to divide the continuous load function into stages. Step 140: Multi-agent deep deterministic policy gradient algorithm is used to optimize multi-agent cooperative scheduling, including: defining the heterogeneous state space and hybrid action space of multi-agents based on the continuous load function of stage division; setting up multi-agent experience replay pool and cooperative evaluator; and designing a dynamic adjustment mechanism that includes local rewards and global cooperative rewards. Step 150: Evaluate the optimization effect by combining multiple indicators; the indicators include at least: load variance reduction rate, cross-regional collaborative efficiency, photovoltaic absorption rate, energy storage health degradation rate, and privacy protection level. Step 160: Perform edge-cloud collaborative verification: upload the action execution records of edge nodes in real time; verify global consistency in the cloud through a digital twin model and send correction instructions for abnormal actions; dynamically adjust node trust and reward allocation coefficients.
[0012] To effectively overcome the challenges of large fluctuations and significant peak-valley load differences in existing photovoltaic power generation technologies, this invention innovatively integrates deep reinforcement learning and Kalman filtering to obtain an optimization method for photovoltaic-storage-DC-flexible systems based on distributed spatiotemporal coordinated control. This aims to achieve precise and efficient solutions to the aforementioned problems. The method fully leverages the advantages of both advanced deep reinforcement learning techniques and Kalman filtering algorithms in data processing, predictive analysis, and decision optimization.
[0013] To achieve the aforementioned objective, specifically, in step 110, data acquisition is performed, including obtaining historical load data; a load prediction model is constructed using a graph neural network; and an adaptive moment estimation algorithm is used to optimize the model parameters to obtain an initial load prediction model, including: First, data acquisition is performed, continuously collecting data on the target area for a set time resolution for the duration of the data acquisition. Historical load data within the time frame This represents the time step (timestamp index) of a time series that is collected with time resolution for the duration. ; It is the duration The total number of time steps (time samples) within the specified time period; synchronously acquire meteorological data and seasonal / periodic information within the corresponding time period; establish an original time-series database based on historical load data, meteorological data, and seasonal / periodic information. The meteorological data includes at least temperature data (e.g., dry-bulb temperature, dew point temperature), humidity data, and light intensity.
[0014] In one embodiment, the duration of data collection of the target area is... Historical load data for the year, with a time resolution of hours, provides typical scenario support for model training. An example of the data format of the original time-series database is shown in Table 1.
[0015] Table 1. Example of data format for the original time-series database
[0016] Then, a load prediction model based on a graph neural network (GNN) is constructed. The GNN consists of multiple graph convolutional layers used to aggregate neighbor node information of the graph structure and capture spatial dependencies. The graph structure of the GNN can be constructed based on geographical associations (e.g., geographical location of regional divisions, geographically based meteorological influences) and / or historical load similarity (e.g., similarity relationships revealed by historical load data). The nodes in the graph structure... Sub-regions in the corresponding target area , ; This represents the total number of nodes in the graph structure, and also the total number of sub-regions divided within the target region. Each graph convolutional layer performs feature propagation and transformation, outputting node embeddings; finally, a fully connected layer maps the node embeddings to load prediction values.
[0017] By encoding the parameter information of the load forecasting model into the node feature vectors of a graph neural network, the model can comprehensively represent each electricity-consuming entity, laying the foundation for analyzing the interaction between parameters.
[0018] In one embodiment, the node feature vectors of the graph neural network are described as follows: ; in, Represents a node The node feature vectors, For nodes The average load, For nodes Temperature sensitivity parameters; The heat transfer coefficient of the building envelope; Encoding the operating cycle of system equipment; For weekdays, use a dummy variable. weekend .
[0019] In one embodiment, the temperature sensitivity parameter The improved quadratic polynomial fitting yielded the following: ; in, It is a node At time (i.e., time step in a time series) Temperature sensitivity parameters; It is a node At any moment The dry bulb temperature; It is a node At any moment The dew point temperature; For nodes At any moment Real-time humidity; For nodes At any moment Light intensity, in units of Fit coefficients The following solution is obtained using the least squares method with L1 (based on L1 norm) regularization: ; In the formula, It is a node At any moment Actual load measurements (obtained from historical load data). It is a node At any moment The load estimate is an intermediate calculated value used to fit the temperature sensitivity parameter, and is not the final system load forecast. It is a fitting coefficient index. ; is the regularization coefficient, used to suppress overfitting, and regularization achieves sparsity of polynomial coefficients.
[0020] In one embodiment, .
[0021] In one embodiment, the building envelope is a reinforced concrete wall, and the heat transfer coefficient of the building envelope structure is taken as... ; workdays weekend .
[0022] The graph convolutional layer propagation rule used for node feature vectors in this invention is as follows: ; in, Indicates the first Nodes in a layered graph convolutional layer Node feature vectors; For nodes The set of neighboring nodes, For nodes The set of neighboring nodes, ; Represents the set of neighboring nodes The base number, Represents the set of neighboring nodes The cardinality; For the first The weight matrix of the layered graph convolutional layer, It is the ReLU activation function. For bias terms; Indicates the first Nodes in a layered graph convolutional layer The node feature vectors.
[0023] The graph convolutional layer propagation rule captures regional and complex spatial dependencies by simulating the flow and transmission of information in the building network. After multiple rounds of graph convolutional propagation, the final representation of each node (i.e., the updated feature vector) not only contains its own information, but also incorporates the contextual information of its neighboring nodes in the graph structure. The representation rich in spatial contextual information is fed into a prediction layer (such as a fully connected layer in a neural network) to generate the final load prediction value.
[0024] Next, using time variables (representing the moments in the time series), temperature variables (representing temperature data), seasonal / periodic dummy variables, and building characteristic parameters as inputs to the load forecasting model, the adaptive moment estimation algorithm is used to optimize the model parameters to obtain the initial load forecasting model.
[0025] The update rule for the adaptive moment estimation algorithm to optimize model parameters is as follows: The first step is to consider any model parameter of the load forecasting model. Calculate its gradient: ; in, For model parameters The gradient of the next iteration; The load forecasting error loss function; model parameters express Any one of the parameters in.
[0026] In one embodiment, the mean squared error (MSE) of the model parameters is selected as the load forecasting error loss function. .
[0027] The second step is to update the first-order moment estimate: ; in, For the first First-moment estimation of the gradient in the next iteration; This is the first-order attenuation coefficient.
[0028] In one embodiment, the first-order attenuation coefficient .
[0029] The third step is to update the second-order moment estimate: ; in, For the first Second-moment estimation of the gradient in the next iteration; It is the second-order attenuation coefficient.
[0030] In one embodiment, the second-order attenuation coefficient .
[0031] Step 4, Deviation Correction: ; in, and These are the corrected first and second moments, respectively; This represents the current iteration number.
[0032] Step 5: Update model parameters: ; in, Indicates the first Any model parameter in the next iteration The value of ; It is the learning rate, used to control the step size for updating model parameters; It is a numerically stable term, taking a smaller non-zero value to prevent division by zero in fractions.
[0033] For example, if the model parameters It is a temperature-sensitive parameter. Then the formula for updating the parameters in step five is expressed as: ; in, Indicates the first Temperature sensitivity parameters in the next iteration The value of .
[0034] In one embodiment, the learning rate Numerical stability term To prevent division by zero.
[0035] Federated Kalman filtering is a distributed estimation technique that combines the results of multiple local Kalman filters. The cloud acts as the fusion center, aggregating the local estimates of all edge nodes to generate a global state estimate. It mainly includes two stages: local filtering and global fusion, and typically involves weight allocation (such as trust weights) to optimize the fusion effect.
[0036] Furthermore, in step 120, based on the initial load forecasting model, long-term load forecasting is performed by fusing local information from edge nodes using federated Kalman filtering, including: The first step is to define the local system state equations and measurement equations for the edge nodes. Specifically, by introducing a regional coupling coefficient matrix, the load correlation strength between regions is quantified and used to construct the state transition matrix of the local system state equations. edge nodes The local state equations and measurement equations are given by the following equations: ; ; in, Represents edge nodes At any moment The state vector (such as load forecast or temperature sensitivity parameter). ; This is the total number of edge nodes; For edge nodes The state transition matrix contains the region coupling coefficients, that is: ; In the above formula, It is the regional coupling coefficient matrix =( The elements in ) represent edge nodes. With nodes The coupling coefficient; For edge nodes The set of neighboring nodes, For edge nodes With nodes Adjacency relationship, It is After vectorization, find the diagonal matrix, where 1 is the dimension. A vector of all 1s; This indicates the creation of diagonal elements. diagonal matrix, It is an edge node Its own state transition coefficient; It is an edge node At any moment The input temperature variable (dry bulb temperature / dew point temperature) can be set to a specific value. This is the control matrix for the temperature variable, used to control the influence of the input temperature variable on the state vector; It is process noise, with a mean of 0 and a covariance matrix of... Gaussian distribution ; Represents edge nodes At any moment Measured values (such as actual load measurements); It is a measurement matrix used to map the state vector to the measurement space; It is measurement noise, which follows a mean of 0 and a covariance matrix of... Gaussian distribution .
[0037] The second step involves performing local prediction and updates at the edge nodes using local Kalman filtering based on the local system state equation and measurement equation, thereby obtaining the state estimates of the edge nodes.
[0038] In this process, the noise covariance update rule is: edge nodes At every moment Update the covariance matrix of the process noise as follows Covariance matrix of measurement noise : ; ; in, For a moment The covariance matrix of the process noise. For a moment The covariance matrix of the measurement noise; It is an edge node At any moment State estimation.
[0039] When a sudden temperature change is detected, the covariance matrix of the measurement noise is temporarily adjusted. ,make It decays linearly to its initial value after a preset duration.
[0040] The temperature jump refers to the rate of change of a temperature variable per unit time. satisfy: ; in, It is a preset temperature change rate threshold.
[0041] In one embodiment, each time step The simulation is on an hourly basis; when a sudden temperature change is detected, the temperature rises by 5°C within 15 minutes during the simulation. It decays linearly after three cycles. After being processed by federated Kalman filtering, the long-term prediction accuracy is significantly improved.
[0042] The third step involves using the fusion learning operation of federated Kalman filtering to define a global fusion equation in the cloud and generate global load prediction data. The weighted average coefficient of the global fusion equation is the trust weight of the edge nodes. The trust weight of the edge nodes is updated in real time based on the local prediction performance.
[0043] The global fusion equation: ; in, It is the cloud that is always there. Global state estimation, used to collect data from all edge nodes in the cloud. State estimation ; For edge nodes Trust weighting.
[0044] In one embodiment, the local prediction performance is measured using the Mean Absolute Percentage Error (MAPE) metric; the trust weights of edge nodes are updated based on MAPE, with the specific update rule being: ; In the above formula, Sensitivity coefficient ( ), For edge nodes At any moment The mean absolute percentage error.
[0045] In one embodiment, the sensitivity coefficient If edge nodes during the experiment At any moment Mean absolute percentage error Then its corresponding trust weight .
[0046] Note that the global load forecast data generated using the global fusion equation is discrete scatter time series data.
[0047] Specifically, in step 130, data preprocessing and stage division are performed, including: The first step is to construct the continuous load function.
[0048] Based on global load forecast data, a continuous load function is constructed using adaptive cubic spline interpolation. The interpolation nodes satisfy the dynamic smoothness condition. The dynamic smoothness condition is used to ensure that the interpolated load curve has higher accuracy and smoothness in areas of drastic load changes (such as high load phases).
[0049] Specifically, the global load forecast data is aggregated according to daily, weekly, or monthly cycles, and a continuous load function is constructed using adaptive cubic spline interpolation.
[0050] In one embodiment, the interpolation nodes that meet the dynamic smoothness condition are encrypted during the high-load phase. For example, one node is set every hour during the high-load phase and one node is set every two hours during the medium-low load phase to ensure curve smoothness.
[0051] The second step is to use the density clustering algorithm (DBSCAN) to divide the continuous load function into load curves for high load, medium load, and low load stages.
[0052] When performing phase division, samples are first taken at fixed time steps on the curve of the continuous load function to obtain sample points for clustering. The total number of sample points is... .
[0053] Develop the continuous load function partitioning rules for the density clustering algorithm, including: Sample points of Neighborhood density is defined as: ; in, Represents sample points of The number of sample points contained in the neighborhood; It is the neighborhood radius; It is an indicator function used to statistically analyze sample points. of The number of sample points in the neighborhood, when the condition in parentheses The value is 1 if the condition is true, and 0 otherwise. Sample points With sample points The weighted Euclidean distance between them Sample points The load function value, This represents the load change over one hour. For sample points Corresponding time period (e.g., the first period of high load) Load variation (hourly); Sample points The load function value, For sample points Corresponding time period (e.g., the first period of high load) The load change (hourly).
[0054] Thresholds are divided by phase: During the high-load phase, sample points exist The number of sample points in the neighborhood satisfies: ,and The continuous crossing time that meets the above conditions is not less than 1 hour; during the medium load phase, the following conditions must be met: ,and During low-load periods, the following conditions must be met: ,and ;in, This is the global load average. The standard deviation of the global load. The average density is defined as follows. The constraint of continuous spanning time (not less than 1 hour) is used to ensure that each phase is continuous in time and to avoid isolated points, such as high-load phases lasting at least 1 hour.
[0055] Furthermore, in step 140, the multi-agent cooperative scheduling optimization using a multi-agent deep deterministic policy gradient algorithm includes: The first step is to define a multi-agent heterogeneous state space based on a continuous load function with phase division. The photovoltaic agent includes irradiance and output prediction error, the energy storage agent includes state of charge (SOC) and state of health (SOH), and the load agent includes flexible load adjustment potential.
[0056] The definition of the multi-agent heterogeneous state space includes the following state definitions: Photovoltaic intelligent agent status: ; in, This is a real-time measurement of light intensity. To predict light intensity, For prediction error, The rate of change of illumination reflects the fluctuation of illumination. Energy storage agent status: ; in, For a moment The state of charge (SOC) indicates the percentage of remaining battery capacity. For a moment The health status of the battery reflects its degree of aging. It is the maximum charging and discharging power; This is the charge / discharge efficiency attenuation coefficient; and They are time points The charging power and discharging power, ; This is the rated capacity.
[0057] Load agent state: ; in, Indicates time The load function value, This demonstrates the potential for flexible load regulation. It is the load change rate, reflecting load fluctuation; The real-time temperature sensitivity coefficient is calculated using temperature sensitivity parameters.
[0058] The second step is to define a hybrid action space, including discrete and continuous actions, to cover a variety of scheduling optimization operations.
[0059] Discrete actions: grid power purchase, direct photovoltaic power supply, selection of energy storage charging and discharging modes, etc.; Continuous operation: energy storage charging and discharging power (0-100%, for example, 50% power is 20MW), continuous adjustment (for example, fine adjustment of photovoltaic inverter output), load shedding ratio (for example, 0-30%).
[0060] The third step involves training the multi-agent cooperative control strategy using the Multi-Agent Deep Deterministic Policy Gradient Algorithm (MADDPG). The aim is to enable the agents to learn how to optimally schedule photovoltaic, energy storage, and load through cooperation, thereby achieving stable, economical, and safe operation of the photovoltaic-storage-DC-flexible system. The training includes setting up a multi-agent experience replay pool and a cooperative evaluator, as well as designing a dynamic adjustment mechanism that incorporates local and global cooperative rewards.
[0061] The multi-agent experience replay pool stores experience samples from agents for offline learning, thereby improving training stability. The collaborative evaluator refers to a critique network equipped with a peer reviewer for each agent, which evaluates the value of actions in the global state, promoting collaboration.
[0062] In the multi-agent deep deterministic policy gradient algorithm, the reward function of the dynamic adjustment mechanism includes at least: Local rewards According to the intelligent agent The values vary depending on the role: If the intelligent agent The role is that of a photovoltaic intelligent agent. Rewards for photovoltaic intelligent agents: ; in, This represents the per-unit value of the actual photovoltaic output. If the intelligent agent Its role is that of an energy storage intelligent agent. Rewards for energy storage agents: ; in, It is in a state of energy storage charge; If the intelligent agent The role is that of a load-bearing intelligent agent. Rewards for the agent with high workload: ; in, This is the maximum load limit.
[0063] Global Collaborative Rewards It is given by the following formula: ; in, To improve cross-domain collaboration efficiency, For photovoltaic power absorption rate, For privacy protection, The power limit for the connecting line is exceeded; The training parameters of the multi-agent deep deterministic policy gradient algorithm include at least: experience replay pool capacity, target network update period, learning rate, and discount factor.
[0064] In one embodiment, the intelligent agent It is a photovoltaic intelligent body. , Prediction error ; , , , Then, the photovoltaic agent reward, which serves as a local reward, can be calculated as follows: ; And global collaborative rewards: ; Training parameters are set as follows: Experience replay pool capacity The batch size is 256, and the number of iterations is 1000.
[0065] Specifically, in step 150, the optimization effect is evaluated by combining multiple indicators; the indicators used to evaluate the optimization effect include at least: load variance reduction rate, cross-regional collaborative efficiency, photovoltaic absorption rate, energy storage health degradation rate, and privacy protection level (local data retention rate).
[0066] The load variance reduction rate is given by the following formula: ; in, The load variance before optimization. This represents the optimized load variance.
[0067] The cross-regional collaboration efficiency is given by the following formula: ; in, This refers to the power transmission volume of inter-regional tie lines.
[0068] The photovoltaic absorption rate reflects the proportion of photovoltaic output that is effectively utilized, thus avoiding curtailment. It is specifically given by the following formula: ; in, This refers to the actual amount of photovoltaic power consumed by the system. This represents the total photovoltaic power generation.
[0069] The energy storage health degradation rate is given by the following formula: ; in, They are time points and The health status of the energy storage smart body at that time.
[0070] The privacy protection level is given by the following formula: .
[0071] By combining the load variance reduction rate, cross-regional collaborative efficiency, photovoltaic absorption rate, energy storage health degradation rate, and privacy protection degree defined in the above manner, a multi-index system for optimizing and evaluating photovoltaic-storage-DC-flexible systems was constructed.
[0072] In one embodiment, a photovoltaic-storage system in a certain region is selected as the research object for scheduling optimization of calculation indicators. The maximum installed capacity of photovoltaic equipment is set to 100MW, which is used in conjunction with a 40MW / 80MWh energy storage plant. The load data adopts the load forecast data for the next year. The reinforcement learning iteration count is 100 times per cycle, and a total of 10 rounds of learning are performed.
[0073] The scheduling optimization results for cross-regional collaborative efficiency are as follows:
[0074] The optimization results for privacy protection are as follows: local data retention rate of 92%; Scheduling optimization results for energy storage health degradation rate:
[0075] Finally, in step 160, edge-cloud collaborative verification is performed, including: The first step is to upload the action execution records of the edge nodes in real time. The cloud uses a digital twin model to calculate and verify global consistency and sends correction instructions for abnormal actions.
[0076] Specifically, the cloud sends correction instructions for dynamic adjustment based on the verification trigger rules of the digital twin model.
[0077] The verification triggering rules for the digital twin model include: Calculate edge nodes Digital twin bias : ; in, Represents edge nodes At any moment The simulated values, Represents edge nodes At any moment The actual value; like If the value exceeds a preset threshold and the duration exceeds a preset period, a correction command for dynamic adjustment will be sent from the cloud. ; in, This is the motion correction amount; The second step is to dynamically adjust the node trust level and reward distribution coefficient.
[0078] Design the dynamic adjustment formula for node trust level: ; in, It is an edge node The level of trust in the nodes.
[0079] In one embodiment, edge nodes upload data every 5 minutes, and the cloud-based digital twin model verifies the deviation between the simulated value and the actual value of a certain node. The node trust level increases by 0.1.
[0080] Formula for dynamic adjustment of reward distribution coefficient: ; in, It is an edge node At any moment The reward distribution coefficient, To reward sensitivity coefficient, This is the baseline reward value.
[0081] In one embodiment, when And for two consecutive cycles, it sends correction commands for dynamic adjustment: .
[0082] Finally, in the experiment of this invention, taking the energy storage of a photovoltaic-storage-DC-flexible system in an industrial park as an example, including 100MW photovoltaic and 40MW / 80MWh energy storage, the results of scheduling optimization according to the method of this invention are as follows: Standard deviation of load for the overall annual data, such as Figure 2 As shown, the load standard deviation decreased from 1.956 before scheduling optimization to 1.7238 after scheduling optimization, a decrease of 11.87%; the load standard deviation during high-load periods, as shown... Figure 3 As shown, the value decreased from 2.5828 before scheduling optimization to 1.9349 after scheduling optimization, a decrease of 25.08%.
[0083] The calculation and evaluation results show that the method provided by this invention can effectively improve the grid load characteristics and enhance system stability and renewable energy utilization efficiency.
[0084] In summary, this invention provides an optimization method for a photovoltaic-storage-DC-flexible system based on distributed spatiotemporal collaborative control. First, a basic load forecasting model is constructed using a graph neural network to capture the load transmission effect between regions. Building envelope parameters and equipment operating cycles are introduced, and wavelet transform is used to decompose load trends and fluctuation characteristics. A federated Kalman filter algorithm is employed, with local filtering at edge nodes and weight aggregation in the cloud, to achieve high-precision long-term load forecasting in distributed scenarios. The average absolute percentage error of this model is controlled below 3.8%. Next, load data is preprocessed. Adaptive spline interpolation is used to construct a continuous load function, and density clustering is used to divide the load curve into high, medium, and low stages (with a newly added medium load stage). Then, a multi-agent collaborative scheduling optimization module is used to define the heterogeneous state space and hybrid action space (including continuous adjustment actions) of photovoltaic, energy storage, and load agents. The Multi-Agent Deep Deterministic Policy Gradient Algorithm (MADDPG) combined with a dynamic adjustment mechanism is used for training to achieve spatiotemporal collaborative scheduling of the "source-storage-load". Finally, a new edge-cloud collaborative verification mechanism is added to verify the optimization effect through multiple dimensions such as load variance reduction rate, improved cross-regional collaborative efficiency, and enhanced privacy protection. The method of this invention is based on a distributed perception-hierarchical decision-spatiotemporal collaborative optimization approach for a direct-flexible optical-storage system. It constructs a basic load prediction model using graph neural networks, combines it with federated Kalman filtering to achieve long-term prediction, and after data preprocessing and stage division, employs multi-agent deep reinforcement learning for scheduling optimization. Ultimately, edge-cloud collaborative verification improves system energy efficiency and stability.
[0085] In one embodiment, the present invention also provides an optimization device for a photovoltaic-storage-direct-flexible system based on distributed spatiotemporal cooperative control, the device comprising: The first module is used for data acquisition, including obtaining historical load data; constructing a load prediction model through a graph neural network; and optimizing the model parameters using an adaptive moment estimation algorithm to obtain the initial load prediction model. The second module is used to perform long-term load forecasting based on the initial load forecasting model, using federated Kalman filtering to fuse local information from edge nodes, and generate global load forecasting data. The third module is used to construct a continuous load function based on global load forecast data using adaptive cubic spline interpolation; and to divide the continuous load function into stages using a density clustering algorithm. The fourth module is used for multi-agent cooperative scheduling optimization using a multi-agent deep deterministic policy gradient algorithm. It includes: defining the heterogeneous state space and hybrid action space of the multi-agent based on a phase-based continuous load function; setting up a multi-agent experience replay pool and a cooperative evaluator; and designing a dynamic adjustment mechanism that includes local rewards and global cooperative rewards.
[0086] The fifth module is used to evaluate the optimization effect by combining multiple indicators; the indicators include at least: load variance reduction rate, cross-regional collaborative efficiency, photovoltaic absorption rate, energy storage health degradation rate, and privacy protection level. The sixth module is used for edge-cloud collaborative verification: real-time uploading of action execution records of edge nodes; verification of global consistency in the cloud through a digital twin model, sending correction instructions for abnormal actions; and dynamic adjustment of node trust and reward allocation coefficients.
[0087] On the other hand, the present invention provides a computer device, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of the optical-storage-direct-flexible system optimization method based on distributed spatiotemporal cooperative control provided in any of the above embodiments. The computer device can be a server. The computer device includes a processor, a memory, a network interface, and a database connected via a system bus. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store sample data. The network interface of the computer device is used for communication with external terminals via a network connection.
[0088] On the other hand, the present invention provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the optical storage direct current flexible system optimization method based on distributed spatiotemporal cooperative control provided in any of the above embodiments.
[0089] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by hardware related to computer program instructions. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0090] Matters not covered in this invention are common knowledge.
[0091] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0092] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these modifications and improvements all fall within the protection scope of this application.
[0093] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. An optimization method for a photovoltaic-storage-direct-drive-flexible system based on distributed spatiotemporal cooperative control, characterized in that, include: Step 110: Perform data collection, including acquiring historical load data; A load prediction model was constructed using a graph neural network. The model parameters are optimized using an adaptive moment estimation algorithm to obtain the initial load prediction model; Step 120: Based on the initial load forecasting model, use federated Kalman filtering to fuse local information from edge nodes for long-term load forecasting and generate global load forecasting data. Step 130: Based on global load forecast data, construct a continuous load function using adaptive cubic spline interpolation; and use density clustering algorithm to divide the continuous load function into stages. Step 140: Multi-agent deep deterministic policy gradient algorithm is used to optimize multi-agent cooperative scheduling, including: defining the heterogeneous state space and hybrid action space of multi-agents based on the continuous load function of stage division; setting up multi-agent experience replay pool and cooperative evaluator; and designing a dynamic adjustment mechanism that includes local rewards and global cooperative rewards. Step 150: Evaluate the optimization effect by combining multiple indicators; the indicators include at least: load variance reduction rate, cross-regional collaborative efficiency, photovoltaic absorption rate, energy storage health degradation rate, and privacy protection level. Step 160: Perform edge-cloud collaborative verification: upload the action execution records of edge nodes in real time; verify global consistency in the cloud through a digital twin model and send correction instructions for abnormal actions; dynamically adjust node trust and reward allocation coefficients.
2. The optimization method for a photovoltaic-storage-direct-drive-flexible system based on distributed spatiotemporal cooperative control according to claim 1, characterized in that, Step 110 involves data collection, which also includes acquiring meteorological data and seasonal / periodic information. The load forecasting model includes: Node feature vectors of a graph neural network: ; in, Represents a node The node feature vectors, ; It is the total number of nodes in the graph structure of a graph neural network; For nodes The average load, This is a temperature-sensitive parameter. The heat transfer coefficient of the building envelope. Encoding the operating cycle of system equipment; For weekdays, use a dummy variable. weekend The nodes in the graph structure Sub-regions in the corresponding target area ; Graph convolutional layer propagation rules for node feature vectors: ; in, Indicates the first Nodes in a layered graph convolutional layer Node feature vectors; For nodes The set of neighboring nodes, For nodes The set of neighboring nodes, ; Represents the set of neighboring nodes The base number, Represents the set of neighboring nodes The cardinality; For the first The weight matrix of the layered graph convolutional layer, It is the first convolutional layer in the graph. Subregion of the layer Node feature vectors; It is the ReLU activation function. For bias terms; Indicates the first Nodes in a layered graph convolutional layer The node feature vectors.
3. The optimization method for a photovoltaic-storage-direct-flexible system based on distributed spatiotemporal cooperative control according to claim 2, characterized in that, In step 110, the step of optimizing the model parameters using an adaptive moment estimation algorithm to obtain the initial load prediction model includes: Time variables, temperature variables, seasonal / periodic dummy variables, and building characteristic parameters are used as inputs to the load forecasting model; When optimizing model parameters using the adaptive moment estimation algorithm, the following update rule is designed: Calculate the gradient: ; in, For model parameters No. The gradient of the next iteration; The load forecasting error loss function is defined as follows: the mean square error of the model parameters is selected as the load forecasting error loss function; the model parameters... express Any one of the parameters in; Update the first-order moment estimate: ; in, For the first First-moment estimation of the gradient in the next iteration; It is the first-order attenuation coefficient; Update the second-order moment estimate: ; in, For the first Second-moment estimation of the gradient in the next iteration; It is the second-order attenuation coefficient; Deviation correction: ; in, and These are the corrected first and second moments, respectively; Update model parameters: ; in, Indicates the first Model parameters in the next iteration The possible values of ; It is the learning rate; It is a numerically stable term.
4. The optimization method for a photovoltaic-storage-direct-flexible system based on distributed spatiotemporal cooperative control according to claim 2, characterized in that, The temperature sensitivity parameter was obtained through an improved quadratic polynomial fitting: ; in, It is a node At any moment Temperature sensitivity parameters; It is a node At any moment The dry bulb temperature; It is a node At any moment The dew point temperature; For nodes At any moment Real-time humidity; For nodes At any moment Light intensity, in units of Fit coefficients The following solution using least squares with L1 regularization is obtained: ; In the formula, It is a node At any moment Actual load measurement value It is a node At any moment The load estimate is used to fit intermediate calculated values of the temperature sensitivity parameters; It is a fitting coefficient index. ; This is the regularization coefficient, used to suppress overfitting.
5. The optimization method for a photovoltaic-storage-direct-drive-flexible system based on distributed spatiotemporal cooperative control according to claim 1, characterized in that, Step 120 includes: Step 121, define the local system state equation and measurement equation for the edge nodes, which are given by the following equations: ; ; in, Represents edge nodes At any moment The state vector, ; This is the total number of edge nodes; For edge nodes The state transition matrix is given by the following equation: ; In the above formula, It is the regional coupling coefficient matrix =( The elements in ) represent edge nodes. With nodes The coupling coefficient; For edge nodes The set of neighboring nodes, For edge nodes With nodes Adjacency relationship, It is After vectorization, find the diagonal matrix, where 1 is the dimension. A vector of all 1s; This indicates the creation of diagonal elements. diagonal matrix, It is an edge node Its own state transition coefficient; It is an edge node At any moment The value of the input temperature variable; The control matrix for the temperature variable; It is process noise, with a mean of 0 and a covariance matrix of... Gaussian distribution ; Represents edge nodes At any moment The measured value; It is a measurement matrix; It is measurement noise, which follows a mean of 0 and a covariance matrix of... Gaussian distribution ; Step 122: Based on the local system state equation and measurement equation, local prediction and update are performed at the edge nodes using local Kalman filtering to obtain the state estimate of the edge nodes; Step 123: Using the fusion learning operation of federated Kalman filtering, a global fusion equation is defined in the cloud to generate global load prediction data; the weighted average coefficient of the global fusion equation is the trust weight of the edge nodes; the trust weight of the edge nodes is updated in real time based on the local prediction performance. The global fusion equation is: ; in, It is the cloud that is always there. Global state estimation, used to collect data from all edge nodes in the cloud. State estimation , ; For edge nodes Trust weighting; Update edge nodes using mean absolute percentage error The trust weight is updated according to the following rules: ; in, This is the sensitivity coefficient. , For edge nodes exist Mean absolute percentage error at any given time.
6. The optimization method for a photovoltaic-storage-direct-drive-flexible system based on distributed spatiotemporal cooperative control according to claim 5, characterized in that, In step 120, the local prediction and update are performed at the edge nodes using local Kalman filtering. The noise covariance update rule used is as follows: edge nodes At any moment Update the covariance matrix of the process noise as follows Covariance matrix of measurement noise : ; ; in, For a moment The covariance matrix of the process noise. For a moment The covariance matrix of the measurement noise; It is an edge node At any moment State estimation; When a sudden temperature change is detected, the covariance matrix of the measurement noise is temporarily adjusted. ,make It decays linearly to its initial value after a preset duration; The temperature abrupt change refers to the rate of change of a temperature variable per unit time. satisfy: ; in, It is a preset temperature change rate threshold.
7. The optimization method for a photovoltaic-storage-direct-drive-flexible system based on distributed spatiotemporal cooperative control according to claim 1, characterized in that, In step 130, the step of using density clustering algorithm to divide the continuous load function into stages includes using density clustering algorithm to divide the continuous load function into load curves for high load, medium load, and low load stages: Sampling is performed at fixed time steps on the curve of the continuous load function to obtain sample points for clustering. The total number of sample points is [number missing]. ; Develop the continuous load function partitioning rules for the density clustering algorithm, including: Define sample points of Neighborhood density: ; in, Represents sample points of The number of sample points contained in the neighborhood; It is the neighborhood radius; It is an indicator function used to statistically analyze sample points. of The number of sample points in the neighborhood, when the condition in parentheses The value is 1 if the condition is true, and 0 otherwise. Sample points With sample points The weighted Euclidean distance between them is given by the following formula: ; In the above formula, Sample points The load function value, This represents the load change over one hour. For sample points The load change during the corresponding time period; Sample points The load function value, For sample points The load change during the corresponding time period; Thresholds are categorized by stage: During high-load phases, the following conditions must be met: ,and The continuous crossing time shall not be less than 1 hour; During the medium load phase, the following conditions must be met: ,and ; The following conditions must be met during the low-load phase: ,and ; in, This is the global load average. The standard deviation of the global load. This represents the average density.
8. The optimization method for a photovoltaic-storage-direct-drive-flexible system based on distributed spatiotemporal cooperative control according to claim 1, characterized in that, Step 140 includes: Step 141, define the multi-agent heterogeneous state space based on the continuous load function of stage partitioning, including: Photovoltaic intelligent agent status: ; in, This is a real-time measurement of light intensity. To predict light intensity, For prediction error, The rate of change of illumination reflects the fluctuation of illumination. Energy storage agent status: ; in, For a moment The state of charge (SOC) indicates the percentage of remaining battery capacity. For a moment ; health status; It is the maximum charging and discharging power; This is the charge / discharge efficiency attenuation coefficient; It is a moment The charging power, It is a moment The discharge power, ; It is the rated capacity; Load agent state: ; in, Indicates time The load function value, This demonstrates the potential for flexible load regulation. It is the load change rate, reflecting load fluctuation; This is the real-time temperature sensitivity coefficient; Step 142: Define a hybrid action space, including discrete and continuous actions, to cover various scheduling optimization operations; The discrete actions include at least the selection of grid power purchase, direct photovoltaic power supply, and energy storage charging and discharging modes; The continuous operation includes at least the energy storage charging and discharging power, continuous adjustment, and load shedding ratio; Step 143: The multi-agent deep deterministic policy gradient algorithm is used to train the cooperative control policy of the multi-agent agents, including: setting up a multi-agent experience replay pool and a cooperative evaluator, and designing a dynamic adjustment mechanism that includes local rewards and global cooperative rewards.
9. The optimization method for a photovoltaic-storage-direct-drive-flexible system based on distributed spatiotemporal cooperative control according to claim 8, characterized in that, In the multi-agent deep deterministic policy gradient algorithm: The reward function of the dynamic adjustment mechanism includes at least: local reward. and global collaborative rewards ; The local reward According to the intelligent agent The values vary depending on the role: If the intelligent agent The role is that of a photovoltaic intelligent agent. Rewards for photovoltaic intelligent agents: ; in, This represents the per-unit value of the actual photovoltaic output. If the intelligent agent Its role is that of an energy storage intelligent agent. Rewards for energy storage agents: ; If the intelligent agent The role is that of a load-bearing intelligent agent. Rewards for the agent with high workload: ; in, This is the maximum load limit; The global collaborative reward It is given by the following formula: ; in, The power limit for the connecting line is exceeded; To improve cross-domain collaboration efficiency, For power transmission of inter-area tie lines; For photovoltaic power absorption rate, This refers to the actual amount of photovoltaic power consumed by the system. This represents the total photovoltaic power generation. For privacy protection; the aforementioned cross-domain collaboration efficiency Photovoltaic absorption rate and privacy protection Used in relation to load variance reduction rate and energy storage health degradation rate Together, we will construct a multi-indicator system for optimizing and evaluating photovoltaic-storage direct-drive-flexible systems; The training parameters of the multi-agent deep deterministic policy gradient algorithm include at least: experience replay pool capacity, target network update period, learning rate, and discount factor.
10. The method according to claim 1, characterized in that, In step 160, the verification triggering rules for the digital twin model include: Calculate edge nodes Digital twin bias : ; in, Represents edge nodes At any moment The simulated values, Represents edge nodes At any moment The actual value; like If the value exceeds a preset threshold and the duration exceeds a preset period, a correction command for dynamic adjustment will be sent from the cloud. ; in, This is the motion correction amount; The dynamic adjustment formula for the node trust level is: ; in, It is an edge node The node's trust level; The dynamic adjustment formula for the reward allocation coefficient is as follows: ; in, It is an edge node At any moment The reward distribution coefficient, To reward sensitivity coefficient, This is the baseline reward value.