Zone area light storage and charging coordinated regulation and control system and autonomous method based on side end self-control

By employing edge-controlled photovoltaic-storage-charging collaborative regulation and control system, lightweight reinforcement learning and federated learning technologies are used to solve the problems of real-time response, privacy protection and physical security integration of photovoltaic-storage-charging system, thus achieving efficient and safe photovoltaic-storage-charging collaborative regulation and control.

CN121965627APending Publication Date: 2026-05-01SICHUAN SIJI TECHNOLOGY CO LTD

Patent Information

Application Number
CN202610416329.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-01
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technologies in photovoltaic, energy storage, and charging systems in power distribution areas suffer from insufficient real-time response and edge autonomy capabilities, contradictions between data privacy protection and collaborative optimization, difficulties in system cold start, and a lack of deep integration of physical security constraints. These issues result in slow system response to rapid disturbances, high risk of privacy leaks, long deployment times, and insufficient equipment security.

Method used

A photovoltaic-storage-charging collaborative control system based on edge-end self-control is constructed, including end-level autonomous units, edge autonomous nodes, and substation-level federated learning centers. Lightweight reinforcement learning, federated learning, secure multi-party computation, and digital twin technologies are adopted to achieve deep integration of millisecond-level reflective control, second-level optimization, personalized collaboration, and physical security constraints.

Benefits of technology

It enhances the system's real-time response capability and autonomy, protects user data privacy, solves the cold start problem of newly commissioned distribution areas, and ensures the unity of AI decision-making and physical security constraints, thus achieving efficient and safe coordinated control of photovoltaic, energy storage, and charging distribution areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121965627A_ABST
    Figure CN121965627A_ABST
Patent Text Reader

Abstract

The invention discloses an end-end self-control-based transformer area light storage and charging coordinated regulation and control system and an autonomous method. The system constructs an end-end-transformer three-level autonomous architecture: an end-level autonomous unit is responsible for local millisecond-level reflection control; the edge autonomous nodes realize intra-zone collaboration through federated learning, and physical security constraints are embedded into AI decisions by using digital twin bodies; the court-level federation learning center realizes knowledge sharing under cross-court privacy protection through secure multi-party computing, the method further integrates transfer learning to solve the cold start problem of a new court, and a three-time scale regulation and control mechanism is adopted to cope with disturbance of different rates; the autonomous capability, the response speed, the safety level and the intelligent degree of the system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power system operation and control technology, specifically to a photovoltaic-storage-charging coordinated control system and autonomous method for distribution transformer areas based on edge-end self-control. Background Technology

[0002] With the advancement of the dual-carbon strategy, the penetration rate of distributed photovoltaic, energy storage systems and electric vehicle charging facilities in distribution network areas continues to increase, forming a complex integrated photovoltaic-energy storage-charging system. How to coordinate these distributed resources to achieve efficient, stable and safe operation at the distribution network level has become a current research hotspot and practical challenge.

[0003] In the prior art, patent CN114977236A discloses a photovoltaic energy storage and charging system, storage medium and photovoltaic energy storage and charging station based on an energy router. It constructs an architecture based on an energy router and realizes energy distribution and management through multi-level cloud-edge collaborative control and multi-objective hierarchical optimization. This solution represents a control idea in the current field, that is, to perform centralized optimization scheduling of the system through the powerful computing capabilities of the cloud or regional center. However, this existing technology, as well as similar centralized or cloud-edge collaborative solutions, have gradually revealed several inherent technical defects and challenges when dealing with the specific scenario of a photovoltaic-storage-charging system in a distribution area: The system lacks real-time response and edge autonomy capabilities. Existing technologies rely on centralized computing in the cloud or energy routers for optimization decisions. When faced with rapid disturbances on the order of seconds or even milliseconds, such as instantaneous fluctuations in photovoltaic output or random start-stop of electric vehicles, this architecture cannot avoid delays caused by communication links because optimization instructions need to be calculated in the cloud before being sent to the edge for execution. At the same time, once communication with the cloud is interrupted, the intelligent decision-making capability of the entire system will be greatly reduced. The edge side lacks sufficient computing power and algorithm models for rapid and autonomous emergency control, and the overall robustness and reliability of the system are challenged. The conflict between data privacy protection and collaborative optimization: In order to achieve collaborative optimization of multiple distribution zones or multiple distributed units, existing solutions usually need to upload detailed operating data or complete model parameters of each unit to a central node, such as a cloud server or energy router, for processing. This approach not only brings huge communication bandwidth pressure, but more importantly, this data contains sensitive information such as users' electricity consumption behavior and geographical location. Its centralized storage and processing poses a risk of privacy leakage. How to achieve effective collaborative optimization and knowledge sharing while protecting the data privacy of all parties is a key problem that existing technologies have not been able to properly solve. Newly commissioned systems face the challenge of a cold start. The performance of data-driven intelligent control models, such as deep learning models, is highly dependent on training with a large amount of high-quality historical data. For a newly built or newly connected distribution area, due to the lack of local historical operating data, the intelligent control model cannot be effectively trained, resulting in a long period of low performance or no intelligence in the early stage of commissioning. It takes several months to accumulate data to gradually improve, which seriously affects the system's plug-and-play and rapid deployment capabilities. The disconnect between physical safety constraints and data-driven optimization models means that many existing intelligent optimization algorithms, including some cloud-edge collaborative solutions, focus primarily on performance indicators such as economy and power quality. However, they fail to deeply embed the key physical safety boundaries of the distribution area, especially the thermal safety constraints of transformers, into the training and real-time decision-making loop of the data-driven model in a computable and quantifiable manner. This may result in optimized strategies that perform well in terms of economy, but inadvertently exacerbate the aging of transformer insulation, such as long-term overheating of windings, which poses a hidden danger to the long-term safe operation of distribution area assets. In summary, existing technologies have significant shortcomings in terms of high real-time edge autonomy, collaborative learning under privacy protection, system cold start, and deep integration of physical information security. Therefore, there is an urgent need in this field for a new scheme for coordinated control of photovoltaic, energy storage, and charging in power distribution areas that can achieve rapid response, autonomous collaboration, and self-learning capabilities while ensuring data privacy and device security. Summary of the Invention

[0004] In order to overcome the problems of the prior art, the present invention discloses a photovoltaic-storage-charging coordinated control system and autonomous method based on edge-end self-control.

[0005] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows: The photovoltaic-storage-charging coordinated control system for distribution areas based on edge-end self-control includes: End-level autonomous units, deployed in photovoltaic inverters, energy storage converters and charging piles, are configured to achieve millisecond-level reflective control based on locally observed electrical state parameters; Edge autonomous nodes are deployed in the intelligent converged terminal of the distribution area and communicate with the end-level autonomous units. They are configured to aggregate the policies of each end-level autonomous unit based on federated learning and generate distribution area-level scheduling instructions. The district-level federated learning center communicates with edge autonomous nodes of multiple districts and is configured to build virtual federated learning centers through secure multi-party computation to exchange encrypted model gradients.

[0006] Preferably, the end-level autonomous unit includes: Lightweight reinforcement learning agents; The local status observation module is configured to collect voltage, power, and state of charge in real time. The reflection control module is configured to instantaneously trigger reactive power support when a voltage over-limit is detected.

[0007] Preferably, the edge autonomous node includes: The prediction model is configured to perform second-level time series prediction and optimization. The regional collaborative controller is configured to aggregate the policies of each end-level autonomous unit based on federated learning. The digital twin adopts a physical-data hybrid driving model, is configured to simulate the operating state of a transformer in real time, and transforms physical constraints into a reinforcement learning reward function penalty term.

[0008] Preferably, the digital twin includes: The physical model uses a transformer thermal circuit model. In the data-driven part, a time-delay neural network is used to learn the residuals.

[0009] Preferably, the district-level federated learning center adopts a personalized federated learning algorithm, where each district retains a personalized output layer and only shares the parameters of the bottom feature extraction layer; The edge autonomous nodes are equipped with a contribution assessment module, which is used to evaluate the data contribution of each participating area before the federated model is aggregated.

[0010] Preferably, the contribution assessment module is configured as follows: Based on the transformer load rate, photovoltaic power output fluctuation rate and electric vehicle charging demand time series similarity of each distribution area, a distribution area feature vector is constructed. Calculate the correlation weights between the feature vectors of the transformer area and the performance improvement of the global model; Based on the relevance weights, the weights of each substation in the federated learning model aggregation are dynamically adjusted.

[0011] Preferably, it further includes a transfer learning module, which includes: The source area knowledge base storage unit is configured to store historical data of typical transformer areas; The meta-learning model training unit communicates with the edge autonomous nodes and is configured to train the meta-learning model based on historical data. The feature mapping processor is configured to map heterogeneous features from different stations to a unified latent space.

[0012] Preferably, it further includes a three-time-scale control module, which includes: The prediction layer is configured to predict future photovoltaic power output and the probability of electric vehicles entering and leaving the site on a minute-level time scale; The correction layer is configured to correct prediction biases on a second-level time scale and uses a model predictive control method to continuously optimize and correct the prediction bias. The emergency layer is configured to trigger end-level agent reflection control on a millisecond timescale during sudden voltage changes.

[0013] Preferably, in the correction layer, model predictive control and rolling optimization objective function embed the joint uncertainty quantification results; The joint uncertainty is obtained by identifying the noise characteristics introduced by the intermittent output of distributed photovoltaic power and the random charging of electric vehicles in the electrical measurement data of the transformer area, and quantifying them using the random discarding method.

[0014] Preferably, a method for coordinated autonomous operation of photovoltaic, energy storage, and charging systems in a distribution area based on the above system includes: Millisecond-level reflective control is achieved through end-level autonomous units based on local observation states; By aggregating the strategies of each end-level autonomous unit through edge autonomous nodes based on federated learning, a dispatching instruction at the station level is generated. Through the district-level federated learning center, encrypted model gradients are exchanged between multiple districts based on secure multi-party computation. The transformer's operating state is simulated in real time using digital twins in edge autonomous nodes, and physical constraints are transformed into penalty terms in a reinforcement learning reward function. By using a district-level federated learning center and employing a personalized federated learning algorithm, each district retains its own output layer and shares only the parameters of the underlying feature extraction layer.

[0015] The beneficial effects of this invention are as follows: Compared with the prior art, the technical solution provided by the present invention has the following significant advantages: To enhance the system's real-time response capability and autonomy, a three-level autonomous architecture of end-edge-platform is constructed, which decentralizes intelligent decision-making. End-level units can achieve millisecond-level local reflective control, effectively suppressing rapid disturbances such as voltage surges. Edge nodes achieve second-level optimization, ensuring that the system can still maintain a high degree of autonomy and stable operation in abnormal situations such as communication interruptions. To achieve efficient collaboration and knowledge sharing under privacy protection, multiple stations can train collaboratively by combining federated learning and secure multi-party computation without sharing the original data, only exchanging encrypted model gradients. This fundamentally protects user data privacy, while personalized federated learning enables the model to both absorb global knowledge and adapt to local characteristics. To overcome the cold start problem of newly commissioned transformer substations, a transfer learning module is introduced. By utilizing the source substation knowledge base and meta-learning model, existing regulatory knowledge from related fields is quickly transferred to the new substation. This enables the substation to obtain effective initial regulatory strategies with only a small amount of new data, shortens the strategy convergence time, and achieves rapid and intelligent deployment of the system. To ensure the deep integration of AI decision-making and physical safety constraints, a lightweight digital twin is deployed to transform key physical constraints such as transformer thermal safety into penalty terms of a reinforcement learning reward function in a computable form. This guides the AI ​​agent to automatically meet the safety operation boundaries of equipment when exploring optimal strategies, thereby unifying optimization goals and safety goals and ensuring the safety of power grid assets. Attached Figure Description

[0016] Figure 1 This is a schematic diagram of the structure of the photovoltaic-storage-charging coordinated control system based on edge self-control provided in Embodiment 1 of the present invention; Figure 2 This diagram illustrates the steps of the autonomous method for the coordinated control system of photovoltaic, energy storage, and charging in a distribution area based on edge self-control, as provided in Embodiment 1 of the present invention. Detailed Implementation

[0017] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, exemplary embodiments will be described in detail below, examples of which are illustrated in the accompanying drawings. In the following description relating to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of methods and systems consistent with some aspects of this application as detailed in the appended claims.

[0018] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a” and “the” as used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0019] The following detailed description of the specific implementation methods, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided in detail.

[0020] Example 1 Please refer to Figure 1 This embodiment provides a photovoltaic-storage-charging coordinated control system for transformer substations based on edge-end self-control, including: End-level autonomous units, deployed in photovoltaic inverters, energy storage converters and charging piles, are configured to achieve millisecond-level reflective control based on locally observed electrical state parameters; Edge autonomous nodes are deployed in the intelligent converged terminal of the distribution area and communicate with the end-level autonomous units. They are configured to aggregate the policies of each end-level autonomous unit based on federated learning and generate distribution area-level scheduling instructions. The district-level federated learning center communicates with edge autonomous nodes of multiple districts and is configured to build virtual federated learning centers through secure multi-party computation to exchange encrypted model gradients.

[0021] Furthermore, end-level autonomous units include: In practical implementation, lightweight reinforcement learning agents can preferably adopt algorithm structures based on Deep Q-Network (DQN), but the network size is specially trimmed. In one embodiment, the lightweight reinforcement learning agent consists of an input layer, a hidden layer containing 64 neurons, and an output layer, with the total number of parameters controlled to within 80KB. This design can be embedded on the main control chip of photovoltaic inverters, energy storage converters or charging pile controllers, such as the ARM Cortex-M7 core, to achieve policy inference without relying on cloud computing power. The local status monitoring module is implemented through the collaboration of sensors deployed on the end device and built-in software; Voltage observation involves real-time acquisition of the effective voltage value at the grid connection point via a voltage transformer and an analog-to-digital converter (ADC). Power observation involves using current transformers and voltage sampling values, and then employing multiplication and low-pass filtering algorithms to calculate real-time active and reactive power. State of charge (SOC) observation is a feature specifically designed for the end-level autonomous unit of the energy storage converter. This module acquires and updates the SOC value of the battery in real time through communication with the battery management system (BMS) using either the coulomb counting method or a combination of Kalman filtering. This module assembles the above observation data into a state vector, and periodically provides it to the lightweight reinforcement learning agent as decision input, such as every 100 milliseconds. The reflection control module is implemented as a high-priority interrupt service routine, and its workflow is as follows: The detection and reflection control module continuously monitors the voltage value provided by the local status observation module; The system determines whether to trigger the control logic instantaneously when the detected voltage value exceeds the preset safety limit, such as ±10% of the rated voltage. Trigger reactive power support; the triggering action is specific to the type of end device. In photovoltaic inverters, this manifests as the immediate execution of control commands, rapidly injecting or absorbing a preset capacity of reactive power depending on the direction of the voltage exceeding the limit, whether it is too high or too low. In energy storage converters, this manifests as adjusting the reactive power output command within milliseconds, prioritizing reactive power support, and temporarily suspending active power scheduling. In charging piles, this manifests as adjusting the internal power factor correction (PFC) circuit to control the reactive power absorbed from or sent to the grid during the charging process. For charging piles with bidirectional power flow capability, this control is also applicable to scheduling electric vehicles to generate or absorb reactive power. This reflective control process is completed entirely in a closed loop on the end device, ensuring that even in abnormal situations such as communication interruption with edge nodes, the most basic power quality guarantee can still be provided. Specifically, by integrating the lightweight reinforcement learning agent, the local state observation module, and the reflection control module into the end device hardware, a collaborative whole with autonomous perception, decision-making, and execution capabilities is formed. This effectively solves the inherent defects of traditional centralized control, such as slow response and reliance on communication, and enables millisecond-level reflective local control of key electrical parameters such as transformer voltage, thereby improving the system's autonomy and operational reliability.

[0022] Furthermore, edge autonomous nodes include: The prediction model is configured to perform second-level time series prediction and optimization. In practical implementation, the prediction model can be deployed on edge autonomous nodes, such as within the computing unit of a smart converged terminal in a distribution area. The prediction model preferably adopts a lightweight long short-term memory network structure. The input is the total load of the transformer area, the total output of photovoltaic power and the total charging power of electric vehicles in the past time series. The model predicts the time series change trend of the above key parameters in the next short period of time, such as the next 30 seconds to 5 minutes, with a period of seconds. Based on this prediction, the module performs optimization calculations with the goal of smoothing the net load curve of the distribution area and suppressing transformer overload. The output is a pre-scheduling instruction for photovoltaic, energy storage and charging equipment in subsequent periods, i.e., the end-level autonomous unit, such as the reference value of the charging and discharging power of the energy storage system. The regional collaborative controller is the core software module for realizing district-level collaborative decision-making. It is configured to aggregate the policies of each autonomous unit at the terminal level based on federated learning, and its workflow is as follows: The regional collaborative controller periodically receives local control strategies or status information uploaded from the end-level autonomous units of various photovoltaic inverters, energy storage converters and smart charging piles through local communication networks, such as HPLC or low-power wireless. The regional collaborative controller uses the federated averaging algorithm as the specific implementation of federated learning. Instead of collecting the raw data from each end, the regional collaborative controller performs a weighted average calculation of the model parameters of multiple end-level autonomous units, such as the weights and biases of the neural network, to generate a collaborative control model with stronger globality and better generalization ability. The updated model obtained after aggregation, or the optimized scheduling instructions calculated based on the model, are sent to each end-level autonomous unit within the distribution area for execution, thereby realizing the collaborative operation of distributed resources; The digital twin adopts a physical-data hybrid driving model, is configured to simulate the operating state of a transformer in real time, and transforms physical constraints into a reinforcement learning reward function penalty term; The physical model uses the thermal circuit model of a transformer, which is the basic simulation engine for the digital twin. In practical implementation, the physical model part adopts the equivalent thermal circuit model of the transformer to simulate its temperature rise dynamics. This model compares the thermal characteristics of the transformer to those of a circuit. The model simulates the thermal dynamics of a transformer using thermal resistance and thermal capacity parameters. Thermal resistance is analogous to resistance in a circuit, representing the obstruction to heat transfer. Thermal capacity is analogous to capacitance in a circuit, representing the ability of a transformer component to store heat; The model takes real-time operating data obtained through edge autonomous nodes as input, which mainly includes the transformer's load current and ambient temperature. Based on these inputs, the model calculates the predicted temperature of the transformer winding hotspots by solving the differential equations corresponding to the thermal circuit model. This physical model provides the temperature change trend under standard operating conditions, forming a benchmark for the digital twin simulation results. The data-driven part is used to correct the inherent biases in the physical model part, and it uses a lightweight feedforward neural network. The input to the lightweight feedforward neural network is a time series window, and its data is the prediction residual of the physical model part, that is, the difference between the predicted hot spot temperature calculated by the physical model and the actual measured hot spot temperature. Lightweight feedforward neural networks are trained to learn the complex dynamic relationships in these residual sequences that are not revealed by physical models, such as errors caused by factors such as changes in cooling efficiency and nonlinear internal losses. The network output is the real-time correction to the physical model's prediction results. Ultimately, the high-precision simulation results output by the digital twin are: the physical model's predicted values ​​plus the corrections provided by the data-driven part. The digital twin drives the hybrid model to run by reading data such as the total power of the transformer area in real time, thereby simulating the transformer's operating status, especially its hot spot temperature, with high accuracy in real time. To achieve safety constraints, the twin transforms physical safety constraints, such as the winding hot spot temperature not exceeding the safety limit that the insulation material can withstand, into a penalty term in the reinforcement learning reward function. Specifically, when training or running end-level or edge-level reinforcement learning agents, the reward function R is designed to introduce a penalty term that is positively correlated with the degree of temperature exceedance:

[0023] In this formula, R pPerformance rewards related to the economic efficiency and stability of system operation; T h Predict hot spot temperatures of transformer windings in real time for digital twin calculations; T s This is the preset upper limit of the safe temperature for transformer windings; λ is a penalty coefficient greater than zero, used to adjust the severity of the penalty for exceeding temperature limits; max(0,T h -T s The function ensures that a penalty is only incurred if the predicted temperature exceeds the safety limit; When the predicted temperature approaches or exceeds the safety limit, this penalty term will increase, guiding the reinforcement learning agent to explore strategies that can both achieve the optimization goal and ensure the safety of the transformer, such as actively reducing the total load or adjusting the equipment operation mode when the temperature is too high. Specifically, edge autonomous nodes achieve rapid optimization and control through forward-looking decision-making by integrating predictive models, distributed intelligent fusion of regional collaborative controllers, and physical security embedding of digital twins. By deeply integrating precise physical simulation with artificial intelligence learning processes, they ensure the safe operation of power equipment and improve the autonomy and intelligence level of the distribution area energy system.

[0024] Furthermore, the district-level federated learning center adopts a personalized federated learning algorithm, where each district retains a personalized output layer and only shares the parameters of the bottom feature extraction layer; In practice, the personalized federated learning algorithm used by the district-level federated learning center is specifically designed for the neural network model structure of the multiple participating districts. The local model of each district, such as the deep learning network used for photovoltaic power output prediction, is divided into two parts: The bottom feature extraction layer, which is shared by all substations participating in federated learning, is responsible for learning and extracting general high-level features from raw, heterogeneous local data, such as voltage sequences and weather information. The personalized output layer is owned and retained independently by each station area and does not participate in sharing or aggregation. Its function is to combine the general features output by the shared feature extraction layer with the unique load characteristics, equipment parameters and operating topology of this distribution area to make the final decision output, such as generating specific control instructions for this distribution area. During the federated learning process, the federated learning center only collects the model parameter updates of the bottom feature extraction layer of each substation, aggregates them through the federated averaging algorithm, generates an improved global feature extraction layer, and then distributes it to each substation. The personalized output layer of each station area is independently trained and fine-tuned using local data to maintain its specificity. This solution effectively solves the problem that a single model is difficult to adapt to all scenarios due to the huge differences in load composition and resource endowment between different stations. The edge autonomous nodes are equipped with a contribution evaluation module, which is used to evaluate the data contribution of each participating area before the federated model is aggregated; The contribution assessment module is a software functional unit integrated within the edge autonomous node of each substation; it is triggered before the start of each federated model aggregation cycle and is configured to execute the following core processes: The data value assessment and contribution assessment module first evaluates the potential value of the local data used by its own substation in this round of training for improving the performance of the global model; In one implementation, this can be achieved by calculating the performance improvement of the current local model update compared to the previous global model on a batch of public benchmark data; The module quantifies the assessed data value into specific contribution scores. In one embodiment, a weighted method based on model performance improvement can be used, and its quantification formula can be expressed as: Contribution score = (Accuracy of the current model - Accuracy of the previous model) × Local dataset size; Here, the accuracy of the model in this round refers to the performance metric obtained by evaluating the model trained locally in this round on the benchmark dataset; The previous round model accuracy refers to the performance metric obtained by evaluating the global model obtained from the previous round of aggregation on the same benchmark dataset; The local dataset size represents the number of local data samples used in this round of training; Finally, the module uploads the calculated contribution score and the updated model parameters of the local underlying feature extraction layer to the regional federated learning center. After receiving information from all the stations, the Federated Learning Center will perform a weighted average of the model parameter updates based on the contribution scores reported by each station. Stations with higher contribution scores will receive greater weight in the aggregation. Furthermore, the contribution assessment module is configured as follows: Based on the transformer load rate, photovoltaic power output fluctuation rate and electric vehicle charging demand time series similarity of each distribution area, a distribution area feature vector is constructed. The contribution evaluation module is configured to first construct a feature vector characterizing the operational characteristics of each participating zone in federated learning. This module obtains real-time and historical operational data of its zone through edge autonomous nodes and performs the following calculations: Transformer load factor is calculated as the ratio of the total active power of the transformer substation in real time to the rated capacity of the transformer. Photovoltaic power output volatility, within an assessment period such as 15 minutes, is calculated as the ratio of the difference between the maximum and minimum photovoltaic power output to the average power output during that period. Electric vehicle charging demand time series similarity: This module extracts the electric vehicle charging power sequence of the local area and adjacent areas in the same historical time period. By calculating the dynamic time warping distance between the sequence of the local area and the sequences of other areas, and taking its reciprocal as the similarity measure, the similarities and differences in charging behavior patterns are quantified. By combining the above three scalar parameters in a predetermined order, a unique feature vector representing the current operating characteristics of this transformer area is constructed. Calculate the correlation weights between the feature vectors of the transformer area and the performance improvement of the global model; The contribution assessment module is further configured to calculate the correlation between the feature vector of the transformer area and the model performance; the specific steps include: The performance improvement is defined as the improvement in accuracy of the global model on the public validation dataset after a new round of federated aggregation compared to the previous round of model. Correlation calculation: This module uses statistical methods, such as Pearson correlation coefficient, to analyze the relationship between the feature vectors of each region and their contribution to the performance improvement of the global model in the current round of historical multi-round federated learning. The stronger the positive correlation between a dimension and the performance improvement in the feature vector, the higher its importance is assigned in the subsequent weight calculation, thus calculating the correlation weight. Based on the relevance weights, the weights of each substation in the federated learning model aggregation are dynamically adjusted; Finally, the contribution evaluation module applies the calculated relevance weights to the model aggregation stage of federated learning. When executing the weighted average aggregation algorithm at the district-level federated learning centers, the weights of the model parameters in each district no longer depend solely on the size of their data, but are dynamically adjusted by the relevance weights. The aggregation weight calculation formula can be expressed as: Aggregate weight = base weight × relevance weight; The basic weight is usually the proportion of the local dataset size in that area to the total dataset size; The correlation weight is an adjustment coefficient provided by this module that reflects the correlation between the characteristics of the transformer area and the improvement of model performance; Through this mechanism, regions whose operational characteristics are represented by feature vectors, which are more conducive to improving the performance of the global model, will have a greater say in federated learning. Specifically, through structural innovation in personalized federated learning algorithms, a balance is achieved between global knowledge sharing and preservation of local characteristics. At the same time, through dynamic weight allocation in the contribution evaluation module, high-quality data is incentivized and selected to participate in federated modeling. The two work together to solve the industry problem of how to train a general and accurate transformer area control model using distributed heterogeneous data under the constraint of data privacy protection.

[0025] Furthermore, it also includes a transfer learning module, which includes: The source area knowledge base storage unit is configured to store historical data of typical transformer areas; In practice, the source area knowledge base storage unit can be deployed on a cloud platform or regional data center, and its configuration is to systematically store historical operating data from multiple typical stations that have been operating stably. These data are categorized and stored according to transformer substation type, primarily including industrial, commercial, and residential substations. The stored data dimensions cover photovoltaic output time-series data, load curve data, electric vehicle charging behavior data, transformer operating parameters, and meteorological correlation data. The source platform knowledge base storage unit provides standardized, high-quality data support for subsequent meta-learning model training by establishing a unified data indexing mechanism; The meta-learning model training unit communicates with the edge autonomous nodes and is configured to train the meta-learning model based on historical data. The meta-learning model training unit establishes a communication connection with the edge autonomous nodes via a power grid wireless private network or fiber optic network. This unit is configured to execute a model-independent meta-learning algorithm, and its training process consists of two phases: In the meta-training phase, using multi-type transformer area data stored in the source transformer area knowledge base storage unit, the model is trained through multiple simulation tasks to obtain initial parameters that can quickly adapt to new tasks. Specifically, each task corresponds to the regulation problem of a specific transformer area, and the goal of model learning is to achieve rapid fitting of the regulation strategy for that transformer area within a few iteration steps. During the model distribution phase, when a newly commissioned transformer substation is connected to the system, the trained meta-learning model is distributed to the edge autonomous node corresponding to the new substation via a communication connection, serving as the initial network parameters for the local reinforcement learning agent. The feature mapping processor is configured to map heterogeneous features from different stations to a unified latent space. The feature mapping processor is deployed in edge autonomous nodes and configured to unify the feature space through deep neural networks; Input processing receives heterogeneous features from local observations, including but not limited to raw features of different dimensions such as load characteristic curves, photovoltaic installed capacity, and electric vehicle penetration rate; Spatial mapping, through a neural network with an encoder structure, transforms the aforementioned heterogeneous features into a unified, low-dimensional latent space representation; Within this potential space, samples with similar operating characteristics from different stations are close to each other, while samples with large differences in characteristics are far apart. To minimize domain differences, the feature map processor employs the maximum mean difference criterion as one of its optimization objectives during training. This aims to reduce the distributional differences between the source and target areas in the latent space. The objective function can be expressed as:

[0026] Among them, L task This is the main loss function of the model on the source area task; MMD(X s ,X t To measure the source area sample set X s Sample set X of the target station area t A measure of the distributional differences in the latent space; λ is a hyperparameter that balances the two optimization objectives; Specifically, the transfer learning module provides a knowledge base through the source area knowledge base storage unit, extracts general control prior knowledge that can be quickly transferred through the meta-learning model training unit, and finally solves the problem of feature heterogeneity between different areas with the help of the feature mapping processor. The three work together to form a complete technical path, effectively overcoming the industry bottleneck of cold start faced by newly commissioned areas due to the lack of historical data, and shortening the convergence time of the control strategy from several months in traditional methods to weekly.

[0027] Furthermore, it also includes a three-time-scale control module, which includes: The prediction layer is configured to predict future photovoltaic power output and the probability of electric vehicles entering and leaving the site on a minute-level time scale; The prediction layer is deployed in edge autonomous nodes and configured to perform minute-level advance prediction functions. Specifically, in implementation: Data collection and prediction are carried out periodically, such as every 5 minutes, to collect historical operating data of the transformer area, including photovoltaic output curves, electric vehicle charging records, weather forecast information and transformer area load data. Predictive execution employs lightweight time-series prediction algorithms, such as attention-based recurrent neural networks, using data from a previous time period, such as 4 hours, as input to predict future time periods, such as photovoltaic power output and electric vehicle entry / exit probabilities within the next 2 hours. The output is passed to the next layer, the correction layer, with a minute-level time resolution, providing forward-looking guidance for optimizing control. The correction layer is configured to correct prediction biases on a second-level time scale and uses a model predictive control method to continuously optimize and correct the prediction bias. The correction layer is also deployed on edge autonomous nodes and configured to perform real-time optimization control at the second level. The model predictive control framework, with a correction layer establishing a rolling optimization framework in seconds, optimizes within each control cycle, such as 10 seconds: A simplified dynamic model of the transformer area is adopted to predict the system behavior in the short term, such as the next 5 minutes, based on the current system state and the long-term prediction provided by the prediction layer. Online solution to finite-time optimal control problems, whose objective function comprehensively considers multiple objectives such as tracking the planned curve, suppressing voltage fluctuations, and reducing network losses; The first control variable of the optimized sequence, such as the energy storage output set value or the charging pile power adjustment command, is sent to the execution device, and then optimized again in the next cycle based on the actual measurement value to form a closed-loop control. When the actual operating data deviates from the prediction results of the prediction layer, the layer automatically corrects the control strategy through a rolling optimization mechanism to ensure that the system always operates in an optimized state. The emergency layer is configured to trigger end-level agent reflection control on a millisecond timescale during sudden voltage changes. The emergency layer is the core defense line for ensuring system security, and is configured as follows: The voltage waveform of key nodes in the transformer area is monitored in real time through a high-speed sampling circuit with a sampling rate of no less than 10kHz. Using hardware logic based on a field-programmable gate array (FPGA), when power quality events such as voltage spikes, drops, or momentary interruptions are detected, event identification and classification can be completed within milliseconds, such as less than 5 milliseconds. Once an emergency is identified, the system immediately bypasses the conventional decision-making process of the prediction and correction layers and directly sends predefined safety control commands to the relevant end-level autonomous units, such as photovoltaic inverters, energy storage converters, and charging piles, to trigger their local protection control modes and achieve the fastest safety response. Specifically, the three-timescale control module constructs a complete control system that balances economic optimization and safety and stability through minute-level forward planning in the prediction layer, second-level rolling optimization in the correction layer, and millisecond-level reflective control in the emergency layer. This multi-timescale collaborative architecture effectively solves the technical problem that single-timescale control cannot simultaneously balance long-term optimization and instantaneous safety, ensuring the economical operation of the system while ensuring rapid protection in the event of emergencies, and improving the overall operational performance of the photovoltaic-storage-charging system in the distribution area.

[0028] Furthermore, the model predictive control in the correction layer incorporates the joint uncertainty quantification results into the rolling optimization objective function; The model predictive control implemented in the correction layer enhances the design of the optimization objective function. Specifically, the rolling optimization objective function, in addition to the traditional consideration of tracking error and control costs, adds an explicit consideration of system uncertainties. The standard form of this objective function can be extended as follows:

[0029] Where J is the total value of the objective function to be minimized; To predict the time domain from the first future time to the Nth future time... p Summing all points at time N, N p To predict the length of the time domain; y(k+i) is the system's predicted output at the i-th time in the future; r(k+i) is the desired reference trajectory; To control the time domain from the current time k to the Nth time... c Summing all points at time -1, N c To control the length of the time domain; u(k+i) is the control variable that needs to be solved; UQ(k+i) represents the quantification result of the joint uncertainty at the i-th time in the future; η and γ are the weighting coefficients for control costs and uncertainty penalties, respectively; By introducing a penalty term containing η·UQ(k+i), the optimization algorithm will automatically tend to select robust control strategies that are insensitive to uncertainty when solving the problem. The joint uncertainty is obtained by identifying the noise characteristics introduced by the intermittent output of distributed photovoltaic power and the random charging of electric vehicles in the electrical measurement data of the transformer area, and quantifying them using the random discarding method; The quantification of joint uncertainty is achieved through the following steps: The correction layer is equipped with a characteristic identification unit, which identifies and establishes a stochastic model introduced by distributed photovoltaic and electric vehicles by analyzing the historical sequence of electrical measurement data of the transformer area. Photovoltaic power output is intermittent, and its randomness is manifested in the rapid fluctuation of the theoretical value of the output power under clear sky conditions. This characteristic is quantified by analyzing the statistical features of the minute-level fluctuation rate of photovoltaic power output. Random charging of electric vehicles is characterized by sudden increases and decreases in charging power and uncertainty in charging time. This characteristic is described by establishing a probability distribution model of charging start time and charging time. After obtaining the current system state, the correction layer uses a random discarding method to quantify uncertainty: The current state is input into the neural network model of the prediction layer for forward propagation calculation multiple times, such as 100 times. In each forward propagation process, some neurons in the neural network are randomly shut down according to a preset probability, thereby introducing randomness into the model and simulating potential random disturbances in the system. The prediction results obtained from multiple forward propagations, such as future photovoltaic power output and load values, are statistically analyzed, and their variance or confidence interval is calculated. This statistic is the quantitative value of the joint uncertainty UQ(k+i), which directly characterizes the degree of uncertainty of the future system behavior given the current state. Specifically, by embedding joint uncertainty quantization results into the rolling optimization of model predictive control, the traditional optimization based on deterministic prediction is transformed into robust optimization that considers stochasticity. This method identifies specific noise characteristics and uses a random dropout method for online quantization, enabling the control system to anticipate and proactively adapt to uncertainties brought about by new energy sources and random loads. Thus, while maintaining optimization performance, it improves the control robustness and operational reliability of the system when facing real-world environmental fluctuations.

[0030] Please refer to Figure 2 Furthermore, this embodiment also provides a method for coordinated autonomous operation of photovoltaic, energy storage, and charging systems in a distribution area based on the above system, including: Millisecond-level reflective control is achieved through end-level autonomous units based on local observation states; The method steps are executed by end-level autonomous units deployed in photovoltaic inverters, energy storage converters, and charging piles. The specific implementation process includes: State observation: Each terminal autonomous unit collects voltage, power and state of charge in real time through local sensors; The reflection detection mechanism triggers a control response within 5 milliseconds when the detected voltage value exceeds ±10% of the rated voltage. Local control involves the photovoltaic inverter adjusting reactive power output, the energy storage converter adjusting charging and discharging power, and the charging pile adjusting charging power, all working together to achieve voltage support. By aggregating the strategies of each end-level autonomous unit through edge autonomous nodes based on federated learning, a dispatching instruction at the station level is generated. This method involves steps executed by edge autonomous nodes deployed in the smart converged terminal of the distribution area: Policy collection: Edge autonomous nodes collect local control policies of each end-level autonomous unit through power line carrier or wireless communication; Federated aggregation uses a federated averaging algorithm to weighted average the model parameters of each end-level unit to generate a coordinated control model. Instruction generation: Based on the aggregated model, calculate the regional-level optical storage and charging collaborative scheduling instructions and issue them to each end-level unit for execution; Through the district-level federated learning center, encrypted model gradients are exchanged between multiple districts based on secure multi-party computation. This method is implemented by the district-level federal learning center: Gradient encryption: The edge autonomous nodes of each transformer area use homomorphic encryption technology to encrypt the gradient of the local model. Secure exchange involves exchanging encrypted model gradients across multiple stations via a secure multi-party computation protocol. Collaborative learning: The Federated Learning Center aggregates encrypted gradients from various regions, updates the global model, and enables cross-regional knowledge sharing. The transformer's operating state is simulated in real time using digital twins in edge autonomous nodes, and physical constraints are transformed into penalty terms in a reinforcement learning reward function. This method's steps are performed by the digital twin in the edge autonomous node: State simulation: The digital twin is based on the transformer thermal circuit model and calculates the winding hot spot temperature in real time. The constraint transformation converts the transformer temperature safety constraint into a penalty term in the reinforcement learning reward function, mathematically expressed as:

[0031] Where R is the total reward value, R 性能 As a reward for system performance, T 热点 The predicted hotspot temperature T calculated for the digital twin 安全 The preset safe temperature limit is α, where α is the penalty coefficient. Through the regional federated learning center, the method of using a personalized federated learning algorithm to ensure that each region retains its own personalized output layer and only shares the parameters of the bottom feature extraction layer is executed by the regional federated learning center: Model segmentation divides the local models of each station area into a shared low-level feature extraction layer and a retained personalized output layer; Parameter sharing: Each station only uploads the parameters of the bottom feature extraction layer to the federated learning center for aggregation. To maintain individuality, each station independently trains a personalized output layer based on local data, thus maintaining its adaptability to local characteristics. Specifically, the collaborative autonomy approach forms a complete autonomous solution for power distribution zones through the organic combination of five core steps: edge-level reflection control, edge federation aggregation, cross-regional secure learning, digital twin security constraints, and personalized model sharing. It fully leverages the advantages of edge-end self-control, and under the premise of ensuring data privacy and security, achieves effective collaboration of distributed resources within the power distribution zone and knowledge sharing between multiple power distribution zones, thereby improving the system's autonomy, learning efficiency, and operational security.

[0032] Although alternative embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make further changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0033] The above specific embodiments further illustrate the purpose, technical solution and beneficial effects of this application. It should be understood that the above are only specific embodiments of this application and are not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made on the basis of the technical solution of this application should be included within the scope of protection of this invention.

Claims

1. A photovoltaic-storage-charging coordinated control system for distribution areas based on edge-end self-control, characterized in that, include: End-level autonomous units, deployed in photovoltaic inverters, energy storage converters and charging piles, are configured to achieve millisecond-level reflective control based on locally observed electrical state parameters; Edge autonomous nodes are deployed in the intelligent converged terminal of the distribution area and communicate with the end-level autonomous units. They are configured to aggregate the policies of each end-level autonomous unit based on federated learning and generate distribution area-level scheduling instructions. The district-level federated learning center communicates with the edge autonomous nodes of multiple districts and is configured to build a virtual federated learning center through secure multi-party computation and exchange encrypted model gradients.

2. The system according to claim 1, characterized in that, The endpoint-level autonomous unit includes: Lightweight reinforcement learning agents; The local status observation module is configured to collect voltage, power, and state of charge in real time. The reflection control module is configured to instantaneously trigger reactive power support when a voltage over-limit is detected.

3. The system according to claim 1, characterized in that, The edge autonomous nodes include: The prediction model is configured to perform second-level time series prediction and optimization. The regional collaborative controller is configured to aggregate the policies of each end-level autonomous unit based on federated learning. The digital twin adopts a physical-data hybrid driving model, is configured to simulate the operating state of a transformer in real time, and transforms physical constraints into a reinforcement learning reward function penalty term.

4. The system according to claim 3, characterized in that, The digital twin includes: The physical model uses a transformer thermal circuit model. In the data-driven part, a time-delay neural network is used to learn the residuals.

5. The system according to claim 1, characterized in that, The district-level federated learning center adopts a personalized federated learning algorithm, with each district retaining a personalized output layer and sharing only the parameters of the bottom feature extraction layer. The edge autonomous node is equipped with a contribution evaluation module, which is used to evaluate the data contribution of each participating area before the federated model is aggregated.

6. The system according to claim 5, characterized in that, The contribution assessment module is configured as follows: Based on the transformer load rate, photovoltaic power output fluctuation rate and electric vehicle charging demand time series similarity of each distribution area, a distribution area feature vector is constructed. Calculate the correlation weight between the feature vector of the transformer area and the performance improvement of the global model; Based on the aforementioned relevance weights, the weights of each substation in the federated learning model aggregation are dynamically adjusted.

7. The system according to claim 1, characterized in that, It also includes a transfer learning module, which includes: The source area knowledge base storage unit is configured to store historical data of typical transformer areas; The meta-learning model training unit is communicatively connected to the edge autonomous node and is configured to train the meta-learning model based on the historical data. The feature mapping processor is configured to map heterogeneous features from different stations to a unified latent space.

8. The system according to claim 1, characterized in that, It also includes a three-time-scale control module, which comprises: The prediction layer is configured to predict future photovoltaic power output and the probability of electric vehicles entering and leaving the site on a minute-level time scale; The correction layer is configured to correct prediction biases on a second-level time scale and uses a model predictive control method to continuously optimize and correct the prediction bias. The emergency layer is configured to trigger end-level agent reflection control on a millisecond timescale during sudden voltage changes.

9. The system according to claim 8, characterized in that, The model predictive control in the correction layer incorporates the joint uncertainty quantification result into the rolling optimization objective function. The joint uncertainty is obtained by identifying the noise characteristics introduced by the intermittent output of distributed photovoltaic power and the random charging of electric vehicles in the electrical measurement data of the transformer area, and quantifying them using a random discarding method.

10. A method for coordinated autonomous operation of photovoltaic, energy storage, and charging systems in a distribution area based on the system described in any one of claims 1 to 9, characterized in that, include: The end-level autonomous unit achieves millisecond-level reflective control based on local observation status; The edge autonomous nodes aggregate the strategies of each end-level autonomous unit based on federated learning to generate area-level scheduling instructions. Through the aforementioned district-level federated learning center, encrypted model gradients are exchanged between multiple districts based on secure multi-party computation. The transformer's operating state is simulated in real time using a digital twin in the edge autonomous node, and physical constraints are transformed into a reinforcement learning reward function penalty term. Through the aforementioned district-level federated learning center, a personalized federated learning algorithm is adopted, enabling each district to retain a personalized output layer and share only the parameters of the underlying feature extraction layer.

Citation Information

Patent Citations

  • Distributed energy aggregation regulation and autonomous regulation collaborative optimization method and system

    CN115149586A

  • Power grid operation state evaluation method

    CN119886533A

  • Hot continuous rolling multi-twin intelligent collaborative optimization system and method

    CN121188812A

Cited By

  • A desulfurization control method based on federated learning and adaptive model predictive control

    CN122345974A