New energy vehicle charging network load balancing control method and corresponding product
Patent Information
- Application Number
- CN202610666465.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-14
- Publication Date
- 2026-08-18
AI Technical Summary
[0004]然而,上述现有技术方案仍存在如下明显缺陷:1)缺乏对充电网络与电网耦合关系的精细化、全局化建模,难以在满足海量差异化用户需求的前提下,精准预判和优化网络整体的负载均衡状态;2)决策过程多依赖于静态规则或简单的优化模型,无法适应电网状态、用户行为、环境因素等多维信息的动态变化与高度不确定性,导致调度策略的实时性、适应性与最优性不足;3)控制指令的生成与下发方式未能充分考虑大规模充电网络控制的实时性、可靠性和协同性要求,可能造成控制滞后或指令冲突
[0020]As can be seen from the technical solution provided in this application, on the one hand, by collecting multi-source heterogeneous data in parallel and performing spatiotemporal alignment and feature fusion to generate a dynamic panoramic data view of the network, information silos are overcome, and a unified, comprehensive, and structured digital representation of the real-time status of the charging network-grid coupled system is achieved, providing reliable and multi-dimensional data support for advanced decision-making. On the other hand, a dynamic digital twin model is constructed and updated in real time based on the dynamic panoramic data view of the network, and this model is used to perform millisecond-level parallel simulation and deduction of various load allocation schemes for the requests to be processed. This enables the system to proactively evaluate the complex impact of different schemes on the overall load balance of the network, the load of key equipment, and the voltage stability of the grid in a low-cost and zero-risk manner before actual scheduling, thereby screening out potential high-performance scheduling schemes. The first aspect is the implementation of a dynamic scheduling strategy, which avoids the grid risks that trial-and-error control may bring. Secondly, the load balancing decision engine based on deep reinforcement learning generates globally optimal dynamic scheduling strategies. This not only allocates specific charging resources and power curves to each request, but also ensures that while meeting all needs, the overall network load fluctuations are controlled and strictly adhere to grid safety constraints. This achieves an intelligent closed loop from perception and prediction to decision-making, improving the automation level and global optimization capabilities of load balancing control. Thirdly, the cloud is responsible for macro-level strategies and collaborative planning, while edge nodes are responsible for micro-level instruction decomposition and emergency handling. This architecture ensures the global consistency of control objectives and improves the real-time performance and reliability of instruction issuance and execution, ensuring that complex load balancing strategies can be reliably and collaboratively implemented in the physical charging network. In summary, the technical solution of this application achieves globally optimal load balancing scheduling by constructing a dynamic digital twin model and a deep reinforcement learning decision engine, effectively optimizing charging network performance and grid stability.
Smart Images

Figure CN122600002A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of new energy charging facilities, and in particular to a load balancing control method and corresponding products for new energy vehicle charging networks. Background Technology
[0002] With the widespread adoption of new energy vehicles, their charging demand has exploded, and the large-scale, high-density charging activity has placed significant load pressure on local power grids. Charging networks are no longer isolated collections of facilities, but rather complex systems deeply coupled with distribution networks. If charging loads are concentrated and disordered, it can easily lead to problems such as power grid overload, voltage exceeding limits, and increased transformer losses, threatening the safe and stable operation of the power grid. It can also result in excessively long charging wait times and a degraded user experience.
[0003] To address these challenges, several charging load management solutions have been proposed in existing technologies. For example, time-of-use pricing strategies can be used to guide users to charge during off-peak hours, or charging stations can be locally controlled based on simple rules (such as "first-come, first-served" combined with power limits). Some solutions attempt to introduce centralized scheduling, which uses a unified approach to start / stop or adjust the power of charging stations based on the total grid load or regional load thresholds.
[0004] However, the existing technical solutions mentioned above still have the following obvious defects: 1) They lack refined and global modeling of the coupling relationship between the charging network and the power grid, making it difficult to accurately predict and optimize the overall load balance of the network while meeting the needs of a large number of differentiated users; 2) The decision-making process relies heavily on static rules or simple optimization models, which cannot adapt to the dynamic changes and high uncertainty of multi-dimensional information such as power grid status, user behavior, and environmental factors, resulting in insufficient real-time performance, adaptability, and optimality of the scheduling strategy; 3) The generation and issuance of control commands fail to fully consider the real-time, reliability, and coordination requirements of large-scale charging network control, which may cause control lag or command conflicts. Summary of the Invention
[0005] This application provides a load balancing control method and corresponding product for new energy vehicle charging networks. By constructing a dynamic digital twin model and a deep reinforcement learning decision engine, it achieves globally optimal load balancing scheduling, effectively optimizing the performance of the charging network and the stability of the power grid.
[0006] On the one hand, this application provides a load balancing control method for a new energy vehicle charging network, the method comprising:
[0007] Receive and parse charging requests from multiple users. The charging request includes at least the user's identity identifier, the target charging pile cluster identifier, the expected charging amount, the expected completion time, and the user-defined charging priority weight. The charging priority weight is used to quantify the user's preference trade-off between charging cost and charging speed.
[0008] The system collects real-time operating status time-series data of all charging piles in the charging network, charging station environmental perception data, and ultra-short-term load forecast data from the power grid dispatch system in parallel. The system then performs spatiotemporal alignment and feature fusion of the real-time operating status time-series data and the ultra-short-term load forecast data to generate a panoramic dynamic data view of the network.
[0009] Based on the network panoramic dynamic data view, a dynamic digital twin model of the charging network is constructed and updated in real time. All current pending charging requests are input into the dynamic digital twin model. The dynamic digital twin model is used to perform millisecond-level parallel simulation and deduction of various possible load distribution schemes to evaluate the impact of each scheme on the overall load balance of the network, transformer peak load and grid node voltage stability.
[0010] Based on the results of the parallel simulation, a load balancing decision engine based on deep reinforcement learning is adopted to generate a globally optimal dynamic scheduling strategy. The dynamic scheduling strategy assigns a specific charging pile, charging start time, and charging power curve to each charging request, and ensures that while meeting the charging needs of all users, the overall load variance is lower than a preset threshold and the grid constraints are not violated.
[0011] The dynamic scheduling strategy is decomposed into a series of time-ordered, charging pile group coordinated control command sequences, which are then distributed to the corresponding edge computing nodes and charging pile controllers through a cloud-edge collaborative architecture to drive the charging network to perform the load balancing control.
[0012] On the other hand, this application provides a load balancing control device for a new energy vehicle charging network, the device comprising:
[0013] The receiving module is used to receive and parse charging requests from multiple users. The charging request includes at least the user's identity identifier, the target charging pile cluster identifier, the expected charging amount, the expected completion time, and the user-defined charging priority weight. The charging priority weight is used to quantify the user's preference trade-off between charging cost and charging speed.
[0014] The fusion module is used to collect real-time operating status time-series data of all charging piles in the charging network, charging station environmental perception data, and ultra-short-term load forecast data from the power grid dispatch system in parallel, and to perform spatiotemporal alignment and feature fusion of the real-time operating status time-series data and the ultra-short-term load forecast data to generate a dynamic panoramic data view of the network.
[0015] The evaluation module is used to construct and update a dynamic digital twin model of the charging network in real time based on the network panoramic dynamic data view. All pending charging requests are input into the dynamic digital twin model. The dynamic digital twin model is used to perform millisecond-level parallel simulation and deduction of various possible load distribution schemes to evaluate the impact of each scheme on the overall load balance of the network, transformer peak load and grid node voltage stability.
[0016] The generation module is used to generate a globally optimal dynamic scheduling strategy based on the results of the parallel simulation and a load balancing decision engine based on deep reinforcement learning. The dynamic scheduling strategy assigns a specific charging pile, charging start time, and charging power curve to each charging request, and ensures that the network meets the charging needs of all users while the overall load variance is lower than a preset threshold and the grid constraints are not violated.
[0017] The driving module is used to decompose the dynamic scheduling strategy into a series of time-ordered, charging pile group coordinated control command sequences, and send them to the corresponding edge computing nodes and charging pile controllers through the cloud-edge collaborative architecture, so as to drive the charging network to perform the load balancing control.
[0018] Thirdly, this application provides an apparatus comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the technical solution of the above-described new energy vehicle charging network load balancing control method.
[0019] Fourthly, this application provides a storage medium storing a computer program, which, when executed by a processor, implements the steps of the technical solution of the above-described new energy vehicle charging network load balancing control method.
[0020] As can be seen from the technical solution provided in this application, on the one hand, by collecting multi-source heterogeneous data in parallel and performing spatiotemporal alignment and feature fusion to generate a dynamic panoramic data view of the network, information silos are overcome, and a unified, comprehensive, and structured digital representation of the real-time status of the charging network-grid coupled system is achieved, providing reliable and multi-dimensional data support for advanced decision-making. On the other hand, a dynamic digital twin model is constructed and updated in real time based on the dynamic panoramic data view of the network, and this model is used to perform millisecond-level parallel simulation and deduction of various load allocation schemes for the requests to be processed. This enables the system to proactively evaluate the complex impact of different schemes on the overall load balance of the network, the load of key equipment, and the voltage stability of the grid in a low-cost and zero-risk manner before actual scheduling, thereby screening out potential high-performance scheduling schemes. The first aspect is the implementation of a dynamic scheduling strategy, which avoids the grid risks that trial-and-error control may bring. Secondly, the load balancing decision engine based on deep reinforcement learning generates globally optimal dynamic scheduling strategies. This not only allocates specific charging resources and power curves to each request, but also ensures that while meeting all needs, the overall network load fluctuations are controlled and strictly adhere to grid safety constraints. This achieves an intelligent closed loop from perception and prediction to decision-making, improving the automation level and global optimization capabilities of load balancing control. Thirdly, the cloud is responsible for macro-level strategies and collaborative planning, while edge nodes are responsible for micro-level instruction decomposition and emergency handling. This architecture ensures the global consistency of control objectives and improves the real-time performance and reliability of instruction issuance and execution, ensuring that complex load balancing strategies can be reliably and collaboratively implemented in the physical charging network. In summary, the technical solution of this application achieves globally optimal load balancing scheduling by constructing a dynamic digital twin model and a deep reinforcement learning decision engine, effectively optimizing charging network performance and grid stability. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a flowchart of the load balancing control method for new energy vehicle charging networks provided in the embodiments of this application;
[0023] Figure 2 This is a schematic diagram of the structure of the new energy vehicle charging network load balancing control device provided in the embodiments of this application;
[0024] Figure 3 This is a schematic diagram of the device provided in the embodiments of this application. Detailed Implementation
[0025] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0026] In this specification, adjectives such as "first" and "second" are used only to distinguish one element or action from another, without necessarily requiring or implying any actual such relationship or order. Where circumstances permit, reference to an element or component or step (etc.) should not be construed as being limited to only one of the elements, components, or steps, but may be one or more of the elements, components, or steps, etc.
[0027] For ease of description, the dimensions of the various parts shown in the accompanying drawings are not drawn to actual scale.
[0028] To address the concentrated and disorderly charging loads of new energy vehicle charging networks, which can easily lead to grid overload, voltage exceeding limits, and increased transformer losses, existing technologies have proposed several charging load management solutions. For example, time-of-use pricing strategies can guide users to charge during off-peak hours, or localized control of charging stations can be implemented based on simple rules (such as "first-come, first-served" combined with power limits). Some solutions attempt to introduce centralized scheduling, uniformly starting, stopping, or adjusting the power of charging stations based on the total grid load or regional load thresholds. However, the existing technical solutions mentioned above still have the following obvious defects: 1) They lack refined and global modeling of the coupling relationship between the charging network and the power grid, making it difficult to accurately predict and optimize the overall load balance of the network while meeting the needs of a large number of differentiated users; 2) The decision-making process relies heavily on static rules or simple optimization models, which cannot adapt to the dynamic changes and high uncertainty of multi-dimensional information such as power grid status, user behavior, and environmental factors, resulting in insufficient real-time performance, adaptability, and optimality of the scheduling strategy; 3) The generation and issuance of control commands fail to fully consider the real-time, reliability, and coordination requirements of large-scale charging network control, which may cause control lag or command conflicts.
[0029] To address the aforementioned problems in the prior art, this application proposes a load balancing control method for new energy vehicle charging networks, the flowchart of which is attached. Figure 1 As shown, it mainly includes steps S101 to S105, which are detailed below:
[0030] Step S101: Receive and parse charging requests from multiple users.
[0031] Traditional charging requests typically only include the target charging station and the desired charging capacity, making it difficult for scheduling systems to make fine-grained trade-offs between user satisfaction and grid efficiency. This application introduces user-defined charging priority weights to quantify users' personalized preferences, providing a key input for subsequent multi-objective optimization. Specifically, the received charging requests from multiple users include at least the user's identity identifier, the target charging station cluster identifier, the desired charging capacity, the desired completion time, and user-defined charging priority weights. These user-defined charging priority weights are used to quantify the user's preference trade-off between charging cost and charging speed.
[0032] Users submit requests via mobile applications, in-vehicle terminals, or the charging pile's interactive interface. The "target charging pile cluster identifier" provides flexibility, allowing it to be a single pile, a charging station, or a geographical area. User-defined charging priority weights are also included. It is a configurable value, for example, continuously adjustable within the range [0, 1]. When When the value approaches 0, it indicates that users are extremely concerned about "cost" and are willing to wait longer for lower electricity prices; when... When the value approaches 1, it indicates that the user is extremely focused on "speed" and is willing to pay higher fees to complete charging as quickly as possible. This weight will directly serve as an important parameter in the reward function of the subsequent deep reinforcement learning decision engine, enabling truly personalized scheduling.
[0033] Considering that a completely passive request reception may lead to load aggregation in time and space, exacerbating the peak-valley difference in the power grid, in order to utilize information prompts to flexibly adjust user demand and optimize load distribution from the source, the above embodiment of receiving and parsing charging requests from multiple users also includes the following steps S1011 to S1013: demand-side response guidance based on digital twins.
[0034] Step S1011: Before the user submits a charging request, push the charging cost-time heatmap of different charging periods and different charging piles based on the dynamic digital twin model to the user's interactive terminal for the current and future periods.
[0035] Traditional methods only display current electricity prices and queuing status, lacking predictions of future trends. This application utilizes a constructed dynamic digital twin model (the details of which will be described in subsequent step S103) to rapidly simulate the future state of the network. Specifically, the system starts with the current network state and drives the digital twin model to perform rolling simulations of multiple discrete time slices in the future (e.g., the next 24 hours, with each time slice being 15 minutes). In the simulation, the system virtually schedules a standardized charging request (e.g., 50kWh of electricity) for each time slice and each charging station, and records the total cost (electricity cost + service fee) and total time (waiting time + charging time) required to complete the charging. Subsequently, the cost and time are normalized and weighted according to the user's possible preference weights to form a comprehensive cost index. For example, for a user who prefers "speed" (higher), the weight of time in the comprehensive cost is greater. Finally, using time and space (charging station location) as two dimensions, the calculated comprehensive cost index is visualized and rendered using color depth to generate an intuitive charging cost-time heatmap. This charging cost-time heatmap clearly shows "when" and "where" charging is more economical or faster.
[0036] Step S1012: Receive the user's adjusted charging demand parameters based on the charging cost-time heatmap, and record the user's adjustment behavior as a successful demand-side response event.
[0037] After observing the heatmap, users may adjust their "expected completion time" or "target charging station cluster." For example, a user might initially want to charge at station A immediately, but the heatmap shows that charging at station B an hour later would have a lower overall cost, so they modify their request. The system compares the user's request before and after the modification. If the modification aligns the request with the system's desired "peak shaving and valley filling" direction (e.g., shifting from peak hours to off-peak hours), the interaction is recorded as a "successful demand-side response event," along with information such as the user ID, adjustment magnitude, and time. This essentially incorporates the user into the collaborative optimization system of load balancing.
[0038] It should be noted that the charging cost-time heatmap of the above embodiment can be generated through the following steps S1011a to S1011c:
[0039] S1011a: Utilizes a dynamic digital twin model to rapidly simulate the current network state over multiple future time slices in response to virtual standardized charging requests.
[0040] S1011b: Records the simulation results of each virtual request at different start times and at different charging stations, including total cost and total time.
[0041] S1011c: After normalizing the cost and time consumption, the cost and time consumption are weighted and merged into a comprehensive cost indicator, and then visualized and rendered in terms of time and space to generate a charging cost-time heat map.
[0042] In the process of generating the charging cost-time heatmap described above, "normalization" aims to eliminate the differences in units and magnitudes between cost and time. For example, a minimum-maximum normalization method can be used to map both cost and time to the [0,1] interval. The weights in "weighted merging" can use a set of preset values to generate a general heatmap by default, or they can be personalized based on the user's historical preferences or real-time selected preference tags to generate a unique charging cost-time heatmap.
[0043] Step S1013: In the subsequent design of the reward function of the deep reinforcement learning decision engine, additional positive rewards are given to the scheduling decision that successfully guides the demand-side response event.
[0044] Step S1013 is crucial for forming the incentive closed loop. This is because the deep reinforcement learning decision engine (details of which will be elaborated in subsequent step S104) learns which scheduling strategy is "good" through the reward function. In this embodiment, the reward function R is designed to include not only conventional indicators such as grid load balancing and user charging costs, but also a special reward item. This is used to incentivize scheduling strategies that successfully prompt users to make beneficial adjustments. Specifically, during the decision engine's simulation, if the scheduling strategy it generates (e.g., by displaying a certain type of heatmap) is evaluated as potentially guiding users to make positive adjustments similar to those in step S1012, then that strategy will receive additional rewards in the reward function. Bonus points. Through repeated learning, the decision engine will gradually master the ability to generate the most effective scheduling strategies and information prompts to guide demand-side responses, thus forming a virtuous cycle of "intelligent guidance - user response - system optimization".
[0045] Through the above step S101, the system not only obtains the user's original needs, but also guides the optimization of needs through intelligent interaction, and feeds back the experience of successful guidance to the AI decision model, providing high-quality and "flexible" demand-side input for the entire load balancing control system.
[0046] Step S102: Collect real-time operating status time-series data of all charging piles in the charging network, charging station environmental perception data, and ultra-short-term load forecast data from the power grid dispatching system in parallel. Then, perform spatiotemporal alignment and feature fusion of the real-time operating status time-series data of all charging piles and the ultra-short-term load forecast data from the power grid dispatching system to generate a panoramic dynamic data view of the network.
[0047] The goal of step S102 is to create a unified, real-time, multi-dimensional, and high-quality "digital mirror" of the system state, providing a reliable data foundation for subsequent simulations and decision-making. Traditional charging network monitoring often involves isolated data sources (separate data from charging piles, the power grid, and the environment) and lacks effective cleaning and fusion, leading to biased, delayed, or even noisy decision-making. This application constructs a panoramic view that truly reflects the coupling relationship between "vehicle-charging pile-station-network" through parallel acquisition, spatiotemporal alignment, and intelligent fusion.
[0048] In the above embodiments, the parallel acquisition of real-time operating status time-series data of all charging piles in the charging network, charging station environmental perception data, and ultra-short-term load forecast data from the power grid dispatching system mainly includes the following steps:
[0049] 1) Collect real-time operating status time-series data, specifically: collect the three-phase voltage of each charging pile at a high frequency (e.g., 1-10 times per second) through the smart meters and sensors built into the charging pile controller. Current Active power P, reactive power Q, power factor and the temperature of the charging gun interface And so on, forming a continuous time series.
[0050] 2) Environmental sensing data for charging stations, specifically: collected through the station's Internet of Things (IoT) sensor network, including: ambient temperature obtained from temperature and humidity sensors. With humidity Parking space occupancy status (Boolean value, 0 for empty, 1 for occupied) obtained from parking space cameras or geomagnetic sensors. And the node's own computing load rate obtained by the system monitoring module of the edge computing node within the site. .
[0051] 3) Ultra-short-term load forecast data, specifically obtained by subscribing to the regional power grid dispatch center's energy management system (EMS) via a secure data gateway according to standard protocols (e.g., IEC 61850, IEC 104). This data typically forecasts the total active power load of each distribution transformer node or feeder on a rolling basis, with a period of 15 minutes to 4 hours. reactive load and voltage reference value It also includes a clear timestamp and node topology identifier.
[0052] As an embodiment of this application, step S102, which involves spatiotemporal alignment and feature fusion of real-time operational status time-series data and ultra-short-term load forecast data to generate a dynamic panoramic data view of the network, can be achieved through the following steps S1021 to S1024:
[0053] Step S1021: The real-time operating status time-series data from the charging piles is processed in real time using a sliding window-based streaming computing framework to calculate the instantaneous power, average power, power change rate, and estimated remaining charging time for each charging pile, thus obtaining the first processed data.
[0054] Traditional batch processing methods cannot meet real-time requirements. This application's embodiments deploy streaming computing engines such as Apache Flink or Apache Spark Streaming at the edge or in regional data centers. This engine continuously consumes the raw data stream using a sliding time window (e.g., window size W = 300 seconds, sliding step S = 10 seconds). Within each window, the following data for each charging station are aggregated and feature extracted in real time: instantaneous power, average power, power change rate, and estimated remaining charging time.
[0055] 1) Instantaneous power : Retrieves the active power value of the latest sampling point within the window;
[0056] 2) Average power : Calculate the arithmetic mean of all active power sampling points within the window, i.e. This reflects the recent average load;
[0057] 3) Power change rate By performing linear fitting on the power sequence within the window, the slope is obtained as the rate of change, which characterizes the severity of load ramping or dropping.
[0058] 4) Estimated remaining charging time Based on the expected charging amount of the current users served by this charging station. Charged amount (This can be obtained by integrating the power) and the current charging efficiency. Estimate, that is, This feature is crucial for predicting the future release time of piles.
[0059] After step S1021, the original second-level data stream is transformed into a feature vector with pile as the dimension and window as the period, denoted as the first processed data, and denoted as Data_pile_feat.
[0060] Step S1022: The charging station environmental perception data and the corresponding real-time operating status time series data are correlated and fused to obtain the second processed data, wherein the charging station environmental perception data includes temperature, humidity and parking space occupancy status.
[0061] Step S1022 aims to establish the association between operational data and the operating environment and resource status. The core of "association fusion" is based on spatial location identifiers. The system maintains a mapping table of "charging pile-charging station-sensor". For each charging pile, the system finds its corresponding charging station according to the mapping table and then associates it with the environmental sensor data within that station. The association fusion is completed in real time within the streaming computing framework as follows: the feature vector of each pile in Data_pile_feat is correlated with the ambient temperature of its corresponding station. ,humidity Perform splicing; introduce parking space occupancy status. If the parking space corresponding to the piling is occupied, but the piling status is "idle," it may indicate that there are vehicles queuing or malfunctioning. This information is of great value for scheduling. Derivative features, such as temperature rise, can be calculated. This is used to more accurately assess the heating status of the pile itself.
[0062] After step S1022, an enhanced feature vector containing runtime features and environmental context is generated, denoted as the second processing data, i.e., Data_pile_env.
[0063] Step S1023: Merge the first processed data and the second processed data to form the processed charging network data.
[0064] This step is the logical result of S1021 and S1022. In fact, within the streaming pipeline, Data_pile_env already naturally contains all the characteristics of the first processed data. Therefore, the processed charging network data Data_network refers to a unified data set containing the operating characteristics, environmental status, and resource occupancy of each charging pile as a basic unit. Each data unit can be represented as: [timestamp, pile ID, location, instantaneous power, average power, power change rate, remaining time, gun temperature, ambient temperature, humidity, parking space status, ...].
[0065] Step S1024: Match and overlay the processed charging network data with the ultra-short-term load forecast data with spatiotemporal labels issued by the power grid dispatching system to form a unified dynamic panoramic data view of the network with power grid nodes and charging piles as basic units.
[0066] This is key to achieving "network-pile" integration. The system maintains an electrical connection topology table for "charging piles - distribution transformers / feeders," and through the following matching, overlaying, and view formation, finally achieves step S1024:
[0067] 1) Matching, specifically: Based on the topology table, aggregate the data of each charging pile in Data_network to its superior grid node. For example, calculate the total real-time power of all charging piles connected to "Transformer T1". Average load rate, etc.
[0068] 2) Overlay, specifically: the aggregated charging load data is combined with the ultra-short-term base load forecast value issued by the power grid EMS for the same node "transformer T1". Align on the timeline (interpolated based on timestamps), and then calculate the total predicted load for that node: This organically combines the unpredictable charging load forecasting with the orderly load planning of the power grid.
[0069] 3) The final generated network panoramic dynamic data view is a hierarchical data structure. The top layer is the power grid node layer, which contains information such as the total predicted load, voltage prediction, and capacity limit of each node. The lower layer is the charging equipment layer, which contains the detailed feature vector of each charging pile. The network panoramic dynamic data view also integrates the time dimension, which can reflect the dynamic evolution trend of the system state. This network panoramic dynamic data view is the only authoritative data source for all subsequent advanced applications (digital twin, AI decision-making).
[0070] Furthermore, after forming a dynamic panoramic data view of the network, this application also includes anomaly detection and resilient disaster recovery mechanisms, the technical solution of which is implemented through the following steps S1025 to S1027:
[0071] Step S1025: Integrate a hybrid anomaly detection model combining isolated forests and autoencoders into the streaming computing framework to monitor each data stream in the network panoramic dynamic data view in real time.
[0072] Charging network data may contain noise or outliers due to sensor malfunctions, communication interference, equipment abnormalities, etc., directly affecting the quality of subsequent decisions. Traditional thresholding methods are ineffective for complex, multivariate joint anomalies. This embodiment employs an unsupervised hybrid model for real-time streaming anomaly detection, mainly including the following steps: model construction, training process, and real-time monitoring:
[0073] 1) Model construction: The constructed models mainly include isolated forests and autoencoders, detailed below:
[0074] Isolation Forest is a fast anomaly detection algorithm suitable for continuous data. Within a streaming computing framework, it maintains a lightweight iForest model for each key data stream (such as the power sequence of a single node or the total load sequence of a node). This model "isolates" data points by randomly partitioning the feature space; outliers, due to their significant differences from normal points, are typically isolated more quickly (with shorter paths).
[0075] An autoencoder is a type of neural network that learns to reconstruct the intrinsic distribution of normal data. We train a lightweight autoencoder for multivariate correlated data, such as the power combination of all the charging piles in a charging station. Its structure is simple (e.g., input layer -> 16-dimensional encoding layer -> output layer) to accommodate edge computing resource constraints. During training, a large amount of historical normal data is used, with the goal of making the output as close to the input as possible.
[0076] 2) Training process: Offline training is conducted using historical data from normal time periods. iForest directly fits the data distribution; the autoencoder minimizes the reconstruction error. Training was conducted, including... It is input. It is a reconstruction output.
[0077] 3) Real-time monitoring: In a stream processing pipeline, each data unit (such as the feature vector of a stub) passes through two models in parallel:
[0078] 3.1) The abnormal score is calculated by iForest. That is, the standardized value of the path length of the data points.
[0079] 3.2) Input the autoencoder and calculate the reconstruction error. .
[0080] Finally, through weighted synthesis and The data point is then compared with a dynamic threshold to determine whether it is an anomaly. The hybrid model combines iForest's sensitivity to global sparse anomalies with the autoencoder's sensitivity to local pattern anomalies, significantly improving detection rate and robustness.
[0081] Step S1026: When the data mode of a certain charging pile is detected as abnormal, the dynamic digital twin model is automatically triggered to perform logical isolation simulation on the abnormal charging pile, assess the impact of its failure on the network, and start the backup charging pile discovery and switching plan.
[0082] Once a charging station is flagged as abnormal (e.g., the power reading remains 0 but the status is "charging", or the power value far exceeds the physical limit), the system will not immediately discard the data. Instead, it will initiate an intelligent handling process, including the following logical isolation simulation, impact assessment, and contingency plan activation:
[0083] The logical isolation simulation includes: the system immediately calls the dynamic digital twin model (the construction of which is detailed in subsequent step S103), marks the abnormal stake as "unavailable" or "operating at reduced capacity" in the twin, and reruns a rapid simulation based on the current network status and pending scheduling requests. This is intended to proactively assess whether the failure of the stake will cause the grid to exceed limits or prevent the completion of critical charging tasks.
[0084] Impact assessment and contingency plan activation: Based on the simulation results, if the impact is significant (e.g., causing an important request to time out), the system will automatically generate a backup charging pile discovery and switching plan. The plan includes: a recommended list of alternative charging piles, load transfer schemes, and the expected impact on the power grid and users.
[0085] Step S1027: Push the anomaly information, impact assessment and switching plan to the load balancing decision engine in real time as one of the inputs for generating or adjusting the dynamic scheduling strategy.
[0086] Anomaly handling and core scheduling decisions are tightly integrated into a closed loop. When calculating the scheduling strategy for the next cycle, the decision engine considers this contingency plan as a hard constraint or optimization objective. For example, the decision engine will prioritize migrating charging tasks originally scheduled for faulty charging stations to backup charging stations recommended in the contingency plan, and will provide a positive reward for successfully completing the migration in the reward function. This ensures that the system can dynamically, self-heal, and seamlessly maintain load balancing services in the face of equipment failures.
[0087] In another embodiment of this application, the anomaly detection and resilient disaster recovery mechanism can also be: using a graph spatiotemporal neural network to jointly model multi-source data in the dynamic panoramic data view of the network, so as to simultaneously capture the correlation of the charging network in spatial topology and the evolution pattern in time, thereby realizing the probabilistic prediction of potential anomalies in the next few minutes; when the probability of an anomaly occurring in a certain node or charging pile in the future exceeds the warning threshold, proactive defense simulation is initiated in advance in the dynamic digital twin model to evaluate and pre-generate a variety of preventive scheduling plans; the optimal preventive scheduling plan is integrated into the current decision cycle of the load balancing decision engine in advance, and by actively adjusting the power of surrounding charging piles, hot backup and load evacuation are achieved before the potential fault actually occurs. Compared to the passive response of the scheme in steps S1025 to S1027 after an anomaly is detected, the above scheme upgrades from anomaly detection to anomaly prediction, and uses graph spatiotemporal neural networks to achieve forward-looking perception. Among them, the active defense simulation and hot backup based on the prediction results is an advanced and predictive system self-healing mechanism. It actively intervenes through scheduling strategies before physical anomalies occur, which greatly improves the reliability and resilience of the system.
[0088] In order to improve the generalization ability of the global model through secure collaboration among multiple networks, as another embodiment of this application, Figure 1 The example method may also include the following steps S102a to S102d, a cross-regional model optimization process based on federated learning:
[0089] Step S102a: Deploy this method in multiple geographically isolated charging networks (belonging to different operators or cities), with each local network independently running its load balancing system and accumulating local data.
[0090] Each local network possesses a complete edge-cloud architecture, independently executes steps S101 to S105, and accumulates locally unique operational data. This data includes user behavior patterns, local power grid characteristics, and differences in equipment models, which are extremely valuable but also involve privacy and trade secrets.
[0091] Step S102b: Each local system periodically uploads the model gradient updates of its local deep reinforcement learning decision engine to a secure aggregation server in encrypted form.
[0092] This is the core of federated learning: data stays within the local domain. Each local network's cloud platform periodically (e.g., weekly) uses its recently accumulated data to perform several training iterations on its local deep reinforcement learning decision engine model (i.e., the Actor network or Critic network), calculating the gradient updates to the model parameters. Instead of the original data or model parameters themselves. Then, Encrypt the data (such as with homomorphic encryption) and upload it to a neutral, secure aggregation server.
[0093] Step S102c: The aggregation server securely aggregates gradient updates from multiple local sources, generates an aggregated global model update, and then distributes it to each local system.
[0094] After receiving the encrypted gradient updates from all participants, the aggregation server performs a secure aggregation operation (e.g., using the Secure Aggregation protocol) to calculate a weighted average of the gradients, i.e.: This process is performed in encrypted form, meaning the aggregation server cannot decrypt the gradient information of individual participants. Subsequently, the aggregated global gradient is updated. Distribute it back to each local system.
[0095] Step S102d: Each local system integrates global model updates to enhance the generalization ability of its own decision engine and its ability to cope with rare scenarios, while protecting the privacy of each local data.
[0096] Each local system received Then, it is applied to its own local model: In this way, each local model absorbs desensitized learning experiences from other networks, enabling it to better cope with rare scenarios that it has never seen but that other networks may encounter (such as extreme weather or special load patterns caused by large-scale events). This achieves secure knowledge sharing and co-evolution, while strictly ensuring the privacy of local data in each network is not leaked.
[0097] As another embodiment of this application, cross-regional model optimization based on federated learning can also be as follows: before securely aggregating model gradient updates, each local system uses secure multi-party computation to evaluate the quality and diversity of its local dataset and generate a verifiable contribution evaluation report; the contribution evaluation report, encrypted gradient updates, and proof of model performance improvement are recorded in the form of transactions on a permissioned blockchain; based on the contribution recorded on the chain, virtual incentive points are automatically allocated to each participant through smart contracts, and their weights in federated aggregation are dynamically adjusted based on their contribution, thereby incentivizing high-quality data holders to actively participate and suppressing malicious or low-quality participation behavior, thus building a sustainable, high-quality, and trustworthy federated learning ecosystem. Compared to the federated learning-based cross-regional model optimization process exemplified in steps S102a to S102d, the above solution further addresses two core challenges in the practical application of federated learning: fair assessment of participant contributions and sustainable participation incentives. By combining secure multi-party computation, blockchain, and smart contracts, a complete governance and incentive system is constructed. Specifically, the above solution ensures the transparency, fairness, and traceability of the collaboration process through technical means, attracting more high-quality participants and fundamentally improving the overall effectiveness and sustainability of federated learning, resulting in significant commercial and ecological value. On the other hand, the deep coupling of privacy computation, blockchain incentives, and federated learning, with its ingenious design, provides an innovative technical solution to the tragedy of the commons and free-rider problems in distributed AI collaboration.
[0098] Step S103: Based on the network panoramic dynamic data view, construct and update the dynamic digital twin model of the charging network in real time. Input all the charging requests to be processed at present into the dynamic digital twin model. Perform millisecond-level parallel simulation and deduction of various possible load distribution schemes through the dynamic digital twin model to evaluate the impact of each scheme on the overall load balance of the network, the peak load of transformers and the voltage stability of grid nodes.
[0099] Step S103 aims to address the fundamental deficiency of existing charging scheduling methods in lacking forward-looking and globally accurate assessment capabilities. Traditional methods typically base scheduling decisions on a snapshot of the system state at the current moment or rely on simple linear extrapolation and empirical rules. This "step-by-step" decision-making model cannot foresee the profound impact of different scheduling schemes on the future state of the entire system, including complex power flow, equipment thermal stress, queuing delays, and other coupled factors. Decision-making is like playing chess blindly, highly unpredictable, and prone to causing unforeseen secondary risks such as local overload, voltage exceeding limits, or a sharp drop in user satisfaction. To address this, this application creatively introduces and constructs a dynamic digital twin model of the charging network, and uses this as a basis for millisecond-level parallel simulation and deduction, providing a high-fidelity, zero-risk "forward-looking telescope" and "strategy testing ground" for intelligent decision-making.
[0100] Specifically, as an embodiment of this application, the construction and real-time updating of a dynamic digital twin model of the charging network can be achieved through the following steps S1031 to S1033:
[0101] Step S1031: Based on the physical connection topology, electrical parameters, historical operating data and real-time reported status data of the charging pile, initialize the dynamic digital twin model, wherein the dynamic digital twin model includes at least the power grid flow calculation sub-model, the charging pile thermal dynamics sub-model and the queuing theory service sub-model.
[0102] A digital twin model is not simply a data mirror, but a computable, simulateable, multiphysics-coupled virtual entity. Its initialization is a process that combines knowledge injection and data-driven approaches, mainly including model architecture and sub-model construction, as well as model initialization, detailed below:
[0103] 1) Model architecture and sub-model construction: The dynamic digital twin model includes the following sub-models: power grid flow calculation, charging pile thermal dynamics, and queuing theory service:
[0104] 1.1) Power Flow Calculation Sub-model: This is the electrical core of the digital twin. Specifically, it can be based on the electrical connection topology of the charging network (already used in the S1024 matching), constructing an equivalent power grid calculation network in the cloud or on a high-performance computing node. Nodes represent distribution transformers and buses, and branches represent feeders and cables. Electrical parameters are assigned to each branch, including resistance R, reactance X, and capacitance to ground B. This sub-model uses either the forward-backward substitution method or the Newton-Raphson method for three-phase power flow calculation. Its input is the net injected power (charging load + grid base load) at each node, and its output is the voltage amplitude U and phase angle of all nodes in the entire network. And the trends of each branch road This sub-model is directly used to assess "grid node voltage stability" and "transformer / line overload".
[0105] 1.2) Charging Pile Thermal Dynamics Sub-model: When a charging pile operates at high power, its internal power devices and connection points generate heat. Overheating can lead to malfunctions or derating. This sub-model establishes a simplified thermal circuit model for each charging pile, commonly described by a first-order RC equivalent circuit. For example, the pile body can be considered as a heat capacity. Its heat exchange with the environment is considered as thermal resistance. Its thermal equilibrium differential equation is: ,in, The thermal time constant, The power loss of the charging pile (related to the output power P) (where k is an empirical coefficient). The ambient temperature is used (from step S102 of the aforementioned embodiment). This model is used to simulate the temperature rise process of the charging pile under continuous high-power operation, assess its thermal safety status, and prevent failures or forced power reduction due to overheating.
[0106] 1.3) Queuing Theory Service Sub-Model: Used to simulate the service process of user requests. Each charging station is considered a "service counter," and the charging request (with power demand) is considered a "customer." This sub-model adopts a G / G / c queuing model (general arrival time, general service time, c parallel service counters). The service time is dynamically calculated based on the requested power and the allocated power curve. This model is used to predict the waiting time, queue length, and completion time of requests, and is key to evaluating user satisfaction indicators.
[0107] • Model initialization: Utilize historical operating data (e.g., typical load curves, ambient temperature changes) to initialize key parameters (e.g., thermal time constant) in the aforementioned sub-models. Average service rate The initial calibration is performed. The real-time reported status data (the network panorama dynamic data view in step S102) serves as the initial status input for each simulation run of the model.
[0108] Step S1032: Establish a real-time data mapping channel between the dynamic digital twin model and the real charging network, and continuously feed the network panoramic dynamic data view.
[0109] To ensure synchronization between the digital twin and the physical entity's state, a low-latency, highly reliable data mapping channel is established. This channel is built based on a message middleware (e.g., Apache Kafka). The dynamic network panorama data view generated in step S102 is encapsulated into messages of a specific format and continuously flows into the data receiving layer of the digital twin model through this channel in real time. Each sub-model in the twin model subscribes to and extracts corresponding real-time data (such as node injection power, ambient temperature, and stake status) from this data stream as needed, updating its own internal state variables to ensure that each simulation begins at an initial point highly synchronized with the real world.
[0110] Step S1033: Using a preset model calibration algorithm, periodically compare the simulated output of the dynamic digital twin model with the monitoring feedback of the real network, and dynamically adjust the key parameters in the model to reduce model mismatch error.
[0111] Due to factors such as model simplification, equipment aging, and environmental changes, the simulated output of a digital twin model will gradually deviate from the real system, i.e., "model mismatch". To address this, this application designs a closed-loop online adaptive calibration mechanism. The step of "periodically comparing the simulated output of the digital twin model with the monitoring feedback of the real network through a preset model calibration algorithm, and dynamically adjusting key parameters in the model to reduce model mismatch error" is achieved through the following steps S1033a to S1033d:
[0112] Step S1033a: In each calibration cycle (e.g., every 6 hours or during off-peak hours each day), obtain a set of actual operating results data from the real network as a calibration benchmark.
[0113] After the calibration period begins, the system selects a recently completed, representative operational period (e.g., a complete peak load period). It then retrieves the actual operational results data of the real network for that period from the database, including the actual power curves of each charging station. Actual voltage curves of each node and the actual completion time of each charging task. And so on. This set of data serves as the gold standard for evaluating the accuracy of the model.
[0114] Step S1033b: Using the same initial conditions as input in the dynamic digital twin model, run the dynamic digital twin model to obtain simulation result data.
[0115] Starting from the beginning of the calibration period, the same initial conditions (the network panorama at that moment) are input into the digital twin model. The model is driven to simulate with the same time length and step size as the actual process, yielding the corresponding simulation result data: simulated power curve. Simulated voltage curve and simulated completion time ,etc.
[0116] Step S1033c: Calculate the multi-objective loss function between the simulation result data and the calibration reference data, wherein the loss function includes at least the load curve fitting error and the charging completion time prediction error.
[0117] Model calibration is structured as a multi-objective optimization problem. A loss function L is defined, which includes several error-measuring terms such as load curve fitting error, charging completion time prediction error, and voltage curve fitting error.
[0118] 1) Load curve fitting error: Calculate the root mean square error (RMSE) between the simulated total active load and the actual total active load sequence for the key transformer or line throughout the entire calibration period.
[0119]
[0120] 2) Charging Completion Time Prediction Error: Calculate the mean absolute percentage error (MAPE) between the simulated completion time and the actual completion time for all tasks that complete charging within this time period.
[0121]
[0122] 3) Voltage curve fitting error (this error term is optional): Calculate the RMSE between the simulated and actual voltage values at key nodes:
[0123]
[0124] The final multi-objective loss function is the weighted sum of the above terms, that is, ,in, , and These are the weighting coefficients.
[0125] S1033d: Gradient descent is used to backpropagate the loss L and update the values of adjustable parameters in the digital twin model to minimize the loss function.
[0126] In this application embodiment, key parameters refer to those parameters in the model that have uncertainties or drift over time, such as the equivalent impedance of the line. Thermal time constant of charging pile and charging efficiency And so on. The system uses automatic differentiation techniques to calculate the loss function L relative to these adjustable parameters. gradient Then, the parameters are updated using gradient descent (or a variant thereof, such as the Adam optimizer):
[0127]
[0128] in, This is the learning rate for the calibration process. This process is iterative until the loss function L converges to its minimum or reaches the predetermined number of iterations. In this way, the key parameters of the dynamic digital twin model are dynamically adjusted, making its simulated behavior continuously approximate the physical behavior of the real system, significantly reducing model mismatch error and ensuring the credibility of the simulation.
[0129] After the model is built and updated, a simulation is performed: all pending charging requests are input into the dynamic digital twin model, and millisecond-level parallel simulations are conducted on various possible load distribution schemes through the dynamic digital twin model.
[0130] When a new charging request arrives or triggers a periodic scheduling, the system injects all pending requests (including new requests and scheduled but not yet started requests) into the digital twin model. "Multiple possible load allocation schemes" are candidate scheduling actions generated by the decision engine in subsequent step S104. For example, a candidate action might be "allocate request A to pile X for immediate charging at 90kW, and postpone request B for 30 minutes to pile Y for charging at 60kW." For each candidate action, the digital twin model simulates the system evolution over the next few hours in parallel at millisecond speeds on a high-performance computing cluster in the cloud. Each simulation instance outputs a series of evaluation metrics, where the overall network load balance is typically measured by the variance or Gini coefficient of all transformer load rates; transformer peak load is the maximum apparent power of the transformers observed during the simulation; and grid node voltage stability is determined by checking whether the voltage of all nodes remains within the national standard allowable range (e.g., ±10% of the nominal voltage) throughout the simulation.
[0131] As an embodiment of this application, the dynamic digital twin model also integrates the impact of distributed energy access during simulation. This solution is an enhancement to address the trend of high-proportion renewable energy access to new power systems, and its technical solution is implemented through the following steps S1034 to S1036:
[0132] Step S1034: In the power flow calculation sub-model of the dynamic digital twin model, input the output prediction data of distributed energy sources such as photovoltaic and energy storage systems and their grid connection point information.
[0133] In modern charging stations, photovoltaic (PV) carports and in-station energy storage systems are becoming increasingly common. This embodiment expands the boundaries of the digital twin model. In the topology diagram of the power flow calculation sub-model, distributed energy nodes such as photovoltaics (PV) and energy storage systems (ESS) are added, and their grid connection locations are marked. Simultaneously, PV output forecast curves for future periods are obtained from weather forecast systems or local forecasting modules; and planned charge / discharge power curves or operating strategies are obtained from the energy storage management system. These data serve as boundary condition inputs for power flow calculation.
[0134] Step S1035: When simulating and extrapolating load distribution schemes, the random output of distributed energy is taken as one of the boundary conditions to evaluate the network robustness of different load distribution schemes under the fluctuation of distributed energy.
[0135] Photovoltaic power output is intermittent and random. This embodiment uses a scenario-based approach or probability distribution to describe this uncertainty. For example, multiple photovoltaic power output scenarios are generated (typical day, weak sunlight day, and day with severe fluctuations), with each scenario corresponding to a... Curves. When simulating load allocation schemes, performance is evaluated not only under the predicted expected scenario but also in parallel under multiple different stochastic output scenarios. A key observation is whether the planned charging load allocation scheme will cause significant grid voltage fluctuations or transformer overload when photovoltaic output suddenly drops. This allows for the assessment of the network robustness of each scheduling scheme under distributed energy fluctuations. Robustness can be quantified as the proportion of scenarios in which key system constraints (voltage, load) are not violated, or the average degree of violation.
[0136] Step S1036: When generating scheduling policies, the load balancing decision engine uses the robustness evaluation results as an optimization objective.
[0137] In the subsequent intelligent decision-making (step S104), the reward function of the decision engine will incorporate a "robustness" metric. A scheduling strategy that performs well in the desired scenario but is vulnerable (frequently exceeding limits) in multiple random scenarios will have its overall reward reduced. Conversely, a "robust" strategy that may not be optimal in the desired scenario but maintains system safety in all random scenarios will receive a higher reward. In this way, the system is driven to generate not only optimized but also robust load balancing scheduling strategies, significantly improving the adaptability and reliability of the charging network when operating in conjunction with a high proportion of new energy sources.
[0138] Step S104: Based on the results of parallel simulation, a load balancing decision engine based on deep reinforcement learning is adopted to generate a globally optimal dynamic scheduling strategy. The dynamic scheduling strategy assigns a specific charging pile, charging start time and charging power curve to each charging request, and ensures that the network meets the charging needs of all users while the overall load variance is lower than a preset threshold and the grid constraints are not violated.
[0139] Step S104 aims to solve the ultra-high-dimensional, nonlinear, and strongly coupled sequential decision-making problem of "finding the globally optimal load allocation scheme in real time while meeting the needs of a large number of differentiated users and complex power grid security constraints".
[0140] Existing technical solutions, whether based on fixed rules or optimizedrs based on traditional mathematical programming (e.g., linear programming, dynamic programming), suffer from fundamental limitations when faced with real-time influx of charging requests, rapidly changing grid conditions, and conflicting optimization objectives (user satisfaction, load balancing, and economy): rule-based engines are rigid and unable to handle unpredictable scenarios; traditional optimization methods have high computational complexity, making it difficult to meet the "second-level / minute-level" real-time response requirements, and exhibit poor adaptability to dynamic environmental changes. Therefore, this application abandons the traditional approach and innovatively adopts a load balancing decision engine based on deep reinforcement learning. This engine, through continuous "trial and error-learning" in a high-fidelity, zero-risk virtual environment constructed using a dynamic digital twin model, autonomously masters the ability to generate scheduling strategies approaching the global optimum, achieving a paradigm shift from "rule-based response" to "agent-learning-based optimization."
[0141] Specifically, as an embodiment of this application, a load balancing decision engine based on deep reinforcement learning is used to generate a globally optimal dynamic scheduling strategy, which can be achieved through the following steps S1041 to S1043, as detailed below:
[0142] Step S1041: Construct the load balancing decision problem as a Markov decision process, where the state space S is jointly defined by the network panoramic dynamic data view and the queue of charging requests to be processed, the action space A is defined as the set of allocation schemes for all charging requests to be scheduled, and the reward function R is designed as a comprehensive evaluation of load balancing degree, user satisfaction and power grid security indicators.
[0143] The following details the state space S, action space A, and reward function R in a Markov decision process:
[0144] 1) State space S, that is: at each decision time t, the state It is a high-dimensional tensor consisting of two parts: a feature representation of the network's panoramic dynamic data view and a feature representation of the queue of charging requests to be processed.
[0145] 1.1) Feature representation of the network panoramic dynamic data view: namely, the coded real-time system status output in step S102, including the load rate and voltage level of each power grid node, the real-time power, status, temperature of each charging pile, and environmental parameters, etc.
[0146] 1.2) Feature representation of the pending charging request queue: including feature vectors of all currently unserved or queued requests. Each request vector contains its user identity, target charging pile cluster identity, expected charging capacity, expected completion time, and user-defined charging priority weight. And the waiting time.
[0147] It fully depicts "how the system is" and "what the task is" at the moment of decision-making.
[0148] 2) Action Space A: Action This is the output of the decision engine, representing a scheduling scheme. It is defined as the allocation set of all pending requests. For a request, the allocation scheme is a triple: (allocated charging station ID, charging start time, charging power curve), where the charging start time can be discretized within a future time window; the charging power curve can be simplified to a set of discrete power levels. Therefore, the action space is a composite action space, which is enormous. To reduce dimensionality, hierarchical decision-making or parameterized representation of actions can be used.
[0149] 3) Reward function R: Reward function It serves as a guiding force for the learning of intelligent agents, and its design must comprehensively reflect the demands of multi-objective optimization:
[0150]
[0151] in:
[0152] As a load balancing incentive, it is often a negative overall network load variance or Gini coefficient to encourage even load distribution.
[0153] Rewards for user satisfaction are based on the deviation between the actual completion time and the user's expected completion time, the actual cost, and the user's priority weight. Quantification is performed. User priority weights are also considered. Users with high speed (high priority) will be penalized more severely for timeouts; this applies to user priority weighting. Users with low (high-cost) electricity bills face a heavier penalty.
[0154] Rewards for grid safety: Imposing large negative rewards (penalties) on any voltage overruns or line / transformer overloads.
[0155] Rewards for demand-side responses: For example, as described in step S101, positive incentives are given for scheduling decisions that successfully guide users to adjust their demands.
[0156] , and These are weighting coefficients used to balance the relative importance of different objectives.
[0157] Step S1042: Use the dynamic digital twin model as an environment simulator for reinforcement learning to evaluate the immediate reward of action A in state S and the next state. .
[0158] Deep reinforcement learning requires interaction with the environment to learn. In this embodiment, the dynamic digital twin model (the high-fidelity model constructed and updated in step S103) acts as an environment simulator. The decision engine (agent) outputs an action. That is, after the scheduling scheme is determined, it does not need to be executed in the real network, but is instead input into the digital twin model. The dynamic digital twin model is based on the current state. and actions The simulation projects the system's evolution over a future period and outputs two key results: 1) the immediate reward calculated using the reward function described above during the simulation. ;2) The new state at the end of the simulation This constitutes a complete "state-action-reward-new state" transition sample. Using digital twins as an environment allows for the low-cost, high-efficiency, and risk-free generation of massive amounts of training data, which is a prerequisite for the application of deep reinforcement learning.
[0159] Step S1043: Deploy a deep Q network in the load balancing decision engine. Through interaction with the environment simulator and training with offline historical data, it learns and outputs the optimal action that can obtain the maximum long-term cumulative reward in the current state. This optimal action constitutes the dynamic scheduling strategy.
[0160] At the core of the decision engine is a deep Q-network, a deep reinforcement learning algorithm that approximates the value function. Specifically, the deep Q-network structure is a deep neural network whose input is a vectorized representation of the state S, and whose output is the Q-value for all possible actions A. Where Q represents the expected long-term cumulative discounted reward obtained after performing action A in state S, and the network parameters are... Deep Q-networks typically consist of convolutional layers (for processing parts of the state that have spatial topological relationships, such as load distribution maps of power grid nodes) and fully connected layers (for processing other features).
[0161] In this embodiment, the deep Q-network employs techniques such as experience replay and target network for stable training, mainly including the following stages: interaction and storage, sampling and learning, and policy generation:
[0162] 1) Interaction and Storage: The intelligent agent interacts with the digital twin environment simulator, generating transfer samples. Stored in an experience playback buffer.
[0163] 2) Sampling and Learning: Periodically sample a small batch of samples randomly from the buffer. Utilize the target Q-network (parameters are...) Calculate the target Q value: ,in, It is the discount factor. Then, by minimizing the current Q-network prediction... The main network parameters are updated using the temporal difference error between the target value y and the target value y. :
[0164]
[0165] 3) Policy Generation: After training, for any real-time state S, the optimal action is... Depend on Give this optimal action. This is the globally optimal dynamic scheduling strategy. The strategy ensures the maximization of long-term cumulative rewards, thus achieving a long-term optimal trade-off among multiple objectives.
[0166] As an embodiment of ensuring decision reliability in this application, the load balancing decision engine based on deep reinforcement learning used in the above embodiment to generate the globally optimal dynamic scheduling strategy may further include an adversarial sample robustness enhancement mechanism. This is because, considering that in actual deployment, state observation S may contain abnormal or malicious disturbances due to sensor noise, communication interference, or data poisoning attacks, this mechanism aims to improve the robustness of the decision engine. Specifically, the adversarial sample robustness enhancement mechanism is implemented through the following steps S1044 to S1046:
[0167] Step S1044: Before inputting state S into the deep Q network, apply a small-range random perturbation that conforms to the historical noise distribution to the key features in state S to generate multiple perturbation state variants. The key features in state S include node load and charging request distribution.
[0168] During the online decision-making phase, the original state S is not directly input into the network. The system maintains a historical noise model that describes the distribution of observation errors for key features (such as node load readings and the geographical distribution density of charging requests) under normal conditions (e.g., mean 0, variance 0). (Gaussian distribution). For the current state S, K perturbation versions are generated by sampling the key feature vectors. ,in, , This simulates possible observational uncertainties.
[0169] Step S1045: Input the state variants into the deep Q network to obtain the corresponding multiple action outputs and Q-value estimates.
[0170] The K perturbation state variants are sequentially input into the trained deep Q-network to obtain K action outputs. and the corresponding Q-value estimate .
[0171] Step S1046: Based on the variance of multiple Q-value estimates, assess the uncertainty of the current decision. If the uncertainty exceeds the threshold, the decision engine will call the rule-based fallback strategy module to generate a conservative but safe scheduling strategy.
[0172] Calculate the variance of these K Q-value estimates. Its magnitude directly reflects the degree to which small perturbations in the current state affect the network's output Q-value (i.e., the judgment of the value of actions). The larger the variance, the more uncertain and fragile the network's decisions are near that state. A variance threshold is set. .like If the decision is deemed reliable, the action with the highest Q value will be adopted. As the final output. If If this occurs, a fallback strategy is triggered. The fallback strategy module is a pre-designed scheduler based on conservative rules (e.g., "all charging piles operate at 80% of their rated power" or "adopt the simplest first-come, first-served scheduling"). Enabling the fallback strategy ensures that when the AI decision engine faces a highly uncertain "unfamiliar" state that may lead to dangerous actions, the system can safely degrade, prioritizing the basic safety of the power grid, reflecting the design principle of "safety first".
[0173] In another embodiment of this application, after the adversarial sample robustness enhancement step, a formal security verification process for the decision strategy is further included, namely: extracting the policy network in the current state from the deep reinforcement learning decision engine and converting its decision logic into a verifiable equivalent neural network abstract model; defining a series of formal specifications based on the electrical safety constraints of the charging network (including node voltage upper and lower limits and line current thermal stability limits) to describe the security attributes that the scheduling strategy should satisfy under any permissible input disturbance; using a formal verification method combining symbolic execution and interval arithmetic to automatically verify the equivalent neural network abstract model, proving or disproving that its output strategy does not violate the security attributes under all possible inputs; if the verification fails, the action space of the strategy is automatically shrunk, and the decision engine is guided to regenerate the strategy within the security boundary. This scheme introduces formal verification methods from the fields of high-reliability software and chip design to provide mathematically rigorous security proofs for AI decision-making strategies. Specifically, the scheme applies verification techniques such as neural network abstraction, formal specification description, and symbolic execution to the verification of real-time charging scheduling strategies, providing theoretically stronger security guarantees that far exceed traditional methods based on probability and thresholds. When verification fails, its mechanism of automatically shrinking the action space and guiding regeneration realizes the dynamic and closed-loop interaction between security constraints and AI decision-making, enabling the system to strictly safeguard the bottom line of security while pursuing optimal performance.
[0174] As an example of how the system achieves continuous self-evolution in this application, Figure 1 The example method may also include policy performance evaluation and knowledge distillation steps, implemented through the following steps S1047a to 1047d:
[0175] Step S1047a: After completing the control of a scheduling cycle, collect the complete running data of the real network after executing the dynamic scheduling strategy within the scheduling cycle, as a training sample.
[0176] After the scheduling policy generated by the decision engine is actually executed on the physical network (via step S105), the system collects the complete data loop within this cycle, including: the input state S, the executed policy (action A), and the subsequent state fed back from the real environment. And the rewards calculated based on the actual results This constitutes a transfer sample originating from the real world. .
[0177] Step S1047b: Add the training samples to the experience replay buffer of the deep reinforcement learning decision engine.
[0178] This real sample is added to the experience replay buffer used for training the decision engine. By mixing it with purely simulated samples, the difference between the digital twin model and the real world (Sim-to-Real Gap) can be gradually corrected, making the learning strategy closer to actual physical laws.
[0179] Step S1047c: Periodically sample data from the experience replay buffer, fine-tune the deep Q network in the decision engine offline, and solidify the policy network parameters with stable performance improvement into a policy "snapshot" and store it in the policy library.
[0180] The system periodically (e.g., during daily off-peak hours) initiates an offline learning task. The deep Q-network is incrementally fine-tuned by sampling from an experience replay buffer containing both new and old samples. When the evaluation shows a consistent improvement in the network's performance (e.g., average reward on the validation set) compared to previous versions, the current network parameters are adjusted. Save it as a policy "snapshot" and store it in a policy library. Each snapshot is associated with metadata describing its strengths in different scenarios (such as "good at handling load fluctuations during midday solar PV peaks").
[0181] Step S1047d: When the online decision engine faces a state of high uncertainty, it has the ability to quickly retrieve and apply high-performance "snapshot" strategies from the strategy library under similar historical scenarios.
[0182] This step is linked to the adversarial robustness enhancement step (i.e., step S1046). When the online engine is about to trigger a fallback strategy due to high uncertainty, it can first try to find a solution from the policy library. The system calculates the similarity between the current state and the scene metadata associated with each snapshot in the policy library. If a highly similar snapshot is found, the network parameters of that snapshot are directly loaded. A temporary alternative to the online network is used, and decisions are generated from this network. This is equivalent to invoking expert strategies from historical experience for handling similar situations, which are often better than conservative fallback strategies, thus maintaining the system's intelligent performance as much as possible while ensuring safety.
[0183] In another embodiment of this application, the above-mentioned policy performance evaluation and knowledge distillation can also be implemented as follows: A hybrid experience pool is constructed, comprising an elite pool storing high-frequency successful experiences, an exploration pool storing exploratory experiences, and a safety pool storing rare but important critical failure experiences; a meta-learner is used to meta-train the data sampled from the hybrid experience pool, enabling the base model of the deep reinforcement learning decision engine to have the meta-learning ability to quickly adapt to new operating scenarios (e.g., holiday patterns, extreme weather); a neural architecture search process is periodically initiated to automatically explore and evaluate miniaturized or sparse policy network architectures more suitable for the current network topology and load characteristics within a limited search space, replacing the original network and achieving co-evolution of model performance and efficiency. Compared to the schemes in steps S1047a to S1047d, where knowledge distillation focuses on preserving and reusing historical policies, this scheme constructs a systematic continuous learning and evolution system. Its hybrid experience pool replay manages and utilizes data more scientifically, and the introduction of meta-learning enables the model to have advanced learning capabilities, actively adapting to unknown scenarios. The introduction of neural architecture search also enables the model structure to evolve adaptively. Therefore, this solution enables the entire decision-making system to learn not only from data, but also to have the self-evolutionary ability to manage the learning process, optimize learning objectives, and evolve learning carriers, greatly reducing the reliance on manual adjustments in long-term system maintenance. On the other hand, it deeply integrates cutting-edge AI methods such as meta-learning and neural architecture search with industrial control systems to solve the problem of continuous model adaptability, pointing to an advanced direction for AI industrial applications.
[0184] As an advanced embodiment of this application addressing ultra-large-scale, distributed charging networks, the deep reinforcement learning decision engine can employ a multi-agent architecture for training and decision-making. This process is implemented through steps S10411 to S10414, as detailed below:
[0185] Step S10411: Model each edge computing node or each key power grid node in the network as an intelligent agent, and each intelligent agent has local observation capabilities.
[0186] For charging networks with a wide geographical area and a large number of nodes, centralized single-agent decision-making faces problems such as state space explosion and training difficulties. This embodiment adopts a multi-agent reinforcement learning framework. The entire charging network is divided into regions, with each region managed by an edge computing node, or the jurisdiction of each key grid node (e.g., a distribution transformer) is considered as a region. Each region is modeled as an independent agent. Local observations of agent i. It is limited to status information within its jurisdiction (e.g., the status of charging piles within the area, local load), and cannot obtain global information. This is consistent with the information structure of a real distributed system.
[0187] Step S10412: The global load balancing decision problem is modeled as a cooperative multi-agent reinforcement learning problem, in which each agent learns a joint policy to maximize the global reward.
[0188] All agents share a global reward. (i.e., the comprehensive reward defined in step S1041). Their goal is to cooperate and learn a joint strategy. This maximizes the long-term global reward. The actions of each agent... It is a scheduling scheme for charging requests within its jurisdiction.
[0189] Step S10413: During the training phase, a centralized training and distributed execution paradigm is adopted, and multi-agent parallel simulation training is performed using a dynamic digital twin model.
[0190] The training employs the CTDE (Centralized Training with Decentralized Execution) paradigm, one of the successful paradigms for multi-agent reinforcement learning, as explained below:
[0191] 1) Centralized training: During training, there is a centralized network of commentators that can access the local observations of all agents. The global state S is used to more accurately evaluate the Q-value of the joint action. The actor networks of individual agents, however, rely only on their own local observations. Output actions. Training takes place in a simulation environment provided by the digital twin model, which allows a large number of simulation instances to be run in parallel, accelerating learning.
[0192] 2) Distributed execution: After training, during the execution phase, each agent relies solely on its own actor network and local observations. Decisions can be made without real-time communication with other intelligent agents or reliance on a central node, which meets the real-time and reliability requirements of distributed control.
[0193] Step S10414: During the execution phase, each agent independently outputs control suggestions for its governed units based on its local observations, and a top-level coordinator conducts final arbitration to form a unified dynamic scheduling strategy.
[0194] In actual operation, each agent independently generates scheduling proposals for its local area. These proposals are uploaded to a lightweight top-level coordinator. The coordinator does not perform complex replanning; it only performs simple conflict detection and arbitration (e.g., if agents in two adjacent areas propose transferring load to the same boundary transformer, it may lead to overload). After resolving conflicts, the coordinator compiles the scheduling schemes from each area into a unified global dynamic scheduling policy and distributes it for execution. This architecture achieves a balance between scalability of decision-making capabilities and execution efficiency in ultra-large-scale networks.
[0195] Step S105: Decompose the dynamic scheduling strategy into a series of time-ordered, charging pile group coordinated control command sequences, and send them to the corresponding edge computing nodes and charging pile controllers through the cloud-edge collaborative architecture to drive the charging network to perform load balancing control.
[0196] Step S105 aims to solve the "last mile" problem of reliably, in real-time, and collaboratively implementing intelligent scheduling strategies in complex, distributed physical systems. Existing purely centralized control architectures suffer from single-point-of-failure risks, high communication latency, and poor scalability, while fully distributed local control lacks a global perspective, making it difficult to achieve system-level load balancing optimization. Therefore, this application adopts a cloud-edge collaborative architecture and decomposes the dynamic scheduling strategy into a time-ordered, cluster-coordinated control command sequence, achieving a perfect combination of global optimization decision-making and precise local control and rapid response, ensuring that the optimal strategy can efficiently and robustly drive the physical charging network.
[0197] Specifically, the dynamic scheduling strategy is decomposed into a series of time-ordered, cluster-coordinated control command sequences, and distributed to the corresponding edge computing nodes and charging pile controllers through a cloud-edge collaborative architecture. This can be achieved through the following steps S1051 to S1053:
[0198] Step S1051: The cloud center server is responsible for generating a global control instruction framework with coarse time granularity. The control instruction framework specifies the total power quota and key control objectives of each edge node in the next time slice.
[0199] In the cloud, the dynamic scheduling strategy generated in step S104 is a global optimization scheme with a long time span (e.g., several hours in the future). The cloud does not directly translate this into microsecond-level instructions sent to each charging station; instead, it decomposes and abstracts the task. Based on the scheduling strategy, the cloud divides the time into larger time granularities (e.g., 15 minutes as a time slice). For each time slice, and for the set of all charging stations managed by the same edge computing node within that time slice, the cloud generates a control instruction framework. This framework is a high-level, goal-oriented instruction, whose content mainly includes total power quota, key control objectives, and time references, etc.
[0200] 1) Total power quota: The upper limit of the total active power of all charging piles within the jurisdiction of this edge node during this time slice. .
[0201] 2) Key control objectives: For example, load balancing objectives in the region (minimizing the load variance of each transformer in the region), a list of specific high-priority charging tasks that must be guaranteed, and grid constraints that need to be avoided (e.g., the voltage of a certain node must not fall below a certain value), etc.
[0202] 3) Time reference: Start time of the framework taking effect and end time .
[0203] This framework does not specify which stub runs at what power, but instead provides a "resource package" and an "optimization target," giving edge nodes ample room for autonomous decision-making.
[0204] Step S1052: After receiving the control command framework, each edge computing node combines the microsecond-level real-time status of the charging piles within its jurisdiction and the local optimization objectives to perform fine-grained decomposition and time synchronization of the control command framework, generating specific power adjustment commands with timestamps.
[0205] Edge computing nodes are deployed in charging stations or power distribution rooms, close to the controlled equipment, and have low-latency communication and fast computing capabilities, including receiving and parsing, local state awareness, fine-grained decomposition, and time synchronization, etc.
[0206] 1) Receiving and parsing: Specifically, edge nodes receive control command frameworks from the cloud through IoT networks (such as 5G and industrial Ethernet) and parse the constraints and objectives within them.
[0207] 2) Local state awareness, specifically including: edge nodes collect the real-time status of each charging pile under their jurisdiction via local buses (such as CAN, RS485) or high-speed local area networks at microsecond intervals, including instantaneous power, on / off status, connection status, and more detailed local sensor data (such as cabinet temperature). The real-time nature of this information is far higher than that of data uploaded to the cloud.
[0208] 3) Fine-grained decomposition, i.e., running a local optimizer within the edge node. This optimizer is constrained by the framework deployed from the cloud, guided by the control objectives within the framework, and combines its unique, millisecond / microsecond-level real-time status information to perform rapid secondary optimization calculations. Its task is to optimize the "total power quota". The system decomposes and assigns power adjustments to each subordinate charging station, generating a specific power adjustment command sequence for each station with timestamps accurate to milliseconds and power values accurate to kilowatts. For example, the command might be: "At absolute timestamp T=2023-10-27 14:30:00.500, set the output power of charging station Pile-001 to 50.0kW." This achieves localized and fine-grained execution of the global framework.
[0209] 4) Time synchronization: This is to ensure that the command execution time of all charging piles in the area is strictly synchronized and to avoid instantaneous power surges caused by asynchrony. A precision clock synchronization protocol (e.g., IEEE 1588 PTP) is used between the edge node and the subordinate charging pile controller to ensure that the clock deviation of all devices is within the microsecond level.
[0210] Step S1053: The charging pile controller executes the power adjustment command and feeds back the execution result and real-time status to the corresponding edge node. The edge node summarizes the data and uploads it asynchronously to the cloud for closed-loop optimization.
[0211] in:
[0212] 1) Command execution: Specifically, the charging pile controller receives timestamped power commands from edge nodes and drives the power module to adjust its output at the specified precise moment. The controller typically has local closed-loop control capabilities to ensure that the actual output power quickly and accurately tracks the set value.
[0213] 2) Feedback and Closed Loop: Specifically, the charging pile controller feeds back the command execution results (actual output power, execution success / failure status) and more detailed real-time status (such as module temperature, error codes) to its subordinate edge nodes in real time. The edge nodes aggregate the feedback information from all subordinate piles for preliminary diagnosis and processing. Simultaneously, the edge nodes asynchronously and non-real-time package and upload this aggregated execution result data to the cloud. The cloud platform utilizes this feedback data from the real world for multiple closed-loop optimizations: 1) verifying the accuracy of the digital twin model; 2) evaluating the actual execution effect of the scheduling strategy as training samples for reinforcement learning (see step S1047); 3) for online learning of the cloud-based prediction model. This constitutes a complete closed loop of "perception-decision-execution-learning".
[0214] As another embodiment of this application, the dynamic scheduling strategy can be decomposed into a series of time-ordered, cluster-coordinated control command sequences, and distributed to the corresponding edge computing nodes and charging pile controllers through a cloud-edge collaborative architecture. Alternatively, the cloud central server can monitor the computing load, pending command queue depth, and network communication quality of each edge computing node in real time. Based on the monitoring information, an online optimization algorithm is used to dynamically decide the computing task offloading strategy and model accuracy-speed trade-off strategy for each edge node: for edge nodes with high load or high latency, some of their fine-grained decomposition computing tasks or high-precision model inference tasks are dynamically migrated to the cloud or nearby lightly loaded edge nodes for execution. When the control command sequence is distributed, appropriate computing resource packages and model version identifiers are synchronously allocated to each control command framework to ensure that edge nodes can still complete the command decomposition on time and with high quality under resource constraints, thereby achieving the global optimal elastic allocation of cloud-edge computing resources. This solution introduces intelligent perception and dynamic orchestration of computing resources, upgrading "cloud-edge collaboration" from a fixed task division to an intelligent system that dynamically schedules computing tasks and model resources based on real-time system status. Specifically, the solution explicitly proposes computing task offloading and precision-speed trade-offs, which directly addresses the core contradiction in the actual deployment of edge AI (i.e., limited resources vs. computing demands). Through online optimization, it achieves elastic resource allocation, ensuring the overall stability and responsiveness of the system under high load or abnormal conditions. Furthermore, by incorporating computing and communication resources into a global optimization framework for load balancing, the system optimization dimension expands from a single power flow to a multi-flow collaboration of "power + computing power + communication," making the system more advanced and complete.
[0215] As an embodiment of this application to ensure the physical safety of the power grid under extreme conditions, when performing fine-grained decomposition at the edge computing node, a local emergency overload protection step is also executed, specifically including steps S1054 to S1056:
[0216] Step S1054: The edge computing node continuously monitors the real-time current and real-time voltage of the local power grid nodes it is connected to.
[0217] This is a local, hardware-based security defense independent of cloud commands. Edge nodes continuously monitor the three-phase current at their grid connection point at millisecond-level frequencies via high-precision electrical measurement units (or data obtained from local protection devices). and phase voltage These measurements are direct, fast, and reliable.
[0218] Step S1055: When the real-time current or real-time voltage exceeds the safety limit, without waiting for instructions from the cloud, immediately reduce the power or suspend the charging pile under its jurisdiction based on the preset priority rules, and simultaneously send an alarm to the cloud.
[0219] When the current in any phase exceeds the instantaneous current protection setting of the line / switch, Or the voltage exceeds the voltage protection setting (e.g., or When this occurs, the protection logic of the edge node is triggered immediately. The key is that "no need to wait for cloud commands," meaning this protection action is completely localized and autonomous, with a response time in the millisecond range, sufficient to suppress potential short-circuit current surges or voltage drops. The following explains the execution of the protection action and the synchronized warnings.
[0220] Protection action execution: Based on preset priority rules (e.g., prioritizing charging for specific vehicles such as fire trucks and emergency vehicles; within the same priority range, following rules such as "last charging vehicle disconnects first" or "highest power vehicle is charged first"), the edge node immediately issues instructions to reduce power or directly suspend charging for the charging piles under its jurisdiction. This action is decisive and mandatory.
[0221] Synchronous alarm: Almost simultaneously with the execution of the protection action, the edge node generates a high-priority alarm message, which is reported to the cloud platform via the fastest path. The alarm message includes the over-limit parameters, action time, list of controlled stakes, action result, etc.
[0222] Step S1056: After receiving the alarm, the cloud will re-coordinate the control commands of other edge nodes in the next scheduling cycle to compensate for the deviation in the global scheduling plan caused by the protection action of the node.
[0223] Upon receiving an emergency alarm, the cloud understands that a certain area has partially withdrawn its load due to a protection action. When the next regular scheduling cycle begins (e.g., a few minutes later), the cloud will consider the load reduction in that area as a new boundary condition when making global optimization decisions (step S104). The cloud may re-coordinate the control command framework of other unaffected edge nodes; for example, it may appropriately increase the charging power quota for other areas, or prioritize users who did not complete charging due to the protection action when scheduling new requests. This compensates for scheduling plan deviations caused by local protection at the global level, maintaining the achievement of the overall optimization goal. This embodies the collaborative concept of "local security first, global optimization later."
[0224] As an embodiment of this application that enables the system to adapt to changes in network structure, Figure 1 The example method performs a topology adaptive optimization step during the initialization phase or when a significant change occurs in the network topology. The specific process is detailed in steps S1057 to S1059 as follows:
[0225] Step S1057: Automatically identify the electrical connection relationships of all charging piles in the charging network, the capacity of the upstream transformer, and the line impedance parameters, and construct or update the network electrical topology diagram.
[0226] The charging network may be expanded or upgraded, and electrical connections may change. This step enables automatic topology discovery and updating, which can be achieved automatically or semi-automatically through various methods, including analysis of outage messages, power flow correlation analysis, and importing planning documents, etc.
[0227] 1. Power outage message analysis: When a circuit breaker trips, analyze which charging piles lose power supply, and thus infer their electrical connection relationship.
[0228] 2. Based on power flow correlation analysis: Long-term monitoring of the correlation between the power of each charging pile and the total power of the upstream transformer / line. Pile with high correlation are likely to belong to the same power supply branch.
[0229] 3. Import based on planning documents: Import standard electrical single-line diagrams and data from the distribution management system (DMS) or engineering design.
[0230] Parameters such as transformer capacity and line impedance can be obtained from equipment nameplates, design databases, or through parameter identification algorithms. The identified "pile-transformer-feeder" connection relationships and electrical parameters are structured and stored as a graph model, i.e., a network electrical topology diagram, which is the foundation for all advanced applications (power flow calculation, area partitioning).
[0231] Step S1058: Based on the network electrical topology diagram, redefine the management scope of edge computing nodes in the cloud-edge collaborative architecture. The goal is to make the electrical coupling of charging piles within the management scope of each node tight, while the electrical coupling between nodes is sparse.
[0232] This is the core of dynamic optimization of the cloud-edge collaborative management scope. The initial management scope division may be geographical or manual, and may not be optimal. This step re-divides the scope based on the latest electrical topology map, using electrical coupling as the criterion. Specifically, the entire charging network is divided into several regions, each managed by an edge node. The goal of this division is to maximize the electrical coupling between charging piles within a region (i.e., strong power interaction and voltage influence between them) and minimize the electrical coupling between different regions. This allows most optimization and control problems to be solved locally within the region, reducing the need for coordination and communication overhead between regions, and improving the overall system operating efficiency. The above problem can be formalized as a graph partitioning problem, i.e., using charging piles and grid nodes as vertices, and using the electrical distance (e.g., impedance magnitude) or power sensitivity between them as edge weights. Graph partitioning algorithms (e.g., spectral clustering, multi-level segmentation algorithms) are used to segment the graph, maximizing the sum of edge weights within the subgraphs and minimizing the sum of edge weights between subgraphs. The partitioning result determines the new management scope of each edge node.
[0233] Step S1059: Based on the new management scope division, reconfigure the allocation relationship of data acquisition, command issuance and edge computing tasks.
[0234] After the management scope is redefined, the system automatically performs a series of configuration updates, including data collection redirection, command issuance path updates, edge task deployment, and metadata synchronization:
[0235] 1. Data Acquisition Redirection: Adjust the data acquisition route to ensure that real-time data from each charging station is sent to its new edge node.
[0236] 2. Command delivery path update: Update the mapping relationship between the cloud and edge nodes to ensure that the control command framework issued by the cloud can be correctly routed to the new responsible edge node.
[0237] 3. Edge task deployment: Migrate or deploy the corresponding local optimizer, protection logic and other software tasks to new edge nodes.
[0238] 4. Metadata synchronization: Synchronize the new topology partitioning results to all relevant modules such as the cloud-based digital twin model and global decision engine.
[0239] Through the above steps, the entire cloud-edge collaborative system can adapt to the expansion and changes in the charging network topology, always maintaining the optimal organizational structure and operational efficiency.
[0240] From the above appendix Figure 1The example of the load balancing control method for new energy vehicle charging networks demonstrates that, on the one hand, by collecting multi-source heterogeneous data in parallel and performing spatiotemporal alignment and feature fusion to generate a dynamic panoramic data view of the network, information silos are overcome, achieving a unified, comprehensive, and structured digital representation of the real-time state of the charging network-grid coupled system, providing reliable and multi-dimensional data support for advanced decision-making. On the other hand, based on the dynamic panoramic data view of the network, a dynamic digital twin model is constructed and updated in real time. This model is then used to perform millisecond-level parallel simulations of various load allocation schemes for the requests to be processed. This allows the system to proactively evaluate the complex impacts of different schemes on the overall load balance of the network, the load of key equipment, and the stability of the grid voltage in a low-cost and zero-risk manner before actual scheduling, thereby screening out potential high-load-balanced systems. The performance scheduling strategy avoids the grid risks that trial-and-error control may bring. Thirdly, the load balancing decision engine based on deep reinforcement learning generates a globally optimal dynamic scheduling strategy that not only allocates specific charging resources and power curves to each request, but also ensures that while meeting all needs, the overall network load fluctuations are controlled and strictly adhere to grid safety constraints. This achieves an intelligent closed loop from perception and prediction to decision-making, improving the automation level and global optimization capability of load balancing control. Fourthly, the cloud is responsible for macro-level strategies and collaborative planning, while edge nodes are responsible for micro-level instruction decomposition and emergency handling. This architecture ensures the global consistency of control objectives and improves the real-time performance and reliability of instruction issuance and execution, ensuring that complex load balancing strategies can be reliably and collaboratively implemented in the physical charging network. In summary, the technical solution of this application achieves globally optimal load balancing scheduling by constructing a dynamic digital twin model and a deep reinforcement learning decision engine, effectively optimizing charging network performance and grid stability.
[0241] Please see the appendix Figure 2 This application provides a load balancing control device for a new energy vehicle charging network. The device may include a receiving module 201, a fusion module 202, an evaluation module 203, a generation module 204, and a driving module 205, as detailed below:
[0242] The receiving module 201 is used to receive and parse charging requests from multiple users. The charging request includes at least the user identity identifier, the target charging pile cluster identifier, the expected charging amount, the expected completion time, and the user-defined charging priority weight. The charging priority weight is used to quantify the user's preference for charging cost and charging speed.
[0243] The fusion module 202 is used to collect real-time operating status time-series data of all charging piles in the charging network, charging station environmental perception data and ultra-short-term load forecast data from the power grid dispatching system in parallel, and to perform spatiotemporal alignment and feature fusion of real-time operating status time-series data and ultra-short-term load forecast data to generate a dynamic panoramic data view of the network.
[0244] Evaluation module 203 is used to build and update a dynamic digital twin model of the charging network in real time based on the network panoramic dynamic data view. All pending charging requests are input into the dynamic digital twin model. The dynamic digital twin model performs millisecond-level parallel simulation and deduction of various possible load distribution schemes to evaluate the impact of each scheme on the overall load balance of the network, transformer peak load and grid node voltage stability.
[0245] The generation module 204 is used to generate a globally optimal dynamic scheduling strategy based on the results of parallel simulation and a load balancing decision engine based on deep reinforcement learning. The dynamic scheduling strategy assigns a specific charging pile, charging start time, and charging power curve to each charging request, and ensures that the network meets the charging needs of all users while the overall load variance is lower than a preset threshold and the grid constraints are not violated.
[0246] The driver module 205 is used to decompose the dynamic scheduling strategy into a series of time-ordered, charging pile group coordinated control command sequences, and send them to the corresponding edge computing nodes and charging pile controllers through the cloud-edge collaborative architecture to drive the charging network to perform load balancing control.
[0247] From the above appendix Figure 2As illustrated by the example of a load balancing control device for a new energy vehicle charging network, on the one hand, by collecting multi-source heterogeneous data in parallel and performing spatiotemporal alignment and feature fusion to generate a dynamic panoramic data view of the network, information silos are overcome. This achieves a unified, comprehensive, and structured digital representation of the real-time state of the charging network-grid coupled system, providing reliable and multi-dimensional data support for advanced decision-making. On the other hand, based on the dynamic panoramic data view of the network, a dynamic digital twin model is constructed and updated in real time. This model is then used to perform millisecond-level parallel simulations of various load allocation schemes for the requests to be processed. This allows the system to proactively evaluate the complex impacts of different schemes on the overall load balance of the network, the load of key equipment, and the stability of the grid voltage in a low-cost and zero-risk manner before actual scheduling is executed, thereby screening out potential high-load balance schemes. The performance scheduling strategy avoids the grid risks that trial-and-error control may bring. Thirdly, the load balancing decision engine based on deep reinforcement learning generates a globally optimal dynamic scheduling strategy that not only allocates specific charging resources and power curves to each request, but also ensures that while meeting all needs, the overall network load fluctuations are controlled and strictly adhere to grid safety constraints. This achieves an intelligent closed loop from perception and prediction to decision-making, improving the automation level and global optimization capability of load balancing control. Fourthly, the cloud is responsible for macro-level strategies and collaborative planning, while edge nodes are responsible for micro-level instruction decomposition and emergency handling. This architecture ensures the global consistency of control objectives and improves the real-time performance and reliability of instruction issuance and execution, ensuring that complex load balancing strategies can be reliably and collaboratively implemented in the physical charging network. In summary, the technical solution of this application achieves globally optimal load balancing scheduling by constructing a dynamic digital twin model and a deep reinforcement learning decision engine, effectively optimizing charging network performance and grid stability.
[0248] Figure 3 This is a schematic diagram of the structure of a device provided in one embodiment of this application. For example... Figure 3 As shown, the device 3 in this embodiment mainly includes: a processor 30, a memory 31, and a computer program 32 stored in the memory 31 and executable on the processor 30, such as a program for a load balancing control method for a new energy vehicle charging network. When the processor 30 executes the computer program 32, it implements the steps described in the above embodiment of the load balancing control method for a new energy vehicle charging network, for example... Figure 1 The steps S101 to S105 are shown. Alternatively, when the processor 30 executes the computer program 32, it implements the functions of each module / unit in the above-described device embodiments, for example... Figure 2 The functions of the receiving module 201, fusion module 202, evaluation module 203, generation module 204, and driving module 205 are shown.
[0249] For example, the computer program 32 of the new energy vehicle charging network load balancing control method mainly includes: receiving and parsing charging requests from multiple users, wherein the charging request includes at least user identification, target charging pile cluster identification, expected charging capacity, expected completion time, and user-defined charging priority weight, the charging priority weight being used to quantify the user's preference trade-off between charging cost and charging speed; parallelly collecting real-time operating status time-series data of all charging piles in the charging network, charging station environmental perception data, and ultra-short-term load forecast data from the power grid dispatching system, and performing spatiotemporal alignment and feature fusion of the real-time operating status time-series data and ultra-short-term load forecast data to generate a network panoramic dynamic data view; based on the network panoramic dynamic data view, constructing and updating a dynamic digital twin model of the charging network in real time, and inputting the current... All pending charging requests are simulated in parallel at millisecond levels using a dynamic digital twin model to evaluate the impact of each scheme on the overall network load balance, transformer peak load, and grid node voltage stability. Based on the results of the parallel simulation, a load balancing decision engine based on deep reinforcement learning is used to generate a globally optimal dynamic scheduling strategy. This strategy assigns a specific charging pile, charging start time, and charging power curve to each charging request, ensuring that the network meets the charging needs of all users while maintaining an overall load variance below a preset threshold and without violating grid constraints. The dynamic scheduling strategy is decomposed into a series of time-ordered, pile-group coordinated control instructions, which are then distributed to the corresponding edge computing nodes and charging pile controllers via a cloud-edge collaborative architecture to drive the charging network to perform load balancing control. The computer program 32 can be divided into one or more modules / units, which are stored in memory 31 and executed by processor 30 to complete this application. One or more modules / units can be a series of computer program instruction segments capable of performing specific functions, describing the execution process of the computer program 32 in device 3.For example, computer program 32 can be divided into the functions of receiving module 201, fusion module 202, evaluation module 203, generation module 204, and driving module 205 (a module in the virtual device). The specific functions of each module are as follows: Receiving module 201 is used to receive and parse charging requests from multiple users. The charging request includes at least the user's identity identifier, the target charging pile cluster identifier, the expected charging amount, the expected completion time, and the user-defined charging priority weight. The charging priority weight is used to quantify the user's preference trade-off between charging cost and charging speed. Fusion module 202 is used to collect real-time operating status time-series data of all charging piles in the charging network, charging station environmental perception data, and ultra-short-term load forecast data from the power grid dispatching system in parallel. It performs spatiotemporal alignment and feature fusion of real-time operating status time-series data and ultra-short-term load forecast data to generate a network panoramic dynamic data view. Evaluation module 203 is used to construct and update in real time based on the network panoramic dynamic data view. The dynamic digital twin model of the charging network takes all pending charging requests as input and performs millisecond-level parallel simulations of various possible load allocation schemes to evaluate the impact of each scheme on the overall load balance of the network, transformer peak load, and grid node voltage stability. The generation module 204, based on the results of the parallel simulations, uses a deep reinforcement learning-based load balancing decision engine to generate a globally optimal dynamic scheduling strategy. This strategy assigns a specific charging pile, charging start time, and charging power curve to each charging request, ensuring that the network meets the charging needs of all users while maintaining an overall load variance below a preset threshold and without violating grid constraints. The driving module 205 decomposes the dynamic scheduling strategy into a series of time-ordered, pile-group coordinated control command sequences and distributes them to the corresponding edge computing nodes and charging pile controllers through a cloud-edge collaborative architecture to drive the charging network to perform load balancing control.
[0250] Device 3 may include, but is not limited to, processor 30 and memory 31. Those skilled in the art will understand that... Figure 3 This is merely an example of device 3 and does not constitute a limitation on device 3. It may include more or fewer components than shown, or combine certain components, or different components. For example, the device may also include input / output devices, network access devices, buses, etc.
[0251] The processor 30 may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0252] The memory 31 can be an internal storage unit of the device 3, such as a hard disk or RAM of the device 3. The memory 31 can also be an external storage device of the device 3, such as a plug-in hard disk, Smart MediaCard (SMC), Secure Digital (SD) card, or Flash Card equipped on the device 3. Furthermore, the memory 31 can include both internal and external storage units of the device 3. The memory 31 is used to store computer programs and other programs and data required by the device. The memory 31 can also be used to temporarily store data that has been output or will be output.
[0253] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed. That is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above-described device can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0254] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0255] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0256] In the embodiments provided in this application, it should be understood that the disclosed apparatus / device and method can be implemented in other ways. For example, the apparatus / device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0257] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0258] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0259] If integrated modules / units are implemented as software functional units and sold or used as independent products, they can be stored in a storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program of the new energy vehicle charging network load balancing control method can be stored in a storage medium. When the computer program is executed by a processor, it can implement the steps of the above method embodiments, namely, receiving and parsing charging requests from multiple users, wherein the charging request includes at least a user identity identifier, a target charging pile cluster identifier, a desired charging amount, a desired completion time, and a user-defined charging priority weight. The charging priority weight is used to quantify the user's preference trade-off between charging cost and charging speed; collecting real-time operating status time-series data of all charging piles in the charging network, charging station environmental perception data, and ultra-short-term load forecast data from the power grid dispatching system in parallel, and performing spatiotemporal alignment and feature fusion of the real-time operating status time-series data and ultra-short-term load forecast data to generate a network panoramic dynamic data view; based on the network... A dynamic data view is constructed and updated in real time to represent the dynamic digital twin model of the charging network. All pending charging requests are input into this model, and millisecond-level parallel simulations of various possible load allocation schemes are performed to evaluate their impact on the overall network load balance, transformer peak load, and grid node voltage stability. Based on the results of the parallel simulations, a load balancing decision engine based on deep reinforcement learning is used to generate a globally optimal dynamic scheduling strategy. This strategy assigns a specific charging pile, charging start time, and charging power curve to each charging request, ensuring that the network meets the charging needs of all users while maintaining an overall load variance below a preset threshold and without violating grid constraints. The dynamic scheduling strategy is decomposed into a series of time-ordered, pile-group coordinated control command sequences, which are then distributed to the corresponding edge computing nodes and charging pile controllers through a cloud-edge collaborative architecture to drive the charging network to perform load balancing control. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or some intermediate form. Storage media can include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the contents of storage media can be appropriately added to or removed according to the requirements of legislation and patent practice in a jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, storage media may not include electrical carrier signals and telecommunication signals.
[0260] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. These modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application. The specific embodiments described above further illustrate the purpose, technical solutions, and beneficial effects of this application. It should be understood that the above descriptions are merely specific embodiments of this application and are not intended to limit the protection scope of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A load balancing control method for a new energy vehicle charging network, characterized in that, The method includes: Receive and parse charging requests from multiple users. The charging request includes at least the user's identity identifier, the target charging pile cluster identifier, the expected charging amount, the expected completion time, and the user-defined charging priority weight. The charging priority weight is used to quantify the user's preference trade-off between charging cost and charging speed. The system collects real-time operating status time-series data of all charging piles in the charging network, charging station environmental perception data, and ultra-short-term load forecast data from the power grid dispatch system in parallel. The system then performs spatiotemporal alignment and feature fusion of the real-time operating status time-series data and the ultra-short-term load forecast data to generate a panoramic dynamic data view of the network. Based on the network panoramic dynamic data view, a dynamic digital twin model of the charging network is constructed and updated in real time. All current pending charging requests are input into the dynamic digital twin model. The dynamic digital twin model is used to perform millisecond-level parallel simulation and deduction of various possible load distribution schemes to evaluate the impact of each scheme on the overall load balance of the network, transformer peak load and grid node voltage stability. Based on the results of the parallel simulation, a load balancing decision engine based on deep reinforcement learning is adopted to generate a globally optimal dynamic scheduling strategy. The dynamic scheduling strategy assigns a specific charging pile, charging start time, and charging power curve to each charging request, and ensures that while meeting the charging needs of all users, the overall load variance is lower than a preset threshold and the grid constraints are not violated. The dynamic scheduling strategy is decomposed into a series of time-ordered, charging pile group coordinated control command sequences, which are then distributed to the corresponding edge computing nodes and charging pile controllers through a cloud-edge collaborative architecture to drive the charging network to perform the load balancing control.
2. The load balancing control method for new energy vehicle charging networks according to claim 1, characterized in that, The load balancing decision engine based on deep reinforcement learning generates a globally optimal dynamic scheduling strategy, including: The load balancing decision problem is constructed as a Markov decision process, where the state space S is jointly defined by the network panoramic dynamic data view and the queue of charging requests to be processed, the action space A is defined as the set of allocation schemes for all charging requests to be scheduled, and the reward function R is designed as a comprehensive evaluation of load balancing degree, user satisfaction and power grid security indicators. The dynamic digital twin model is used as an environment simulator for reinforcement learning to evaluate the immediate reward of action A in state S and the next state S'. A deep Q-network is deployed in the load balancing decision engine. Through interaction with the environment simulator and training with offline historical data, it learns and outputs the optimal action that can obtain the maximum long-term cumulative reward in the current state. The optimal action constitutes the dynamic scheduling strategy.
3. The load balancing control method for new energy vehicle charging networks according to claim 1, characterized in that, The step of performing spatiotemporal alignment and feature fusion of the real-time operating status time-series data and the ultra-short-term load forecast data to generate a dynamic panoramic data view of the network includes: The real-time operating status time-series data from charging piles is processed in real time using a sliding window-based streaming computing framework to calculate the instantaneous power, average power, power change rate, and estimated remaining charging time for each charging pile, thus obtaining the first processed data. The charging station environmental perception data is correlated and fused with the corresponding real-time operating status time series data to obtain the second processed data. The charging station environmental perception data includes temperature, humidity and parking space occupancy status. The first processed data and the second processed data are merged to form processed charging network data; The processed charging network data is matched and overlaid with the ultra-short-term load forecast data with spatiotemporal labels issued by the power grid dispatching system to form a unified panoramic dynamic data view of the network with power grid nodes and charging piles as basic units.
4. The load balancing control method for new energy vehicle charging networks according to claim 3, characterized in that, After forming a unified dynamic data view of the network panoramic view with grid nodes and charging piles as basic units, the following anomaly detection and resilient disaster recovery steps are also included: A hybrid anomaly detection model combining isolated forests and autoencoders is integrated into the streaming computing framework to monitor each data stream in the network panoramic dynamic data view in real time. When the data mode of a charging pile is detected as abnormal, the dynamic digital twin model is automatically triggered to perform logical isolation simulation on the abnormal charging pile, assess the impact of its failure on the network, and start the backup charging pile discovery and switching plan. The abnormal information, impact assessment, and switching plan are pushed to the load balancing decision engine in real time as one of the inputs for generating or adjusting the dynamic scheduling strategy.
5. The load balancing control method for new energy vehicle charging networks according to claim 2, characterized in that, The method of using a load balancing decision engine based on deep reinforcement learning to generate a globally optimal dynamic scheduling strategy also includes the following steps to enhance robustness against adversarial samples: Before inputting state S into the deep Q network, a small-range random perturbation conforming to the historical noise distribution is applied to the key features in state S to generate multiple perturbation state variants. The key features in state S include node load and charging request distribution. The state variants are input into a deep Q-network to obtain the corresponding multiple action outputs and Q-value estimates. Based on the variance of multiple Q-value estimates, the uncertainty of the current decision is assessed. If the assessed uncertainty exceeds a threshold, the decision engine will call the rule-based fallback strategy module to generate a conservative but safe scheduling strategy.
6. The load balancing control method for new energy vehicle charging networks according to claim 4, characterized in that, The dynamic scheduling strategy is decomposed into a series of time-ordered, charging pile-group coordinated control command sequences, and distributed to the corresponding edge computing nodes and charging pile controllers through a cloud-edge collaborative architecture to drive the charging network to perform the load balancing control, including: The cloud-based central server is responsible for generating a global control instruction framework with coarse time granularity. This control instruction framework specifies the total power quota and key control objectives for each edge node in the next time slice. After receiving the control command framework, each edge computing node combines the microsecond-level real-time status of the charging piles within its jurisdiction with the local optimization objectives to perform fine-grained decomposition and time synchronization of the control command framework, generating specific power adjustment commands with timestamps. The charging pile controller executes the power adjustment command and feeds back the execution result and real-time status to its respective edge node. The edge node summarizes the data and uploads it asynchronously to the cloud for closed-loop optimization.
7. The load balancing control method for new energy vehicle charging networks according to claim 1, characterized in that, After completing the control of a scheduling cycle, the following policy performance evaluation and knowledge distillation steps are also included: Collect complete operational data of the real network after executing the dynamic scheduling strategy within the scheduling period, and use it as a training sample; The training samples are added to the experience replay buffer of the deep reinforcement learning decision engine. Data is periodically sampled from the experience replay buffer to fine-tune the deep Q-network in the decision engine offline, and the policy network parameters with stable performance improvement are solidified as policy "snapshots" and stored in the policy library; When the online decision engine faces a state of high uncertainty, it has the ability to quickly retrieve and apply high-performance "snapshot" strategies from the strategy library based on similar historical scenarios.
8. A load balancing control device for a new energy vehicle charging network, characterized in that, The device includes: The receiving module is used to receive and parse charging requests from multiple users. The charging request includes at least the user's identity identifier, the target charging pile cluster identifier, the expected charging amount, the expected completion time, and the user-defined charging priority weight. The charging priority weight is used to quantify the user's preference trade-off between charging cost and charging speed. The fusion module is used to collect real-time operating status time-series data of all charging piles in the charging network, charging station environmental perception data, and ultra-short-term load forecast data from the power grid dispatch system in parallel, and to perform spatiotemporal alignment and feature fusion of the real-time operating status time-series data and the ultra-short-term load forecast data to generate a dynamic panoramic data view of the network. The evaluation module is used to construct and update a dynamic digital twin model of the charging network in real time based on the network panoramic dynamic data view. All pending charging requests are input into the dynamic digital twin model. The dynamic digital twin model is used to perform millisecond-level parallel simulation and deduction of various possible load distribution schemes to evaluate the impact of each scheme on the overall load balance of the network, transformer peak load and grid node voltage stability. The generation module is used to generate a globally optimal dynamic scheduling strategy based on the results of the parallel simulation and a load balancing decision engine based on deep reinforcement learning. The dynamic scheduling strategy assigns a specific charging pile, charging start time, and charging power curve to each charging request, and ensures that the network meets the charging needs of all users while the overall load variance is lower than a preset threshold and the grid constraints are not violated. The driving module is used to decompose the dynamic scheduling strategy into a series of time-ordered, charging pile group coordinated control command sequences, and send them to the corresponding edge computing nodes and charging pile controllers through the cloud-edge collaborative architecture, so as to drive the charging network to perform the load balancing control.
9. An apparatus comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 7.
10. A storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 7.