Charging pile power distribution method and system based on dynamic load balancing
Through multi-agent reinforcement learning and game coordination mechanism, the problems of traffic dynamics and electricity price fluctuations in charging pile power distribution are solved, intelligent and adaptive charging power distribution is achieved, battery health and grid load are optimized, operating costs are reduced, and charging system efficiency is improved.
Patent Information
- Application Number
- CN202510750632.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-09-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing charging pile power allocation methods are unable to effectively cope with traffic dynamics, electricity price fluctuations and battery health management, resulting in long charging waiting times, severe grid load fluctuations, shortened battery life and high operating costs.
By adopting multi-agent reinforcement learning, multi-objective optimization and game coordination mechanism, the local state information of the charging pile is obtained to predict the vehicle entry and exit frequency, electricity price trend and power demand in the future period. A game agent model and multi-objective reward function are constructed to achieve intelligent and adaptive dynamic allocation of charging power.
It achieves intelligent and adaptive charging power distribution, reduces charging waiting time, optimizes battery health and grid load, reduces operating costs, and improves the overall efficiency of the charging system and user experience.
Smart Images

Figure CN120645762A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of electric vehicle charging infrastructure, and in particular to a charging pile power distribution method and system based on dynamic load balancing, belonging to the cross-technical field of intelligent charging scheduling and energy management. Background Art
[0002] With the rapid growth of new energy vehicles, infrastructure such as charging stations and charging stacks faces increasing pressure from high concurrent access and energy allocation. Existing power allocation methods, which mostly rely on static strategies, are unable to effectively address multi-dimensional objectives such as traffic dynamics, electricity price fluctuations, and battery health management. This leads to long charging wait times, severe grid load fluctuations, shortened battery life, and high operating costs. Therefore, a dynamic load balancing method and system with adaptive and multi-objective optimization capabilities is urgently needed. Summary of the Invention
[0003] In view of this, the present invention provides a charging stack power distribution method and system based on dynamic load balancing. Through multi-agent reinforcement learning, multi-objective optimization and game coordination mechanism, it can perceive key parameters such as traffic flow, power grid, battery status, etc. in real time, and realize intelligent, adaptive and low-cost dynamic distribution of charging power to solve or alleviate the technical problems existing in the existing technology, and at least provide a beneficial option.
[0004] The technical solution of the embodiment of the present invention is achieved as follows:
[0005] In a first aspect, the present invention provides a charging stack power distribution method based on dynamic load balancing, comprising the following steps:
[0006] Obtaining local status information for each charging station, including at least: charging queue length, vehicle current state of charge, current load power, battery health status, expected stay time, electricity price information, and grid load status;
[0007] Based on historical charging behavior, electricity price fluctuations, and traffic flow data, a time series neural network is used to predict the vehicle entry and exit frequency, electricity price trends, and power demand of each charging station in the future period; and a predictive model is constructed.
[0008] Constructing a game agent model, based on each charging stack as a game agent, and determining that the game agents coordinate power allocation through a game collaboration method, wherein the game collaboration method includes neighborhood-wide power coordination and a target utility function between the game agents;
[0009] The central control and dispatching platform constructs a multi-objective reward function based on charging demand response, vehicle waiting time, battery health, electricity price forecast, and grid load level, and obtains the optimal power dispatch path through reinforcement learning algorithm training.
[0010] The optimized power scheduling path is sent to the local collaborative controller of each charging stack.
[0011] Further preferably: the target utility function is defined as U i , the specific formula is as follows:
[0012]
[0013] Among them, R i (a i ):Charging stack i in power allocation strategy a i The profit function under
[0014] λ: collaborative weight factor;
[0015] represents the power difference between charging pile i and its neighboring charging pile j.
[0016] Further preferably: the multi-objective reward function is defined as R t The specific formula is as follows:
[0017]
[0018] Among them, α1: weight coefficient, representing the vehicle waiting time optimization target; α2: weight coefficient, representing the charging stack power load balancing optimization target; α3: weight coefficient, representing the battery health optimization target; α4: weight coefficient, representing the grid load optimization target; α5: weight coefficient, representing the operating cost optimization target; T wait,i : Current waiting time of vehicle i; Average charging power; D battery,i : The impact factor of current fast charging on battery aging; G overload : Power grid pressure index; C i : Estimated unit charging cost under the current allocation strategy.
[0019] Further preferably, the reinforcement learning algorithm adopts a multi-agent deterministic policy gradient algorithm, each of the charging stack agents independently trains the multi-agent deterministic policy gradient algorithm, and shares global state information for policy enhancement.
[0020] Further preferably: the prediction model includes:
[0021] Traffic flow prediction network based on Transformer structure;
[0022] Electricity price prediction model based on attention mechanism;
[0023] State of charge evolution predictor based on gradient boosted tree model.
[0024] Further preferably, the dispatch target weight coefficients α1 to α5 of the central control dispatch platform support dynamic adjustment according to different time periods, grid load peaks and valleys, electricity price sensitivity or user priority strategies.
[0025] Further preferably, the step of sending the optimized power scheduling path to the local collaborative controller of each charging stack specifically includes:
[0026] Each charging stack sends a power request to the local power scheduler based on the weight coefficient α1. The power control module maps the request to the actual power supply value and fine-tunes the power based on the battery BMS feedback.
[0027] A second aspect: A charging stack power distribution system based on dynamic load balancing, comprising:
[0028] An acquisition module configured to acquire local status information of each charging pile, wherein the status information includes at least: charging queue length, current state of charge of the vehicle, current load power, battery health status, expected stay time, electricity price information, and grid load status;
[0029] Build a predictive modeling module; based on historical charging behavior, electricity price fluctuations, and traffic flow data, use a time series neural network to predict the vehicle entry and exit frequency, electricity price trends, and power demand of each charging station in the future time period; and build a predictive model;
[0030] A game agent model building module is configured to build the game agent model, based on each charging stack being a game agent, determine that the game agents coordinate power allocation through a game collaboration method, wherein the game collaboration method includes neighborhood-wide power coordination and target utility function between the game agents;
[0031] Centralized control and dispatch module: This module is deployed on the central control and dispatch platform. It builds a multi-objective reward function based on charging demand response, vehicle waiting time, battery health, electricity price forecast, and grid load level, and obtains the optimal power dispatch path through reinforcement learning algorithm training.
[0032] Dynamic load balancing module; configured to send the optimized power scheduling path to the local collaborative controller of each charging pile.
[0033] A third aspect: A readable medium having instructions stored thereon, which, when executed on an electronic device, causes the electronic device to execute the charging stack power distribution method based on dynamic load balancing.
[0034] A fourth aspect: An electronic device, characterized in that the electronic device comprises:
[0035] a memory for storing instructions to be executed by one or more processors of the electronic device, and
[0036] The processor is one of the processors of the electronic device and is used to execute the charging stack power distribution method based on dynamic load balancing.
[0037] The embodiment of the present invention adopts the above technical solution, which has the following advantages:
[0038] The present invention realizes forward-looking power planning by introducing a multi-dimensional prediction mechanism;
[0039] Improve power allocation efficiency by adopting local cooperative games;
[0040] Achieve self-adaptation and generalization capabilities through reinforcement learning, enhancing the robustness of scheduling strategies in multiple scenarios;
[0041] Through multi-objective scheduling indicators, electricity price, battery life, waiting time and grid load are optimized simultaneously.
[0042] The above summary is for illustrative purposes only and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments and features described above, further aspects, embodiments and features of the present invention will be readily apparent by reference to the accompanying drawings and the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0044] Figure 1 This is a flow chart of a charging stack power distribution method based on dynamic load balancing according to the present invention;
[0045] Figure 2 This is a schematic diagram of the computer device structure of Example 3 of the present invention.
[0046] In the figure, 10 is a computer device; 1002 is a processor; 1004 is a memory; and 1006 is a transmission device. DETAILED DESCRIPTION
[0047] Hereinafter, only certain exemplary embodiments are briefly described. As will be appreciated by those skilled in the art, the described embodiments may be modified in various ways without departing from the spirit or scope of the present invention. Therefore, the drawings and description are to be considered as illustrative in nature and not restrictive.
[0048] The embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0049] Example 1
[0050] A charging stack power distribution method based on dynamic load balancing, such as Figure 1 As shown, the following steps are included:
[0051] Step S100: Acquire local status information of each charging pile, wherein the status information includes at least: charging queue length, current state of charge of the vehicle, current load power, battery health status, expected stay time, electricity price information, and grid load status;
[0052] Step S200: Based on historical charging behavior, electricity price fluctuations, and vehicle flow data, a time series neural network is used to predict the vehicle entry and exit frequency, electricity price trends, and power demand of each charging station in the future time period; and a prediction model is constructed;
[0053] Step S300: Constructing a game agent model, based on each charging stack being a game agent, determining that the game agents coordinate power allocation through a game collaboration method, wherein the game collaboration method includes neighborhood-wide power coordination and a target utility function between the game agents;
[0054] The objective utility function is defined as U i , the specific formula is as follows:
[0055]
[0056] Among them, R i (a i ):Charging stack i in power allocation strategy a i The profit function under
[0057] λ: collaborative weight factor;
[0058] represents the power difference between charging pile i and its neighboring charging pile j.
[0059] Step S400: The central control and dispatching platform constructs a multi-objective reward function based on charging demand response, vehicle waiting time, battery health, electricity price forecast, and grid load level, and obtains the optimal power dispatch path based on reinforcement learning algorithm training;
[0060] The multi-objective reward function is defined as R t The specific formula is as follows:
[0061]
[0062] Among them, α1: weight coefficient, representing the vehicle waiting time optimization target; α2: weight coefficient, representing the charging stack power load balancing optimization target; α3: weight coefficient, representing the battery health optimization target; α4: weight coefficient, representing the grid load optimization target; α5: weight coefficient, representing the operating cost optimization target; T wait,i : Current waiting time of vehicle i; Average charging power; D battery,i : The impact factor of current fast charging on battery aging; G overload : Power grid pressure index; C i : Estimated unit charging cost under the current allocation strategy.
[0063] The reinforcement learning algorithm adopts a multi-agent deterministic policy gradient algorithm. Each charging stack agent independently trains the multi-agent deterministic policy gradient algorithm and shares global state information for policy enhancement.
[0064] Step S500: Send the optimized power scheduling path to the local collaborative controller of each charging stack.
[0065] In this embodiment, specifically: the prediction model includes:
[0066] Traffic flow prediction network based on Transformer structure;
[0067] Electricity price prediction model based on attention mechanism;
[0068] State of charge evolution predictor based on gradient boosted tree model.
[0069] In this embodiment, specifically: the central control and dispatching platform dispatch target weight coefficients α1 to α5 support dynamic adjustment according to different time periods, grid load peaks and valleys, electricity price sensitivity or user priority strategies.
[0070] In this embodiment, specifically, sending the optimized power scheduling path to the local collaborative controller of each charging stack includes:
[0071] Each charging stack sends a power request to the local power scheduler based on the weight coefficient α1. The power control module maps the request to the actual power supply value and fine-tunes the power based on the battery BMS feedback.
[0072] Specifically, the key status information of each charging stack in the system is collected in real time through the sensors embedded in the charging stack and the network communication system. This information forms a multi-dimensional state space vector that represents the current state of the charging system. This state space includes the following core elements:
[0073] Charging stack load rate:
[0074] Indicates the current load level of each charging station, that is, whether the charging station is close to full capacity. This is closely related to "Current Load Power" and "Expected Dwell Time", because a higher load means more power needs to be deployed, which also leads to longer waiting times.
[0075] Number of vehicles in queue:
[0076] The number of vehicles currently waiting in line at each charging station. This reflects the congestion of charging demand and directly affects the "charging queue length." During high-demand periods, when there are many vehicles in the queue, scheduling strategies are needed to avoid overloading.
[0077] Battery state of charge:
[0078] Indicates the current state of charge (SOC) of each vehicle being charged. This data is directly related to the vehicle's current state of charge (SOC) and can help the system understand which vehicles need to be charged quickly and which can be charged later, further optimizing power distribution.
[0079] Grid power:
[0080] The total power currently available on the grid. This value is affected by the current grid load and external supply conditions, and is highly correlated with the "grid load status." Real-time grid power acquisition is the core of dynamic load balancing, ensuring that power distribution remains within a controllable range and preventing grid overload.
[0081] Historical power allocation trends:
[0082] This shows the historical power distribution trend for each charging stack. This provides a reference for past load changes and a basis for predicting future power distribution. It is closely related to actual power output and forms the historical data foundation for calculating power scheduling optimization.
[0083] Electricity price factor:
[0084] The current electricity price information, especially the electricity price under the time-of-use electricity price model. Electricity price fluctuations will affect the choice of charging strategy, especially during periods of low electricity prices. The system will tend to adjust the charging strategy to optimize economic benefits.
[0085] It reflects the multi-dimensional construction of the charging stack status perception system.
[0086] For example:
[0087] “Charging queue length” corresponds to the number of vehicles in the queue
[0088] The "vehicle's current state of charge (SOC)" corresponds to the battery's state of charge. The "current load power" is related to the charging stack load rate and historical power distribution trends.
[0089] The “battery health status” can be indirectly inferred through battery state of charge (SOC) and charging history analysis;
[0090] The “expected dwell time” can be estimated by combining the number of vehicles in the queue, the battery charging progress, and the grid power.
[0091] This information will be updated in real time and aggregated into a multi-dimensional tensor to form a state vector for further analysis and processing by the system, thereby achieving dynamic load balancing and power scheduling optimization.
[0092] By introducing a multi-dimensional perception state space model, not only can more comprehensive and accurate charging stack status information be obtained in real time, but this information can also be combined for efficient prediction and dynamic adjustment. Traditional solutions often only focus on single-dimensional status information, while this invention uses multi-dimensional data fusion, which has the following advantages:
[0093] High-dimensional spatiotemporal correlation modeling: The status information of each charging pile includes multiple dimensions (load rate, queue number, SOC, electricity price, grid power, etc.), making state modeling more accurate and supporting high-dimensional spatiotemporal correlation modeling.
[0094] Adaptive capability: This method can dynamically perceive environmental changes and respond quickly based on factors such as different loads, electricity price fluctuations, and traffic volume, and has the ability of adaptive scheduling.
[0095] Optimizing power distribution: By comprehensively considering multiple factors, the system can more efficiently distribute and dispatch power, avoid system overload or uneven battery charging, and improve the charging experience and grid utilization.
[0096] Predictive modeling mainly models and predicts the future power demand of charging piles, electricity price trends, traffic flow, etc., in order to provide a basis for subsequent scheduling optimization.
[0097] Predictive Modeling with Time Series Neural Networks:
[0098] Data collection and feature selection;
[0099] Real-time status information of the charging stack is collected, and this data will serve as the basis for model training and prediction.
[0100] Core data includes:
[0101] Historical charging behavior data: vehicle entry and exit frequency, charging time, battery SOC (state of charge), historical power distribution, etc. of each charging pile.
[0102] Electricity price fluctuation data: time-of-use electricity prices, historical electricity price fluctuations, etc.
[0103] Traffic flow data: vehicle entry and departure times, vehicle type (such as battery capacity, charging requirements, etc.).
[0104] Time series model design
[0105] In order to predict future traffic flow, electricity price trends, power demand, etc., a model based on time series neural networks (Time Series Neural Networks) is used for training and prediction.
[0106] Specific modeling methods include:
[0107] LSTM (Long Short-Term Memory): LSTM is highly effective at processing time series data and is particularly well-suited for capturing long-term dependencies between charging behavior, electricity prices, and vehicle traffic. LSTM can be used to predict future charging pile loads, vehicle entry and exit frequencies, and electricity price trends.
[0108] Transformer model: To further improve prediction accuracy, the Transformer architecture is combined to enhance the model's ability to process long time series. The advantage of the Transformer is its ability to process data at multiple time steps in parallel, thereby better capturing complex dependencies over long periods of time.
[0109] Time2Vec: Building on LSTM and Transformer, Time2Vec enhances the representation of temporal features. By learning the periodic and trend characteristics of time series, Time2Vec can further improve the model's predictive capabilities.
[0110] Prediction tasks and outputs
[0111] The main task of the model is to predict the following key parameters in the future period:
[0112] Frequency of vehicles entering and exiting the charging pile in the future: By learning from historical traffic data, the frequency of vehicles entering and exiting each charging pile in the future is predicted.
[0113] Electricity price fluctuation trend: Based on historical electricity prices and grid demand in future periods, predict the changing trend of electricity prices.
[0114] Power demand forecasting: Based on factors such as vehicle state of charge, battery health, and electricity prices, the future power demand of each charging stack is predicted.
[0115] Combined with multi-objective scheduling optimization
[0116] The predicted traffic volume, electricity price trends, and power demand will become the input for subsequent multi-objective scheduling optimization. Based on the predicted results, the power distribution of the charging pile will be adjusted to optimize the following objectives:
[0117] Minimize waiting time: Dynamically adjust power allocation by predicting future traffic flow to reduce user waiting time.
[0118] Load balancing: Based on power demand forecasts, load balancing is performed among multiple charging piles to avoid overloading a single charging pile.
[0119] Battery life optimization: Predict battery life based on the vehicle's SOC and battery health to avoid battery loss caused by fast charging.
[0120] Grid load balancing: Through grid load forecasting, power distribution is dispatched to smooth out grid fluctuations.
[0121] Operation cost optimization: Based on electricity price forecasts, choose charging time when electricity prices are low to reduce costs.
[0122] Specifically:
[0123] Historical charging behavior, electricity price fluctuations, and traffic data are all raw data collected and used.
[0124] Time series neural networks are used to predict future charging pile power demand, vehicle flow and electricity price trends.
[0125] The prediction results serve as the basis for multi-objective scheduling optimization, helping the system make more accurate and efficient power allocation decisions, reducing unnecessary charging waiting time, and reducing grid load fluctuations, thereby improving the operational efficiency of the entire charging system.
[0126] Multi-model fusion: Combining the advantages of models such as LSTM, Transformer, and Time2Vec, it improves the modeling capabilities of complex time series data. In particular, it can capture long-term dependencies in the data when predicting charging pile power demand and electricity price fluctuations.
[0127] Real-time prediction and scheduling: Through accurate prediction of future time periods, the charging system can adjust the scheduling strategy according to actual conditions, avoid overload or battery damage during peak hours, optimize power distribution, and improve user experience.
[0128] Strong adaptability: The model can adaptively update prediction results based on new traffic flow, electricity prices and power data to adapt to the dynamically changing charging environment.
[0129] By combining multiple cutting-edge neural network models, the predictive and adaptive capabilities of the charging stack power distribution system have been improved, enabling it to more efficiently respond to complex charging demands and grid load changes.
[0130] By introducing the intelligent agent model in game theory, the optimal power distribution is achieved between charging piles through collaborative games.
[0131] Local collaborative power coordination:
[0132] 1. Game Agent Modeling
[0133] In this step, each charging station is treated as an independent agent, and the charging stations coordinate power allocation through game collaboration. Each agent (charging station) exchanges information with the charging stations in its neighborhood, using collaborative game play to optimize local power allocation while ensuring overall system efficiency.
[0134] Agent definition: Each charging pile is considered an agent, defined as i. Each charging pile decides how to allocate its own power based on its own charging needs, battery status, grid load, and other information.
[0135] Neighborhood Set: Each charging station will play a cooperative game with the charging stations within its neighborhood. This includes charging piles that are adjacent to charging pile i in physical space or grid topology. These charging piles affect each other's power demand and load distribution, so collaborative decision-making is required.
[0136] 2. Target utility function definition
[0137] The goal of each charging pile is to maximize its revenue function while minimizing the power difference between itself and its neighboring charging piles. To this end, the objective utility function U is defined as i :
[0138]
[0139] in:
[0140] R i (a i ):Charging stack i in power allocation strategy a i The benefit function is usually related to factors such as the charger's power utilization, battery state of charge, and electricity price. The charger optimizes its charging efficiency or benefit by adjusting power distribution.
[0141] λ: A synergistic weighting factor that controls the impact of power differences on overall utility. A larger λ means that the power differences between charging stacks must be kept small, while a smaller λ means that the charging stacks can tolerate more power fluctuations.
[0142] represents the power difference between charging station i and its neighboring charging station j. This penalty term allows the charging station to consider the power coordination with neighboring charging stations when allocating power, avoiding local overload or resource waste.
[0143] 3. Game Optimization Process
[0144] Each charging station optimizes according to its own target utility function in the local collaborative game. The specific steps are as follows:
[0145] Strategy update: Each charging station i is updated according to its own benefit function R i (a i ) and neighborhood power difference The goal of the charging stack is to find an optimal power allocation strategy a i *, so that the utility function U i maximize.
[0146] Game equilibrium: Through the concept of Nash equilibrium in game theory, the charging stacks will repeatedly adjust until a stable power allocation strategy is reached, in which the choice of each charging stack will no longer change. That is, each charging stack will no longer unilaterally change its own strategy when considering the strategies of other charging stacks.
[0147] Local load balancing: During the game, the power difference is controlled by the penalty term, and the power distribution between the charging piles tends to be balanced, thereby ensuring that the load of each charging pile is not too high and avoiding the problem of insufficient power in some charging piles.
[0148] 4. Overall revenue improvement
[0149] While each charging station performs local optimization, the collaboration and competition between charging stations ultimately lead to improved overall system profitability. Specifically, by coordinating power distribution among charging stations, the system improves resource utilization, reduces queue times, and optimizes grid load while ensuring local load balancing, ultimately maximizing overall system profitability.
[0150] Game agent modeling: The charging pile participates in power allocation decisions as an independent game agent. Each charging pile makes adjustments based on its own goals and the status of neighboring charging piles, which is consistent with the definition of game agent in the claims.
[0151] Objective utility function: This is a utility function that includes a revenue function and a penalty term for power differences, specifically expressing how the charging stack strikes a balance between its own objectives and neighboring power coordination.
[0152] Local load balancing and global benefit improvement: Through the game optimization process, the charging stack can not only optimize its own power distribution, but also improve the operating efficiency and benefits of the entire system through collaboration.
[0153] Applying game theory to charger power coordination, we achieve load balancing through local game optimization while ensuring overall system profitability. This approach transcends the limitations of traditional centralized scheduling models and offers greater flexibility and adaptability.
[0154] By solving the Nash equilibrium of the game, each charger's power allocation strategy can ensure its own interests while coordinating with neighboring chargers, ensuring stable system operation despite fluctuations in charging demand. This method fully leverages the advantages of local cooperative games, promoting global benefits while ensuring local load balancing, thereby achieving more efficient power allocation and grid load management.
[0155] By utilizing the intelligent agent collaboration mechanism of game theory, the power distribution is optimized through dynamic adjustment and local game, thus solving the problems of inefficiency and local overload in traditional scheduling methods.
[0156] The power scheduling strategy of the charging stack is optimized by constructing a reward function that includes multiple objectives, each of which is closely related to different system requirements.
[0157] 1. Comprehensively optimize charging demand, vehicle waiting time, battery health, electricity price forecasts, and grid load levels.
[0158] The reward function has the following form:
[0159]
[0160] Among them, each item represents a scheduling optimization goal:
[0161] α1: Weight coefficient, representing the optimization target of vehicle waiting time. The goal is to minimize the waiting time T of the vehicle. wait,i , maximizing the utilization rate and charging efficiency of the charging stack.
[0162] α2: Weight coefficient, representing the optimization goal of charging pile power load balancing. The goal is to reduce the power difference between charging piles and make the power distribution of each charging pile close to the average value. This prevents overloading of certain charging stacks.
[0163] α3: Weight coefficient, representing the battery health optimization target. The goal is to avoid overcharging or fast charging that will reduce battery life and try to maintain the battery's health. battery,i .
[0164] α4: Weight coefficient, which represents the grid load optimization target. The goal is to avoid grid overload and balance the grid load by adjusting the power distribution of the charging stack.
[0165] α5: Weight coefficient, representing the operating cost optimization goal. The goal is to reduce the system's operating costs by selecting charging periods with lower electricity prices.
[0166] 2. Application of reinforcement learning (DDPG) algorithm
[0167] To achieve this multi-objective optimization, the Deep Deterministic Policy Gradient (DDPG) algorithm in reinforcement learning is used. DDPG is a reinforcement learning method based on an actor-critic architecture that is suitable for handling problems in continuous action spaces, and is particularly well-suited for tasks such as power scheduling, which require real-time feedback and decision-making.
[0168] The core idea of the DDPG algorithm is to continuously optimize the decision-making strategy through reinforcement learning, so that the system gradually learns the optimal power scheduling strategy through interaction with the environment.
[0169] The specific steps of the DDPG algorithm are as follows:
[0170] Environment modeling and state definition: State S t It contains real-time information of the charging stack, such as the vehicle's waiting time, current power load, vehicle state of charge (SOC), battery health, grid load, etc.
[0171] Action α t Indicates the power allocation decision of the charging stack, such as adjusting the power output P of the charging stack i .
[0172] Reward function R t It is calculated based on the various objectives in the multi-objective scheduling optimization and used as an evaluation indicator of system performance.
[0173] Strategy Optimization:
[0174] Actor network: learns the power scheduling strategy through neural network, that is, according to the state S t Given a continuous power output α t .
[0175] Critic network: Learn to estimate each state-action pair (S t , α t ), which is the long-term reward of the current power allocation strategy under a specific state.
[0176] Target network: Ensure training stability by updating the target network (delayed update).
[0177] Training process:
[0178] At each time step t, the Actor network of the charging stack selects a power scheduling action α according to the current statet And feed the result back to the Critic network to calculate the reward R t
[0179] The DDPG algorithm optimizes the strategy so that the Actor network can select the optimal power scheduling strategy in the shortest possible time and maximize the multi-objective reward function.
[0180] Through multiple rounds of interaction, the charging stack learns the optimal power allocation strategy under different conditions and ultimately achieves global optimization.
[0181] 3. Scheduling strategy updates and system feedback
[0182] After each dispatch decision, the system provides feedback based on the actual power dispatch results. The charger adjusts its strategy based on this feedback, aiming to further optimize power allocation in future dispatches. This feedback mechanism ensures that the charger can flexibly adapt and continuously optimize in the face of uncertain grid loads, electricity price fluctuations, and changes in traffic flow.
[0183] Multi-objective reward function: The multi-objective reward function is designed in detail in the steps, covering multiple objectives such as charging pile waiting time, power balance, battery health, grid load, etc. Reinforcement learning-based scheduling optimization: The claims clearly state that the optimal power scheduling strategy is trained using reinforcement learning (such as the DDPG algorithm), and the reinforcement learning application in the steps fully meets this requirement.
[0184] Through the DDPG algorithm, the system can dynamically adjust the power scheduling strategy according to actual conditions, and has high adaptability and flexibility.
[0185] Simultaneously optimizing multiple objectives (such as waiting time, battery health, grid load, etc.) not only improves user experience but also ensures the stability and economy of system operation.
[0186] Reinforcement learning enables the charging stack to find the optimal power scheduling strategy through repeated learning when facing variable grid loads and traffic changes, thereby achieving long-term stable optimization of the system.
[0187] Efficiently distribute the optimized power allocation strategy to each charging stack local controller;
[0188] Ensure that the system can quickly respond and self-update optimization strategies when faced with real-time state changes and environmental disturbances, keeping the overall system stable and running efficiently.
[0189] Example 2
[0190] A charging stack power distribution system based on dynamic load balancing, comprising:
[0191] An acquisition module configured to acquire local status information of each charging pile, wherein the status information includes at least: charging queue length, current state of charge of the vehicle, current load power, battery health status, expected stay time, electricity price information, and grid load status;
[0192] Build a predictive modeling module; based on historical charging behavior, electricity price fluctuations, and traffic flow data, use a time series neural network to predict the vehicle entry and exit frequency, electricity price trends, and power demand of each charging station in the future time period; and build a predictive model;
[0193] A game agent model building module is configured to build the game agent model, based on each charging stack being a game agent, determine that the game agents coordinate power allocation through a game collaboration method, wherein the game collaboration method includes neighborhood-wide power coordination and target utility function between the game agents;
[0194] Centralized control and dispatch module: This module is deployed on the central control and dispatch platform. It builds a multi-objective reward function based on charging demand response, vehicle waiting time, battery health, electricity price forecast, and grid load level, and obtains the optimal power dispatch path through reinforcement learning algorithm training.
[0195] Dynamic load balancing module; configured to send the optimized power scheduling path to the local collaborative controller of each charging pile.
[0196] The foregoing Figure 1 The various variations and specific examples of the charging stack power distribution method based on dynamic load balancing in Example 1 are also applicable to the charging stack power distribution system based on dynamic load balancing in this embodiment. Through the above detailed description of the charging stack power distribution method based on dynamic load balancing, those skilled in the art can clearly know the implementation method of the charging stack power distribution system based on dynamic load balancing in this embodiment, so for the sake of brevity of the specification, it will not be described in detail here.
[0197] Example 3
[0198] An embodiment of the present application provides a computer device, which includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement a charging stack power distribution method based on dynamic load balancing as provided in the above method embodiment.
[0199] Figure 2 The hardware structure diagram of a device for implementing a charging stack power distribution method based on dynamic load balancing provided in an embodiment of the present application is shown. The device may participate in or include the apparatus or system provided in an embodiment of the present application. Figure 2As shown, the computer device 10 may include one or more processors 1002 (the processor may include but is not limited to a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 1004 for storing data, and a transmission device 1006 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that Figure 2 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 2 More or fewer components than shown, or with Figure 2 Different configurations shown.
[0200] It should be noted that the one or more processors and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry". The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuitry may be a single independent processing module, or may be incorporated in whole or in part into any of the other components of the computer device 10 (or mobile device). As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).
[0201] The memory 1004 can be used to store software programs and modules of application software, such as a program instruction / data storage device corresponding to a charging stack power distribution method based on dynamic load balancing in an embodiment of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory 1004, that is, implementing one of the above methods. The memory 1004 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 1004 may further include a memory remotely located relative to the processor, and these remote memories may be connected to the computer device 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0202] The transmission device 1006 is used to receive or send data via a network. Specific examples of the aforementioned network may include a wireless network provided by the communications provider of the computer device 10. In one embodiment, the transmission device 1006 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 1006 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0203] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer device 10 (or mobile device).
[0204] Example 4
[0205] An embodiment of the present application also provides a computer-readable storage medium, which can be set in a server to store at least one instruction or at least one program related to a charging stack power distribution method based on dynamic load balancing in an embodiment of the method. The at least one instruction or the at least one program is loaded and executed by the processor to implement a charging stack power distribution method based on dynamic load balancing provided in the above-mentioned embodiment of the method.
[0206] Optionally, in this embodiment, the storage medium may be located in at least one of a plurality of network servers in a computer network. Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing program code, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0207] Example 5
[0208] An embodiment of the present invention further provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform a charging stack power allocation method based on dynamic load balancing provided in the various optional embodiments described above.
[0209] It should be noted that the order of the embodiments of the present application described above is for descriptive purposes only and does not represent the superiority or inferiority of the embodiments. The above description is of specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or advantageous.
[0210] The various embodiments in this application are described in a progressive manner. Similar portions between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the device, equipment, and storage medium embodiments are generally similar to the method embodiments, so their descriptions are relatively simple. For relevant portions, refer to the descriptions of the method embodiments.
[0211] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.
[0212] With the above-described preferred embodiments of the present invention as a guide, and with reference to the above description, relevant personnel are fully capable of making various changes and modifications without departing from the technical scope of this invention. The technical scope of this invention is not limited to the contents of the specification and must be determined according to the scope of the claims.
[0213] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various modifications and substitutions within the technical scope disclosed in the present invention, and such modifications and substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.
Claims
1. A charging stack power distribution method based on dynamic load balancing, characterized in that: The following steps are involved: Obtaining local status information for each charging station, including at least: charging queue length, vehicle current state of charge, current load power, battery health status, estimated dwell time, electricity price information, and grid load status; Based on historical charging behavior, electricity price fluctuations, and vehicle flow data, a time series neural network is used to predict vehicle entry and exit frequency, electricity price trends, and power demand at each charging station in the future; and a predictive model is constructed. Constructing a game agent model, based on each charging stack as a game agent, and determining that the game agents coordinate power allocation through a game collaboration method, wherein the game collaboration method includes neighborhood-wide power coordination and a target utility function between the game agents; The central control and dispatching platform constructs a multi-objective reward function based on charging demand response, vehicle waiting time, battery health, electricity price forecast, and grid load level, and obtains the optimal power dispatch path through reinforcement learning algorithm training. The optimized power scheduling path is sent to the local collaborative controller of each charging stack.
2. The method according to claim 1, wherein: The objective utility function is defined as U i , the specific formula is as follows: Among them, R i (a i ):Charging stack i in power allocation strategy a i The profit function under λ: collaborative weight factor; represents the power difference between charging pile i and its neighboring charging pile j.
3. The method according to claim 1, wherein: The multi-objective reward function is defined as R t The specific formula is as follows: Among them, α1: weight coefficient, representing the vehicle waiting time optimization target; α2: weight coefficient, representing the charging stack power load balancing optimization target; α3: weight coefficient, representing the battery health optimization target; α4: weight coefficient, representing the grid load optimization target; α5: weight coefficient, representing the operating cost optimization target; T wait,i : Current waiting time of vehicle i; Average charging power; D battery,i : The impact factor of current fast charging on battery aging; G overload : Power grid pressure index; C i : Estimated unit charging cost under the current allocation strategy.
4. The method according to claim 1, wherein: The reinforcement learning algorithm adopts a multi-agent deterministic policy gradient algorithm. Each charging stack agent independently trains the multi-agent deterministic policy gradient algorithm and shares global state information for policy enhancement.
5. The method according to claim 1, wherein: The prediction model includes: Traffic flow prediction network based on Transformer structure; Electricity price prediction model based on attention mechanism; State of charge evolution predictor based on gradient boosted tree model.
6. The method according to claim 3, wherein: The central control and dispatching platform dispatch target weight coefficients α1 to α5 support dynamic adjustment according to different time periods, grid load peaks and valleys, electricity price sensitivity or user priority strategies.
7. The method according to claim 1, wherein: The optimized power scheduling path is sent to the local collaborative controller of each charging stack, specifically including: Each charging stack sends a power request to the local power scheduler based on the weight coefficient α1. The power control module maps the request to the actual power supply value and fine-tunes the power based on the battery BMS feedback.
8. A charging stack power distribution system based on dynamic load balancing, characterized in that: include: Get module; configured to obtain local status information of each charging pile, the status information including at least: charging queue length, current state of charge of the vehicle, current load power, battery health status, expected stay time, electricity price information and grid load status; Build a predictive modeling module; based on historical charging behavior, electricity price fluctuations, and traffic flow data, use a time series neural network to predict the vehicle entry and exit frequency, electricity price trends, and power demand of each charging station in the future time period; and build a predictive model; A game agent model building module is configured to build the game agent model, based on each charging stack being a game agent, determine that the game agents coordinate power allocation through a game collaboration method, wherein the game collaboration method includes neighborhood-wide power coordination and a target utility function between the game agents; Centralized control and dispatch module: This module is deployed on the central control and dispatch platform. It builds a multi-objective reward function based on charging demand response, vehicle waiting time, battery health, electricity price forecast, and grid load level, and obtains the optimal power dispatch path based on reinforcement learning algorithm training. Dynamic load balancing module; configured to send the optimized power scheduling path to the local collaborative controller of each charging pile.
9. A readable medium, characterized in that The readable medium stores instructions, which, when executed on an electronic device, enable the electronic device to execute the charging stack power distribution method based on dynamic load balancing according to any one of claims 1 to 7.
10. An electronic device, characterized in that: The electronic device comprises: a memory for storing instructions to be executed by one or more processors of the electronic device, and The processor is one of the processors of the electronic device, and is used to execute the charging stack power distribution method based on dynamic load balancing according to any one of claims 1 to 7.
Citation Information
Cited By
Method for jointly determining taxi scheduling strategy and charging station pricing strategy
CN121391344A
Battery management method and system for intelligent shared battery replacement cabinet
CN121882650A
OCPP-based charging pile group dynamic load balancing method and system and storage medium
CN122463721A