Reinforcement learning based scalable optimization control method for charging and discharging of electric vehicle cluster
By combining Markov chain simulation and deep reinforcement learning, the problem of low computational efficiency in the optimal control of electric vehicle clusters is solved, and individual flexibility constraints and optimal cluster scheduling are realized, thereby improving the accuracy and efficiency of the optimal control of electric vehicle clusters.
Patent Information
- Application Number
- CN202510846484.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2045-06-24
AI Technical Summary
Existing electric vehicle cluster optimization control technologies suffer from low computational efficiency due to the complexity of individual electric vehicle constraints, making it difficult to balance computational accuracy and efficiency. Furthermore, they lack methods that consider individual flexibility constraints and optimal scheduling of the cluster as a whole.
By employing Markov chain simulation and deep reinforcement learning, a dynamic power boundary is constructed through state dimensionality reduction and action decomposition. Combined with a greedy linear annealing strategy, a safety control strategy is designed to achieve flexible constraint and optimal control of electric vehicle clusters.
It effectively solves the curse of dimensionality problem, improves computational efficiency and optimization accuracy, ensures system stability and reliability, and forms an optimal strategy that meets actual needs.
Smart Images

Figure CN120792569B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of electric vehicle charging and discharging optimization control, in particular to an electric vehicle cluster charging and discharging scalable optimization control method and device based on reinforcement learning. BACKGROUND
[0002] The electric vehicle (EV) provides a huge technical possibility for coping with global energy crisis and power grid fluctuations due to its electric characteristics. Under the background of increasingly tight energy supply, electric vehicles significantly improve energy utilization efficiency compared to traditional internal combustion engine vehicles through efficient energy conversion mechanisms, reduce dependence on fossil fuels, and electric vehicles have the technical characteristics of two-way energy flow and can be used as distributed energy storage units to provide adjustable load resources for power supply and demand matching. With the increasing penetration rate of electric vehicles, the impact of electric vehicle load on the power grid is becoming more and more prominent. Large-scale electric vehicle charging may exacerbate power grid load peaks, making the power grid face greater pressure; reasonable regulation and control of electric vehicle charging and discharging behavior can enhance the flexibility and stability of the power grid and reduce the operating cost of the power grid.
[0003] The schedulable charging and discharging capacity of a single electric vehicle is small and highly uncertain. Electric vehicle aggregators (EVA) act as a bridge between electric vehicle users and the power grid, ensuring user travel needs and economic benefits through smart contracts; and through advanced optimization and control algorithms, the provision of power grid ancillary services is maximized, achieving a win-win situation for all parties. The scheduling of EVA to the electric vehicle cluster needs to be carried out under the premise of strictly meeting the user's energy demand, and if the user's expected power cannot be met before the electric vehicle leaves, it will seriously damage the satisfaction and willingness of the electric vehicle user to participate in the scheduling. The scheduling of the electric vehicle cluster is also affected by the user's random travel demand, and the parking time and location of the electric vehicle are uncertain and exhibit different aggregation effects in different time periods and spatial ranges, which will lead to dynamic changes in the schedulable flexibility of the corresponding period.
[0004] After the EVA obtains the overall charging and discharging optimization strategy of the cluster, the EVA allocates the charging and discharging quota and determines the specific charging and discharging control strategy of the individual electric vehicle. The sorting method based on time laxity plays a key role in this process and provides a scientific basis for the EVA to make individual electric vehicle control decisions. Time laxity is defined as the difference between the remaining time available for charging before the electric vehicle leaves and the shortest time required to complete charging. The smaller the time laxity of the electric vehicle, the higher the charging urgency, and the electric vehicle should be charged as soon as possible. Based on the concept of time laxity, multi-dimensional sorting indicators can be constructed, such as calculating time urgency and energy demand urgency by combining energy laxity and time laxity.
[0005] In the optimization problem of the charging and discharging control of the electric vehicle cluster, as the number of electric vehicles increases and the model setting becomes closer to reality, the optimization model becomes more complex, and the optimization solving efficiency also decreases. The existing optimization control technology of electric vehicle cluster is often based on a simplified electric vehicle demand model, which makes the assumption that the dispatchable time and the required power of individual electric vehicles are homogeneous, which does not conform to the real-world electric vehicle usage pattern. Some studies use Markov chain method to simulate more realistic individual electric vehicle level uncertainty demand, but only stay in the flexibility evaluation of aggregated load, without trying to optimize the control of aggregated load.
[0006] In the existing optimization control technology, the problem of declining computational efficiency caused by the complexity of individual electric vehicle constraints has not been overcome. Centralized optimization methods (such as mixed integer linear programming and quadratic programming, etc.) in solving EVA charging and discharging problems, focusing on individual electric vehicle level uncertainty will lead to the number of decision variables of the optimization problem growing in proportion to the size of the electric vehicle cluster, and the solving time also grows exponentially; mathematical programming method is a static solving method, which needs to solve the problem again in each iteration of the day-ahead scheduling, resulting in high computational cost. When facing the dynamic changes of the electric vehicle cluster state, the solving speed does not meet the requirements of real-time control, and it takes a long time to complete an optimization iteration, which is difficult to deal with frequent fluctuations of power grid and sudden changes of user behavior.
[0007] The rolling optimization framework integrates the future energy consumption prediction of the electric vehicle cluster, but the computational burden increases dramatically with the increase of prediction time domain, which cannot guarantee the optimization quality while considering the computational efficiency. Some studies use mixed integer programming-based rolling optimization method to solve the V2G potential estimation problem of large-scale electric vehicles in the electricity market, which quantitatively describes the energy and power constraints of electric vehicles through the establishment of aggregated model, and realizes the accurate estimation of V2G potential of vehicle fleet. But this method focuses on the estimation of V2G potential of vehicle fleet in the past few hours, and it is difficult to realize real-time scheduling control of V2G.
[0008] Deep reinforcement learning (DRL) can better capture the dynamic characteristics of the electric vehicle cluster, and quickly complete the charging and discharging optimization solution. In the scenario of electric vehicle group providing auxiliary services to the power grid, deep reinforcement learning algorithm aggregates and captures various uncertainties related to scheduling. In this method, a feature-based linear function approximator is used to handle the time-varying state-action space in the aggregated electric vehicle group scheduling problem.
[0009] Existing reinforcement learning methods also face the challenge of balancing optimization control accuracy and computational efficiency in describing the state and action spaces of large-scale electric vehicle (EV) clusters. As the size of the EV cluster grows, the action and state spaces of the agents may increase exponentially, leading to a severe "curse of dimensionality." To address this issue, a flexible EV control model based on uncertain departure times is proposed by modeling the aggregation flexibility of the EV cluster. This model focuses on allocating the overall capacity of scheduled EVs to individual response capacities by optimizing the charging power peaks caused by load rebound effects. However, these methods fail to fully reflect the complexity of EV-level constraints in the overall scheduling optimization problem of EV clusters. Currently, there is a lack of a scalable optimization control method for EV cluster charging and discharging that considers the complex flexibility constraints of individual EVs and the overall optimal scheduling of the EV cluster, based on deep reinforcement learning. Summary of the Invention
[0010] To address the technical problems of existing electric vehicle demand modeling, such as simplistic assumptions that are seriously divorced from reality, difficulty in balancing computational efficiency and model accuracy, and the lack of an optimization control framework that integrates the complex constraints of individual electric vehicle flexibility and the overall optimal scheduling of electric vehicle clusters, this invention provides a scalable optimization control method and apparatus for charging and discharging electric vehicle clusters based on reinforcement learning. The technical solution is as follows:
[0011] On the one hand, a reinforcement learning-based scalable optimization control method for charging and discharging electric vehicle clusters is provided. This method is implemented by a scalable optimization control device for charging and discharging, and includes:
[0012] S1. Based on the Markov chain simulation method, the uncertain travel and energy demand of electric vehicles are simulated to obtain the parking demand and battery information of electric vehicles.
[0013] S2. Based on parking demand and battery information, calculate the charging and discharging priority and reference charging power of electric vehicles, construct the state variables of electric vehicles, and calculate the charging relaxation and discharging relaxation.
[0014] S3. Obtain the electricity price signal; perform state dimensionality reduction calculation on the state variables of all electric vehicles in the cluster to obtain the aggregated state vector of the electric vehicle aggregator, and calculate the dynamic power boundary of the cluster.
[0015] S4. Based on the linear annealing-based greedy strategy, the deep reinforcement learning agent selects charging and discharging control actions according to the aggregation state vector of the electric vehicle aggregator to obtain the initial control action vector.
[0016] S5. Based on the dynamic power boundary, filter the over-limit actions of the preliminary control action vector to obtain the safe control action vector, and impose penalties on the over-limit action decisions of the deep reinforcement learning agent.
[0017] S6. Based on the dynamic power boundary and the charging and discharging priority and relaxation of electric vehicles, the charging and discharging task actions of electric vehicles are decomposed according to the safety control action vector to obtain the execution action vector of electric vehicles.
[0018] S7. Based on the electric vehicle's action vector, the distributed intelligent charging pile performs charging and discharging control actions.
[0019] On the other hand, a reinforcement learning-based scalable optimization control device for charging and discharging of electric vehicle clusters is provided. This device is applied to a reinforcement learning-based scalable optimization control method for charging and discharging of electric vehicle clusters. The device includes:
[0020] The information acquisition module is used to simulate the uncertain travel and energy demand of electric vehicles based on the Markov chain simulation method, and obtain the parking demand and battery information of electric vehicles.
[0021] The electric vehicle state construction module is used to calculate the charging and discharging priority and reference charging power of electric vehicles based on parking demand and battery information, construct electric vehicle state variables, and calculate charging relaxation and discharging relaxation.
[0022] The electric vehicle aggregator state aggregation module is used to acquire electricity price signals; perform state dimensionality reduction calculations on the state variables of all electric vehicles in the cluster to obtain the aggregated state vector of the electric vehicle aggregator, and calculate the dynamic power boundary of the cluster.
[0023] The control action selection module is used for linear annealing-based... - Greedy strategy: The deep reinforcement learning agent selects charging and discharging control actions based on the aggregation state vector of electric vehicle aggregators to obtain the initial control action vector;
[0024] The charging and discharging safety control module is used to filter out excessive actions from the initial control action vector based on the dynamic power boundary, obtain the safety control action vector, and impose penalties on the excessive action decisions of the deep reinforcement learning agent.
[0025] The control action decomposition module is used to decompose the electric vehicle charging and discharging task actions based on the dynamic power boundary and the charging and discharging priority and relaxation of the electric vehicle, and obtain the electric vehicle execution action vector according to the safety control action vector.
[0026] The charging and discharging action execution module is used to control the charging and discharging actions of the distributed smart charging piles based on the action vector of the electric vehicle.
[0027] On the other hand, a scalable optimization control device for charging and discharging is provided, the scalable optimization control device for charging and discharging includes: a processor; a memory, the memory storing computer-readable instructions, which, when executed by the processor, implement any of the methods in the above-described reinforcement learning-based scalable optimization control method for charging and discharging electric vehicle clusters.
[0028] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction is stored in the storage medium, the at least one instruction being loaded and executed by a processor to implement any of the methods in the above-described reinforcement learning-based scalable optimization control method for charging and discharging electric vehicle clusters.
[0029] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following:
[0030] This invention proposes a scalable optimization control method for charging and discharging electric vehicle (EV) clusters based on reinforcement learning. By introducing an EV behavior generation method based on Markov chains, it overcomes the limitations of oversimplifying EV usage patterns and accurately captures the time and energy constraints of individual EV uncertainties. The EV aggregator charging and discharging agent based on deep reinforcement learning exhibits good adaptability to uncertainties in real-time scheduling and can learn the dynamic changes in the flexibility of the EV cluster through interaction with the environment, forming an optimal strategy that meets the actual user needs.
[0031] By integrating the state vectors of individual electric vehicles, the state of the electric vehicle aggregator is represented in a dimensionality-reduced manner. The proposed dynamic power boundary model and action decomposition method effectively overcome the challenge of the exponential increase in the dimension of the action space with the growth of the electric vehicle cluster size. By decoupling the state and action vector dimensions from the cluster size, the problem is solved from... Dimensional reduction This achieves a lightweight design of the state and action space of deep reinforcement learning agents, effectively solving the dimensionality curse problem in the optimization process.
[0032] The safety control strategy designed in this invention ensures the stability and reliability of the system from two aspects. Firstly, a dynamic power boundary for the electric vehicle aggregator is established based on the flexibility constraints of the electric vehicles in the cluster. Penalties are applied to actions by the agent that violate the dynamic power boundary, thereby accelerating the convergence of the agent's policy. Secondly, safety filtering is performed on the actions exceeding the limits given by the agent, projecting the control actions onto the dynamic power boundary to ensure the safety of actual control. This invention is a scalable optimization control method for charging and discharging of electric vehicle clusters based on deep reinforcement learning, considering the complex flexibility constraints of individual electric vehicles and the overall optimal scheduling of the electric vehicle cluster. Attached Figure Description
[0033] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0034] Figure 1 This is a flowchart of a scalable optimization control method for charging and discharging electric vehicle clusters based on reinforcement learning, provided in an embodiment of the present invention.
[0035] Figure 2 This is a schematic diagram of an overall framework provided by an embodiment of the present invention;
[0036] Figure 3 This is a schematic diagram illustrating the trend of an input electricity price signal provided in an embodiment of the present invention;
[0037] Figure 4 This is a schematic diagram of the state of charge and upper and lower limits of an electric vehicle cluster under different charging and discharging strategies, provided by an embodiment of the present invention.
[0038] Figure 5 This is a schematic diagram of the charging and discharging power of an electric vehicle cluster under different charging and discharging strategies, provided by an embodiment of the present invention.
[0039] Figure 6 This is a block diagram of a scalable and optimized control device for charging and discharging electric vehicle clusters based on reinforcement learning, provided in an embodiment of the present invention.
[0040] Figure 7 This is a schematic diagram of the structure of a charge and discharge scalable optimization control device provided in an embodiment of the present invention. Detailed Implementation
[0041] The technical solution of the present invention will now be described with reference to the accompanying drawings.
[0042] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.
[0043] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.
[0044] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.
[0045] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.
[0046] This invention provides a reinforcement learning-based scalable optimization control method for charging and discharging electric vehicle clusters. This method can be implemented by a scalable optimization control device, which can be a terminal or a server. Figure 1 The flowchart shown is for a scalable optimization control method for charging and discharging electric vehicle clusters based on reinforcement learning. One feasible implementation is as follows: Figure 2 As shown, this method can be divided into four levels: electric vehicle level, electric vehicle aggregator level, safety control level, and optimization decision level. For each time step of the electric vehicle cluster charging and discharging optimization control period, at the electric vehicle level, the information acquisition module simulates the battery parameters and demand data of each electric vehicle, and the electric vehicle state construction module constructs the electric vehicle flexibility constraints, calculating the electric vehicle state variables and relaxation information. The electric vehicle state variables and relaxation information are input into the electric vehicle aggregator level, and the electric vehicle aggregator state aggregation module obtains the dynamic power boundary and aggregated state vector. The aggregated state vector is input into the control action selection module of the optimization decision level, and the deep reinforcement learning agent outputs a preliminary action vector based on the aggregated state vector. The preliminary action vector is input into the safety control level, and the charging and discharging safety control module outputs a safety control action vector. The safety control action vector is input into the electric vehicle aggregator level, and the control action decomposition module obtains the electric vehicle execution action vector. The execution action vector is input into the electric vehicle level, and the charging and discharging action execution module performs the actual charging and discharging control.
[0047] The processing flow of this method may include the following steps:
[0048] S1. Based on the Markov chain simulation method, the uncertain travel and energy demand of electric vehicles are simulated to obtain the parking demand and battery information of electric vehicles.
[0049] Among them, the uncertain travel and energy demand of electric vehicles include: parking time, parking duration, travel duration, driving distance, travel purpose and parking location. All of these parameters follow the corresponding probability distribution. The travel mode and energy use preference of electric vehicles can be set by adjusting the probability distribution of the Markov chain model.
[0050] Specifically, upon the arrival of an electric vehicle, the system uses distributed smart charging stations to read the individual electric vehicle's parking needs data, including arrival time. Expected departure time Initial state of charge and minimum charge upon departure Read individual electric vehicle battery information, including battery capacity. Rated charging power Rated discharge power Charging efficiency and discharge efficiency .
[0051] Based on the parking needs of electric vehicles and battery information, the electric vehicle access status at the smart charging station is updated at each time step. .
[0052] S2. Based on parking demand and battery information, calculate the charging and discharging priority and reference charging power of electric vehicles, construct the state variables of electric vehicles, and calculate the charging relaxation and discharging relaxation.
[0053] Optionally, the specific implementation process of S2 includes S21-S25:
[0054] S21. Based on parking demand and battery information, calculate the flexibility constraints for a single electric vehicle, including the maximum capacity constraint. Minimum discharge constraint Constraints on electric vehicle user satisfaction ;in, and It can be determined based on the performance of the electric vehicle battery and can be set as a reasonable constant. To ensure that electric vehicles reach their parking time limit at the end of the parking period The minimum charge level that should be reached at time t during this parking period; based on the minimum charge level that should be reached at time t during this parking period, calculate the user satisfaction constraint. This can be expressed by the following formula (1):
[0055] (1);
[0056] in, This represents the current time step.
[0057] S22. Based on the flexibility constraints of electric vehicles, perform charging and discharging priority analysis to obtain the charging and discharging priority of a single electric vehicle. ;
[0058] In one feasible implementation, based on the electric vehicle's state of charge... Electric vehicles are classified into four categories based on their charging and discharging priorities. :
[0059] achieve At the same time greater than and greater than Electric vehicles only offer the flexibility of discharging; they cannot be recharged due to battery capacity limitations. .
[0060] Not achieved ,at the same time equal Electric vehicles only offer charging flexibility. .
[0061] Less than or Electric vehicles lack flexibility and must be recharged during the remaining parking time. .
[0062] In other cases, electric vehicles offer charging and discharging flexibility. .
[0063] S23, Based on the rated charging power of electric vehicles Based on the current state of charge and user satisfaction constraints Determine the baseline charging power for a single electric vehicle;
[0064] Among them, when individual electric vehicles do not reach It should be continuously charged beforehand; the baseline charging power for an individual electric vehicle is its rated power. When it reaches... Charging can then be stopped; the baseline charging power for an individual electric vehicle is 0.
[0065] S24. Based on the classification results of electric vehicle charging and discharging priority Reference charging power, state of charge at current step size t and battery capacity Construct state variables for a single electric vehicle ;
[0066] S25, Based on electric vehicle battery information and user satisfaction constraints According to the state of charge of electric vehicles Obtain the charging slack of a single electric vehicle. and discharge relaxation .
[0067] Among them, the relaxation of electric vehicles is used as a control ranking reference, and the charging relaxation is used as a reference. and discharge relaxation The calculations are shown in formulas (2) and (3) respectively:
[0068] (2);
[0069] (3);
[0070] S3. Obtain the electricity price signal; perform state dimensionality reduction calculations on the state variables of all electric vehicles in the cluster to obtain the aggregated state vector of the electric vehicle aggregator, and calculate the dynamic power boundary of the cluster. The trend of the input electricity price signal is as follows: Figure 3 As shown.
[0071] In one feasible implementation, at each time step, the electric vehicle aggregator accepts the state variables of all electric vehicles. Charging slack for all electric vehicles and discharge relaxation At each time step, receive the safety control action vector sent by the electric vehicle charging and discharging safety control module. Acquire electricity price signals at each time step. At each time step, the deep reinforcement learning optimization module sends the aggregation state vector of the electric vehicle aggregator. At each time step, a control action vector is sent to the distributed electric vehicle charging / discharging flexibility calculation module. .
[0072] Electric vehicle (EV) aggregators can optimize the charging and discharging of EVs connected to the charging infrastructure by signing contracts with users, provided that the flexibility constraints of the EVs are met. The maximum fleet size that an EV aggregator can accommodate is set as N, which can be determined by the charging infrastructure that the aggregator can control. Since the time when an EV connects to a charging station is uncertain, and the time an EV stays and the amount of electricity it requires are also uncertain, the size of the EV cluster controlled by the aggregator is dynamically changing, and the dispatchability flexibility of different EVs varies significantly.
[0073] The state and action decision variables of an electric vehicle swarm, composed of variables related to individual electric vehicles, are highly correlated with the number of controlled vehicles. Therefore, this invention proposes a state dimensionality reduction and action decomposition method. The aim is to reduce the dimensionality of the state vectors acquired by the deep reinforcement learning agent and the dimensionality of the control action vectors required, while ensuring the accuracy of the optimized control of electric vehicle charging and discharging.
[0074] At each time step, the electric vehicle aggregator receives the electricity price signal input from the external power grid. Simultaneously receive usage status information from all charging and discharging infrastructure locations. and the state vectors of all controlled individual electric vehicles. This includes: classification results based on electric vehicle charging and discharging priority. Reference charging power, state of charge at current step size t and battery capacity The electric vehicle aggregator updates its aggregation status based on the electric vehicle status information. The flexibility of all electric vehicles controlled by the aggregator is considered as a virtual energy storage system.
[0075] Optionally, the specific implementation process of S3 includes S31-S33:
[0076] S31, Maximum control capacity based on electric vehicle aggregators Based on the state variables of all electric vehicles in the cluster Calculate the state of charge of electric vehicle aggregators and reference charging power Among them, the benchmark charging power of electric vehicle aggregators The sum of the baseline charging power of all controlled individual electric vehicles;
[0077] In one feasible implementation, the state of charge of the electric vehicle aggregator The calculation formula is as follows (4):
[0078] (4);
[0079] in, This represents the maximum capacity that the virtual energy storage system of an electric vehicle aggregator can control.
[0080] S32. Based on the flexibility constraints of all electric vehicles in the cluster, calculate the dynamic power boundary of the electric vehicle aggregator, including the maximum charging power limit. Minimum charging power limit and maximum discharge power limit ;
[0081] Among them, the charging and discharging power constraints of electric vehicle aggregators are based on the flexibility constraints of individual electric vehicles, and the maximum charging power limit of electric vehicle aggregators. Minimum charging power limit Maximum discharge power limit The calculation formulas are shown in equations (5), (6) and (7) below:
[0082] (5);
[0083] (6);
[0084] (7);
[0085] S33, Based on electricity price signals State of charge of electric vehicle aggregators The benchmark charging power of electric vehicle aggregators Constructing the aggregation state vector of the electric vehicle aggregator from the dynamic power boundary of the electric vehicle aggregator .
[0086] Among them, the final aggregation state vector of the electric vehicle aggregator The following formula (8) represents:
[0087] (8);
[0088] S4, based on linear annealing - Greedy strategy: The deep reinforcement learning agent selects charging and discharging control actions based on the aggregation state vector of electric vehicle aggregators to obtain the initial control action vector.
[0089] In one feasible implementation, at each time step, a dual-depth Q-network agent observes the aggregation state vector of the electric vehicle aggregator. ,right Normalization is performed. The action vector output by the agent is represented as follows: , This represents the total charging and discharging power at the electric vehicle aggregator level, with positive values indicating charging and negative values indicating discharging. The agent's reward function is the sum of the energy cost of charging and discharging and the penalty for exceeding the charging and discharging power limit, expressed as... ,in, Indicates the weighting coefficient. This indicates a penalty for exceeding the limits.
[0090] In one feasible implementation, the agent, based on linear annealing... Greedy strategy for selecting actions In a probability of 1- In this case, the agent uses the following parameters: Q network selection With probability of In this situation, the agent randomly selects an action. As the number of training iterations decreases linearly, it eventually reaches a minimum value. In this way, the agent can improve convergence by conducting extensive exploration at the beginning and leveraging experience.
[0091] In one feasible implementation, the dual deep Q-network agent uses a deep neural network to approximate the action-value function and passes it through a target Q-network. The target Q-value is calculated to stabilize the training process, thereby indirectly optimizing the strategy to maximize the cumulative reward.
[0092] S5. Based on the dynamic power boundary, filter out excessive actions from the initial control action vector to obtain a safe control action vector, and impose penalties on the excessive action decisions of the deep reinforcement learning agent.
[0093] Optionally, the specific implementation process of S5 includes S51-S53:
[0094] S51. Based on the over-limit action protection rules, and according to the dynamic power boundary information of the electric vehicle aggregator, the initial control action vector is... Perform verification to obtain the security verification result;
[0095] Among them, the present invention proposes an over-limit action protection rule, based on dynamic power boundary information. As a verification standard for whether the output action of an intelligent agent exceeds the limit, there may be two types of over-limit actions, including overcharging or over-discharging caused by the decision power exceeding the maximum charging power limit or maximum discharging power limit of the electric vehicle.
[0096] S52. Based on the security verification results, the initial control action vector will be... Truncate to dynamic power boundary to obtain safety control action vector ;
[0097] Among them, based on the safety verification results, when the initial control action vector When the dynamic power limit of the electric vehicle aggregator is exceeded, the action is truncated to the executable range; otherwise, no filtering is performed. Based on the over-limit action protection rules, a safety control action vector is obtained. The calculation process is shown in formula (9); the safety control action vector Output to electric vehicle aggregators.
[0098] (9)
[0099] S53. Based on the security verification results, obtain the penalty items for exceeding the limits. .
[0100] Among them, the penalty items for exceeding the limit. Calculated using formula (10):
[0101] (10);
[0102] The calculated With preset weighting factors Multiply by the product to obtain the over-limit penalty value for the action. The data is then sent to the deep reinforcement learning agent to help it avoid giving actions that exceed the limits and to accelerate policy convergence.
[0103] S6. Based on the dynamic power boundary and the charging and discharging priority and relaxation of electric vehicles, the charging and discharging task actions of electric vehicles are decomposed according to the safety control action vector to obtain the execution action vector of electric vehicles.
[0104] Optionally, the specific implementation process of S6 includes:
[0105] S61, Charging and discharging priority based on electric vehicles The electric vehicle clusters are classified to obtain priority classifications;
[0106] The priority classification for individual electric vehicles includes: .
[0107] S62. Based on the priority classification results, according to the charging relaxation... and discharge relaxation The electric vehicles in each category are sorted in ascending order to obtain the priority order of charging and discharging of electric vehicles in each priority category;
[0108] In this process, the electric vehicle aggregator collects the charging slack of all controlled electric vehicles at each time step. and discharge relaxation The tasks are sorted as a quantitative indicator to determine their priority. The smaller the slack, the greater the urgency. Electric vehicle aggregators allocate charging or discharging tasks in ascending order of slack.
[0109] S63, Reference charging power based on electric vehicle aggregator and safety control action vector To obtain electric vehicle charging and discharging tasks categorized by different priorities;
[0110] In one feasible implementation, when At that time, all individual electric vehicles are charged according to the benchmark charging power, and the electric vehicle aggregator does not change the charging behavior of individual electric vehicles; when At that time, electric vehicle aggregators will assign charging tasks exceeding the baseline power to priority categories. ;when First, classify the priority as Electric vehicle delayed charging, if delayed charging is insufficient to achieve Electric vehicle aggregators assign electric vehicle discharge tasks to priority categories. Electric vehicles. Among them, to protect user comfort, for any The charging and discharging control of electric vehicle aggregators meets the priority classification as follows: The charging requirements will remain unchanged.
[0111] S64. Based on the priority order of electric vehicle charging and discharging, assign charging and discharging tasks of each priority category to electric vehicles to obtain the execution action vector of electric vehicles. .
[0112] To mitigate battery degradation risks, electric vehicle aggregators allocate demand response tasks across priority categories. For electric vehicles within the same flexibility category, the scheduling order is determined by the minimum slack priority principle. The total charging and discharging capacity for each priority category is decomposed into individual charging and discharging tasks achievable by each electric vehicle, thus obtaining the execution action vector of the electric vehicle. The electric vehicle aggregator will send action vectors to the corresponding electric vehicles, combined with electricity price signals. Calculate the cost of charging and discharging .
[0113] Optionally, after obtaining the electric vehicle's action vector at this time step, in order to iteratively optimize the charging and discharging control at the next time step, the method further includes:
[0114] Update the state vector of the electric vehicle at the next time step t+1 based on the action vector executed by the electric vehicle. Then, dimensionality reduction calculations are performed to obtain the aggregation state vector of the electric vehicle aggregator. ;
[0115] Based on the cost of exceeding the limit of the action and the energy cost of charging and discharging Obtain reward function ;
[0116] Based on the aggregation state vector of electric vehicle aggregators at the current time step and the next time step and Electric vehicle aggregator safety control action vector and reward function Build training experience data It is stored in the playback buffer 𝐵;
[0117] After the amount of training experience data reaches a preset minimum threshold, a small batch of training experience is extracted at each time step to train the deep reinforcement learning agent, thereby updating the parameters of the neural network of the deep reinforcement learning agent.
[0118] In one feasible implementation, when When the capacity reaches the minimum threshold, from A small batch of experience is extracted from the training data to train the agent. The parameters of the Q-network and the target Q-network are then updated. During training, the learning rate is set to 1. 10 -3 The training batch size and minimum replay buffer size are set to 32 and 1, respectively. 10 3 .
[0119] S7. Based on the electric vehicle's action vector, the distributed intelligent charging pile performs charging and discharging control actions.
[0120] in, Figure 4 This is a schematic diagram illustrating the state of charge and upper and lower limits of an electric vehicle cluster under different charging and discharging strategies, provided by an embodiment of the present invention; it demonstrates the electric vehicle cluster in a disordered charging scenario. Figure 4 (a) Intelligent charging scenario ( Figure 4 (b) and intelligent charging and discharging scenarios ( Figure 4 (c) The curves showing the changes in the aggregated state of charge and its upper and lower limits. As shown in the figure, the random arrival and departure of electric vehicles lead to a step-like change in the aggregated state of charge. By observing the real-time changes in the aggregated state of charge, the deep reinforcement learning agent can learn the uncertain demand patterns of electric vehicles. At the same time, the method of this invention can follow the upper and lower limits of the electric vehicle cluster's power, thus avoiding any compromise to user comfort.
[0121] Correspondingly, Figure 5 This demonstrates the situation of electric vehicles charging in a disorderly charging scenario. Figure 5 (a) Intelligent charging scenario ( Figure 5 (b) and intelligent charging and discharging scenarios ( Figure 5 (c) Charging and discharging power. In the unordered charging scenario, the peak charging period occurs in the evening when electricity prices are high. In contrast, the intelligent charging and discharging strategy controlled by the deep reinforcement learning agent makes full use of the low electricity price period for charging and the high electricity price period for discharging. The method of this invention can concentrate the discharging and charging power during the peak and low electricity price periods, respectively, which can effectively reduce the energy cost of electric vehicles.
[0122] This invention proposes a scalable optimization control method for charging and discharging electric vehicle (EV) clusters based on reinforcement learning. By introducing an EV behavior generation method based on Markov chains, it overcomes the limitations of oversimplifying EV usage patterns and accurately captures the time and energy constraints of individual EV uncertainties. The EV aggregator charging and discharging agent based on deep reinforcement learning exhibits good adaptability to uncertainties in real-time scheduling and can learn the dynamic changes in the flexibility of the EV cluster through interaction with the environment, forming an optimal strategy that meets the actual user needs.
[0123] By integrating the state vectors of individual electric vehicles, the state of the electric vehicle aggregator is represented in a dimensionality-reduced manner. The proposed dynamic power boundary model and action decomposition method effectively overcome the challenge of the exponential increase in the dimension of the action space with the growth of the electric vehicle cluster size. By decoupling the state and action vector dimensions from the cluster size, the problem is solved from... (N) Dimensional reduction to (1) A lightweight design of the state and action space of a deep reinforcement learning agent was achieved, effectively solving the dimensionality curse problem in the optimization process.
[0124] The safety control strategy designed in this invention ensures the stability and reliability of the system from two aspects. Firstly, a dynamic power boundary for the electric vehicle aggregator is established based on the flexibility constraints of the electric vehicles in the cluster. Penalties are applied to actions by the agent that violate the dynamic power boundary, thereby accelerating the convergence of the agent's policy. Secondly, safety filtering is performed on the actions exceeding the limits given by the agent, projecting the control actions onto the dynamic power boundary to ensure the safety of actual control. This invention is a scalable optimization control method for charging and discharging of electric vehicle clusters based on deep reinforcement learning, considering the complex flexibility constraints of individual electric vehicles and the overall optimal scheduling of the electric vehicle cluster.
[0125] Figure 6 This is a block diagram illustrating a reinforcement learning-based scalable optimization control device for charging and discharging electric vehicle clusters, according to an exemplary embodiment. The device is used in a reinforcement learning-based scalable optimization control method for charging and discharging electric vehicle clusters. (Refer to...) Figure 6 The device includes an information acquisition module 610, an electric vehicle state construction module 620, an electric vehicle aggregator state aggregation module 630, a control action selection module 640, a charging and discharging safety control module 650, a control action decomposition module 660, and a charging and discharging action execution module 670. Among them:
[0126] The information acquisition module 610 is used to simulate the uncertain travel and energy demand of electric vehicles based on the Markov chain simulation method, and obtain the parking demand and battery information of electric vehicles.
[0127] The electric vehicle state construction module 620 is used to calculate the charging and discharging priority and reference charging power of the electric vehicle based on parking demand and battery information, construct the electric vehicle state variables, and calculate the charging relaxation and discharging relaxation.
[0128] The electric vehicle aggregator state aggregation module 630 is used to acquire electricity price signals; perform state dimensionality reduction calculation on the state variables of all electric vehicles in the cluster to obtain the aggregated state vector of the electric vehicle aggregator, and calculate the dynamic power boundary of the cluster.
[0129] Control action selection module 640, for linear annealing-based... - Greedy strategy: The deep reinforcement learning agent selects charging and discharging control actions based on the aggregation state vector of electric vehicle aggregators to obtain the initial control action vector;
[0130] The charge and discharge safety control module 650 is used to filter out excessive actions from the preliminary control action vector based on the dynamic power boundary, obtain the safety control action vector, and impose penalties on the excessive action decisions of the deep reinforcement learning agent.
[0131] The control action decomposition module 660 is used to decompose the electric vehicle charging and discharging task actions based on the dynamic power boundary and the charging and discharging priority and relaxation of the electric vehicle, and obtain the electric vehicle execution action vector.
[0132] The charging and discharging action execution module 670 is used to control the charging and discharging actions of the distributed smart charging pile according to the action vector of the electric vehicle.
[0133] Optionally, the uncertain travel and energy demand of the electric vehicle include the parking time, parking duration, travel duration, driving distance, travel purpose and parking location of the electric vehicle;
[0134] The parking requirements for the electric vehicles include arrival time. Expected departure time Initial state of charge and minimum charge upon departure ;
[0135] The battery information includes battery capacity. Rated charging power Rated discharge power Charging efficiency and discharge efficiency .
[0136] Optionally, the electric vehicle state construction module 620 is used for:
[0137] Based on parking demand and battery information, calculate the flexibility constraints for a single electric vehicle, including the maximum capacity constraint. Minimum discharge constraint and user satisfaction constraints ;
[0138] Based on the flexibility constraints of electric vehicles, a charging and discharging priority analysis is performed to obtain the charging and discharging priority of a single electric vehicle. ;
[0139] Based on the rated charging power of electric vehicles Based on the current state of charge and user satisfaction constraints Determine the baseline charging power for a single electric vehicle;
[0140] Based on the classification results of electric vehicle charging and discharging priorities Reference charging power, state of charge at current step size t and battery capacity Construct state variables for a single electric vehicle ;
[0141] Based on electric vehicle battery information and user satisfaction constraints According to the state of charge of electric vehicles Obtain the charging slack of a single electric vehicle. and discharge relaxation .
[0142] Optionally, the electric vehicle aggregator status aggregation module 630 is used for:
[0143] Based on the maximum control capacity of electric vehicle aggregators Based on the state variables of all electric vehicles in the cluster Calculate the state of charge of electric vehicle aggregators and reference charging power ;
[0144] Based on the flexibility constraints of all electric vehicles in the cluster, calculate the dynamic power boundary of the electric vehicle aggregator, including the maximum charging power limit. Minimum charging power limit and maximum discharge power limit ;
[0145] According to electricity price signals State of charge of electric vehicle aggregators The benchmark charging power of electric vehicle aggregators Constructing the aggregation state vector of the electric vehicle aggregator from the dynamic power boundary of the electric vehicle aggregator .
[0146] Optionally, the charge / discharge safety control module 650 is used for:
[0147] Based on the over-limit protection rules, and according to the dynamic power boundary information of the electric vehicle aggregator, the initial control action vector is... Perform verification to obtain the security verification result;
[0148] Based on the security verification results, the initial control action vector will be... Truncate to dynamic power boundary to obtain safety control action vector ;
[0149] Based on the security verification results, obtain the penalty items for exceeding the limits. .
[0150] Optionally, the control action decomposition module 660 is used for:
[0151] Based on the charging and discharging priority of electric vehicles The electric vehicle clusters are classified to obtain priority classifications;
[0152] Based on the priority classification results, according to the charging relaxation and discharge relaxation The electric vehicles in each category are sorted in ascending order to obtain the priority order of charging and discharging of electric vehicles in each priority category;
[0153] Based on the benchmark charging power of electric vehicle aggregators and safety control action vector To obtain electric vehicle charging and discharging tasks categorized by different priorities;
[0154] Based on the charging and discharging priority of electric vehicles, charging and discharging tasks of each priority category are assigned to electric vehicles to obtain the execution action vectors of electric vehicles. .
[0155] Optionally, after the control action decomposition module 660, the system further includes:
[0156] Update the state vector of the electric vehicle at the next time step t+1 based on the action vector executed by the electric vehicle. Then, dimensionality reduction calculations are performed to obtain the aggregation state vector of the electric vehicle aggregator. ;
[0157] Based on the cost of exceeding the limit of the action and the energy cost of charging and discharging Obtain reward function ;
[0158] Based on the aggregation state vector of electric vehicle aggregators at the current time step and the next time step and Electric vehicle aggregator safety control action vector and reward function Build training experience data Stored in the playback buffer middle;
[0159] After the amount of training experience data reaches a preset minimum threshold, a small batch of training experience is extracted at each time step to train the deep reinforcement learning agent, thereby updating the parameters of the neural network of the deep reinforcement learning agent.
[0160] This invention proposes a scalable optimization control method for charging and discharging electric vehicle (EV) clusters based on reinforcement learning. By introducing an EV behavior generation method based on Markov chains, it overcomes the limitations of oversimplifying EV usage patterns and accurately captures the time and energy constraints of individual EV uncertainties. The EV aggregator charging and discharging agent based on deep reinforcement learning exhibits good adaptability to uncertainties in real-time scheduling and can learn the dynamic changes in the flexibility of the EV cluster through interaction with the environment, forming an optimal strategy that meets the actual user needs.
[0161] By integrating the state vectors of individual electric vehicles, the state of the electric vehicle aggregator is represented in a dimensionality-reduced manner. The proposed dynamic power boundary model and action decomposition method effectively overcome the challenge of the exponential increase in the dimension of the action space with the growth of the electric vehicle cluster size. By decoupling the state and action vector dimensions from the cluster size, the problem is solved from... (N) Dimensional reduction to (1) A lightweight design of the state and action space of a deep reinforcement learning agent was achieved, effectively solving the dimensionality curse problem in the optimization process.
[0162] The safety control strategy designed in this invention ensures the stability and reliability of the system from two aspects. Firstly, a dynamic power boundary for the electric vehicle aggregator is established based on the flexibility constraints of the electric vehicles in the cluster. Penalties are applied to actions by the agent that violate the dynamic power boundary, thereby accelerating the convergence of the agent's policy. Secondly, safety filtering is performed on the actions exceeding the limits given by the agent, projecting the control actions onto the dynamic power boundary to ensure the safety of actual control. This invention is a scalable optimization control method for charging and discharging of electric vehicle clusters based on deep reinforcement learning, considering the complex flexibility constraints of individual electric vehicles and the overall optimal scheduling of the electric vehicle cluster.
[0163] Figure 7 This is a schematic diagram of the structure of a charge / discharge scalable optimization control device provided in an embodiment of the present invention, as shown below. Figure 7As shown, the charge / discharge scalable optimization control device may include the above-mentioned Figure 6 The illustrated device is a reinforcement learning-based scalable optimization control device for charging and discharging electric vehicle clusters. Optionally, the scalable optimization control device 710 may include a first processor 2001.
[0164] Optionally, the charge / discharge scalable optimization control device 710 may also include a memory 2002 and a transceiver 2003.
[0165] The first processor 2001, memory 2002, and transceiver 2003 can be connected via a communication bus.
[0166] The following is combined with Figure 7 A detailed description of each component of the charge / discharge scalable optimization control device 710 is provided below:
[0167] The first processor 2001 is the control center of the charge / discharge scalable optimization control device 710. It can be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 can be one or more central processing units (CPUs), application-specific integrated circuits (ASICs), or one or more integrated circuits configured to implement embodiments of the present invention, such as one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs).
[0168] Optionally, the first processor 2001 can perform various functions of the charge / discharge scalable optimization control device 710 by running or executing software programs stored in the memory 2002 and calling data stored in the memory 2002.
[0169] In a specific implementation, as one example, the first processor 2001 may include one or more CPUs, for example... Figure 7 CPU0 and CPU1 are shown in the diagram.
[0170] In a specific implementation, as one example, the charge / discharge scalable optimization control device 710 may also include multiple processors, for example... Figure 7The first processor 2001 and the second processor 2004 are shown in the diagram. Each of these processors can be a single-core processor or a multi-core processor. Here, a processor can refer to one or more devices, circuits, and / or processing cores used to process data (such as computer program instructions).
[0171] The memory 2002 is used to store the software program that executes the present invention, and is controlled by the first processor 2001 to execute it. The specific implementation method can be referred to the above method embodiment, and will not be repeated here.
[0172] Optionally, the memory 2002 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. The memory 2002 may be integrated with the first processor 2001 or exist independently, and its interface circuitry can be expanded and optimized through charging and discharging. Figure 7 (Not shown in the image) is coupled to the first processor 2001, and this embodiment of the invention does not specifically limit this.
[0173] The transceiver 2003 is used to communicate with network devices or with terminal devices.
[0174] Alternatively, transceiver 2003 may include a receiver and a transmitter. Figure 7 (Not shown separately). The receiver is used to implement the receiving function, and the transmitter is used to implement the transmitting function.
[0175] Optionally, the transceiver 2003 can be integrated with the first processor 2001 or exist independently, and the interface circuit of the control device 710 can be expanded and optimized through charging and discharging. Figure 7 (Not shown in the image) is coupled to the first processor 2001, and this embodiment of the invention does not specifically limit this.
[0176] It should be noted that, Figure 7 The structure of the charge / discharge scalable optimization control device 710 shown does not constitute a limitation on the router. Actual knowledge structure identification devices may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0177] Furthermore, the technical effects of the charge-discharge scalable optimization control device 710 can be referred to the technical effects of the reinforcement learning-based electric vehicle cluster charge-discharge scalable optimization control method described in the above method embodiments, and will not be repeated here.
[0178] It should be understood that the first processor 2001 in the embodiments of the present invention may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0179] It should also be understood that the memory in the embodiments of the present invention can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0180] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.
[0181] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.
[0182] In this invention, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be a single item or multiple items.
[0183] It should be understood that, in various embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0184] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0185] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0186] In the several embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0187] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0188] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0189] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0190] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A reinforcement learning based scalable optimization control method for charging and discharging of an electric vehicle cluster, characterized in that, The method comprises: S1, based on the Markov chain simulation method, simulating the uncertain travel and energy demand of the electric vehicle, obtaining the parking demand and battery information of the electric vehicle; S2, based on the parking demand and battery information, calculating the charging and discharging priority and reference charging power of the electric vehicle, constructing the state variable of the electric vehicle, and calculating the charging and discharging relaxation; S3, obtaining the electricity price signal; performing state dimension reduction calculation on the state variable of all electric vehicles in the cluster to obtain the aggregation state vector of the electric vehicle aggregator, and calculating the dynamic power boundary of the cluster; S4, based on the linear annealing epsilon-greedy strategy, the deep reinforcement learning agent selects the charging and discharging control action according to the aggregation state vector of the electric vehicle aggregator, and obtains the preliminary control action vector; S5, based on the dynamic power boundary, filtering the preliminary control action vector to obtain the safe control action vector, and imposing a penalty on the out-of-limit action decision of the deep reinforcement learning agent; S6, based on the dynamic power boundary and the charging and discharging priority and relaxation of the electric vehicle, performing electric vehicle charging and discharging task decomposition according to the safe control action vector to obtain the electric vehicle execution action vector; S7, according to the electric vehicle execution action vector, the distributed intelligent charging pile performs charging and discharging control action.
2. The reinforcement learning-based scalable optimization control method for charging and discharging of an electric vehicle cluster according to claim 1, characterized in that, The S1 uncertain travel and energy demand of the electric vehicle includes the parking time, parking time, travel time, driving distance, travel purpose and parking location of the electric vehicle; The parking requirements of the electric vehicle include an arrival time , an estimated departure time , an initial state of charge and a minimum amount of charge at departure ; The battery information includes a battery capacity , a rated charge power , a rated discharge power , a charge efficiency , and a discharge efficiency . 3.The reinforcement learning based scalable optimization control method for charging and discharging of an electric vehicle cluster according to claim 1, characterized in that, The S2 based on the parking demand and battery information, calculating the charging and discharging priority and reference charging power of the electric vehicle, constructing the state variable of the electric vehicle, and calculating the charging and discharging relaxation, includes: S21, calculate the single electric vehicle flexibility constraints, including maximum capacity constraint, minimum discharge constraint and user satisfaction constraint, according to the parking demand and battery information ; S22, based on the electric vehicle flexibility constraint, performing charge-discharge priority analysis to obtain a single electric vehicle charge-discharge priority ; S23, determining a reference charging power for each electric vehicle based on the rated charging power of the electric vehicle , the reference charging power for each electric vehicle is determined based on the current state of charge and the user satisfaction constraint S24, based on the electric vehicle charging and discharging priority classification result , reference charging power, state of charge of the current step t and battery capacity , construct a single electric vehicle state variable ; S25, based on the electric vehicle battery information and the user satisfaction constraint , obtaining the charging relaxation and the discharging relaxation of the single electric vehicle , according to the state of charge of the electric vehicle and the state of discharge of the electric vehicle . 4.The reinforcement learning based scalable optimization control method for charging and discharging of an electric vehicle cluster according to claim 1, characterized in that, The S3 obtains the electricity price signal; performing state dimension reduction calculation on the state variable of all electric vehicles in the cluster to obtain the aggregation state vector of the electric vehicle aggregator, and calculating the dynamic power boundary of the cluster, including: S31, maximum control capacity of the electric vehicle aggregator based on state variables of all electric vehicles in the cluster calculating state of charge of the electric vehicle aggregator and reference charging power ; S32. Calculate the dynamic power boundary of the electric vehicle aggregator including the maximum charging power limit, the minimum charging power limit and the maximum discharging power limit according to the flexibility constraints of all electric vehicles in the cluster ; S33. constructing an aggregation state vector of the electric vehicle aggregator based on the power price signal , state of charge of the electric vehicle aggregator , reference charging power of the electric vehicle aggregator and dynamic power boundary of the electric vehicle aggregator .
5. The reinforcement learning-based scalable optimization control method for charging and discharging of an electric vehicle cluster according to claim 1, characterized in that, The S5 based on the dynamic power boundary, filtering the preliminary control action vector to obtain the safe control action vector, and imposing a penalty on the out-of-limit action decision of the deep reinforcement learning agent, including: S51. Based on the over-limit action protection rules, and according to the dynamic power boundary information of the electric vehicle aggregator, the initial control action vector is... Perform verification to obtain the security verification result; S52、based on the security check result, the preliminary control action vector cut off to the dynamic power boundary, obtain the security control action vector ; S53、based on the security check result, obtain an out-of-limit action penalty term .
6. The reinforcement learning-based scalable optimization control method for charging and discharging of an electric vehicle cluster according to claim 1, characterized in that, The S6 based on the dynamic power boundary and the charging and discharging priority and relaxation of the electric vehicle, performing electric vehicle charging and discharging task decomposition according to the safe control action vector to obtain the electric vehicle execution action vector, including: S61、based on the electric vehicle charging and discharging priority Classify the electric vehicle cluster to obtain a priority classification; S62, based on the priority classification result, according to the charging relaxation degree and the discharging relaxation degree , the electric vehicles in each category are respectively sorted in order from small to large, and the charging and discharging priority order of the electric vehicles in each priority category is obtained. S63, benchmark charging power based on electric vehicle aggregator and safety control action vectors obtaining electric vehicle charging and discharging tasks of different priority classifications; S64, based on the electric vehicle charging and discharging priority order, assigning the charging and discharging tasks of each priority category to the electric vehicles, obtaining an electric vehicle execution action vector .
7. The reinforcement learning-based scalable optimization control method for charging and discharging of an electric vehicle cluster according to claim 1, characterized in that, After obtaining the electric vehicle execution action vector in S6, the method further comprises: updating the state vector of the electric vehicle at the next time step t+1 according to the electric vehicle performing the action vector and performing dimension reduction calculation to obtain the aggregated state vector of the electric vehicle aggregator ; According to the action cost exceeding the limit And the energy cost of charging and discharging Obtain the reward function ; According to the aggregation state vector of the electric vehicle aggregator at the current time step and the next time step and , the electric vehicle aggregator security control action vector and the reward function , the training experience data ( , , , ) is stored in the replay buffer B; After the amount of training experience data reaches the preset minimum threshold, a small batch of training experience is extracted at each time step to train the deep reinforcement learning agent, so that the neural network of the deep reinforcement learning agent updates the parameters.
8. A reinforcement learning based scalable optimization control device for charging and discharging of an electric vehicle cluster, the reinforcement learning based scalable optimization control device for charging and discharging of an electric vehicle cluster being configured to implement the reinforcement learning based scalable optimization control method for charging and discharging of an electric vehicle cluster according to any one of claims 1 to 7. The device comprises: An information acquisition module is configured to simulate the uncertain travel and energy demand of the electric vehicle based on the Markov chain simulation method, and obtain the parking demand and battery information of the electric vehicle; An electric vehicle state construction module is configured to calculate the charging and discharging priority and reference charging power of the electric vehicle based on the parking demand and battery information, construct the state variable of the electric vehicle, and calculate the charging and discharging relaxation; An electric vehicle aggregator state aggregation module is configured to acquire a power price signal, perform state dimensionality reduction calculation on state variables of all electric vehicles in a cluster, obtain an aggregated state vector of the electric vehicle aggregator, and calculate a dynamic power boundary of the cluster; A control action selection module is configured to perform charging and discharging control action selection of a deep reinforcement learning agent according to the aggregated state vector of the electric vehicle aggregator based on a linear annealing e-greedy strategy, and obtain a preliminary control action vector; A charging and discharging safety control module is configured to perform over-limit action filtering on the preliminary control action vector based on the dynamic power boundary, obtain a safety control action vector, and impose a penalty on over-limit action decision of the deep reinforcement learning agent; A control action decomposition module is configured to perform electric vehicle charging and discharging task action decomposition according to the safety control action vector based on the dynamic power boundary and charging and discharging priorities and slackness of the electric vehicles, and obtain an electric vehicle execution action vector; A charging and discharging action execution module is configured to perform charging and discharging control action of a distributed intelligent charging pile according to the electric vehicle execution action vector.
9. A charge-discharge scalable optimization control device characterized by comprising: The charging and discharging scalable optimization control device comprises: a processor; a memory having computer readable instructions stored thereon, the computer readable instructions being executed by the processor to implement the method of any one of claims 1 to 7.
10. A computer readable storage medium, characterized in that, The computer readable storage medium stores program code, which can be called and executed by the processor to implement the method of any one of claims 1 to 7.
Citation Information
Patent Citations
Cluster electric vehicle charging behavior optimization method based on deep reinforcement learning
CN111934335A
Optimal scheduling method for electric vehicles to participate in energy storage market based on reinforcement learning
CN119362543A