Electric vehicle cluster charging and discharging extensible optimization control method based on reinforcement learning
By combining Markov chains and deep reinforcement learning, the problems of insufficient computational efficiency and accuracy in the optimal control of electric vehicle clusters are solved. This approach enables flexible constraints and overall optimal scheduling of electric vehicle clusters, thereby improving the real-time scheduling capability and system stability of electric vehicle clusters.
Patent Information
- Application Number
- CN202510846484.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-06-24
AI Technical Summary
Existing electric vehicle cluster optimization control technologies have failed to effectively address the problems of low computational efficiency and insufficient model accuracy caused by the complex constraints of individual electric vehicles, especially in large-scale electric vehicle clusters where real-time scheduling and optimization control are difficult to achieve.
By employing Markov chain simulation and deep reinforcement learning, and considering the uncertain travel and energy demands of electric vehicles, dynamic power boundaries and safety control strategies are constructed through state dimensionality reduction and action decomposition, thereby achieving flexibility constraints and overall optimal scheduling of electric vehicle clusters.
It improves the computational efficiency and accuracy of optimized charging and discharging control of electric vehicle clusters, can adapt to real-time dynamic changes, ensures system stability and reliability, reduces the risk of dimensional disasters, and achieves a balance between power grid and user needs.
Smart Images

Figure CN120792569A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of electric vehicle charging and discharging optimization control, in particular to an electric vehicle cluster charging and discharging scalable optimization control method and device based on reinforcement learning. BACKGROUND
[0002] The electric vehicle (EV) provides a huge technical possibility for coping with global energy crisis and power grid fluctuations due to its electric characteristics. Under the background of increasingly tight energy supply, electric vehicles significantly improve energy utilization efficiency compared to traditional internal combustion engine vehicles through efficient energy conversion mechanisms, reduce dependence on fossil fuels, and electric vehicles have the technical characteristics of two-way energy flow and can be used as distributed energy storage units to provide adjustable load resources for power supply and demand matching. With the increasing penetration rate of electric vehicles, the impact of electric vehicle load on the power grid is becoming more and more prominent. Large-scale electric vehicle charging may exacerbate power grid load peaks, making the power grid face greater pressure; reasonable regulation and control of electric vehicle charging and discharging behavior can enhance the flexibility and stability of the power grid and reduce the operating cost of the power grid.
[0003] The schedulable charging and discharging capacity of a single electric vehicle is small and highly uncertain. Electric vehicle aggregators (EVA) act as a bridge between electric vehicle users and the power grid, ensuring user travel needs and economic benefits through smart contracts; and through advanced optimization and control algorithms, the provision of power grid ancillary services is maximized, achieving a win-win situation for all parties. The scheduling of EVA to the electric vehicle cluster needs to be carried out under the premise of strictly meeting the user's energy demand, and if the user's expected power cannot be met before the electric vehicle leaves, it will seriously damage the satisfaction and willingness of the electric vehicle user to participate in the scheduling. The scheduling of the electric vehicle cluster is also affected by the user's random travel demand, and the parking time and location of the electric vehicle are uncertain and exhibit different aggregation effects in different time periods and spatial ranges, which will lead to dynamic changes in the schedulable flexibility of the corresponding period.
[0004] After the EVA obtains the overall charging and discharging optimization strategy of the cluster, the EVA allocates the charging and discharging quota and determines the specific charging and discharging control strategy of the individual electric vehicle. The sorting method based on time laxity plays a key role in this process and provides a scientific basis for the EVA to make individual electric vehicle control decisions. Time laxity is defined as the difference between the remaining time available for charging before the electric vehicle leaves and the shortest time required to complete charging. The smaller the time laxity of the electric vehicle, the higher the charging urgency, and the electric vehicle should be charged as soon as possible. Based on the concept of time laxity, multi-dimensional sorting indicators can be constructed, such as calculating time urgency and energy demand urgency by combining energy laxity and time laxity.
[0005] In the optimization problem of the charging and discharging control of the electric vehicle cluster, as the number of electric vehicles increases and the model setting becomes closer to reality, the optimization model becomes more complex, and the optimization solving efficiency also decreases. The existing optimization control technology of the electric vehicle cluster is often based on a simplified electric vehicle demand model, which makes the assumption that the dispatchable time and the required power of individual electric vehicles are homogeneous, which does not conform to the real-world electric vehicle usage pattern. Some studies use Markov chain method to simulate more realistic individual electric vehicle level uncertainty demand, but only stay in the flexibility evaluation of the aggregated load, without trying to optimize the control of the aggregated load.
[0006] In the existing optimization control technology, the problem of decreased calculation efficiency caused by the complexity of individual electric vehicle constraints has not been overcome. Centralized optimization methods (such as mixed integer linear programming and quadratic programming, etc.) focus on individual electric vehicle level uncertainty when solving the EVA charging and discharging problem, which leads to the number of decision variables of the optimization problem growing in proportion to the size of the electric vehicle cluster, and the solving time also grows exponentially. Mathematical programming method is a static solving method, which needs to solve the problem again in each iteration of the day-ahead scheduling, resulting in high calculation cost. When facing the dynamic changes of the electric vehicle cluster state, the solving speed does not meet the requirements of real-time control, and it takes a long time to complete an optimization iteration, making it difficult to respond to frequent fluctuations in the power grid and sudden changes in user behavior.
[0007] The rolling optimization framework integrates the future energy consumption prediction of the electric vehicle cluster, but the calculation burden increases dramatically with the increase of the prediction time domain, which cannot guarantee the optimization quality while considering the calculation efficiency. Some studies use mixed integer programming-based rolling optimization method to solve the V2G potential estimation problem of large-scale electric vehicles in the electricity market, which quantitatively describes the energy and power constraints of electric vehicles through the establishment of an aggregated model, and realizes the accurate estimation of the V2G potential of the vehicle fleet. However, this method focuses on the V2G potential estimation of the vehicle fleet in the past few hours, and it is difficult to realize real-time scheduling control of V2G.
[0008] Deep reinforcement learning (DRL) can better capture the dynamic characteristics of the electric vehicle cluster and quickly complete the charging and discharging optimization solution. In the scenario where the electric vehicle cluster provides auxiliary services to the power grid, the deep reinforcement learning algorithm aggregates and captures various uncertainties related to scheduling. In this method, a feature-based linear function approximator is used to handle the time-varying state-action space in the aggregated electric vehicle fleet scheduling problem.
[0009] In the aspect of describing the state and action space of large-scale electric vehicle cluster, the existing reinforcement learning method also faces the problem of balancing the optimization control precision and the calculation efficiency. With the growth of the electric vehicle cluster scale, the action and state space of the agent may grow exponentially, thereby causing serious "dimension disaster". In order to solve this problem, by modeling the aggregation flexibility of the electric vehicle cluster, an electric vehicle control flexibility model based on uncertain departure time is proposed, and by optimizing the charging power peak caused by the load rebound effect, the allocation control from the overall capacity of the scheduled electric vehicle to the individual response capacity is focused on. However, the above method fails to fully reflect the complexity of the electric vehicle level constraint in the overall scheduling optimization problem of the electric vehicle cluster. In the prior art, there is a lack of an electric vehicle cluster charging and discharging scalable optimization control method based on deep reinforcement learning, which considers the complex flexibility constraint of individual electric vehicles and the overall optimal scheduling of the electric vehicle cluster. SUMMARY
[0010] In order to solve the technical problems of the prior art that the electric vehicle demand modeling exists serious simplification assumption deviating from reality, it is difficult to balance the calculation efficiency and the model precision, and there is a lack of optimization control framework integrating the individual electric vehicle flexibility complex constraint and the overall scheduling optimality of the electric vehicle cluster, the embodiment of the present application provides an electric vehicle cluster charging and discharging scalable optimization control method and device based on reinforcement learning. The technical solution is as follows:
[0011] On the one hand, an electric vehicle cluster charging and discharging scalable optimization control method based on reinforcement learning is provided, which is realized by a charging and discharging scalable optimization control device, and the method comprises the following steps:
[0012] S1, based on the Markov chain simulation method, the uncertainty of the electric vehicle travel and the energy demand are simulated to obtain the parking demand and the battery information of the electric vehicle;
[0013] S2, based on the parking demand and the battery information, the charging and discharging priority and the reference charging power of the electric vehicle are calculated, the electric vehicle state variable is constructed, and the charging relaxation degree and the discharging relaxation degree are calculated;
[0014] S3, the electricity price signal is obtained; the state variable of all electric vehicles in the cluster is calculated, the aggregation state vector of the electric vehicle aggregator is obtained, and the dynamic power boundary of the cluster is calculated;
[0015] S4, based on the linear annealing e-greedy strategy, the deep reinforcement learning agent selects the charging and discharging control action according to the aggregation state vector of the electric vehicle aggregator to obtain the preliminary control action vector;
[0016] S5, based on the dynamic power boundary, the preliminary control action vector is filtered for over-limit action, a safe control action vector is obtained, and a punishment is imposed on the over-limit action decision of the deep reinforcement learning agent;
[0017] S6, based on the dynamic power boundary and the charging and discharging priority and relaxation degree of the electric vehicle, the charging and discharging task action of the electric vehicle is decomposed according to the safe control action vector, and an execution action vector of the electric vehicle is obtained;
[0018] S7, according to the execution action vector of the electric vehicle, the distributed intelligent charging pile performs charging and discharging control action.
[0019] On the other hand, a reinforcement learning-based electric vehicle cluster charging and discharging scalable optimization control device is provided, which is applied to the reinforcement learning-based electric vehicle cluster charging and discharging scalable optimization control method, and the device comprises:
[0020] An information acquisition module is configured to simulate the uncertainty of travel and energy demand of the electric vehicle based on a Markov chain simulation method, and obtain parking demand and battery information of the electric vehicle;
[0021] An electric vehicle state construction module is configured to calculate charging and discharging priority and reference charging power of the electric vehicle based on the parking demand and the battery information, construct state variables of the electric vehicle, and calculate charging relaxation degree and discharging relaxation degree;
[0022] An electric vehicle aggregator state aggregation module is configured to obtain a price signal, perform state dimension reduction calculation on state variables of all electric vehicles in the cluster, obtain an aggregated state vector of the electric vehicle aggregator, and calculate a dynamic power boundary of the cluster;
[0023] A control action selection module is configured to select a control action based on a linear annealing greedy strategy, the deep reinforcement learning agent selects a charging and discharging control action based on the aggregated state vector of the electric vehicle aggregator, and obtains a preliminary control action vector;
[0024] A charging and discharging safety control module is configured to filter the preliminary control action vector for over-limit action based on the dynamic power boundary, obtain a safe control action vector, and impose a punishment on the over-limit action decision of the deep reinforcement learning agent;
[0025] A control action decomposition module is configured to decompose the charging and discharging task action of the electric vehicle based on the dynamic power boundary and the charging and discharging priority and relaxation degree of the electric vehicle according to the safe control action vector, and obtain an execution action vector of the electric vehicle;
[0026] A charging and discharging action execution module is configured to perform charging and discharging control action of the distributed intelligent charging pile according to the execution action vector of the electric vehicle.
[0027] In another aspect, there is provided a charge-discharge scalable optimization control device, comprising: a processor; a memory having computer readable instructions stored thereon, the computer readable instructions, when executed by the processor, implement any one of the above-mentioned reinforcement learning-based electric vehicle cluster charge-discharge scalable optimization control methods.
[0028] In another aspect, there is provided a computer readable storage medium having at least one instruction stored therein, the at least one instruction being loaded and executed by a processor to implement any one of the above-mentioned reinforcement learning-based electric vehicle cluster charge-discharge scalable optimization control methods.
[0029] The technical scheme provided by the embodiments of the present application has at least the following beneficial effects:
[0030] The present application provides a reinforcement learning-based electric vehicle cluster charge-discharge scalable optimization control method, which introduces a Markov chain-based electric vehicle behavior generation method, breaks through the limitation of oversimplified electric vehicle usage mode, and accurately captures the time and energy constraints of individual electric vehicles. The electric vehicle aggregator charge-discharge agent based on deep reinforcement learning has good adaptability to the uncertainty in real-time scheduling, can learn the dynamic changes of the flexibility of the electric vehicle cluster through interaction with the environment, and form an optimal strategy that meets the actual user demand.
[0031] By integrating the state vector of individual electric vehicles, the state of the electric vehicle aggregator is represented in a reduced dimension, and the proposed dynamic power boundary model and action decomposition method effectively overcomes the challenge of exponential increase in action space dimension with the growth of the electric vehicle cluster size. By decoupling the state and action vector dimension from the size of the vehicle fleet, the problem is reduced from to , realizing the lightweight design of the state and action space of the deep reinforcement learning agent, and effectively solving the dimension disaster problem in the optimization process.
[0032] The safety control strategy designed in the present application guarantees the stability and reliability of the system from two aspects. On the one hand, the dynamic power boundary of the electric vehicle aggregator is established according to the flexibility constraints of the electric vehicles in the cluster, and a penalty term is imposed on the actions of the agent that violate the dynamic power boundary, thereby accelerating the convergence of the agent's strategy. On the other hand, the safety filter is applied to the over-limit actions given by the agent, and the control actions are projected on the dynamic power boundary to ensure the safety of the actual control. The present application is a reinforcement learning-based electric vehicle cluster charge-discharge scalable optimization control method that considers the complex flexibility constraints of individual electric vehicles and the overall optimal scheduling of the electric vehicle cluster. BRIEF DESCRIPTION OF DRAWINGS
[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiments will be briefly introduced as follows. Obviously, the drawings in the following description only constitute some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained from these drawings without creative effort.
[0034] Figure 1 is a flow chart of a scalable optimization control method for charging and discharging of an electric vehicle cluster based on reinforcement learning provided by an embodiment of the present application.
[0035] Figure 2 is a schematic diagram of a general framework provided by an embodiment of the present application.
[0036] Figure 3 is a schematic diagram of a trend of an input electricity price signal provided by an embodiment of the present application.
[0037] Figure 4 is a schematic diagram of state of charge and upper and lower limits of an electric vehicle cluster under different charging and discharging strategies provided by an embodiment of the present application.
[0038] Figure 5 is a schematic diagram of charging and discharging power of an electric vehicle cluster under different charging and discharging strategies provided by an embodiment of the present application.
[0039] Figure 6 is a block diagram of a scalable optimization control device for charging and discharging of an electric vehicle cluster based on reinforcement learning provided by an embodiment of the present application.
[0040] Figure 7 is a schematic diagram of a structure of a charging and discharging scalable optimization control device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0041] The technical solutions in the present application will be described below with reference to the drawings.
[0042] In the embodiments of the present application, the words such as “example”, “for example” and the like are used to represent an example, illustration or description. Any embodiment or design scheme described as “example” in the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, the word “example” is intended to present a concept in a specific manner. In addition, in the embodiments of the present application, the meaning expressed by “and / or” can be both, or can be one of the two optionally.
[0043] In the embodiments of the present application, the terms "image" and "picture" can be used interchangeably, and it should be pointed out that the meanings expressed are consistent when the distinction is not emphasized.
[0044] In the embodiments of the present application, the subscript such as W1 may be written in the form of non-subscript such as W1, and the meanings expressed are consistent when the distinction is not emphasized.
[0045] To make the technical problems, technical solutions and advantages of the present application clearer, the following will be described in detail with reference to the drawings and specific embodiments.
[0046] The embodiments of the present application provide a reinforcement learning-based scalable charging and discharging optimization control method for an electric vehicle cluster, which can be implemented by a charging and discharging scalable optimization control device, which can be a terminal or a server. Figure 1 As shown in the flowchart of the reinforcement learning-based scalable charging and discharging optimization control method for an electric vehicle cluster. In a feasible implementation manner, as shown in Figure 2 The method can be divided into four levels: an electric vehicle level, an electric vehicle aggregator level, a safety control level and an optimization decision level. For each time step of the electric vehicle cluster charging and discharging optimization control period, at the electric vehicle level, the information acquisition module simulates the electric vehicle battery parameter and demand data, the electric vehicle state construction module constructs the electric vehicle flexibility constraint, and the electric vehicle state variable and slack information are calculated; the electric vehicle state variable and slack information are input into the electric vehicle aggregator level, the dynamic power boundary and aggregated state vector are obtained through the electric vehicle aggregator state aggregation module; the aggregated state vector is input into the control action selection module of the optimization decision level, and the preliminary action vector is output by the deep reinforcement learning agent according to the aggregated state vector; the preliminary action vector is input into the safety control level, and the safety control action vector is output through the charging and discharging safety control module; the safety control action vector is input into the electric vehicle aggregator level, and the electric vehicle execution action vector is obtained through the control action decomposition module; and the execution action vector is input into the electric vehicle level, and the actual charging and discharging control is performed through the charging and discharging action execution module.
[0047] The processing flow of the method can include the following steps:
[0048] S1, based on the Markov chain simulation method, the uncertainty of the electric vehicle and the energy demand are simulated, and the parking demand and battery information of the electric vehicle are obtained.
[0049] Among them, the uncertain travel and energy demand of electric vehicles include: parking time, parking duration, travel duration, driving mileage, travel purpose and parking location of electric vehicles. The above parameters all obey the corresponding probability distribution. The travel mode and energy usage preference of electric vehicles can be set by adjusting the probability distribution of the Markov chain model.
[0050] Among them, when the electric vehicle arrives, the distributed intelligent charging pile reads the parking demand data of the individual electric vehicle. The parking demand of the electric vehicle includes the arrival time , Estimated departure time , initial state of charge and minimum charge when leaving ; Read individual electric vehicle battery information including battery capacity , rated charging power , rated discharge power , charging efficiency and discharge efficiency .
[0051] According to the parking demand and battery information of electric vehicles, the electric vehicle access status at the smart charging pile is updated at each time step .
[0052] S2. Based on parking demand and battery information, the charging and discharging priority and benchmark charging power of the electric vehicle are calculated, the electric vehicle state variables are constructed, and the charging slack and discharging slack are calculated.
[0053] Optionally, the specific implementation process of S2 includes S21-S25:
[0054] S21. Calculate the flexibility constraints of a single electric vehicle, including the maximum capacity constraint, based on parking requirements and battery information. , minimum discharge constraint and electric vehicle user satisfaction constraints ;in, and It can be determined based on the performance of the electric vehicle battery and can be set as a reasonable constant; To ensure that electric vehicles reach the , and the minimum charge that should be reached at time t during this parking period; According to the minimum charge that should be reached at time t during this parking period, calculate the user satisfaction constraint It is expressed by the following formula (1):
[0055] (1);
[0056] in, Represents the current time step.
[0057] S22, based on the electric vehicle flexibility constraint, the charging and discharging priority analysis is carried out, and the charging and discharging priority of a single electric vehicle is obtained ;
[0058] In an available implementation, according to the state of charge of the electric vehicle The charging and discharging priority of the electric vehicle is classified into four categories :
[0059] reach , and greater than and greater than , the electric vehicle only has discharging flexibility, and cannot be charged due to battery capacity limitation, .
[0060] Not reaching , and equal to , the electric vehicle only has charging flexibility, .
[0061] Less than or , the electric vehicle has no flexibility and must be charged within the remaining parking time, .
[0062] In other cases, the electric vehicle has charging and discharging flexibility, .
[0063] S23, based on the rated charging power of the electric vehicle , the reference charging power of a single electric vehicle is determined according to the current state of charge and the user satisfaction constraint ;
[0064] Among them, when the individual electric vehicle should continue to charge before reaching , the reference charging power of the individual electric vehicle is the rated power. When reaching , the charging can be stopped, and the reference charging power of the individual electric vehicle is 0.
[0065] S24, based on the charging and discharging priority classification result of the electric vehicle, the reference charging power, the state of charge and the battery capacity of the current step t, the state variable of a single electric vehicle is constructed ;
[0066] S25, based on the battery information of the electric vehicle and the user satisfaction constraint , according to the state of charge of the electric vehicle , the charging relaxation of a single electric vehicle and the discharging relaxation .
[0067] wherein the relaxation of the electric vehicle is taken as a control ranking reference, the calculation of the charging relaxation and the discharging relaxation is shown in formula (2) and formula (3) respectively:
[0068] (2);
[0069] (3);
[0070] S3, an electricity price signal is obtained; state dimension reduction calculation is performed on the state variables of all electric vehicles in the cluster to obtain the aggregated state vector of the electric vehicle aggregator, and the dynamic power boundary of the cluster is calculated. Wherein, the trend of the input electricity price signal is shown in formula (4). Figure 3
[0071] In a feasible implementation, at each time step, the electric vehicle aggregator receives the state variables of all electric vehicles , the charging relaxation of all electric vehicles and the discharging relaxation of all electric vehicles. At each time step, the safety control action vector sent by the electric vehicle charging and discharging safety control module is received ; at each time step, the electricity price signal is obtained ; at each time step, the aggregated state vector of the electric vehicle aggregator is sent to the deep reinforcement learning optimization module ; at each time step, the control action vector is sent to the distributed electric vehicle charging and discharging flexibility calculation module .
[0072] The electric vehicle aggregator can sign a contract with the user, and the electric vehicle aggregator can perform charging and discharging optimization control on the electric vehicles connected to the charging and discharging infrastructure, provided that the flexibility constraints of the electric vehicles are met. The maximum number of vehicle fleets that the electric vehicle aggregator can accommodate is set to N, which can be determined by the charging infrastructure that the electric vehicle aggregator can control. Since the time of connecting the electric vehicle at the charging pile is uncertain, and the time of staying and the amount of electricity required by the connected electric vehicle are also uncertain, the size of the electric vehicle cluster controlled by the electric vehicle aggregator is dynamically changing, and the dispatchable flexibility of different electric vehicles has a large difference.
[0073] The state and action decision variables of the electric vehicle group composed of individual electric vehicle related variables are highly related to the number of controlled vehicles. Therefore, the application proposes a state dimension reduction and action decomposition method, aiming to reduce the dimension of the state vector obtained by the deep reinforcement learning agent and the dimension of the control action vector given by the agent under the premise of ensuring the accuracy of the electric vehicle charging and discharging optimization control.
[0074] At each time step, the electric vehicle aggregator receives the price signal input by the external power grid , receives the usage state at all charging and discharging infrastructure , and receives the state vector of all controlled individual electric vehicles , including: based on the electric vehicle charging and discharging priority classification result , the reference charging power, the state of charge at the current step t and the battery capacity . The electric vehicle aggregator updates the aggregation state of the electric vehicle aggregator based on the electric vehicle state information. Among them, the flexibility of all electric vehicles controlled by the electric vehicle aggregator is regarded as a virtual energy storage system.
[0075] Optionally, the specific implementation process of S3 includes S31-S33:
[0076] S31, based on the maximum control capacity of the electric vehicle aggregator , according to the state variables of all electric vehicles in the cluster , calculate the state of charge of the electric vehicle aggregator and the reference charging power ; wherein the reference charging power of the electric vehicle aggregator is the sum of the reference charging powers of all controlled individual electric vehicles;
[0077] In a feasible implementation manner, the state of charge of the electric vehicle aggregator The calculation formula is as follows formula (4):
[0078] (4);
[0079] Wherein, represents the maximum capacity that the electric vehicle aggregator virtual energy storage system can control.
[0080] S32, according to the flexibility constraints of all electric vehicles in the cluster, calculate the dynamic power boundary of the electric vehicle aggregator, including the maximum charging power limit , the minimum charging power limit and the maximum discharging power limit ;
[0081] wherein the charging and discharging power constraints of the electric vehicle aggregator are based on the flexibility constraints of individual electric vehicles, the maximum charging power limit of the electric vehicle aggregator , the minimum charging power limit , the maximum discharging power limit , the calculation formula is shown in the following formulas (5), (6) and (7):
[0082] (5);
[0083] (6);
[0084] (7);
[0085] S33, constructing an aggregation state vector of the electric vehicle aggregator according to the electricity price signal , the state of charge of the electric vehicle aggregator , the reference charging power of the electric vehicle aggregator and the dynamic power boundary of the electric vehicle aggregator .
[0086] wherein the finally formed aggregation state vector of the electric vehicle aggregator is shown in the following formula (8):
[0087] (8);
[0088] S4, a linear annealing-based -greedy strategy, a deep reinforcement learning agent selects a charging and discharging control action according to the aggregation state vector of the electric vehicle aggregator to obtain a preliminary control action vector.
[0089] In a feasible implementation, at each time step, a double deep Q network agent observes the aggregation state vector of the electric vehicle aggregator , and performs normalization processing on . The action vector output by the agent is represented as , representing the total charging and discharging power of the electric vehicle aggregator level, a positive value representing charging and a negative value representing discharging. The reward function of the agent is the sum of the energy cost of charging and discharging and the over-limit punishment of charging and discharging power, represented as , wherein represents a weight coefficient, represents an over-limit action punishment item.
[0090] In a feasible implementation, the agent selects an action according to a linear annealing-based -greedy strategy, and the probability of selecting the action is 1- In the case of , the agent selects an action using a parameterized Q-network with probability , and selects a random action with probability linearly decreasing with the number of training iterations, and eventually reaching a minimum value
[0091] . In this way, the agent is able to improve convergence by exploring a lot at the beginning and using experience. In a feasible implementation, the double deep Q-network agent uses a deep neural network to approximate the action value function, and calculates the target Q value through a target Q network
[0092] to stabilize the training process, thereby indirectly optimizing the policy to maximize the cumulative reward.
[0093] Optionally, the specific implementation process of S5 includes S51-S53:
[0094] S51, based on the over-limit action protection rule, the preliminary control action vector is verified according to the dynamic power boundary information of the electric vehicle aggregator, and a safe verification result is obtained;
[0095] Among them, the over-limit action protection rule is proposed, and the dynamic power boundary information is used as the verification standard for whether the agent output action is over-limit, and two over-limit action situations may occur, including overcharging or over-discharging caused by decision power exceeding the maximum charging power limit or the maximum discharging power limit of the electric vehicle.
[0096] S52, based on the safe verification result, the preliminary control action vector is truncated to the dynamic power boundary, and a safe control action vector is obtained.
[0097] Among them, according to the safe verification result, when the preliminary control action vector exceeds the dynamic power boundary of the electric vehicle aggregator, the action is truncated to the executable range, otherwise no filtering is performed. Based on the over-limit action protection rule, the safe control action vector is obtained, and the calculation process is shown in formula (9); the safe control action vector is output to the electric vehicle aggregator.
[0098] (9)
[0099] S53, obtaining an out-of-limit action penalty term based on the security check result .
[0100] wherein the out-of-limit action penalty term is obtained by formula (10):
[0101] (10);
[0102] The calculated is multiplied by a preset weight factor to obtain an out-of-limit penalty value of the action and is sent to the deep reinforcement learning agent to help the agent avoid giving out-of-limit actions and accelerate policy convergence.
[0103] S6, based on the dynamic power boundary and the charging and discharging priority and relaxation degree of the electric vehicle, performing electric vehicle charging and discharging task action decomposition according to the safety control action vector to obtain an electric vehicle execution action vector.
[0104] Optionally, the specific implementation process of S6 includes:
[0105] S61, based on the charging and discharging priority of the electric vehicle classifying the electric vehicle cluster to obtain a priority classification;
[0106] wherein the priority classification of a single electric vehicle includes .
[0107] S62, based on the priority classification result, respectively sorting the electric vehicles in each category in order from small to large according to the charging relaxation degree and the discharging relaxation degree to obtain the charging and discharging priority order of the electric vehicles in each priority category;
[0108] wherein the electric vehicle aggregator collects the charging relaxation degree and the discharging relaxation degree of all controlled electric vehicles at each time step to sort as quantitative indicators for determining task allocation priority. The smaller the relaxation degree, the greater the urgency, and the electric vehicle aggregator allocates charging or discharging tasks in order from small to large relaxation degree.
[0109] S63, based on the benchmark charging power of the electric vehicle aggregator and the safety control action vector obtaining charging and discharging tasks for electric vehicles of different priority classifications;
[0110] In a feasible implementation, when When all individual electric vehicles are charging according to the benchmark charging power, the electric vehicle aggregator does not change the charging behavior of individual electric vehicles; when the electric vehicle aggregator allocates the charging task exceeding the baseline power to the priority category ; when , first, the electric vehicle with the priority category is delayed charging, if the delayed charging is not enough to achieve , the electric vehicle aggregator allocates the discharging task to the electric vehicle with the priority category . Wherein, in order to protect the user comfort, for any , the charging and discharging control of the electric vehicle aggregator meets the charging demand of the priority category without change.
[0111] S64, based on the charging and discharging priority order of the electric vehicle, the charging and discharging tasks of each priority category are allocated to the electric vehicle, and the execution action vector of the electric vehicle is obtained .
[0112] Wherein, in order to reduce the risk of battery degradation, when the electric vehicle aggregator allocates the demand response task within each priority category, for the electric vehicles in the same flexibility category, the scheduling order is determined according to the principle of priority of minimum relaxation, the total charging and discharging amount of each priority category is decomposed into the charging and discharging tasks that can be realized by individual electric vehicles, and the execution action vector of the electric vehicle is obtained . The electric vehicle aggregator sends the execution action vector to the corresponding electric vehicle, and calculates the charging and discharging cost combined with the electricity price signal .
[0113] Optionally, after obtaining the electric vehicle execution action vector at this time step, in order to iteratively optimize the charging and discharging control at the next time step, the method further comprises:
[0114] According to the electric vehicle execution action vector, the state vector of the electric vehicle at the next time step t+1 is updated and the dimension reduction calculation is performed, and the aggregated state vector of the electric vehicle aggregator is obtained ;
[0115] According to the action out-of-limit cost and the charging and discharging energy cost , the reward function is obtained;
[0116] According to the aggregated state vectors of the electric vehicle aggregator at the current time step and the next time step and , the safe control action vector of the electric vehicle aggregator and the reward function , build training experience data is stored in the replay buffer 𝐵;
[0117] After the amount of training experience data reaches the preset minimum threshold, a small batch of training experience is extracted at each time step to train the deep reinforcement learning agent, so as to update the parameters of the neural network of the deep reinforcement learning agent.
[0118] In one feasible implementation, when When the capacity reaches the minimum threshold, A small batch of experience is extracted from the training data to train the agent. And the parameters of the Q network and the target Q network are updated. During the training process, the learning rate is set to 1 10 -3 The training batch size and the minimum replay buffer capacity are set to 32 and 1 respectively. 10 3 .
[0119] S7. According to the action vector executed by the electric vehicle, the distributed intelligent charging pile performs charging and discharging control actions.
[0120] in, Figure 4 This is a schematic diagram of the state of charge and upper and lower limits of an electric vehicle cluster under different charging and discharging strategies provided by an embodiment of the present invention; it shows the state of charge and upper and lower limits of an electric vehicle cluster in a disorderly charging scenario ( Figure 4 (a)), smart charging scenario ( Figure 4 (b)) and smart charging and discharging scenarios ( Figure 4 (c) Aggregate SOC and upper and lower limits. As shown in the figure, the random arrival and departure of electric vehicles leads to a step-like change in the aggregate SOC. By observing the real-time changes in the aggregate SOC, the deep reinforcement learning agent can learn the uncertain demand patterns of electric vehicles. At the same time, the method of the present invention can adhere to the upper and lower limits of the electric vehicle cluster's power consumption, avoiding compromising user comfort.
[0121] Correspondingly, Figure 5 Shows the disorderly charging scenario of electric vehicle groups ( Figure 5 (a)), smart charging scenario ( Figure 5 (b)) and smart charging and discharging scenarios ( Figure 5 (c) In a disorderly charging scenario, peak charging occurs in the evening when electricity prices are high. In contrast, the intelligent charging and discharging strategy controlled by a deep reinforcement learning agent takes full advantage of charging during low electricity prices and discharging during high electricity prices. This method can concentrate the discharging and charging power during peak and low electricity price periods, respectively, effectively reducing the energy costs of electric vehicles.
[0122] This paper proposes a scalable optimization control method for charging and discharging electric vehicle clusters based on reinforcement learning. By introducing a Markov chain-based electric vehicle behavior generation method, this method overcomes the limitations of oversimplified electric vehicle usage patterns and accurately captures the uncertain time and energy constraints of individual electric vehicles. The electric vehicle aggregator charging and discharging agent, based on deep reinforcement learning, is highly adaptable to the uncertainties of real-time scheduling. It can learn from the dynamic changes in the flexibility of electric vehicle clusters through interaction with the environment and form an optimal strategy that meets actual user needs.
[0123] By integrating the state vectors of individual electric vehicles and reducing the dimension of the state of the electric vehicle aggregator, the proposed dynamic power boundary model and action decomposition method effectively overcome the challenge of the exponential increase of the action space dimension as the scale of the electric vehicle cluster grows. By decoupling the state and action vector dimensions from the scale of the vehicle cluster, the problem is reduced from (N) Dimensionality reduction to (1) We achieve a lightweight design of the state and action space of deep reinforcement learning agents, effectively solving the curse of dimensionality problem in the optimization process.
[0124] The safety control strategy designed in this invention ensures system stability and reliability from two perspectives. First, a dynamic power boundary for the electric vehicle aggregator is established based on the flexibility constraints of the electric vehicles in the cluster. Penalties are imposed on actions taken by the agent that violate the dynamic power boundary, thereby accelerating the convergence of the agent's strategy. Second, over-limit actions given by the agent are safely filtered, and control actions are projected onto the dynamic power boundary to ensure the safety of actual control. This invention is a scalable optimization control method for charging and discharging electric vehicle clusters based on deep reinforcement learning, taking into account the complex flexibility constraints of individual electric vehicles and the overall optimal scheduling of the electric vehicle cluster.
[0125] Figure 6 This is a block diagram of a device for scalable optimization control of electric vehicle cluster charging and discharging based on reinforcement learning according to an exemplary embodiment. The device is used in a scalable optimization control method for electric vehicle cluster charging and discharging based on reinforcement learning. Figure 6 The device includes an information acquisition module 610, an electric vehicle state construction module 620, an electric vehicle aggregator state aggregation module 630, a control action selection module 640, a charge and discharge safety control module 650, a control action decomposition module 660, and a charge and discharge action execution module 670. Among them:
[0126] An information acquisition module 610 is used to simulate the uncertain travel and energy demand of the electric vehicle based on a Markov chain simulation method to obtain the parking demand and battery information of the electric vehicle;
[0127] The electric vehicle state construction module 620 is configured to calculate the charging and discharging priority and the benchmark charging power of the electric vehicle based on the parking demand and the battery information, construct the electric vehicle state variable, and calculate the charging relaxation and the discharging relaxation.
[0128] The electric vehicle aggregator state aggregation module 630 is configured to obtain the electricity price signal, perform state dimension reduction calculation on the state variables of all the electric vehicles in the cluster to obtain the aggregation state vector of the electric vehicle aggregator, and calculate the dynamic power boundary of the cluster.
[0129] The control action selection module 640 is configured to select the charging and discharging control action of the deep reinforcement learning agent based on the linear annealing - greedy strategy according to the aggregation state vector of the electric vehicle aggregator, to obtain a preliminary control action vector.
[0130] The charging and discharging safety control module 650 is configured to perform over-limit action filtering on the preliminary control action vector based on the dynamic power boundary to obtain a safety control action vector, and impose a penalty on the over-limit action decision of the deep reinforcement learning agent.
[0131] The control action decomposition module 660 is configured to perform electric vehicle charging and discharging task action decomposition according to the safety control action vector based on the dynamic power boundary and the charging and discharging priority and the relaxation of the electric vehicle, to obtain an electric vehicle execution action vector.
[0132] The charging and discharging action execution module 670 is configured to perform charging and discharging control action of the distributed intelligent charging pile according to the electric vehicle execution action vector.
[0133] Optionally, the uncertainty of the electric vehicle and the energy demand include the parking time, the parking duration, the travel duration, the driving mileage, the travel purpose and the parking location of the electric vehicle.
[0134] The parking demand of the electric vehicle includes the arrival time , the estimated departure time , the initial state of charge and the minimum amount of electricity at the time of departure .
[0135] The battery information includes the battery capacity , the rated charging power , the rated discharging power , the charging efficiency and the discharging efficiency .
[0136] Optionally, the electric vehicle state construction module 620 is configured to:
[0137] According to the parking demand and the battery information, the flexibility constraints of a single electric vehicle are calculated, including the maximum capacity constraint , the minimum discharge constraint and the user satisfaction constraint ;
[0138] Based on the flexibility constraints of the electric vehicle, the charging and discharging priority analysis is carried out to obtain the charging and discharging priority of a single electric vehicle ;
[0139] Based on the rated charging power of the electric vehicle , the reference charging power of a single electric vehicle is determined according to the current state of charge and the user satisfaction constraint ;
[0140] Based on the classification result of the charging and discharging priority of the electric vehicle , the reference charging power, the state of charge at the current step t and the battery capacity , the state variable of a single electric vehicle is constructed ;
[0141] Based on the battery information of the electric vehicle and the user satisfaction constraint , the charging relaxation degree and the discharging relaxation degree of a single electric vehicle are obtained according to the state of charge of the electric vehicle .
[0142] Optionally, the electric vehicle aggregator state aggregation module 630 is configured to:
[0143] Based on the maximum control capacity of the electric vehicle aggregator , the state of charge and the reference charging power of the electric vehicle aggregator are calculated according to the state variables of all electric vehicles in the cluster ;
[0144] According to the flexibility constraints of all electric vehicles in the cluster, the dynamic power boundary of the electric vehicle aggregator is calculated, including the maximum charging power limit , the minimum charging power limit and the maximum discharging power limit ;
[0145] According to the electricity price signal , the state of charge of the electric vehicle aggregator , the reference charging power of the electric vehicle aggregator and the dynamic power boundary of the electric vehicle aggregator, the aggregation state vector of the electric vehicle aggregator is constructed .
[0146] Optionally, the charge and discharge safety control module 650 is configured to:
[0147] Based on the over-limit action protection rule, the initial control action vector is Perform verification and obtain security verification results;
[0148] Based on the safety verification results, the initial control action vector Truncating to the dynamic power boundary to obtain a safe control action vector ;
[0149] Based on the safety check results, obtain the penalty item for exceeding the limit action .
[0150] Optionally, the control action decomposition module 660 is configured to:
[0151] Based on charging and discharging priority of electric vehicles Classify electric vehicle clusters and obtain priority classification;
[0152] Based on the priority classification results, according to the charging slack and discharge relaxation , sort the electric vehicles in each category from small to large, and obtain the charging and discharging priority of electric vehicles in each priority category;
[0153] Based on benchmark charging power of EV aggregators and safety control action vector ,obtain EV charging and discharging tasks with different priority classifications;
[0154] Based on the charging and discharging priority of electric vehicles, the charging and discharging tasks of each priority category are assigned to electric vehicles, and the electric vehicle execution action vector is obtained. .
[0155] Optionally, after the control action decomposition module 660, the following further comprises:
[0156] According to the electric vehicle's action vector, the state vector of the electric vehicle at the next time step t+1 is updated And perform dimensionality reduction calculation to obtain the aggregated state vector of the electric vehicle aggregator ;
[0157] According to the action excess cost and the energy cost of charging and discharging , and obtain the reward function ;
[0158] According to the aggregation state vector of the electric vehicle aggregator at the current time step and the next time step And , the electric vehicle aggregator safety control action vector And the reward function , the training experience data is constructed Stored in the replay buffer ;
[0159] After the amount of training experience data reaches the preset minimum threshold, a small batch of training experience is extracted at each time step to train the deep reinforcement learning agent, so that the neural network of the deep reinforcement learning agent is updated in parameters.
[0160] The application provides an electric vehicle cluster charging and discharging scalable optimization control method based on reinforcement learning, which introduces an electric vehicle behavior generation method based on Markov chain, breaks through the limitation of oversimplified electric vehicle usage mode, and accurately captures the time and energy constraints of individual electric vehicles. The electric vehicle aggregator charging and discharging agent based on deep reinforcement learning has good adaptability to the uncertainty in real-time scheduling, can learn the dynamic changes of the flexibility of the electric vehicle cluster through interaction with the environment, and form an optimal strategy that meets the actual user demand.
[0161] By integrating the state vector of individual electric vehicles, the state of the electric vehicle aggregator is represented in a reduced dimension, and the proposed dynamic power boundary model and action decomposition method effectively overcomes the challenge of exponential increase of the action space dimension with the growth of the electric vehicle cluster size. By decoupling the state and action vector dimension from the cluster size, the problem is reduced from (N) to (1), realizing the lightweight design of the state and action space of the deep reinforcement learning agent, and effectively solving the dimension disaster problem in the optimization process.
[0162] The safety control strategy designed in the application guarantees the stability and reliability of the system from two aspects. On the one hand, the dynamic power boundary of the electric vehicle aggregator is established according to the flexibility constraints of the electric vehicles in the cluster, and a penalty term is imposed on the actions of the agent that violate the dynamic power boundary, thereby accelerating the convergence of the agent strategy. On the other hand, the safety filtering is performed on the over-limit actions given by the agent, and the control actions are projected on the dynamic power boundary to ensure the safety of the actual control. The application is an electric vehicle cluster charging and discharging scalable optimization control method based on deep reinforcement learning, which considers the complex flexibility constraints of individual electric vehicles and the overall optimal scheduling of the electric vehicle cluster.
[0163] Figure 7 is a structural schematic view of a charging and discharging scalable optimization control device provided by an embodiment of the application, as Figure 7As shown, the charge-discharge scalable optimization control device can include the above Figure 6 The reinforcement learning-based electric vehicle cluster charge-discharge scalable optimization control device is shown. Optionally, the charge-discharge scalable optimization control device 710 can include a first processor 2001.
[0164] Optionally, the charge-discharge scalable optimization control device 710 can further include a memory 2002 and a transceiver 2003.
[0165] The first processor 2001, the memory 2002, and the transceiver 2003 can be connected through a communication bus.
[0166] The specific implementation of the charge-discharge scalable optimization control device 710 will be described below. Figure 7 The specific implementation of the charge-discharge scalable optimization control device 710 will be described below.
[0167] The first processor 2001 is the control center of the charge-discharge scalable optimization control device 710, which can be one processor or a plurality of processing elements. For example, the first processor 2001 is one or more central processing units (CPUs), which can also be application specific integrated circuits (ASICs), or one or more integrated circuits configured to implement embodiments of the present application, such as one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGA).
[0168] Optionally, the first processor 2001 can execute various functions of the charge-discharge scalable optimization control device 710 by running or executing software programs stored in the memory 2002 and calling data stored in the memory 2002.
[0169] In a specific implementation, as an embodiment, the first processor 2001 can include one or more CPUs, such as the CPU0 and CPU1 shown in Figure 7
[0170] In a specific implementation, as an embodiment, the charge-discharge scalable optimization control device 710 can also include a plurality of processors, such as the CPU0 and CPU1 shown in Figure 7 The first processor 2001 and the second processor 2004 shown in the foregoing embodiments can be implemented by using a single-CPU or a multi-CPU. The processor herein can refer to one or more devices, circuits, and / or processing cores for processing data (for example, computer program instructions).
[0171] The memory 2002 is configured to store a software program for implementing the scheme of the present application, and the first processor 2001 is configured to control execution of the software program. For details, refer to the foregoing method embodiments, which will not be described herein again.
[0172] Alternatively, the memory 2002 can be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disk storage, a magneto-optical disk, a magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer, but is not limited to this. The memory 2002 can be integrated with the first processor 2001 or exist independently, and is coupled to the first processor 2001 through an interface circuit (not shown in the foregoing embodiments) of the charge-discharge scalable optimization control device 710. The embodiments of the present application are not limited in this regard. Figure 7
[0173] The transceiver 2003 is configured to communicate with a network device or a terminal device.
[0174] Alternatively, the transceiver 2003 can include a receiver and a transmitter (not shown in the foregoing embodiments). The receiver is configured to implement a receiving function, and the transmitter is configured to implement a transmitting function. Figure 7
[0175] Alternatively, the transceiver 2003 can be integrated with the first processor 2001 or exist independently, and is coupled to the first processor 2001 through an interface circuit (not shown in the foregoing embodiments) of the charge-discharge scalable optimization control device 710. The embodiments of the present application are not limited in this regard. Figure 7
[0176] It should be noted that, Figure 7 The structure of the charge-discharge scalable optimization control device 710 shown in the figure does not constitute a limitation on the router, and the actual knowledge structure identification device can include more or fewer components than the figure, or combine certain components, or different component arrangements.
[0177] In addition, the technical effects of the charge-discharge scalable optimization control device 710 can refer to the technical effects of the electric vehicle cluster charge-discharge scalable optimization control method based on reinforcement learning described in the above method embodiments, which will not be described here.
[0178] It should be understood that the first processor 2001 in the embodiments of the present application can be a central processing unit (CPU), and the processor can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0179] It should also be understood that the memory in the embodiments of the present application can be volatile or nonvolatile memory, or can include both volatile and nonvolatile memory. The nonvolatile memory can be read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically EPROM (EEPROM), or flash memory. The volatile memory can be random access memory (RAM) used as external cache. By way of example, and not limitation, many forms of random access memory (RAM) are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0180] The above-described embodiments can be implemented in whole or in part by software, hardware (such as a circuit), firmware, or any combination thereof. When implemented in software, the above-described embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are wholly or partially generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center through a wired (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. containing one or more available medium collections. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state disk.
[0181] It should be understood that the term "and / or" herein merely describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can represent the following three cases: A exists alone, A and B exist together, and B exists alone, where A and B can be singular or plural. In addition, the character " / " herein generally represents that the associated objects before and after it are in an "or" relationship, but it can also represent an "and / or" relationship, which can be understood according to the context before and after it.
[0182] In the present application, "at least one" means one or more, and "multiple" means two or more. "At least one of the following" or the like means any combination of the items, including any combination of single or multiple items. For example, at least one of a, b, or c can represent a, b, c, a-b, a-c, b-c, or a-b-c, where a, b, and c can be single or multiple.
[0183] It should be understood that in various embodiments of the present application, the size of the sequence number of the above-described processes does not mean the order of execution, and the execution order of the processes should be determined according to their functions and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0184] Those skilled in the art can clearly understand that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0185] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working processes of the devices, apparatuses and units described above can refer to the corresponding processes in the foregoing method embodiments, which will not be repeated here.
[0186] In several embodiments provided by the present application, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely schematic, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0187] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.
[0188] In addition, each functional unit in each embodiment of the present application can be integrated into a processing unit, or each unit can exist physically independently, or two or more units can be integrated into one unit.
[0189] If the functions are realized in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application or the parts of the present application that essentially contribute to the prior art or the parts of the technical solutions can be embodied in the form of software products. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the various embodiments of the present application. The aforementioned storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0190] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A scalable optimization control method for charging and discharging of electric vehicle clusters based on reinforcement learning, characterized in that: The method comprises: S1. Based on the Markov chain simulation method, the uncertain travel and energy demand of electric vehicles are simulated to obtain the parking demand and battery information of electric vehicles; S2. Based on parking demand and battery information, calculate the charging and discharging priority and benchmark charging power of the electric vehicle, construct the electric vehicle state variable, and calculate the charging slack and discharging slack; S3. Obtain electricity price signals; perform state dimensionality reduction calculations on the state variables of all electric vehicles in the cluster, obtain the aggregated state vector of the electric vehicle aggregator, and calculate the dynamic power boundary of the cluster; S4. Based on the linear annealing 𝜀-greedy strategy, the deep reinforcement learning agent selects the charging and discharging control action according to the aggregated state vector of the electric vehicle aggregator and obtains the preliminary control action vector; S5. Based on the dynamic power boundary, the initial control action vector is filtered out of the limit to obtain a safe control action vector, and penalties are imposed on the deep reinforcement learning agent's out-of-limit action decisions. S6. Based on the dynamic power boundary and the charging and discharging priority and slack of the electric vehicle, the electric vehicle charging and discharging task action is decomposed according to the safety control action vector to obtain the electric vehicle execution action vector; S7. According to the action vector executed by the electric vehicle, the distributed intelligent charging pile performs charging and discharging control actions.
2. The scalable optimization control method for charging and discharging of electric vehicle clusters based on reinforcement learning according to claim 1 is characterized in that: The uncertain travel and energy demand of the electric vehicle in S1, including the parking time, parking duration, travel duration, driving mileage, travel purpose and parking location of the electric vehicle; The parking demand of the electric vehicle includes the arrival time , Estimated departure time , initial state of charge and minimum charge when leaving ; The battery information includes battery capacity , rated charging power , rated discharge power , charging efficiency and discharge efficiency .
3. The scalable optimization control method for charging and discharging of electric vehicle clusters based on reinforcement learning according to claim 1 is characterized in that: The S2 calculates the charging and discharging priority and the reference charging power of the electric vehicle based on the parking demand and the battery information, constructs the electric vehicle state variable, and calculates the charging slack and the discharging slack, including: S21. Calculate the flexibility constraints of a single electric vehicle, including the maximum capacity constraint, based on parking requirements and battery information. , minimum discharge constraint and user satisfaction constraints ; S22. Based on the flexibility constraints of electric vehicles, perform charge and discharge priority analysis to obtain the charge and discharge priority of a single electric vehicle. ; S23, based on the rated charging power of electric vehicles , according to the current state of charge and user satisfaction constraints Determine the benchmark charging power for a single electric vehicle; S24. Classification results based on electric vehicle charging and discharging priority , reference charging power, current state of charge at step t and battery capacity , construct the state variables of a single electric vehicle ; S25, based on electric vehicle battery information and user satisfaction constraints , according to the state of charge of electric vehicles , obtain the charging slack of a single electric vehicle and discharge relaxation .
4. The scalable optimization control method for charging and discharging of electric vehicle clusters based on reinforcement learning according to claim 1 is characterized in that: The S3 acquires the electricity price signal; performs state dimension reduction calculation on the state variables of all electric vehicles in the cluster, obtains the aggregated state vector of the electric vehicle aggregator, and calculates the dynamic power boundary of the cluster, including: S31. Maximum control capacity based on electric vehicle aggregators , according to the state variables of all electric vehicles in the cluster , calculate the state of charge of electric vehicle aggregators and benchmark charging power ; S32. Calculate the dynamic power boundary of the EV aggregator based on the flexibility constraints of all EVs in the cluster, including the maximum charging power limit , minimum charging power limit and maximum discharge power limit ; S33, according to the electricity price signal , State of Charge of Electric Vehicle Aggregators , Benchmark charging power for electric vehicle aggregators The aggregated state vector of the EV aggregator is constructed based on the dynamic power boundary of the EV aggregator. .
5. The scalable optimization control method for charging and discharging of electric vehicle clusters based on reinforcement learning according to claim 1 is characterized in that: The S5 performs over-limit action filtering on the preliminary control action vector based on the dynamic power boundary to obtain a safe control action vector and imposes penalties on the over-limit action decisions of the deep reinforcement learning agent, including: S51, based on the over-limit action protection rule, according to the dynamic power boundary information of the electric vehicle aggregator, the preliminary control action vector Perform verification and obtain security verification results; S52, based on the safety check result, the preliminary control action vector Truncating to the dynamic power boundary to obtain a safe control action vector ; S53. Based on the safety check result, obtain the penalty item for exceeding the limit action .
6. The scalable optimization control method for charging and discharging of electric vehicle clusters based on reinforcement learning according to claim 1 is characterized in that: The S6 decomposes the electric vehicle charging and discharging tasks based on the dynamic power boundary and the charging and discharging priority and slack of the electric vehicle according to the safety control action vector to obtain the electric vehicle execution action vector, including: S61. Charging and discharging priority based on electric vehicles Classify electric vehicle clusters and obtain priority classification; S62, based on the priority classification result, according to the charging slack and discharge relaxation , sort the electric vehicles in each category from small to large, and obtain the charging and discharging priority of electric vehicles in each priority category; S63, Benchmark charging power based on electric vehicle aggregators and safety control action vector ,obtain EV charging and discharging tasks with different priority classifications; S64, based on the charging and discharging priority of the electric vehicle, assign the charging and discharging tasks of each priority category to the electric vehicle, and obtain the electric vehicle execution action vector .
7. The scalable optimization control method for charging and discharging of electric vehicle clusters based on reinforcement learning according to claim 1 is characterized in that: After obtaining the electric vehicle execution action vector in S6, the method further includes: According to the electric vehicle's action vector, the state vector of the electric vehicle at the next time step t+1 is updated And perform dimensionality reduction calculation to obtain the aggregated state vector of the electric vehicle aggregator ; According to the action excess cost and the energy cost of charging and discharging , and obtain the reward function ; According to the aggregated state vector of the electric vehicle aggregator at the current time step and the next time step and , Electric Vehicle Aggregator Safety Control Action Vector and the reward function , construct training experience data ( , , , ) is stored in the playback buffer 𝐵; After the amount of training experience data reaches the preset minimum threshold, a small batch of training experience is extracted at each time step to train the deep reinforcement learning agent, so as to update the parameters of the neural network of the deep reinforcement learning agent.
8. A scalable optimization control device for charging and discharging a cluster of electric vehicles based on reinforcement learning, wherein the scalable optimization control device for charging and discharging a cluster of electric vehicles based on reinforcement learning is used to implement the scalable optimization control method for charging and discharging a cluster of electric vehicles based on reinforcement learning as claimed in any one of claims 1 to 7, characterized in that: The device comprises: An information acquisition module is used to simulate the uncertain travel and energy demand of electric vehicles based on the Markov chain simulation method, and obtain the parking demand and battery information of electric vehicles; An electric vehicle state construction module is used to calculate the charging and discharging priority and benchmark charging power of the electric vehicle based on parking requirements and battery information, construct electric vehicle state variables, and calculate charging slack and discharging slack; The electric vehicle aggregator state aggregation module is used to obtain electricity price signals; perform state dimension reduction calculations on the state variables of all electric vehicles in the cluster to obtain the aggregated state vector of the electric vehicle aggregator and calculate the dynamic power boundary of the cluster; A control action selection module is used to select the charging and discharging control actions based on the 𝜀-greedy strategy based on linear annealing. The deep reinforcement learning agent selects the charging and discharging control actions according to the aggregated state vector of the electric vehicle aggregator to obtain the preliminary control action vector; The charging and discharging safety control module is used to filter out-of-limit actions of the preliminary control action vector based on the dynamic power boundary, obtain the safe control action vector, and impose penalties on the out-of-limit action decisions of the deep reinforcement learning agent; A control action decomposition module is used to decompose the electric vehicle charging and discharging task actions according to the safety control action vector based on the dynamic power boundary and the charging and discharging priority and slack of the electric vehicle, and obtain the electric vehicle execution action vector; The charging and discharging action execution module is used to execute the action vector according to the electric vehicle, and the distributed intelligent charging pile performs the charging and discharging control action.
9. A charge and discharge scalable optimization control device, characterized in that: The charge and discharge scalable optimization control device includes: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program code, which can be called by a processor to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Cluster electric vehicle charging behavior optimization method based on deep reinforcement learning
CN111934335A
Optimal scheduling method for electric vehicles to participate in energy storage market based on reinforcement learning
CN119362543A
Intelligent charging of multiple vehicles through learned experience
US20230196090A1
Electric vehicle charging system having reinforcement learning-based autonomous mobile charger, and operation method thereof
WO2025105559A1