A vehicle-network collaborative cluster orderly scheduling method and system

Through deep reinforcement learning technology, using user travel rules parameters and reward and punishment functions, the agent is trained to realize orderly scheduling of vehicle-network collaborative clusters, solving the problem of unstable load in the power grid, and achieving the minimum variance of load and win-win situation of user interests.

CN114881387BActive Publication Date: 2025-06-17XI AN JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111547330.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-16
Publication Date
2025-06-17
Estimated Expiration
2041-12-16

AI Technical Summary

Technical Problem

The existing time-sharing electricity price strategy is difficult to effectively reduce the valley peak load difference of the power grid, resulting in unstable load load, and the minimum load variance and peak cutting and valley filling effect cannot be achieved.

Method used

By obtaining user travel rules parameters, building reward and punishment functions, and using deep reinforcement learning to train the agent, the orderly scheduling of vehicle-network collaborative clusters is realized and the grid load is optimized.

Benefits of technology

It minimizes the variance of the grid load, stabilizes the grid load, meets users' travel needs, and reduces users' charging costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114881387B_ABST
    Figure CN114881387B_ABST
Patent Text Reader

Abstract

The present invention discloses a vehicle-grid collaborative cluster orderly scheduling method and system, including: obtaining user travel pattern parameters and using the user travel pattern parameters as training samples; constructing a reward and punishment function; based on the reward and punishment function, training an agent with the training samples, and using the trained agent to perform vehicle-grid collaborative cluster orderly scheduling. This method and system minimize the variance of the power grid load and stabilize the power grid load at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a scheduling method and system, and more particularly to a vehicle-grid collaborative cluster orderly scheduling method and system. Background Art

[0002] Vehicle-grid collaboration (V2G) can effectively promote the efficient use of energy and reduce energy loss. At the same time, the current energy network and power generation facilities cannot support the unordered charging of a large number of electric vehicles. Therefore, an effective electric vehicle charging and discharging scheduling strategy is of positive significance for improving the safety and stability of the power system.

[0003] Regarding the orderly charging and discharging strategy for electric vehicles accessing the distribution network, a large number of studies have been carried out in the industry. Implementing the time-of-use electricity price policy is an effective method for controlling the charging load of electric vehicles. Its main idea is that the grid side formulates different electricity prices at different times to guide users to transfer the charging time to the valley period of the grid load, so as to achieve the effect of peak shaving and valley filling.

[0004] With the in-depth study of vehicle to grid (V2G), in view of the formulation problem of time-of-use electricity price during charging and discharging, Yang H et al. constructed an optimized time-of-use electricity price model with penalties through demand response modeling in the paper "Reliability Evaluation of Power System Considering Time-of-Use Electricity Pricing[J].IEEE Transactions on Power Systems,2019,34(3):1991-2002." Dubey A et al. determined a time-of-use electricity price plan that can maximize the benefits of the grid and customers when the electric vehicle load is charged under the worst conditions in the paper "Determining time-of-use schedules for electric vehicle loads:A practical perspective[J].IEEE Power and Energy Technology Systems Journal,2015,2(1):12-20."

[0005] The time-of-use electricity price policy encourages electric vehicle users to charge during the valley period to balance the grid load. However, relying solely on the time-of-use electricity price strategy may not be sufficient to effectively reduce the valley-peak load difference of the grid. In some cases, another load peak may be generated in the local distribution network, so it is impossible to minimize the grid load variance, achieve peak shaving and valley filling, and stabilize the grid load. Summary of the Invention

[0006] The object of the present invention is to overcome the disadvantages of the above-mentioned prior art, and provides a vehicle-grid collaborative cluster orderly scheduling method and system, which minimize the variance of the power grid load and smooth the power grid load at the same time.

[0007] To achieve the above object, the vehicle-grid collaborative cluster orderly scheduling method described in the present invention includes:

[0008] Obtain the user travel pattern parameters and use them as training samples;

[0009] Construct a reward and punishment function;

[0010] Based on the reward and punishment function, train the agent with the training samples, and use the trained agent for vehicle-grid collaborative cluster orderly scheduling.

[0011] The user travel pattern parameters include the user's return time, travel time, state of charge of the electric vehicle battery at the return moment, power grid load including 24 scheduling time periods, SOC sum of the vehicle group, and variance of the power grid load.

[0012] Simulate the user travel pattern parameters according to Monte Carlo.

[0013] The reward and punishment function is:

[0014]

[0015] where r is the reward of the agent, α is the reward coefficient of the power grid load variance, β is the penalty coefficient of the power grid peak-valley difference, χ is the constraint coefficient of the SOC sum of the vehicle group, P max 、P min are respectively the maximum and minimum values of the power grid load after the V2G vehicle group scheduling, Q(t) bounary is the boundary value of the SOC sum of the V2G vehicle group at time t, Q(t) actual is the actual value of the SOC sum of the V2G vehicle group at time t.

[0016] The power grid load variance σ 2 is:

[0017]

[0018]

[0019]

[0020] where N is the number of time periods, is the residential electricity load except for the electric vehicle load in the nth time period, is the charging and discharging power of the i-th electric vehicle in the n-th time period, p Avg is the average value of the total grid load in a day, a n is the sum of the charging and discharging powers of all electric vehicles in the n-th time period.

[0021] Construct a reward and punishment function according to equations (2) to (4).

[0022] The actual SOC of the vehicle group in the n-th time period and is:

[0023]

[0024] Among them, is the sum of the actual SOCs of the vehicle group at the (n - 1)-th moment, is the sum of the SOCs of the vehicle group connected to the grid in the n-th time period, is the charging and discharging power of the i-th vehicle in the n-th time period, I is the number of electric vehicles, and C is the battery capacity of the electric vehicle, is the SOC of the i-th electric vehicle when it is connected to the grid in the n-th time period.

[0025] The vehicle-grid collaborative cluster orderly scheduling system described in the present invention includes:

[0026] An acquisition module for acquiring user travel pattern parameters and using the user travel pattern parameters as training samples;

[0027] A construction module for constructing a reward and punishment function;

[0028] A control module for training an agent based on the reward and punishment function through the training samples, and using the trained agent for vehicle-grid collaborative cluster orderly scheduling.

[0029] The present invention has the following beneficial effects:

[0030] When the vehicle-grid collaborative cluster orderly scheduling method and system described in the present invention are specifically operated, user travel pattern parameters are acquired and used as training samples; a reward and punishment function is constructed; an agent is trained based on the reward and punishment function through the training samples, and the trained agent is used for vehicle-grid collaborative cluster orderly scheduling. It should be noted that the present invention is based on deep reinforcement learning for scheduling, minimizing the variance of the grid load, which can not only meet the travel needs of users but also stabilize the grid load, enabling mutual benefit between the grid and users. Description of the Drawings

[0031] Figure 1 is a schematic diagram of the charging and discharging boundary model of the electric vehicle cluster;

[0032] Figure 2Schematic diagram of the charging and discharging boundary model for an electric vehicle cluster;

[0033] Figure 3 Schematic diagram of the interaction between the reinforcement learning agent and the environment;

[0034] Figure 4 Daily base load diagram;

[0035] Figure 5 Time-of-use electricity price diagram;

[0036] Figure 6 Reward curve diagram during the training process;

[0037] Figure 7 Grid load variance convergence curve diagram during the training process;

[0038] Figure 8 Charging cost convergence curve diagram during the training process;

[0039] Figure 9 Schematic diagram of the sum of the SOC of the vehicle group during the scheduling time;

[0040] Figure 10 Schematic diagram of the charging and discharging power during the scheduling time;

[0041] Figure 11 Schematic diagram of the grid load variance during the scheduling time at the 30,000th training episode;

[0042] Figure 12 Grid load curve diagram after scheduling at the 30,000th training episode;

[0043] Figure 13 Grid load variance convergence curve diagram during the training process under the condition of new energy access;

[0044] Figure 14 Charging cost convergence curve diagram during the training process under the condition of new energy access;

[0045] Figure 15 Reward curve diagram during the training process under the condition of new energy access;

[0046] Figure 16 Schematic diagram of the sum of the SOC of the vehicle group during the scheduling time at the 30,000th training episode under the condition of new energy access;

[0047] Figure 17 Grid load curve diagram after scheduling at the 30,000th training episode under the condition of new energy access;

[0048] Figure 18 Schematic diagram of the grid load variance during the scheduling time at the 30,000th training episode under the condition of new energy access;

[0049] Figure 19 Schematic diagram of charging and discharging power within the scheduling time at the 30,000th training round when new energy is connected. Detailed implementation manners

[0050] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, not all of the embodiments, and are not intended to limit the scope of the present invention disclosure. In addition, in the following description, the description of known structures and technologies is omitted to avoid unnecessarily confusing the concepts disclosed in the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.

[0051] The structural schematic diagrams according to the disclosed embodiments of the present invention are shown in the accompanying drawings. These figures are not drawn to scale, and for the purpose of clear expression, some details are enlarged, and some details may be omitted. The shapes of various regions and layers shown in the figures, as well as their relative sizes and positional relationships, are only exemplary. In practice, there may be deviations due to manufacturing tolerances or technical limitations, and those skilled in the art can design regions / layers with different shapes, sizes, and relative positions according to actual needs.

[0052] Embodiment 1

[0053] The vehicle-grid collaborative cluster orderly scheduling method of the present invention includes the following steps:

[0054] 1) According to the Monte Carlo simulation of user travel patterns, obtain the user's return time, travel time, state of charge of the electric vehicle battery at the return moment, grid load including 24 scheduling time periods, SOC sum of the vehicle group, and load variance of the grid to obtain training samples, as shown in Table 1;

[0055] Table 1

[0056]

[0057] Among them, the SOC sum of the vehicle group is the algebraic sum of the SOCs of electric vehicles within a certain scheduling time period, and the actual SOC of the vehicle group at the nth time period and is:

[0058]

[0059] Among them, is the actual SOC sum of the vehicle group at the (n - 1)th moment, is the sum of the SOCs of the vehicle cluster connected to the power grid in the nth time period. is the charging and discharging power of the ith vehicle in the nth time period, I is the number of electric vehicles, and C is the battery capacity of the electric vehicle. is the SOC of the ith electric vehicle when it is connected to the power grid in the nth time period.

[0060] According to the user travel pattern and the electric vehicle charging and discharging boundary model, an electric vehicle cluster charging and discharging boundary model is established based on the constraint conditions to constrain the sum of the SOCs of the electric vehicle cluster. Specifically, as Figure 1 and Figure 2 shown.

[0061] Figure 1 and Figure 2 are the boundary values of the sum of the SOCs of the electric vehicle cluster within 24 hours a day and the scheduling time respectively. The user starts to connect to the power grid for charging at 15:00 (corresponding to Figure 2 time period 1), and all users finish charging and leave at 10:00 (corresponding to Figure 2 time period 20). The SOC of each vehicle at the leaving moment should be between 0.8 and 0.9, that is, in time period 20, the sum of the SOCs of the vehicle cluster is between [407.2, 458.1]. To extend the cycle life of the battery, the SOC of each vehicle is greater than 0.2, that is, the lower bound of the sum of the SOCs of the vehicle cluster is 101.8. The sum of the SOCs of all vehicles constitutes the electric vehicle cluster charging and discharging boundary model.

[0062] Based on this model, with the minimum variance of the grid load as the optimization goal, the proximal policy optimization (PPO) algorithm in deep reinforcement learning is used for solution.

[0063] In the context of deep reinforcement learning, a computerized agent learns to take actions at discrete time steps t to maximize the numerical reward obtained from the environment. If, in a given state, an action chosen by the agent gives a lower reward, then in subsequent actions in that state, the agent can choose an action that produces a higher reward. In this framework, there is a learning and decision-making entity, namely the agent, and an environment, which is everything outside the agent, such as Figure 3 shown.

[0064] Normally, the actions taken not only affect the immediate reward signal but also all subsequent reward signals. Initially, the agent receives the environmental state s t , and based on this, it performs an action that causes the environmental state to change to s t+1 . In addition to the observation, the environment also provides the agent with feedback related to the behavior in the form of a reward signal r t+1 . Moving the state from s t to st+1 The transition function usually contains a random component and cannot be determined solely by the action of a t The set of rules that an agent follows when mapping states to a probability distribution of actions is called a policy (π). The goal of reinforcement learning is to provide an agent with a policy that maximizes its total future reward.

[0065] The variance of the grid load is an indicator reflecting the stability of the grid load. The smaller its value, the more stable the grid load. Its specific model is:

[0066]

[0067]

[0068]

[0069] where σ 2 is the variance of the grid load. The smaller its value, the more stable the grid operation. N is the number of time periods, is the residential electricity load except for the electric vehicle load in the nth time period, is the charging and discharging power of the ith electric vehicle in the nth time period, with a range of [-6, 6], and p Avg is the average value of the total grid load in a day, and a n is the action of the agent in the nth time period, that is, the sum of the charging and discharging powers of all electric vehicles in the nth time period.

[0070] To avoid overly sparse rewards during the agent's learning process, rewards and penalties are added after each action. Through equations (2) to (4), the reward and penalty function is obtained as:

[0071]

[0072] where r is the reward of the agent, α is the reward coefficient of the grid load variance, β is the penalty coefficient of the grid peak-valley difference, χ is the constraint coefficient of the sum of the vehicle group SOCs, P max and P min are the maximum and minimum values of the grid load after the V2G vehicle group scheduling respectively, Q(t) bounary is the boundary value of the sum of the SOCs of the V2G vehicle group at time t, and Q(t) actual is the actual value of the sum of the SOCs of the V2G vehicle group at time t.

[0073] 2) Based on the reward and penalty function, the agent is trained using training samples to obtain the trained agent, and then the trained agent is used for coordinated and orderly scheduling of the vehicle-grid cluster.

[0074] Example Two

[0075] The vehicle-grid collaborative cluster orderly scheduling system described in the present invention includes:

[0076] An acquisition module, configured to acquire user travel pattern parameters and use the user travel pattern parameters as training samples;

[0077] A construction module, configured to construct a reward and punishment function;

[0078] A control module, configured to train an agent based on the reward and punishment function using the training samples, and perform vehicle-grid collaborative cluster orderly scheduling using the trained agent.

[0079] Verification test

[0080] To verify the correctness of the present invention, the battery capacity of each electric vehicle is taken as 24 kW·h, the maximum charge and discharge power is 6 kW, the duration of the scheduling time period is 1 h, and the number of electric vehicles is 509 obtained according to the peak load proportionally. The load curve is as Figure 4 shown.

[0081] The time-of-use electricity price is 1.2 yuan / (kW·h) during peak hours, 0.8 yuan / (kW·h) during normal hours, and 0.4 yuan / (kW·h) during valley hours. The 24 hours of a day are divided into 24 time periods. 0:00 - 1:00 is charging period 1, and so on. The electricity price distribution of each charging period is as Figure 5 shown.

[0082] From Figure 4 Figure 5 it can be seen that in the case of unordered charging, the variance of the grid load is 751941.4, and the charging cost of the vehicle group is 9074.4 yuan; when wind power and photovoltaic are connected to the grid and the vehicle group charges disorderly, the variance after the new energy is connected is 1798859.1, and the charging cost of the vehicle group is 9074.4 yuan.

[0083] The hyperparameters of PPO are set as shown in Table 2. It should be noted that the state dimension is 26, including the grid load of 24 scheduling time periods, the SOC of the vehicle group, and the variance of the grid load, a total of 26 parameters.

[0084] Table 2

[0085]

[0086] The optimization result of the vehicle group scheduling without wind and light access is as Figure 6 shown.

[0087] From the reward curve, it can be seen that the present invention has good convergence and finally converges near 2100. The large fluctuations at the beginning are due to the overstepping behavior during the exploration process and the large penalty term, which is a normal phenomenon.

[0088] Figure 7 is the convergence curve of the grid load variance. Refer to Figure 7 . The grid load variance decreases continuously with training and finally stabilizes at around 9500. Compared with the unordered charging load variance of 751941.4, it has decreased by 98.7%.

[0089] Refer to Figure 8 . The user's charging cost decreases continuously with training and finally stabilizes at around 2100 yuan. Compared with the unordered charging cost of 9074.4 yuan, it has decreased by 76.9%.

[0090] Refer to Figure 9 and Figure 10 . Taking the 30000th training round as an example for illustration, the vehicle group starts to connect to the grid for scheduling at 15:00, which corresponds to the first period of the scheduling time. The vehicle group leaves completely at 10:00 in the morning, corresponding to the 20th period of the scheduling time. Charging and discharging are carried out according to the peak-valley state of the grid, discharging during peak hours and charging during valley hours. The SOC sum of the vehicle group within the scheduling time meets the constraints, and the final SOC sum is 422.571991, meeting the termination condition of the scheduling. And within the scheduling time, the grid load variance decreases from 90269.23223 to 11012.51375. After the scheduling ends, the grid load curve is as Figure 12 shown. It can be seen from Figure 12 that the effect of the scheduling strategy in shaving peaks and filling valleys is relatively obvious. The load during peak hours decreases significantly, and the load during valley hours increases significantly.

[0091] The optimization result after the integration of wind and solar power is as Figures 13 to 19 shown. It can be seen from the grid load variance convergence curve that the grid load variance decreases continuously with training and finally stabilizes at around 55000. Compared with the unordered charging load of 1798859.1, it has decreased by 96.4%.

[0092] It can be seen from the charging cost convergence curve that the user's charging cost decreases continuously with training and finally stabilizes at around 1600 yuan. Compared with the unordered charging cost of 9074.4 yuan, it has decreased by 82.4%.

[0093] It can be seen from the reward curve that the present invention has good convergence and finally converges near 2100. At the beginning, the algorithm fluctuates greatly because there are out-of-bounds behaviors during the exploration process and the penalty term is relatively large, which is a normal phenomenon.

[0094] Taking the 30,000th training round as an example for illustration, the vehicle group starts connecting to the power grid for scheduling at 15:00. This corresponds to the first period of the scheduling time. The vehicle group all leaves at 10:00 am, corresponding to the 20th period of the scheduling time. Charging and discharging are carried out according to the peak-valley state of the power grid, discharging during peak hours and charging during valley hours. The SOC sum of the vehicle group within the scheduling time meets the constraints, and the final SOC sum is 440.9595317, meeting the termination condition of the scheduling. Moreover, within the scheduling time, the variance of the power grid load decreases from 90269.23223 to 68341.39581. After the scheduling ends, the power grid load curve is as Figure 17 shown, and it can be seen from Figure 17 that the effect of the scheduling strategy on peak shaving and valley filling is relatively obvious. The load during peak hours has a significant decrease, and the load during valley hours significantly increases.

Claims

1. A vehicle-network collaborative cluster orderly scheduling method, characterized in that, Including: Obtain user travel pattern parameters and use them as training samples; Construct a reward and punishment function; Based on the reward and punishment function, train the agent with the training samples, and use the trained agent for orderly scheduling of vehicle-grid collaborative clusters; The reward and punishment function is: Among them, r is the reward of the agent, α is the reward coefficient of the grid load variance, β is the penalty coefficient of the grid peak-valley difference, χ is the constraint coefficient of the sum of the SOCs of the vehicle fleet, P max and P min are respectively the maximum and minimum values of the grid load after the V2G vehicle fleet is dispatched, and Q(t) bounary is the boundary value of the sum of the SOCs of the V2G vehicle fleet at time t, and Q(t) actual is the actual value of the sum of the SOCs of the V2G vehicle fleet at time t; Grid load variance σ 2 is as follows: where N is the number of time periods, is the residential electricity load except for the electric vehicle load in the nth time period, is the charging and discharging power of the ith electric vehicle in the nth time period, p Avg is the average value of the total grid load in a day, a n is the sum of the charging and discharging powers of all electric vehicles in the nth time period; Construct the reward and punishment function according to Equations (2) to (4).

2. The vehicle-network collaborative cluster orderly scheduling method according to claim 1, characterized in that, The user travel pattern parameters include the user's return time, travel time, state of charge of the electric vehicle battery at the return moment, grid load with 24 scheduling time periods, SOC of the vehicle group, and load variance of the grid.

3. The vehicle-network collaborative cluster orderly scheduling method according to claim 1, characterized in that, Simulate user travel pattern parameters based on Monte Carlo.

4. The vehicle-network collaborative cluster orderly scheduling method according to claim 1, characterized in that, The actual SOC of the vehicle group in the nth time period and is as follows: Among them, is the actual sum of SOC of the vehicle group at the (n - 1)-th moment, is the sum of SOC of the vehicle group connected to the power grid in the n-th time period, is the charging and discharging power of the i-th vehicle in the n-th time period, I is the number of electric vehicles, and C is the battery capacity of the electric vehicle, is the SOC of the i-th electric vehicle when it is connected to the power grid in the n-th time period.

5. A vehicle-network collaborative cluster orderly scheduling system, characterized in that, Including: An acquisition module for obtaining user travel pattern parameters and using them as training samples; A construction module for constructing a reward and punishment function; A control module for training the agent with the training samples based on the reward and punishment function and using the trained agent for orderly scheduling of vehicle-grid collaborative clusters; The reward and punishment function is: Among them, r is the reward of the agent, α is the reward coefficient of the grid load variance, β is the penalty coefficient of the grid peak-valley difference, χ is the constraint coefficient of the sum of the vehicle fleet's SOC, P max and P min are respectively the maximum and minimum values of the grid load after the V2G vehicle fleet is dispatched, Q(t) bounary is the boundary value of the sum of the SOC of the V2G vehicle fleet at time t, Q(t) actual is the actual value of the sum of the SOC of the V2G vehicle fleet at time t; Grid load variance σ 2 is as follows: where N is the number of time periods, is the residential electricity load except for the electric vehicle load in the nth time period, is the charging and discharging power of the ith electric vehicle in the nth time period, p Avg is the average value of the total grid load in a day, a n is the sum of the charging and discharging powers of all electric vehicles in the nth time period; Construct the reward and punishment function according to Equations (2) to (4).

Citation Information

Patent Citations

  • Electric vehicle charging load prediction method based on Monte Carlo and deep learning

    CN110570014A

  • Cluster electric vehicle charging behavior optimization method based on deep reinforcement learning

    CN111934335A