Power Data Scheduling Method, System and Device Based on Deep Reinforcement Learning
Through the power data scheduling method based on deep reinforcement learning, the problem of balance between user needs and charging station benefits in V2G technology is solved, and the intelligent and efficient charging and discharging decisions of charging stations are realized, which improves the economic benefits and user satisfaction of charging stations.
Patent Information
- Application Number
- CN202411700879.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-26
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2044-11-26
AI Technical Summary
It is difficult for existing V2G technology to effectively balance user needs and the economic benefits of charging stations in charge and discharge scheduling solutions.
Using the power data scheduling method based on deep reinforcement learning, a mathematical model of the charging and discharging scheduling optimization problem is established by establishing a mathematical model of the charging and discharging scheduling optimization problem of charging stations, and using the DQN algorithm to build a scheduling optimization network, define actions, states and rewards, and train the scheduling optimization network to adaptively generate the best charging and discharging scheme.
It achieves the maximization of the economic benefits and user service satisfaction of the charging station while meeting user needs, dynamically adapts to the operating status of the charging station, and improves the profit and user experience of the charging station.
Smart Images

Figure CN119623981B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of vehicle charging, and particularly relates to a power data scheduling method, system, and device based on deep reinforcement learning. Background Art
[0002] An electric vehicle (EV) is a means of transportation powered by electricity, and has received increasing attention due to its high efficiency, zero pollution, and low emissions. With the continuous growth of the market share of electric vehicles, the charging demand of vehicles is also increasing. The large-scale and disorderly access of electric vehicles will have adverse effects on the power grid, such as an increase in the peak-valley difference, a decline in power quality, and an increase in the difficulty of optimizing and controlling the operation of the power grid.
[0003] In response to the burden brought by the charging demand of electric vehicles to the power grid, technicians have proposed the vehicle-to-grid (V2G) technology. V2G technology is a new type of power grid technology that can achieve two-way, real-time, controllable, and high-speed energy flow between vehicles and the power grid. In this technology, the EV serves as a kind of energy storage device, which can not only obtain electric energy from the power grid but also transmit electric energy to the power grid, having the dual nature of a power source and a load, and providing an energy storage backup function for the power grid.
[0004] Applying V2G technology to control the charging and discharging process of electric vehicles can effectively alleviate the negative impact of the large-scale access of electric vehicles on the stability of the power grid operation; it plays a role in shaving peaks and filling valleys and suppressing renewable energy. Therefore, it is very necessary and urgent to develop V2G technology and implement V2G technology in the charging and discharging scheduling strategy of EVs. However, when applying V2G technology to generate the charging and discharging scheduling scheme of charging stations, how to balance user needs and the economic benefits of charging stations is still a complex problem. Summary of the Invention
[0005] In order to solve the problem that the existing charging and discharging scheduling schemes using V2G technology are difficult to effectively balance customer needs and the benefits of charging stations, the present invention provides a power data scheduling method, system, and device based on deep reinforcement learning.
[0006] The present invention is implemented by the following technical solutions:
[0007] A power data scheduling method based on deep reinforcement learning, which includes:
[0008] S1: Establish a mathematical model representing the charging and discharging scheduling optimization problem of a charging station.
[0009] S2: Model the mathematical problem as a Markov decision process, and construct a scheduling optimization network for solving the mathematical model based on the DQN algorithm model. In the scheduling optimization network, define the actions, states, and rewards as follows:
[0010] (1) The action \(a\) of the control center of the charging station at time \(t\) t includes the charging price, the discharging price, as well as the charging power and discharging power of each charging pile.
[0011] (2) The state \(s\) of the control center of the charging station at time \(t\) t includes the global state information of the charging station and the user state information of the charging piles Both satisfy:
[0012] Among them, and respectively represent the charging power and discharging power of the \(m\)-th charging pile at time \(t\); and respectively represent the charging price and discharging price of the charging station at time \(t\); \(m\in M\), where \(M\) represents the set of charging piles in the charging station; represents the expected SOC information of the \(i\)-th user; represents the expected departure time of the \(i\)-th user; represents the energy demand of the \(i\)-th user, \(i\in N\), where \(N\) represents the set of users.
[0013] (3) The reward \(r\) of the control center of the charging station at time \(t\) t satisfies the following formula:
[0014]
[0015] In the above formula, represents the economic benefit of the charging station; represents the service satisfaction of the \(i\)-th user; \(\alpha\) and \(\beta\) are respectively and the weights of; represents the power purchased by the charging station from the distribution network at time \(t\); represents the real-time electricity price in the power market; represents the arrival time of the \(i\)-th user at the station; represents the shortest charging time of the vehicle of the \(i\)-th user.
[0016] S3: Train the scheduling optimization network according to the operation data of the charging station, so that it can adaptively obtain a better control strategy in continuous decision-making and learning to maximize the economic benefit of the charging station and the service satisfaction of users.
[0017] S4: When any vehicle arrives at the charging station and requests charging, the control center of the charging station obtains the demand information of the current vehicle, including: arrival time, expected departure time, state of charge (SoC) when the vehicle arrives, and expected SoC when leaving. Then, it uses the trained scheduling optimization network to generate the best charging and discharging plan that meets the current user's demand.
[0018] As a further improvement of the present invention, in the mathematical model established in step S1 for characterizing the charging and discharging scheduling optimization problem of the charging station, the optimization objective is to maximize the overall revenue F of the charging station; the decision variables are the charging and discharging power corresponding to each charging and discharging behavior, and the charging and discharging electricity price including service fees; the overall constraints are the battery capacity of the vehicle itself, the charging and discharging power constraints of the vehicle and the charging pile, and the service electricity price constraint.
[0019] As a further improvement of the present invention, in the mathematical model, the optimization objective is:
[0020]
[0021]
[0022] In the above formula, T represents the operation duration of the charging station; represents the basic load of the interaction between the charging station and the power grid.
[0023] As a further improvement of the present invention, in the mathematical model, the expression of the battery capacity constraint for each vehicle is as follows:
[0024]
[0025] In the above formula, Cap represents the battery capacity of the current vehicle; SoC t represents the real-time state of charge of the current vehicle; η chg and η dhg respectively represent the charging energy efficiency and discharging energy efficiency between the vehicle and the charging pile; SoC min and SoC max respectively represent the lower limit and upper limit of the vehicle's state of charge; SoC d represents the expected state of charge of the current vehicle.
[0026] As a further improvement of the present invention, in the mathematical model, the expression of the charging and discharging power constraints of the vehicle and the charging pile is as follows:
[0027]
[0028] In the above formula, and respectively represent the maximum charging power and maximum discharging power allowed by the charging pile connected to the current vehicle; Represents the maximum instantaneous power that the charging station can withstand.
[0029] As a further improvement of the present invention, in the mathematical model, the expression of the service electricity price constraint is as follows:
[0030]
[0031] In the above formula, and are the upper and lower limits of the charging electricity price; E t sc is the slack variable of the charging electricity price; and are the upper and lower limits of the discharging electricity price; is the slack variable of the discharging electricity price.
[0032] As a further improvement of the present invention, in step S2, the shortest charging time T of any vehicle need is used to ensure that the state of charge of the vehicle can meet the expectation before the expected departure time, and its calculation formula is as follows:
[0033]
[0034] The energy demand P of any vehicle req has the following calculation formula:
[0035] P req =(SoC d -SoC a )·Cap;
[0036] In the above formula, SoC a represents the state of charge of the vehicle when it arrives at the station.
[0037] As a further improvement of the present invention, in step S3, the training process of the scheduling optimization network is as follows:
[0038] S31: Initialize the training parameters of the DQN algorithm, including the learning rate, reward discount factor, exploration probability, training step size, and number of iterations; define the experience replay pool and the size of the mini-batch data; initialize the system environment and obtain the initial state of the charging station.
[0039] S32: According to the output of the Online neural network and the greedy strategy, perform charge and discharge behavior control during the operation of the charging pile, and then generate the charge and discharge power and electricity price at the current moment, and obtain the corresponding action a t .
[0040] S33: Determine the state s t of the environment at the next moment according to the determined action a t+1 , and calculate the reward r t of the action.
[0041] S34: Store the correlation parameters of the iterative process as empirical data in the experience replay pool. The empirical data includes: the current state s t+1 , the action a t , the reward r t and the state s at the next moment t+1 .
[0042] S35: Randomly select a small batch of data from the experience replay pool and transfer it to the Online neural network and the Target neural network;
[0043] S36: The Target network takes the small batch of data as input and outputs the target Q value. The Online network takes the small batch of data as input and outputs the estimated Q value; Calculate the error loss based on the target Q value and the estimated Q value;
[0044] S37: Iteratively execute steps S32 - S36, and perform backpropagation of the gradient according to the error loss to update the model parameters of the Online neural network and the Target neural network until the preset number of iterations is reached.
[0045] The present invention further includes a power data scheduling system based on deep reinforcement learning, which includes a demand acquisition unit, a decision-making unit, and an execution unit.
[0046] Among them, the demand acquisition unit is used to obtain the demand information of any arriving vehicle, including: the arrival time, the expected departure time, the state of charge when the vehicle arrives, and the expected state of charge when leaving.
[0047] There is a scheduling optimization network running in the decision-making unit, which is trained by using the power data scheduling method based on deep reinforcement learning as described above. The scheduling optimization network is used to generate a charge-discharge plan that can meet the user's needs and maximize the economic benefits of the charging station and the service satisfaction of the user according to the operating states of each charging pile in the charging station and the user's demand information.
[0048] The execution unit is used to allocate an idle charging pile to the user, and after the user's vehicle is connected to the designated charging pile, provide services to the vehicle according to the charge-discharge plan generated by the decision-making unit.
[0049] The present invention further includes a charging station scheduling device, which includes a memory, a processor, and a computer program stored on the memory and running on the processor. When the processor executes the computer program, it creates a power data scheduling system based on deep reinforcement learning as described above, and then realizes interaction with the user and provides vehicle charge-discharge services that meet the needs of customers.
[0050] The technical solution provided by the present invention has the following beneficial effects:
[0051] The present invention models the charging station scheduling problem in a V2G charging environment as a single-objective optimization model, and then uses the DQN algorithm to design and train a scheduling optimization network. The scheduling optimization network automatically generates a vehicle charging and discharging plan that can meet customer needs and maximize the interests and environmental benefits of the charging station according to the differentiated needs of vehicles and the real-time operating status of the charging station.
[0052] The charging pile control center adopting the solution of the present invention uses a control strategy of adaptively adjusting the charging and discharging price and the charging and discharging power, and has the characteristics of autonomous learning, dynamic adaptation and fast convergence. Compared with the current mathematical optimization methods such as convex optimization, which have problems such as high computational complexity and poor convergence, the present invention has great advantages in terms of convergence speed and convergence performance. The proposed method models the charging station control center as an intelligent agent, constructs a linear weighted reward function related to the EV charging and discharging behavior, the charging and discharging power, the charging and discharging price, and the charging time, so that the intelligent agent can adaptively improve the strategy, realize the intelligence and high efficiency of the EV charging and discharging decision-making plan, thereby achieving the goal of maximizing the profit of the charging station and ensuring the performance of the system. Description of the Drawings
[0053] Figure 1 It is a system architecture diagram of a charging station control system adopting the V2G technology in Embodiment 1 of the present invention.
[0054] Figure 2 It is a step flow chart of a power data scheduling method based on deep reinforcement learning provided in Embodiment 1 of the present invention.
[0055] Figure 3 It is a schematic diagram of the training stage of the scheduling optimization network adopted in Embodiment 1 of the present invention.
[0056] Figure 4 It is a system topology diagram of a power data scheduling system based on deep reinforcement learning provided in Embodiment 2 of the present invention. Detailed Embodiment
[0057] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0058] Embodiment 1
[0059] This embodiment provides a V2G charging and discharging scheduling method based on Deep Reinforcement Learning (DRL). This solution is applied to a charging station as shown in Figure 1 shown, inFigure 1 As can be seen from the architecture, the control center of the charging station serves as the control center of the entire system. The control center interacts with the power grid upward and can achieve the downward consumption and upward grid connection of electric energy. Downward, it interacts with each charging pile and the connected vehicles, and then controls each charging pile to charge or discharge the connected vehicles. In the solution of this embodiment, each EV connected to the charging station defaults to accept the V2G charging and discharging protocol. Under this protocol, the charging station can flexibly control the charging and discharging behaviors of EVs. While accessing the national time-of-use (TOU) electricity market, a dynamic charging and discharging price pricing strategy considering service fees is introduced, which helps users obtain more favorable charging prices and more discharging benefits, so as to encourage EV users to participate in the V2G dispatching process. At the same time, the solution of this embodiment realizes the flexible regulation of the charging and discharging power of EVs by combining customer needs and the operation status of the charging station, and maximizes the economic benefits of the charging station on the premise of meeting the charge state requirements of customers when leaving the station. In addition, in addition to maximizing the interests of users and the charging station, the finally generated charging and discharging scheduling plan of this embodiment can also fully release the energy storage backup function of EVs, contribute to peak shaving and valley filling of the power grid and stabilizing renewable energy, and thus achieve greater social benefits.
[0060] Specifically, as Figure 2 shown, the power data scheduling method based on deep reinforcement learning provided in this embodiment includes the following steps:
[0061] S1: Establish a mathematical model representing the charging and discharging scheduling optimization problem of the charging station.
[0062] The technical problem to be solved by the technical solution provided in this embodiment is to scientifically schedule the charging and discharging capabilities of the charging station to meet the charging needs of vehicles at the charging station. Under the V2G protocol, this scheduling work involves the adaptive adjustment of the charging and discharging power and charging and discharging prices at different times in the charging behavior of users. From the perspective of the charging station management, the essence of solving this problem is to maximize the benefits of the charging station on the premise of ensuring the interests of customers (lower charging costs, charging quantity and time meet the requirements). Therefore, this embodiment regards this problem as a single-objective optimization (maximizing the benefits of the charging station) problem and conducts mathematical modeling.
[0063] The mathematical model of a typical single-objective optimization problem consists of three elements: decision variables, optimization objectives, and constraint conditions. In the mathematical model constructed in this embodiment to represent the charging and discharging scheduling optimization problem of a charging station, the decision variables are the charging and discharging powers corresponding to each charging and discharging behavior, and the charging and discharging electricity prices including service fees are the decision variables. This part of the decision variables can be flexibly adjusted by the charging station according to its operating conditions. Among them, the charging and discharging powers and charging and discharging electricity prices corresponding to the charging and discharging behaviors can be adjusted through the designated charging piles used by users. That is, the charging station can dynamically adjust the charging and discharging states and their powers of the charging piles used by each user according to the actual needs of each user and the overall operating state of the charging station, and provide different electricity prices for different users after the guiding price range when the users settle their accounts.
[0064] I. Optimization Objectives
[0065] In the mathematical model established in this embodiment to represent the charging and discharging scheduling optimization problem of a charging station, the optimization objective is to maximize the overall revenue F of the charging station. The overall revenue F of the charging station is obtained by summing up three parts: the fees (income) charged for providing charging services to each user by all the charging piles in the station during the operating period, the incentive revenue (expenditure) paid for providing discharging services to users by the charging piles, and the electricity fees (expenditure) paid for obtaining electric energy from the power grid. Therefore, the expression of the optimization objective is as follows:
[0066]
[0067] In the above formula, represents the charging power of the m-th charging pile at time t; represents the charging price of the charging station at time t; represents the discharging power of the m-th charging pile at time t; represents the discharging price of the charging station at time t. Among them, m ∈ M, and M represents the set of charging piles in the charging station; t ∈ T, and T represents the operating duration of the charging station; represents the real-time electricity price (TOU) of the power market; represents the power purchase from the distribution network by the charging station at time t.
[0068] In the charging station based on V2G technology in this embodiment, the power purchase from the distribution network by the charging station is related to the real-time charging and discharging powers of each charging pile in the charging station. The charging station should ensure that the charging and discharging powers of the entire grid are balanced with the energy supply of the grid. Therefore, the power purchase from the distribution network by the charging station at time t
[0069]
[0070] In the above formula, Represents the basic load of the charging station interacting with the power grid Represents the sum of the charging powers of all charging piles in the station at time t Represents the sum of the discharging powers of all charging devices in the station at time t
[0071] II. Constraints
[0072] In the mathematical model established in this embodiment to represent the charging and discharging scheduling optimization problem of the charging station, the constraints are divided into three categories, namely: (1) The capacity constraint that the vehicle's own battery needs to meet when enjoying charging and discharging services; (2) The charging and discharging power constraints that the charging pile and the entire charging station need to meet when providing charging and discharging services to the vehicle; (3) The service electricity price constraint that the time-of-use electricity price adopted by the system needs to meet when the charging station settles the service fees with the user. The specific contents of the three categories of constraints are as follows:
[0073] 2.1. Battery capacity constraint
[0074] When in the driving state, the vehicle (EV) needs to consume electric energy, so the state of charge SoC of the electric vehicle gradually decreases. After entering the charging station, the battery stops driving, and the state of charge of the vehicle is related to the charging or discharging behavior of the vehicle. In the charging mode, the state of charge of the vehicle gradually increases with the charging time, and in the discharging mode, the state of charge of the vehicle gradually decreases with the discharging time. Therefore, considering that the V2G type charging station can provide both charging services and discharging services, the SoC update expression of the EV is as follows:
[0075]
[0076] Among them, Cap represents the battery capacity of the current vehicle; SoC t-1 represents the state of charge of the vehicle before entering the station; SoC t represents the state of charge of the vehicle when leaving the station; Δt represents the residence time of the vehicle in the charging station. η chg represents the charging energy efficiency between the vehicle and the charging pile; η dhg represents the discharging energy efficiency between the vehicle and the charging pile.
[0077] In addition, in order to avoid overcharging and overdischarging of the EV's battery pack, the state of charge of the EV also needs to meet the following constraints:
[0078]
[0079] Among them, SoC min represents the lower limit of the vehicle's state of charge; SoC max represents the upper limit of the vehicle's state of charge.
[0080] This constraint means that when the charging pile provides charging service for the vehicle, the charging service should stop when the real-time state of charge SoC of the vehicle t reaches the upper limit of the state of charge of the vehicle SoC max When the charging pile provides discharging service for the vehicle, the discharging service should stop when the real-time state of charge SoC of the vehicle t reaches the lower limit of the state of charge of the vehicle SoC min at this time.
[0081] Finally, after each EV arrives at the charging station, it needs to provide demand information to the control center of the charging station, including the progress time t a of the vehicle, the expected departure time t d the state of charge SoC when the vehicle arrives a the expected state of charge SoC when leaving d . These information are used by the control center to schedule the charging and discharging process between the charging piles and EVs. According to the battery capacity Cap of the EV, the energy demand P req after the EV enters the station can be obtained, and the calculation formula is as follows:
[0082]
[0083] At this time, the control center of the charging station can add the EVs to the waiting service queue according to the order of their arrival, and then provide services to the vehicles in a timely manner according to different demands. For example, for a vehicle entering the station, the earlier the departure time t d (the more urgent the demand) and the more the energy demand P req (the higher the generated income), the higher the priority of the EV when obtaining the charging service.
[0084] Therefore, in order to ensure that the state of charge of each vehicle entering the station can reach the expectation when leaving the station, the constraints on the state of charge of the EV when arriving at and leaving the charging station can be expressed as: assuming that when t = t a , SoC t = SoC a ; then when t = t d , SoC t = SoC d . On this basis, in order to ensure that when t = t d , SoC t = SoC d , this embodiment sets a minimum charging time T need for the charging and discharging service of each vehicle, which satisfies the following formula:
[0085]
[0086] Based on the determined minimum charging time Tneed , the charging station can be used before the EV leaves d -T need At all times, the EV only needs to keep charging without participating in the V2G scheduling of the charging station, thereby ensuring that the vehicle's charge state always meets expectations before leaving.
[0087] 2.2 Charging power constraints
[0088] When a charging station provides charging and discharging services to vehicles, the charging piles and other equipment in the station have their own electrical characteristics. For example, for each charging pile and the vehicle connected to it, there is a maximum charging and discharging power limit when charging and discharging.
[0089] Therefore, the charging pile has the following maximum power constraints:
[0090]
[0091]
[0092] In the above formula, Indicates the maximum charging power allowed by the charging pile to which the vehicle is currently connected; Indicates the maximum discharge power allowed by the charging pile to which the vehicle is currently connected.
[0093] For each charging pile, it can only provide one of the charging and discharging services for the same connected vehicle, and the charging and discharging behaviors of the charging pile cannot exist at the same time. Therefore, this embodiment adds the following charging and discharging behavior constraints of the charging pile based on the characteristic that the charging power and the generated power are not positive at the same time:
[0094]
[0095] This constraint is used to force the product of the charging power and discharging power of each charging pile to be 0, indicating that the same charging pile cannot be charged and discharged at the same time.
[0096] Each charging station is equipped with multiple charging piles. In order to protect the charging facilities of the charging station and the battery life of the EV, the instantaneous power of the charging station should be less than the maximum value it can withstand. Therefore, for the entire charging station, it also needs to meet the following total requirements:
[0097]
[0098]
[0099] In the above formula, Indicates the maximum instantaneous power that the charging station can withstand.
[0100] 2.3 Service electricity price constraint
[0101] In the solution of this embodiment, the charging station motivates EVs to participate in V2G services by adjusting the charging and discharging electricity prices, so as to achieve the purpose of promoting the consumption of new energy and peak shaving and valley filling. The charging and discharging behaviors of EV users are scheduled by the charging station. Therefore, the charging station needs to charge users a part of service fees. Considering market factors and user affordability, this embodiment specifies differentiated charging and discharging electricity prices.
[0102] In actual operation, the charging station should not unilaterally increase the charging service electricity price or decrease the discharging service electricity price in order to pursue corporate profits. Therefore, this embodiment sets the following service electricity price constraints:
[0103] For the charging behavior of users, the expression of the corresponding service electricity price constraint is:
[0104]
[0105] In the above formula, and are the upper and lower limits of the charging electricity price; is the slack variable of the charging electricity price.
[0106] For the discharging behavior of users, the expression of the corresponding service electricity price constraint is:
[0107]
[0108]
[0109] In the above formula, and are the upper and lower limits of the discharging electricity price; is the slack variable of the discharging electricity price.
[0110] S2: Model the mathematical problem as a Markov decision process, and construct a scheduling optimization network for solving the mathematical model based on the DQN algorithm model.
[0111] Based on the mathematical model established above for characterizing the charging and discharging scheduling optimization problem of the charging station, in this embodiment, an intelligent method based on the deep Q-network (DQN) algorithm of reinforcement learning is proposed to find the optimal solution to this problem. Specifically, in this embodiment, the optimization problem is modeled as a Markov decision process. At the beginning of each time slot, the charging station control center determines the charging and discharging behavior decisions (including charging and discharging behavior control and power magnitude determination) and charging and discharging price setting for all charging piles of the charging station according to the observed state at the previous moment. The solution process of the optimization problem is a series of time series decisions, and the decision at the previous moment will affect the behavior of the agent at the next moment. Therefore, the optimal decision problem is transformed into the solution of the Markov decision process. Specifically, the control center is modeled as an agent, and its actions include charging and discharging behavior control (the timing of charging and discharging behavior), charging and discharging power determination (charging and discharging power at different moments), and charging and discharging price setting (time-of-use service price), a total of six parts. The agent optimizes the actions in different states by maximizing the long-term average return, and the reward consists of two parts: the revenue of the charging station and the satisfaction degree of EV users' services. In particular, the timing of charging and discharging behavior can actually be managed by dynamically adjusting the charging and discharging power. For example, when the timing of charging and discharging is not reached, it means that the charging power or power generation power at this time is 0.
[0112] In summary, in the scheduling optimization network of this embodiment, the actions, states, and rewards of the agent are defined as follows:
[0113] (1) Actions
[0114] The action \(a\) of the charging station control center at time \(t\) t includes the charging price, the discharging price, as well as the charging power and discharging power of each charging pile.
[0115] Specifically, the action \(a\) of the control center at time \(t\) t is composed of the charging and discharging price vector and the charging and discharging power vector of the charging pile, that is The two vectors respectively satisfy:
[0116]
[0117]
[0118] where represents the charging price vector; represents the discharging price vector; represents the charging power vector; represents the discharging power vector.
[0119] In addition, considering that the DQN algorithm is applicable to discrete action space tasks, the above actions all need to be discretized.
[0120] (2) Status
[0121] The status s of the control center of the charging station at time t t Includes the global status information of the charging station And the user status information of the charging piles Among them, the global status information of the charging station Is reflected in four aspects: the total charging power, the total discharging power of the charging piles, and the real-time charging price and discharging price. And the user status information of the charging piles Is reflected in three aspects: the expected state of charge of the vehicle, the expected departure time, and the energy demand. Specifically, the global status information And the user status information Respectively satisfy the following formulas:
[0122]
[0123]
[0124] Among them, And Respectively represent the charging power and discharging power of the m-th charging pile at time t; And Respectively represent the charging price and discharging price of the charging station at time t; m ∈ M, M represents the set of charging piles in the charging station; Represents the expected SOC information of the i-th user; Represents the expected departure time of the i-th user; Represents the energy demand of the i-th user, i ∈ N, N represents the set of users.
[0125] (3) Reward
[0126] In the solution of this embodiment, in order to make the optimized scheduling scheme provide high-quality services for users while enabling the charging piles to obtain higher benefits and taking into account the ecological value of the V2G system. Therefore, the reward is divided into the economic benefits of the charging station And the service satisfaction of users Two aspects. That is, the optimized scheduling scheme should not only provide the benefits of the charging station but also improve the user experience.
[0127] In particular, in order to quantify the abstract index of user satisfaction, this embodiment uses the waiting time of users in the charging station as an index to evaluate user satisfaction, that is, the earlier the user completes the service and leaves, the better the user experience, which also conforms to the actual feelings of most users.
[0128] In summary, in this embodiment, the reward r of the control center of the charging station at time t t satisfies the following formula:
[0129]
[0130] In the above formula, represents the economic benefit of the charging station; represents the service satisfaction of the i-th user; α and β are respectively and weights; represents the power purchase power of the charging station from the distribution network at time t; represents the real-time electricity price in the electricity market; represents the arrival time of the i-th user at the charging station; represents the shortest charging time of the vehicle of the i-th user.
[0131] S3: Train the scheduling optimization network according to the operation data of the charging station, so that it can adaptively obtain a better control strategy in continuous decision-making and learning to maximize the economic benefit of the charging station and the service satisfaction of users.
[0132] As Figure 3 shown, in this embodiment, the training process of the scheduling optimization network is as follows:
[0133] S31: Initialize the training parameters of the DQN algorithm, including the learning rate, reward discount factor, exploration probability, training step size and number of iterations; define the experience replay pool and the scale of the mini-batch data; initialize the system environment and obtain the initial state of the charging station.
[0134] S32: According to the output of the Online neural network and the greedy strategy, perform charge and discharge behavior control during the operation of the charging pile, and then generate the charge and discharge power and electricity price at the current moment to obtain the corresponding action a t .
[0135] S33: Determine the state s t of the environment at the next moment according to the formulated action a t+1 , and calculate the reward r t of the action.
[0136] S34: Store the associated parameters of the iterative process as experience data in the experience replay pool. The experience data includes: the current state s t+1 , the action a t , the reward r t and the state s t+1 at the next moment.
[0137] S35: Randomly select a small batch of data from the experience replay pool and transfer it to the Online neural network and the Target neural network;
[0138] S36: The Target network takes the small batch of data as input and outputs the target Q value. The Online network takes the small batch of data as input and outputs the estimated Q value. Calculate the error loss based on the target Q value and the estimated Q value;
[0139] S37: Iteratively execute steps S32 - S36, and perform backpropagation of the gradient according to the error loss to update the model parameters of the Online neural network and the Target neural network until the preset number of iterations is reached.
[0140] S4: When any vehicle arrives at the charging station and requests charging, the control center of the charging station obtains the demand information of the current vehicle, including: the arrival time, the expected departure time, the state of charge when the vehicle arrives, and the expected state of charge when leaving. Then, use the trained scheduling optimization network to generate the best charging and discharging plan that meets the current user's needs.
[0141] Embodiment 2
[0142] Based on the solution in Embodiment 1, this embodiment also provides a power data scheduling system based on deep reinforcement learning, as Figure 4 shown, which includes a demand acquisition unit, a decision-making unit, and an execution unit.
[0143] Among them, the demand acquisition unit is used to obtain the demand information of any arriving vehicle, including: the arrival time, the expected departure time, the state of charge when the vehicle arrives, and the expected state of charge when leaving.
[0144] There is a scheduling optimization network trained by using the power data scheduling method based on deep reinforcement learning as in Embodiment 1 running in the decision-making unit. The scheduling optimization network is used to generate a charging and discharging plan that can meet the user's needs and maximize the economic benefits of the charging station and the service satisfaction of the user according to the operating states of each charging pile in the charging station and the user's demand information.
[0145] The execution unit is used to allocate an idle charging pile to the user, and after the vehicle of the user is connected to the designated charging pile, provide services to the vehicle according to the charging and discharging plan generated by the decision-making unit.
[0146] Embodiment 3
[0147] Based on the solutions in Embodiments 1 and 2, this embodiment further provides a charging station scheduling device, which includes a memory, a processor, and a computer program stored on the memory and running on the processor. When the processor executes the computer program, a power data scheduling system based on deep reinforcement learning as in Embodiment 2 is created, thereby realizing interaction with users and providing vehicle charging and discharging services that meet the needs of customers.
[0148] The charging station scheduling device provided in this embodiment is essentially a computer device. This computer device can be an embedded computer device or a general-purpose computer device. For example, it can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a rack server, a blade server, a tower server, or a cabinet server (including an independent server or a server cluster composed of multiple servers) that can execute programs.
[0149] In one typical application of the solution in this embodiment, the charging station scheduling device can serve as the management background of the entire charging pile and run on the background server. At this time, when a user makes a reservation request through a cloud APP or the like, the device provides corresponding scheduling solutions and services. In another typical application, the charging station scheduling device can also be embedded in each charging pile to communicate with the background service when each charging pile receives a user request and provide services according to the generated scheduling solutions.
[0150] The computer device in this embodiment at least includes, but is not limited to, a memory and a processor that can communicate with each other through a system bus. In this embodiment, the memory (i.e., the readable storage medium) includes flash memory, a hard disk, a multimedia card, a card-type memory (such as an SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory can be an internal storage unit of the computer device, such as the hard disk or memory of the computer device. In other embodiments, the memory can also be an external storage device of the computer device, such as a plug-in hard disk equipped on the computer device, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Of course, the memory can also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the memory is usually used to store the operating system and various application software installed on the computer device. In addition, the memory can also be used to temporarily store various data that have been output or will be output.
[0151] In some embodiments, the processor may be a Central Processing Unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor is generally used to control the overall operation of the computer device.
[0152] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A power data scheduling method based on deep reinforcement learning, characterized in that, It includes: S1: Establish a mathematical model representing the charging and discharging scheduling optimization problem of the charging station; S2: Model the charging and discharging scheduling optimization problem as a Markov decision process, and construct a scheduling optimization network for solving the mathematical model based on the DQN algorithm model; in the scheduling optimization network, define: (1) The action a of the control center of the charging station at time t t includes the charging price, the discharging price, as well as the charging power and discharging power of each charging pile; (2) The state s of the control center of the charging station at time t t includes the global state information of the charging station and the user state information of the charging piles The two satisfy: Among them, and respectively represent the charging power and discharging power of the m-th charging pile at time t; and respectively represent the charging price and discharging price of the charging station at time t; m ∈ M; represents the expected SOC information of the i-th user; represents the expected departure time of the i-th user; represents the energy demand of the i-th user, i ∈ N; (3) The reward r of the control center of the charging station at time t t satisfies the following formula: In the above formula, represents the economic benefit of the charging station; represents the service satisfaction of the i-th user; α and β are the weights of and respectively; represents the power purchase from the distribution network by the charging station at time t; represents the real-time electricity price in the electricity market; represents the arrival time of the i-th user at the charging station; represents the shortest charging time of the vehicle of the i-th user. S3: Train the scheduling optimization network according to the operation data of the charging station, so that it adaptively obtains corresponding control strategies in continuous decision-making and learning to maximize the economic benefits of the charging station and the service satisfaction of users; S4: When any vehicle arrives at the charging station and puts forward a charging demand, the control center of the charging station obtains the demand information of the current vehicle, including: arrival time, expected departure time, state of charge when the vehicle arrives, and expected state of charge when leaving, and then uses the trained scheduling optimization network to generate the best charging and discharging plan that meets the current user's demand.
2. The power data scheduling method based on deep reinforcement learning according to claim 1, wherein: In the mathematical model established in step S1, the optimization objective is to maximize the overall revenue F of the charging station; the charging and discharging power corresponding to each charging and discharging behavior and the charging and discharging electricity price including service fees are used as decision variables; the battery capacity of the vehicle itself, the charging and discharging power constraints of the vehicle and the charging pile, and the service electricity price constraint are used as overall constraints.
3. The power data scheduling method based on deep reinforcement learning according to claim 2, wherein: In the mathematical model, the optimization objective is: In the above formula, T represents the operation duration of the charging station; represents the basic load of the interaction between the charging station and the power grid.
4. The power data scheduling method based on deep reinforcement learning according to claim 3, characterized in that: In the mathematical model, the expression of the battery capacity constraint of each vehicle is as follows: In the above formula, Cap represents the battery capacity of the current vehicle; SoC t represents the real-time state of charge of the current vehicle; η chg and η dhg represent the charging energy efficiency and the discharging energy efficiency between the vehicle and the charging pile respectively; SoC min and SoC max represent the lower limit and the upper limit of the state of charge of the vehicle respectively; SoC d represents the desired state of charge of the current vehicle.
5. The power data scheduling method based on deep reinforcement learning according to claim 4, wherein: In the mathematical model, the expression of the charging and discharging power constraint of the vehicle and the charging pile is as follows: In the above formula, and respectively represent the maximum charging power and the maximum discharging power allowed by the charging pile connected to the current vehicle; represents the maximum instantaneous power that the charging station can withstand.
6. The power data scheduling method based on deep reinforcement learning according to claim 5, characterized in that: In the mathematical model, the expression of the service electricity price constraint is as follows: In the above formula, and are the upper and lower limits of the charging electricity price; is the slack variable of the charging electricity price; and are the upper and lower limits of the discharging electricity price; is the slack variable of the discharging electricity price.
7. The power data scheduling method based on deep reinforcement learning according to claim 6, characterized in that: In step S2, the shortest charging time T of any vehicle need is used to ensure that the state of charge of the vehicle can meet the expectation before the expected departure time, and its calculation formula is as follows: The energy demand P of any vehicle req is calculated as follows: P req = (SoC d - SoC a ) · Cap; In the above formula, SoC a represents the state of charge when the vehicle arrives at the station.
8. The power data scheduling method based on deep reinforcement learning according to claim 1, characterized in that: In step S3, the training process of the scheduling optimization network is as follows: S31: Initialize the training parameters of the DQN algorithm, including learning rate, reward discount factor, exploration probability, training step size and number of iterations; define the experience replay pool and the scale of the mini-batch data; Initialize the system environment and obtain the initial state of the charging station; S32: According to the output of the Online neural network and the greedy strategy, during the operation of the charging pile, charge and discharge behavior control is performed, thereby generating the charging and discharging power and electricity price at the current moment, and obtaining the corresponding action a t ; S33: Determine the state s of the environment at the next moment according to the defined action a t and calculate the reward r of the action t+1 ; t ; S34: Store the associated parameters of the iterative process as empirical data in the experience replay pool, where the empirical data includes: the current state s t+1 , action a t , reward r t and the state s at the next moment t+1 ; S35: Randomly select a mini-batch of data from the experience replay pool and transfer it to the Online neural network and the Target neural network; S36: The Target network takes the mini-batch of data as input and outputs the target Q value, and the Online network takes the mini-batch of data as input and outputs the estimated Q value; and calculate the error loss according to the target Q value and the estimated Q value; S37: Iteratively execute steps S32 - S36, and perform backpropagation of the reverse gradient according to the error loss to update the model parameters of the Online neural network and the Target neural network until the preset number of iterations is reached.
9. A power data scheduling system based on deep reinforcement learning, characterized in that, It includes: A demand acquisition unit, which is used to obtain the demand information of any arriving vehicle, including: arrival time, expected departure time, state of charge when the vehicle arrives, and expected state of charge when leaving; A decision-making unit, in which a scheduling optimization network trained by using the power data scheduling method based on deep reinforcement learning described in any one of claims 1-8 is running; the scheduling optimization network is used to generate a charging and discharging plan that can meet the user's needs and maximize the economic benefits of the charging station and the service satisfaction of the user according to the operating states of each charging pile in the charging station and the user's demand information; An execution unit, which is used to allocate an idle charging pile to the user, and after the vehicle of the user is connected to the designated charging pile, provide services to the vehicle according to the charging and discharging plan generated by the decision-making unit.
10. A charging station scheduling device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and running on the processor, and is characterized in that: when the processor executes the computer program, it creates the power data scheduling system based on deep reinforcement learning described in claim 9, thereby realizing interaction with the user and providing vehicle charging and discharging services that meet the needs of customers.
Citation Information
Patent Citations
Electric vehicle charging and discharging strategy optimization method based on interior point strategy optimization
CN114997935A
Electric vehicle demand response and charging scheduling method based on deep reinforcement learning
CN116384845A