Control method, system and equipment for virtual power plant containing electric vehicle and medium

Through the multi-agent deep reinforcement learning method, the Markov decision-making model is used to use the distributed partial observable Markov decision-making model to solve the problem of user information leakage and poor control effects in virtual power plants, and more efficient electric vehicle charging and discharging control is achieved, reducing electricity consumption costs and reducing user anxiety.

CN120073816APending Publication Date: 2025-05-30STATE GRID HEBEI ELECTRIC POWER CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510088004.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing virtual power plant control methods containing electric vehicles have problems such as user information leakage and poor control effects, especially in large-scale scenarios, it is difficult to effectively coordinate multiple electric vehicles.

Method used

The multi-agent deep reinforcement learning method is adopted, and the Markov decision model can be observed through the distributed partially. Each charging pile in the virtual power plant is used as an agent. Reinforcement learning is used for centralized training, the optimal control strategy is obtained, and user privacy is protected during distributed execution.

Benefits of technology

It achieves better charging and discharging control performance, reduces the electricity cost of electric vehicles, reduces user mileage anxiety, and improves control effect in large-scale scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120073816A_ABST
    Figure CN120073816A_ABST
Patent Text Reader

Abstract

The invention discloses a virtual power plant control method, system and device containing an electric vehicle and a medium, and relates to the field of virtual power plant control, and the method comprises the steps: modeling a virtual power plant control problem as a distributed partial observable Markov decision model; the distributed partial observable Markov decision process comprises a plurality of intelligent agents, the global state of a virtual power plant, joint observation of all the intelligent agents, local observation of all the intelligent agents, joint actions of all the intelligent agents, actions of all the intelligent agents and a reward function; constructing a value network and a strategy network; one agent corresponds to one policy network; performing centralized training on the value network and the strategy networks by using a reinforcement learning method to obtain an optimal control strategy corresponding to the value network and each strategy network; according to the method, user information leakage is prevented, better charging and discharging control performance is realized, the power utilization cost of the electric vehicle is reduced, and the mileage anxiety of the user is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of virtual power plant control, and particularly to a virtual power plant control method, system, device and medium including electric vehicles based on deep reinforcement learning. Background Art

[0002] In recent years, with the rapid development of distributed energy, virtual power plants, as an aggregate that can coordinate the control of distributed energy, contribute to improving the efficiency and reliability of the power grid. Among them, electric vehicles are becoming an important part of virtual power plants due to their advantages in carbon emission reduction and have broad development prospects. The progress of V2G technology enables electric vehicles to have energy storage functions and can realize the reverse power discharge of electric vehicles to the power grid. By aggregating electric vehicle resources in a certain area, virtual power plants can help the power grid adjust the supply-demand balance and improve the stability of power grid operation. For electric vehicle owners, on the one hand, they hope to adjust the charging and discharging power according to the demand response mechanism to save electricity costs, and on the other hand, they are worried about insufficient battery power during travel, which will cause anxiety about insufficient mileage. Therefore, how to achieve the orderly control of virtual power plants including electric vehicles has become a current research hotspot and difficult problem.

[0003] The existing control methods for virtual power plants including electric vehicles are mainly divided into two categories, including mathematical optimization methods and deep reinforcement learning methods. Traditional mathematical optimization methods have a rigorous mathematical theory derivation basis and can give theoretically optimal solutions, but it is usually difficult to obtain accurate model parameters. At the same time, the charging and discharging behaviors of electric vehicles have strong randomness and uncertainty. Therefore, mathematical optimization methods are not applicable to the actual application scenarios of virtual power plant control. As a model-free data-driven method, deep reinforcement learning methods have received extensive attention in recent years. Deep reinforcement learning methods combine the feature extraction ability of deep learning and the search learning ability of reinforcement learning, can obtain experience from historical data, and can learn the optimal control strategy through continuous exploration and update, which is applicable to complex and changeable decision-making scenarios. Therefore, it can be applied to the control problem of virtual power plants including electric vehicles.

[0004] However, the current applied research on deep reinforcement learning is still in its initial stage and there are still many problems. Existing methods usually assume that the environment inside the virtual power plant is completely observable, regard the entire virtual power plant as an intelligent agent, and control the charging and discharging actions of all electric vehicles. However, for electric vehicle users, information such as their own vehicle parameters, commuting time, and charging habits involves personal privacy and is not convenient to share with other users. This leads to partial observability of the environment, and the current single-intelligent-agent method is restricted. Summary of the Invention

[0005] The objective of this application is to provide a virtual power plant control method, system, device, and medium for electric vehicles, which has high applicability, can prevent user information leakage, achieve better charging and discharging control performance, reduce the electricity cost of electric vehicles, and reduce user mileage anxiety.

[0006] To achieve the above objective, the following solutions are provided in this application.

[0007] In the first aspect, this application provides a virtual power plant control method for electric vehicles, including the following steps.

[0008] Construct a virtual power plant control problem and model the virtual power plant control problem as a distributed partially observable Markov decision model; the distributed partially observable Markov decision process includes several agents, the global state of the virtual power plant, the joint observation of all agents, the local observation of each agent, the joint action of all agents, the action of each agent, and the reward function; one agent corresponds to one charging pile in the virtual power plant; the action is the charging and discharging power of the electric vehicle.

[0009] Construct a value network and a policy network; each agent corresponds to one policy network.

[0010] Use the reinforcement learning method to centrally train the value network and the policy network to obtain the optimal control strategy corresponding to the value network and each policy network.

[0011] Based on the optimal control strategy, determine the control scheme of the virtual power plant.

[0012] Optionally, the local observation of the agent includes the current time, the electricity price at the current time, the positions of all users' electric vehicles at the current time, the battery energy of the electric vehicle at the current time, and the household load power at the current time.

[0013] Optionally, the reward function is expressed as follows:

[0014]

[0015] where r j,t is the single-step reward of agent j at time t, C t is the electricity price at time t, P EV,j,t is the charging and discharging power of agent j at time t, α represents the mileage anxiety penalty cost corresponding to the unit of unmet battery energy, E tar,j is the target value to which the battery energy of the expected electric vehicle is to be charged, E j,t is the actual battery energy of electric vehicle j when leaving the charging pile, t arr,j is the time when electric vehicle j arrives home, t dep,j is the time when electric vehicle j leaves home.

[0016] Optionally, the value network includes a first splicing module, a first linear layer, a self-attention encoder, a second splicing module, a second linear layer, a first normalization layer, a first ReLU activation layer, and a third linear layer connected in sequence; the first splicing module is used to splice the local observations and actions of all agents.

[0017] Optionally, the policy network includes a fourth linear layer, a second normalization layer, a second ReLU activation layer, and a fifth linear layer connected in sequence.

[0018] Optionally, using the reinforcement learning method, the value network and the policy network are centrally trained to obtain the optimal control policies corresponding to the value network and each of the policy networks, specifically including:

[0019] Step 301: In each time step, for each agent, using the policy network corresponding to the agent, based on the local observation of the agent, output the action of the agent; the actions of all agents form a joint action;

[0020] Step 302: Based on the actions of the agents, the agents perform charging and discharging actions, calculate the rewards of the agents based on the reward function, and transfer to the next moment state;

[0021] Step 303: Repeat Step 301 - Step 302 to obtain an experience replay array; the experience replay array includes multiple experiences; the global state at the current moment, the actions of the agents at the current moment, the rewards of the agents at the current moment, and the global state at the next moment constitute one experience;

[0022] Step 304: Use the sample set to centrally train the value network and the policy network to obtain the optimal control policies corresponding to the value network and each of the policy networks; the sample set is selected from the experience replay array.

[0023] Optionally, using the Monte Carlo algorithm, several experiences are selected from the experience replay array to obtain the sample set.

[0024] In a second aspect, the present application provides a virtual power plant control system including an electric vehicle, including the following modules.

[0025] A virtual power plant control problem transformation module, configured to: construct a virtual power plant control problem and model the virtual power plant control problem as a distributed partially observable Markov decision model; the distributed partially observable Markov decision process includes several agents, the global state of the virtual power plant, the joint observations of all agents, the local observations of each agent, the joint actions of all agents, the actions of each agent, and a reward function; one agent corresponds to one charging pile in the virtual power plant; the action is the charging and discharging power of the electric vehicle.

[0026] A value network and a policy network construction module, configured to: construct a value network and a policy network; each of the agents corresponds to one of the policy networks.

[0027] An optimal control strategy determination module, configured to: use a reinforcement learning method to centrally train the value network and the policy networks, and obtain the optimal control strategy corresponding to the value network and each of the policy networks.

[0028] A control scheme determination module, configured to: determine the control scheme of the virtual power plant based on the optimal control strategy.

[0029] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the above virtual power plant control method including electric vehicles.

[0030] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above virtual power plant control method including electric vehicles is implemented.

[0031] According to the specific embodiments provided by the present application, the following technical effects are disclosed in the present application:

[0032] The present application provides a virtual power plant control method, system, device and medium including electric vehicles. Through training by a reinforcement learning method, which is a model-free data-driven method, different from existing mathematical optimization methods, it does not need to rely on a complete model and accurate parameters, and can learn the optimal control strategy in continuous interaction with the environment, and obtain user rules and scheduling experience from historical data. Therefore, it has high applicability in actual virtual power plant control scenarios; adopting a multi-agent modeling method, each charging pile in the virtual power plant is regarded as a separate agent, and a centralized training and distributed execution architecture is adopted. In this way, it can learn the globally optimal coordinated control strategy during the training process, avoiding a single agent falling into a local optimum, and does not require device communication and data sharing during the execution process, preventing the leakage of information such as user vehicle parameters, commuting time, and charging habits, and ensuring the personal privacy security of users. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0034] Figure 1 This is an application environment diagram of a virtual power plant control method including electric vehicles in an embodiment of the present application;

[0035] Figure 2 This is a schematic flowchart of a virtual power plant control method including electric vehicles provided in an embodiment of the present application;

[0036] Figure 3 This is a schematic diagram of a value network structure provided in an embodiment of the present application;

[0037] Figure 4 This is a schematic diagram of a policy network structure provided in an embodiment of the present application;

[0038] Figure 5 This is a schematic diagram of a centralized training and distributed execution structure provided in an embodiment of the present application;

[0039] Figure 6 This is a schematic diagram of functional modules of a virtual power plant control system including electric vehicles provided in an embodiment of the present application;

[0040] Figure 7 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. Detailed implementation manners

[0041] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0042] When existing mathematical optimization methods are used to solve the virtual power plant control problem including electric vehicles, a complete optimization model needs to be established and solved based on accurate parameters. Although a theoretical optimal solution can be given, it is difficult to apply in practice. On the one hand, it is difficult to accurately and completely count the parameter indicators of the power grid and electric vehicles in the virtual power plant; on the other hand, the charging and discharging behaviors of users have strong uncertainties, and other loads in the virtual power plant will also fluctuate randomly. These dynamic and random changes are difficult to be processed by the mathematical optimization methods based on fixed models, so their practical applicability is poor.

[0043] Existing deep reinforcement learning methods only use a single agent to train a virtual power plant to control all electric vehicles. When making decision control, the virtual power plant needs to obtain the observation information of all electric vehicles before calculating the charge and discharge action values of each vehicle. However, for electric vehicle owners, information such as their vehicle parameters, commuting time, and charging habits belongs to personal privacy and is not convenient to share with other users in the virtual power plant, resulting in only partial observability of the virtual power plant environment. The single-agent deep reinforcement learning method is no longer applicable.

[0044] The policy network and value network structures in existing deep reinforcement learning methods are relatively simple. Usually, a fully connected neural network is adopted, which is only composed of stacked linear layers and activation layers. However, in a large-scale virtual power plant scenario, the number of electric vehicles increases. Relying solely on a simple fully connected network, it is difficult to effectively extract the high-dimensional features of the input vector, resulting in difficult algorithm training, insufficient learning ability, poor scalability, and the virtual power plant's inability to effectively coordinate numerous electric vehicles, with poor control effects and an inability to meet the demands of electric vehicle owners for electricity costs and electricity consumption. The satisfaction of users with electricity costs and electricity consumption decreases.

[0045] To address the above problems, this application uses a deep reinforcement learning method to obtain a virtual power plant control strategy through model-free data-driven training, thereby overcoming the dependence on precise model parameters. A multi-agent deep reinforcement learning method is adopted, and through a centralized training and distributed execution framework, the protection of user privacy information is achieved. The policy network and value network structures introduce self-attention encoders in the neural network to solve the problem of poor control effects in large-scale scenarios, thereby improving the performance of the algorithm, enhancing the effectiveness of high-dimensional feature extraction of the input vector, reducing the training difficulty, achieving the effective coordination of the virtual power plant for numerous electric vehicles, improving the control effect, and increasing the satisfaction of users with electricity costs and electricity consumption.

[0046] To make the above objects, features, and advantages of this application more obvious and understandable, the following further details this application in conjunction with the accompanying drawings and specific embodiments.

[0047] The virtual power plant control method containing electric vehicles provided by the embodiments of this application can be applied to, for example Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be set separately, integrated on the server 104, placed on the cloud or other servers. The terminal 102 can send the basic information of the virtual power plant to the server 104. After receiving the basic information of the virtual power plant, the server 104 models the virtual power plant control problem as a distributed partially observable Markov decision model, and uses the reinforcement learning method to centrally train the value network and the policy network to obtain the optimal control policy corresponding to the value network and each policy network. Based on the optimal control policy, the control scheme of the virtual power plant is determined. The server 104 can feedback the obtained control scheme for the virtual power plant to the terminal 102. In addition, in some embodiments, the virtual power plant control method including electric vehicles can also be implemented separately by the server 104 or the terminal 102. For example, the terminal 102 can directly perform virtual power plant control on the basic information of the virtual power plant, or the server 104 can obtain the basic information of the virtual power plant from the data storage system and perform virtual power plant control on the video to be processed.

[0048] Among them, the terminal 102 can be, but is not limited to, various desktop computers and laptop computers. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers, and can also be a cloud server.

[0049] In an exemplary embodiment, as Figure 2 shown, a virtual power plant control method including electric vehicles is provided. This method is executed by a computer device, and can be specifically executed alone by a computer device such as a terminal or a server, or jointly executed by a terminal and a server. In the embodiments of the present application, taking this method applied to Figure 1 the server 104 in as an example for illustration, it includes the following steps 201 to step 204. Among them:

[0050] Step 201: Construct a virtual power plant control problem and model the virtual power plant control problem as a distributed partially observable Markov decision model; the distributed partially observable Markov decision process includes several agents, the global state of the virtual power plant, the joint observation of all agents, the local observation of each agent, the joint action of all agents, the action of each agent and the reward function; one agent corresponds to one charging pile in the virtual power plant; the action is the charging and discharging power of the electric vehicle.

[0051] The virtual power plant control problem consists of charging cost and user mileage anxiety cost; the virtual power plant includes several charging piles; each charging pile is used for an electric vehicle to charge. At each time step, each agent obtains its own local observation.

[0052] Step 202: Construct a value network and a policy network; each of the agents corresponds to one of the policy networks.

[0053] Step 203: Use the reinforcement learning method to centrally train the value network and the policy networks to obtain the optimal control policies corresponding to the value network and each of the policy networks.

[0054] To optimize the charging cost, the user range anxiety cost, and the reward, use the reinforcement learning method for centralized training. During the centralized training process, each agent uploads its own observations and actions to the virtual power plant, and the value network of the virtual power plant guides the parameter update of the policy network of each agent.

[0055] Step 204: Based on the optimal control policies, determine the control scheme of the virtual power plant. The control scheme includes the charging and discharging power of each electric vehicle in the virtual power plant.

[0056] Implementing the above Steps 201 to 204 solves the problems of privacy protection and multi-user coordination in the control scenario of a virtual power plant with electric vehicles, so as to reduce the electricity cost of electric vehicles and relieve the user's range anxiety. It provides an innovative solution for the efficient aggregated control of electric vehicles in the virtual power plant. In addition, aiming at the problems of difficult training and poor control effect of existing methods in large-scale scenarios, this application introduces a self-attention encoder into the value network, which can efficiently extract global features, reasonably allocate weight relationships, and thus more accurately evaluate the quality of the agent's actions, urge it to continuously improve the charging and discharging strategy, so that each electric vehicle in the virtual power plant can coordinate and cooperate more effectively, achieve better charging and discharging control performance, reduce the electricity cost of electric vehicles and reduce the user's range anxiety.

[0057] Consider regarding a certain residential area as a virtual power plant. Each household in the area is equipped with a charging pile, which can charge the household electric vehicle and also provide the reverse discharging service of V2G. As an aggregate of distributed electric vehicle resources in the area, the virtual power plant coordinates and controls the charging and discharging actions of each charging pile to reduce the overall electricity cost and ensure the electricity demand for the travel of all users. Assume that the user set in the virtual power plant is J, the user index number is j, the total cycle is T, and the time index number is t. For household users, electric vehicles usually arrive in the evening and leave in the early morning, and use the night time for charging. Assume that the time when user j drives the electric vehicle home is t arr,j , and the time to leave home is t dep,j . The position identifier of the electric vehicle is represented by the Boolean variable F j,t , where 1 represents at home and 0 represents not at home. During the period when the electric vehicle is at home, the battery energy of the electric vehicle of user j at time t is E j,tIt is indicated that the charging and discharging power of the electric vehicle is represented by P EV,j,t It is indicated that the power of the household load other than the electric vehicle is represented by P L,j,t It is indicated that the initial battery energy of the electric vehicle when user j arrives home is The battery energy when leaving home is

[0058] The goal of the above virtual power plant control problem is to reduce the charging cost of the electric vehicle while reducing the user's range anxiety. Considering the intra-day electricity market fluctuations, assuming that the electricity price at time t is C t , then the charging cost of user j during the scheduling period is That is Assume that when user j leaves home, it is expected that the battery energy of the electric vehicle is charged to the target value E tar,j . If the battery energy of the electric vehicle reaches the target value, the user's range anxiety is 0; if the battery energy of the electric vehicle does not reach the expected value, the range anxiety is the difference between the actual battery energy when the electric vehicle leaves and the target battery energy. Therefore, the user's range anxiety can be represented by It is indicated.

[0059] This application adopts the multi-agent deep reinforcement learning method to solve the above virtual power plant control problem containing electric vehicles and models it as a distributed partially observable Markov decision model. The distributed partially observable Markov decision model mainly includes agents, states, observations, actions, rewards, etc. In this control scenario, its specific physical meanings are as follows.

[0060] The basic information of the virtual power plant includes the data of the charging piles.

[0061] Agent: Each charging pile in the virtual power plant is regarded as a separate agent. The charging pile is used to collect information and make action decisions. The set and index of the agents are also represented by and j.

[0062] State: At time t, the global state s of the virtual power plant t is defined as the set where t, C t , F j,t , E j,t , P L,j,t represent the current time, the electricity price at the current time, the location of the electric vehicle of user j at the current time, the battery energy of the electric vehicle of user j at the current time, and the household load power of user j at the current time, respectively. Then the global state of the virtual power plant includes the current time, the electricity price, the locations, battery energies, and household load powers of all users' electric vehicles.

[0063] Observation: represents the joint observation of all agents at time t, while the local observation o of agent j at time t j,t is defined as the set {t, C t , F j,t , E j,t , P L,j,t}}. Except for time and electricity price, information such as the location of electric vehicles, battery energy, and household load is protected by personal privacy and is not shared during actual execution. Therefore, a single charging pile can only obtain its own household information and cannot observe the overall information of the entire virtual power plant. Then the local observation of the agent includes the current time, the electricity price at the current time, the locations of electric vehicles of all users at the current time, the battery energy of the electric vehicle at the current time, and the household load power at the current time.

[0064] Action: represents the joint action of all agents at time t, a j,t = P EV,j,t is the action of agent j at time t, that is, the charging and discharging power of the electric vehicle. The charging and discharging actions are restricted by the rated energy of the battery and the maximum charging and discharging power.

[0065] Reward: represents the global reward of the virtual power plant at time t, that is, the sum of the rewards of all agents. The reward function is shown in Equation (1).

[0066]

[0067] Among them, r j,t is the single-step reward of agent j at time t, C t is the electricity price at time t, P EV,j,t is the charging and discharging power of agent j at time t, α represents the mileage anxiety penalty cost corresponding to the unit of unmet battery energy, E tar,j is the target value of the battery energy that the expected electric vehicle is to be charged to, E j,t is the actual battery energy of electric vehicle j when leaving the charging pile, t arr,j is the time when electric vehicle j arrives home, t dep,j is the time when electric vehicle j leaves home.

[0068] The goal of the above virtual power plant control Markov model is to find the optimal joint control strategy so that the charging piles in the virtual power plant can be effectively coordinated and the electricity consumption costs and mileage anxiety of all users' electric vehicles can be reduced as much as possible.

[0069] The role of the value network is to evaluate the actions of the agents based on all observed information, thereby learning the advantages and disadvantages of the executed strategy and making corresponding improvements. In this application, a self-attention encoder is introduced into the value network to enhance its learning ability and scalability. The self-attention mechanism, whose encoder consists of two sub-layers: a multi-head self-attention network and a position-wise feed-forward network, performs weighted summation on multiple groups of queries, keys, and values, and then, after processes such as concatenation of the same dimension, residual connection, and layer normalization, can adaptively focus on the important features of the input sequence and assign higher weights to key variables. The structure of the value network is as Figure 3 shown.

[0070] The value network includes a first concatenation module, a first linear layer, a self-attention encoder, a second concatenation module, a second linear layer, a first layer normalization layer, a first ReLU activation layer, and a third linear layer connected in sequence; the first concatenation module is used to concatenate the local observations and actions of all agents.

[0071] The first concatenation module concatenates the local observation o j,t and action a j,t of each agent together, and then performs a linear projection of the first linear layer to obtain a sequence of vectors, denoted as the first concatenation feature, and the sequence length is the total number of agents. Then the obtained sequence of vectors is fed into the self-attention encoder, and after the weighted processing of the self-attention encoder, the vector features are fully extracted to obtain a sequence of weighted feature vectors. The second concatenation module concatenates the sequence of weighted feature vectors to obtain a second concatenation feature, and then successively passes through the second linear layer, the first layer normalization, the first ReLU activation layer, and finally linearly maps to the global value Q glo,t . Suppose represents the value network, whose network parameters are ω, and the expression of the value network is shown in Equation (2).

[0072]

[0073] The role of the policy network is to generate corresponding actions according to the local observations of the agents. Each agent has its own separate policy network, and the parameters are not shared. The structure of the policy network is as Figure 4 shown.

[0074] The policy network includes a fourth linear layer, a second layer normalization layer, a second ReLU activation layer, and a fifth linear layer connected in sequence.

[0075] The observation o j,t successively passes through the linear projection of the fourth linear layer, the second layer normalization, the second ReLU activation layer, and finally linearly maps to the action a j,t . Suppose μ jDenote the policy network of agent j with network parameters θ j , and the expression of the policy network is shown in Equation (3):

[0076]

[0077] The virtual power plant control method with electric vehicles proposed in this application adopts a centralized training and distributed execution architecture, and its architecture diagram is as Figure 5 shown.

[0078] In another exemplary embodiment of this application, the above step 203 specifically includes the following steps 301 to 304.

[0079] Step 301: In each time step, for each agent, use the policy network corresponding to the agent to output the action of the agent according to the local observation of the agent; the actions of all agents form a joint action.

[0080] Step 302: Based on the actions of the agents, the agents perform charging and discharging actions, calculate the rewards of the agents based on the reward function, and transfer to the next moment state.

[0081] Step 303: Repeat steps 301 - 302 to obtain an experience replay array; the experience replay array includes multiple experiences; the global state at the current moment, the actions of the agents at the current moment, the rewards of the agents at the current moment, and the global state at the next moment constitute an experience.

[0082] Step 304: Use the sample set to centrally train the value network and the policy network to obtain the optimal control strategies corresponding to the value network and each policy network; the sample set is selected from the experience replay array, and one sample is an experience.

[0083] The centralized training process is completed in the control master station of the virtual power plant. During the training, the agents upload their own observation and action information to the virtual power plant, and the global value network uniformly guides the improvement of each agent's policy network. When the training is over, the virtual power plant will obtain the parameters θ of each agent's policy network jIt is sent to the corresponding charging pile, and the parameters of the value network do not need to be sent. During the step-by-step execution, the charging pile can observe personal privacy information such as the location of the electric vehicle, battery energy, and household load in the household, and obtain public information such as time and electricity price from the virtual power plant. Based on the above observed data, the charging pile can quickly obtain the magnitude of the control action only by the feedforward operation of its own policy network, output the corresponding charging and discharging power, provide G2V or V2G services for the household electric vehicle, help users earn price difference profits by taking advantage of different electricity prices at night, reduce electricity costs, and at the same time ensure the electricity satisfaction when the user drives away and reduce the user's anxiety about insufficient mileage.

[0084] The process of centralized training is specifically described below.

[0085] The centralized training process includes two steps: collecting experience and updating network parameters. During the process of the virtual power plant collecting experience, at each time step of each episode, the agent j obtains the local observation o t from the global state s j,t , sends it into the policy network μ j , and adds the action noise ξ to the output of the policy network μ j to obtain the action a j,t , as shown in Equation (4).

[0086]

[0087] In the formula, the action noise ξ is randomly drawn from a normal distribution with a mean of 0 and a standard deviation of σ .

[0088] The actions a j,t of all agents form the joint action a t , and the charging pile executes the charging and discharging actions. After fluctuations in electricity price and load, arrival and departure of electric vehicles, adjustment of battery energy, etc., the corresponding reward r t is calculated and transferred to the next moment state s t+1 . After one exploration of the behavior policy, an experience is generated, which is represented by the quadruple (s t , a t , r t , s t+1 ) and this experience is stored in the experience replay array.

[0089] After a certain amount of exploration warm-up, when the number of collected experiences reaches a certain value, the Monte Carlo algorithm is used to select several experiences from the experience replay array to obtain the sample set. Specifically, the Monte Carlo algorithm is used to randomly extract small batch samples from the experience replay array to update the neural network parameters of the value network and the policy network. Assume that L samples are randomly extracted from the experience replay array, and the lth quadruple sample is (s l ,a l ,r l ,s l '),in The superscript ' represents the variable at the next moment. In addition, this application introduces the target network corresponding to the strategy network and the value network, respectively represented by μ tar,j , Represents that the target policy network μ tar,j and target value network The network parameters are θ tar,j ,ω tar , in order to alleviate the problem of overestimation. The target network has the same structure as the original network and the same parameters during initialization, but the parameter update process is different.

[0090] First, the time difference algorithm is used to update the value network Parameter ω. For the lth sample, based on the observation sequence o at the next moment j ' ,l , through the target policy network μ tar,j and target value network Calculate the action a at the next moment in sequence j ' ,l and value Q g ' lo,l , thus obtaining the time difference target TD tar,l , as shown in equations (5)-(7):

[0091]

[0092] TD tar,l =r l +γQ g ' lo,l (7).

[0093] Based on the local observation o of all agents in the lth sample j,l and action a j,l , calculate the current global value Q based on the value network glo,l , thus obtaining the time difference error TD err,l , as shown in equations (8)-(9).

[0094]

[0095] TD err,l = Q glo,l - TD tar,l (9).

[0096] The goal of the temporal difference algorithm is to make the current value evaluation Q glo,l closer to the temporal difference target TD tar,l . Therefore, the mean squared error of L samples is used to replace the expectation, and the value loss lossQ is obtained. Then, the network parameters ω are updated by gradient descent, as shown in Eqs. (10)-(11).

[0097]

[0098] where η ω is the learning rate of the value network.

[0099] The policy network μ is updated using the policy gradient algorithm j parameters θ j . For the l-th sample, based on the observation o j,l of agent j, the action j is calculated according to the current policy network μ and the value network is used to evaluate this action Perform scoring to obtain the value corresponding to the l th sample This process is as shown in Equation (12) - (13) as shown.

[0100]

[0101] The goal of policy learning is to maximize the expected value of the score given by the value network for the action. Therefore, the mean of L samples is used to replace the expectation, and its negative value is taken to obtain the policy loss lossμ. Then, the network parameters θ j of the policy network are updated by gradient descent, as shown in Eqs. (14)-(15).

[0102]

[0103] where η θ is the learning rate of the policy network.

[0104] The target policy network μ tar,j and the target value network parameters θ tar,j , ω tar are updated using the soft update algorithm. The soft update factor τ is introduced, and the update formulas are shown in Eqs. (16)-(17).

[0105]

[0106] ω tar ←τω+(1 - τ)ω tar (17).

[0107] After the above training process, the optimal control strategies corresponding to the value network and each of the policy networks can be obtained. The optimal control strategies include the optimal network parameters of the value network and the optimal network parameters of each of the policy networks, and the trained value network and the trained policy networks are obtained.

[0108] Based on the above trained value network and trained policy networks, the control scheme of the virtual power plant is determined.

[0109] This application also provides an application scenario, which applies the virtual power plant control method with electric vehicles described above. Specifically: The virtual power plant control method with electric vehicles provided in this embodiment can be applied in the virtual power plant control scenario. The virtual power plant control scenario includes an information acquisition link and a virtual power plant control link; the basic information of the virtual power plant enters the virtual power plant control link from the information acquisition link to obtain the corresponding control scheme. The virtual power plant control method with electric vehicles provided in this embodiment belongs to the virtual power plant control link. Specifically, in the process of the virtual power plant control link for the virtual power plant, the virtual power plant control problem can be modeled as a distributed partially observable Markov decision model, and the value network and the policy network are centrally trained by using the reinforcement learning method to obtain the optimal control strategies corresponding to the value network and each policy network, and based on the optimal control strategies, the control scheme of the virtual power plant is determined.

[0110] Based on the same inventive concept, the embodiments of this application also provide a virtual power plant control system with electric vehicles for implementing the virtual power plant control method with electric vehicles involved above. The implementation solutions provided by this system to solve problems are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more embodiments of the virtual power plant control system with electric vehicles provided below can refer to the limitations on the virtual power plant control method with electric vehicles in the above text, and will not be repeated here.

[0111] In an exemplary embodiment, as Figure 6 shown, a virtual power plant control system with electric vehicles is provided, which includes the following modules.

[0112] The virtual power plant control problem transformation module T1 is used for: constructing a virtual power plant control problem and modeling the virtual power plant control problem as a distributed partially observable Markov decision model; the distributed partially observable Markov decision process includes several agents, the global state of the virtual power plant, the joint observations of all agents, the local observations of each agent, the joint actions of all agents, the actions of each agent, and a reward function; one agent corresponds to one charging pile in the virtual power plant; the action is the charging and discharging power of the electric vehicle.

[0113] The value network and policy network construction module T2 is used for: constructing a value network and a policy network; each of the agents corresponds to one of the policy networks.

[0114] The optimal control strategy determination module T3 is used for: using the reinforcement learning method to centrally train the value network and the policy networks to obtain the optimal control strategies corresponding to the value network and each of the policy networks.

[0115] The control scheme determination module T4 is used for: determining the control scheme of the virtual power plant based on the optimal control strategy.

[0116] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal, and its internal structure diagram can be as Figure 7 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store virtual power plant control data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a virtual power plant control method including electric vehicles.

[0117] Those skilled in the art can understand that Figure 7 the structure shown in

[0118] In an exemplary embodiment, a computer device is provided, which includes a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0119] In an exemplary embodiment, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0120] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0121] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0122] In each of the embodiments provided in the present application, the database involved may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on blockchain, etc., without limitation. In each of the embodiments provided in the present application, the processor may be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without limitation.

[0123] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0124] Specific examples are used in this article to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A method for controlling a virtual power plant containing electric vehicles, characterized in that: include: A virtual power plant control problem is constructed and modeled as a distributed partially observable Markov decision model; the distributed partially observable Markov decision process includes several agents, the global state of the virtual power plant, the joint observation of all agents, the local observation of each agent, the joint action of all agents, the action of each agent and the reward function; one agent corresponds to one charging pile in the virtual power plant; the action is the charging and discharging power of the electric vehicle; Constructing a value network and a strategy network; each of the intelligent agents corresponds to a strategy network; Using a reinforcement learning method, the value network and the policy network are centrally trained to obtain the optimal control strategy corresponding to the value network and each of the policy networks; Based on the optimal control strategy, a control scheme of the virtual power plant is determined.

2. The control method of a virtual power plant containing electric vehicles according to claim 1, characterized in that: The local observations of the agent include the current time, the electricity price at the current time, the electric vehicle locations of all users at the current time, the battery energy of the electric vehicle at the current time, and the household load power at the current time.

3. The control method of a virtual power plant containing electric vehicles according to claim 1, characterized in that: The reward function is expressed as follows: Among them, r j,t is the single-step reward of agent j at time t, C t is the electricity price at time t, P EV,j,t is the charging and discharging power of agent j at time t, α represents the mileage anxiety penalty cost corresponding to the unsatisfied battery energy per unit, E tar,j E is the target value of the battery energy of the electric vehicle. j,t is the actual battery energy of electric vehicle j when it leaves the charging station, t arr,j is the time when electric car j arrives home, t dep,j Time to leave home for electric car j.

4. The control method of a virtual power plant containing electric vehicles according to claim 1, characterized in that: The value network includes a first splicing module, a first linear layer, a self-attention encoder, a second splicing module, a second linear layer, a first normalization layer, a first ReLU activation layer and a third linear layer connected in sequence; the first splicing module is used to splice the local observations and actions of all intelligent agents.

5. The control method of a virtual power plant containing electric vehicles according to claim 1, characterized in that: The policy network includes a fourth linear layer, a second normalization layer, a second ReLU activation layer and a fifth linear layer which are connected in sequence.

6. The control method of a virtual power plant containing electric vehicles according to claim 1, characterized in that: The value network and the policy network are centrally trained by using a reinforcement learning method to obtain the optimal control strategy corresponding to the value network and each policy network, specifically including: Step 301: In each time step, for each agent, using the policy network corresponding to the agent, output the action of the agent according to the local observation of the agent; the actions of all agents constitute a joint action; Step 302: Based on the action of the agent, the agent performs charging and discharging actions, calculates the reward of the agent based on the reward function, and transfers to the next moment state; Step 303: Repeat steps 301-302 to obtain an experience replay array; the experience replay array includes multiple experiences; the global state at the current moment, the action of the agent at the current moment, the reward of the agent at the current moment and the global state at the next moment constitute an experience; Step 304: using a sample set to centrally train the value network and the policy network to obtain the optimal control strategy corresponding to the value network and each of the policy networks; the sample set is selected from the experience playback array.

7. The control method of a virtual power plant containing electric vehicles according to claim 6, characterized in that: A Monte Carlo algorithm is used to select a number of experiences from the experience playback array to obtain the sample set.

8. A virtual power plant control system containing electric vehicles, characterized in that: The virtual power plant control system including electric vehicles includes: A virtual power plant control problem conversion module is used to: construct a virtual power plant control problem, and model the virtual power plant control problem as a distributed partially observable Markov decision model; the distributed partially observable Markov decision process includes several intelligent agents, the global state of the virtual power plant, the joint observation of all intelligent agents, the local observation of each intelligent agent, the joint action of all intelligent agents, the action of each intelligent agent and the reward function; one intelligent agent corresponds to one charging pile in the virtual power plant; the action is the charging and discharging power of the electric vehicle; The value network and strategy network construction module is used to: construct a value network and a strategy network; each of the intelligent agents corresponds to one strategy network; The optimal control strategy determination module is used to: use the reinforcement learning method to centrally train the value network and the strategy network to obtain the optimal control strategy corresponding to the value network and each of the strategy networks; The control scheme determination module is used to determine the control scheme of the virtual power plant based on the optimal control strategy.

9. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement a virtual power plant control method containing electric vehicles as described in any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the virtual power plant control method containing electric vehicles described in any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Dynamic pricing and scheduling method for electric vehicle charging station

    CN121526676A