An orderly charging control method, device and system for a distribution network
By establishing an orderly charging target model and Q learning algorithm in the distribution station area, and optimizing the charging load access strategy, the load fluctuation problem of electric vehicle charging on the distribution network is solved, and the stability and efficiency of the distribution network are improved.
Patent Information
- Application Number
- CN202411353257.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-26
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-09-26
AI Technical Summary
With the increase in the number of electric vehicles connected to the grid, the charging demand for distribution networks has increased, resulting in line overload and the stability of distribution network operations declining. It is difficult for traditional adjustment methods to cope with the randomness and impact of large-scale charging loads.
By obtaining electricity consumption and charging load data in the distribution station area, an orderly charging target model is established, and a charging scheduling model is constructed using Markov decision-making process and Q learning algorithms to optimize the access strategy of charging load to reduce load fluctuations and improve the stability of the distribution network.
It has achieved the reduction of the impact of charging load on the load fluctuations of the distribution network, and improved the operating stability of the distribution network and the power transmission efficiency of the power grid.
Smart Images

Figure CN119231508B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the technical field of orderly charging, and in particular, to an orderly charging control method, device and system for a distribution network. Background Art
[0002] With the continuous increase in the number of grid-connected electric vehicles, the demand for charging power provided by the distribution network has increased, while the carrying capacity of the existing distribution network is limited and it is difficult to support the newly increased high charging volume. The existing distribution network has faced the risks of line overload, distribution transformer overload and decline in the operation stability of the distribution network within the region.
[0003] Resources such as distributed power sources and energy storage can, to a certain extent, improve the acceptance capacity of the distribution network for charging loads such as electric vehicles. However, various charging loads, as important adjustable and controllable flexible resources on the power consumption side, have become one of the largest loads in the power grid, with randomness, simultaneity and impact. The difficulty of the traditional means of the power grid to adjust and maintain safe and stable operation will be greatly increased, and it is not applicable to the power regulation scenarios of new distribution systems with large-scale charging load access. Summary of the Invention
[0004] The present invention provides an orderly charging control method, device and system for a distribution network, which reduces the impact of the access charging load on the load fluctuation of the distribution network, realizes the peak shaving and valley filling of the distribution network load, and improves the operation stability of the distribution network.
[0005] In a first aspect, an embodiment of the present invention provides an orderly charging control method for a distribution network, including:
[0006] Obtain data of the power consumption load and charging load of the distribution substation area, and establish a target model for orderly charging of the distribution substation area with the minimum load fluctuation of the distribution substation area as the goal;
[0007] Convert the target model into a Markov decision process, and construct a Q-learning algorithm model for orderly charging scheduling;
[0008] Obtain the access strategies of each charging load according to the Q-learning algorithm model and the target model.
[0009] Optionally, obtaining data of the power consumption load and charging load of the distribution substation area, and establishing a target model for orderly charging of the distribution substation area with the minimum load fluctuation of the distribution substation area as the goal includes:
[0010] Obtain the predicted power consumption load at each moment of the distribution substation area, the number of charging loads at each moment, and the power of a single charging load;
[0011] Obtain the total load of the distribution substation area at each moment according to the predicted power consumption load, the number of charging loads, and the power of a single charging load at each moment;
[0012] The minimum load fluctuation is obtained based on the difference between the total load and the average load of the distribution transformer area at each moment, thereby establishing the target model.
[0013] Optionally, the total load of the distribution transformer area at each moment is expressed as:
[0014] P i =I i ·P c +P load (i)
[0015] I i is the number of charging loads at time Ti, P c is the power of a single charging load, P load (i) is the predicted power consumption load of the distribution transformer area at time Ti;
[0016] The target model is expressed as:
[0017]
[0018] where P i is the total load of the distribution transformer area at time Ti, P avg is the average load, P MAX is the upper limit of the distribution transformer area load.
[0019] Optionally, the target model is converted into a Markov decision process, and a Q-learning algorithm model for ordered charging scheduling is constructed, including:
[0020] According to the Markov decision process, a triple including a state space, an action space, and a reward function is set. Among them, the state space represents the total load of the distribution transformer area at each moment, the action space represents the action of scheduling newly added charging loads to a preset time period for charging, and the reward function represents the total number of charging loads after the action;
[0021] According to the state space and the action space, a Q-value function representing the action taken by the distribution transformer area in the current system state is obtained.
[0022] Optionally, the preset time period is the low load period of the total load of the distribution transformer area.
[0023] Optionally, an access strategy for each charging load is obtained according to the Q-learning algorithm model and the target model, including:
[0024] An equation relationship is established between the Q-value function and the target model;
[0025] Taking the minimum value of the target model as the goal, the Q-value function is solved according to the currently input state space, thereby obtaining the access strategy when the target model is at the minimum value.
[0026] Optionally, after converting the target model into a Markov decision process and constructing a Q-learning algorithm model for orderly charging scheduling, the method further includes:
[0027] Training the Q-learning algorithm model;
[0028] Wherein, training the Q-learning algorithm model includes:
[0029] S1. Observing the current system state within a training period;
[0030] S2. Selecting an action according to a random policy to obtain a reward function for the next system state;
[0031] S3. Obtaining the next system state according to the action space of the current system state after the action;
[0032] S4. Updating the Q-value function of the current system state according to the next system state and the reward function of the next system state;
[0033] S5. Repeating steps S2 to S4 until the training period ends.
[0034] Optionally, updating the Q-value function of the current system state according to the next system state and the reward function R of the next system state includes:
[0035] Updating the Q-value of the current system state according to an update function, where the update function is:
[0036]
[0037] where α is the learning rate and β is the discount rate; Q old (s (t) ,a (t) ) represents the Q-value of taking an action in the current system state; Q new (s (t) ,a (t) ) represents the updated Q-value in the current system state; R (t+1) is the reward function of the next system state after executing the action; s (t+1) represents the next system state entered after executing the current action; represents the minimum Q-value of taking an action in the next system state.
[0038] In a second aspect, an embodiment of the present invention provides an orderly charging control device for a distribution network, including:
[0039] A first model establishment module, configured to obtain data of the electricity consumption load and the charging load of a distribution transformer area, and establish a target model for orderly charging of the distribution transformer area with the minimum load fluctuation of the distribution transformer area as the target;
[0040] The second model building module is used to convert the target model into a Markov decision process and construct a Q-learning algorithm model for orderly charging scheduling;
[0041] The calculation module is used to obtain the access strategies of each charging load according to the Q-learning algorithm model and the target model.
[0042] In a third aspect, an embodiment of the present invention provides an orderly charging control system for a distribution network, including an edge computing intelligent device and a communication module. Among them, the edge computing intelligent device includes an electricity consumption information collection unit, an electricity load collection unit, a charging load collection unit, an orderly charging control device for the distribution network, and a database;
[0043] The communication module is used to establish a communication connection between the edge computing intelligent device and external intelligent meters and charging piles;
[0044] The electricity consumption information collection unit is used to collect the historical data and real-time data of the intelligent meters in the distribution transformer area and store them in the database;
[0045] The electricity load collection unit is used to predict the electricity load of the distribution transformer area on the current day according to the historical data and real-time data of the electricity load of the charging piles, and obtain the predicted electricity load at each moment;
[0046] The charging load collection unit is used to collect the quantity of the charging load at each moment and the power of a single charging load in the distribution transformer area of the charging piles and store them in the database;
[0047] The orderly charging control device for the distribution network is used to obtain the data of the electricity load and charging load in the distribution transformer area from the database, establish a target model for orderly charging in the distribution transformer area with the minimum load fluctuation of the distribution transformer area as the target; convert the target model into a Markov decision process, construct a Q-learning algorithm model for orderly charging scheduling; obtain the access strategies of each charging load according to the Q-learning algorithm model and the target model.
[0048] The technical solution provided by the embodiment of the present invention obtains the data of the electricity load and charging load in the distribution transformer area, establishes a target model for orderly charging in the distribution transformer area with the minimum load fluctuation of the distribution transformer area as the target; converts the target model into a Markov decision process, constructs a Q-learning algorithm model for orderly charging scheduling, considers calculating with the minimum load fluctuation of the distribution network as the target to reduce the complexity of the target operation, uses the Q-learning algorithm model to obtain the state-action mapping, and guides the access strategy of the charging load according to the current state, so as to control the access of the charging load to reduce the impact of the access of the charging load on the load fluctuation of the distribution network, and realize the peak shaving and valley filling of the distribution network load to improve the operation stability of the distribution network. Description of the Drawings
[0049] Figure 1 This is a schematic flowchart of an orderly charging control method for a distribution network provided by an embodiment of the present invention;
[0050] Figure 2 This is a schematic flowchart of another orderly charging control method for a distribution network provided by an embodiment of the present invention;
[0051] Figure 3 This is a schematic structural diagram of an orderly charging control device for a distribution network provided by an embodiment of the present invention;
[0052] Figure 4 This is a schematic structural diagram of an orderly charging control system for a distribution network provided by an embodiment of the present invention;
[0053] Figure 5 This is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0054] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0055] With the continuous increase in the number of grid-connected electric vehicles, the required charging volume provided by the distribution network increases. However, the carrying capacity of the existing distribution network is limited and it is difficult to support the newly increased high charging volume. The existing distribution network already faces risks such as overload of lines and distribution transformers within the region and a decline in the operating stability of the distribution network. Resources such as distributed power sources and energy storage can, to a certain extent, improve the acceptance capacity of the distribution network for electric vehicles. However, the charging behavior of electric vehicles has a high randomness, and the output of distributed photovoltaics has strong volatility, with complex dynamic characteristics. Therefore, traditional distribution transformers are not applicable to the power regulation scenarios of a new distribution system facing the access of various distributed resources and large-scale electric vehicles.
[0056] In view of this, Figure 1 This is a schematic flowchart of an orderly charging control method for a distribution network provided by an embodiment of the present invention. This embodiment is applicable to the control of orderly charging of charging loads. This method can be executed by an orderly charging control device, and the device can be implemented in a hardware and / or software manner. The method specifically includes the following steps:
[0057] S110. Obtain the data of the power consumption load and charging load in the distribution substation area, and establish an objective model for orderly charging in the distribution substation area with the minimum load fluctuation of the distribution substation area as the goal.
[0058] Specifically, the distribution substation area supplies power to power users. Among them, power users include regular power consumption users and charging users composed of charging piles. The power consumption of regular power consumption users can be called the power consumption load, and the power consumption of charging piles connected to charging users can be called the charging load. The historical data and / or real-time data related to the power consumption load and charging load can be collected through the acquisition unit equipped in the distribution substation area. Considering the volatility of various distributed resources and the load volatility generated by electric vehicles accessing the charging piles, an objective model of the minimum load fluctuation in the distribution substation area is established. With the minimum load fluctuation as the limit, the access of the charging load is controlled to reduce the change range between the load values of the distribution network and improve the transmission efficiency of the power grid. Exemplarily, the minimum load fluctuation can be used to reflect the change range of the distribution network load by the difference between the total load and the average load of the distribution network at each moment. Further, the standard deviation of the total load and the average load of the distribution network at each moment can be used to reflect the change range of the distribution network load. The value of the standard deviation is closer to the change range of the original data, and the size of the standard deviation can directly reflect the size of the data dispersion degree, which helps to more intuitively evaluate the volatility of the data.
[0059] S120. Convert the objective model into a Markov decision process and construct a Q-learning algorithm model for orderly charging scheduling.
[0060] Specifically, the Markov decision process (MDP) is a mathematical model that describes the decision-making process of a decision-maker in an uncertain environment. The current state of the MDP depends only on the current state and action, and is independent of the historical state and action. Therefore, the objective model can be converted into an MDP problem to construct a Q-learning algorithm model. The Q-learning algorithm model evaluates the expected utility of taking a certain action in a given state by learning an action Q-value function, and then solves the MDP problem. Exemplarily, according to the MDP, a triple including a state space, an action space, and a reward function can be represented as {S, A, R}, where S is the state space, A is the action space, and R is the reward function.
[0061] State space S: Represents the total load of the distribution substation area at each moment. The system state at time slot t is set as:
[0062]
[0063] Action space A: Represents the actions that affect the system performance. Schedule the newly added charging load to a preset time period [T j , T j+LPerform the charging action. The action for time slot t is set as:
[0064]
[0065] That is to say, the newly added charging load is scheduled to the preset time period [T j , T j+L for charging. Then the number of charging loads at the corresponding moment in this time period increases by one, and the number of charging loads at other moments remains unchanged. Therefore, if the preset time period is set as the low-load period of the distribution network, the low-load period can be fully utilized for charging in the action result, and the power distribution utilization rate of the distribution network can be improved.
[0066] Reward function R: Set the reward as selecting an action in the current system state S, obtaining the reward function for the next time slot t + 1, and entering the next system state S (t+1) . The reward function R can represent the total number of charging loads after the action:
[0067]
[0068] In the state space S, take the action space A to obtain the Q-value function, that is, the Q-learning algorithm model, denoted as Q(s, a).
[0069] S130. Obtain the access strategies for each charging load according to the Q-learning algorithm model and the target model.
[0070] Specifically, establish a relationship between the Q-value function and the target model, and solve with the goal of the Q-value function set as the minimum value of the target model to obtain the Q-value. That is to say, a system state and an action correspond to a Q-value. The Q-value maps each state-action pair to a real number, and this real number represents the expected minimum load fluctuation obtained after executing this action from this state. In the Q-Learning algorithm, the Q-value interacts with the environment continuously, observes the rewards obtained by taking different actions in different states, and updates the Q-value accordingly to gradually obtain the optimal value of the target model. Therefore, according to the Q-value that meets the minimum load fluctuation, that is, the target model is the minimum value, the corresponding state-action mapping is determined, and the access strategy of the charging load can be guided according to the current state.
[0071] The technical solution provided by the embodiments of the present invention is to obtain the data of the power consumption load and the charging load in the distribution transformer area, and establish an objective model for orderly charging in the distribution transformer area with the minimum load fluctuation of the distribution transformer area as the goal; convert the objective model into a Markov decision process, construct a Q-learning algorithm model for orderly charging scheduling, consider calculating with the minimum load fluctuation of the distribution network as the goal to reduce the complexity of objective operation, use the Q-learning algorithm model to obtain the state-action mapping, and guide the access strategy of the charging load according to the current state, so as to control the access of the charging load and reduce the impact of the access charging load on the load fluctuation of the distribution network, and realize the peak shaving and valley filling of the distribution network load to improve the operation stability of the distribution network.
[0072] Figure 2 It is a schematic flowchart of another method for controlling orderly charging of a distribution network provided by the embodiments of the present invention. Refer to Figure 2 , including:
[0073] S210. Obtain the predicted power consumption load at each moment in the distribution transformer area, the number of charging loads at each moment, and the power of a single charging load;
[0074] Specifically, the power consumption load in the distribution transformer area includes the power consumption load of regular users, and the power consumption information acquisition unit obtains relevant data by reading the electricity meters in the distribution transformer area. The distribution transformer area is usually equipped with a power consumption load prediction function. Through the real-time data and historical data of the power consumption load in the distribution transformer area, the time series method, neural network prediction method, etc. are used to predict the power consumption load of the day, so as to obtain the predicted power consumption load at each moment in the distribution transformer area. Among them, it should be noted that the above power consumption load prediction method is related prior art and will not be elaborated here. The above moment can be a time point, for example, 24 hours a day, and each hour corresponding time point is used as a moment. The moment can also be a time period, for example, 24 hours a day, and each hour time period is used as a moment. There is no specific limitation here, and the corresponding moment definition can be set according to application requirements. In the embodiments of the present invention, taking the moment corresponding to the time period as an example, if the charging load is mainly the load when new energy vehicles are connected to the charging piles, the number of charging loads at each moment can be the number of charging piles connected within a time period. The power of a single charging load can be the power demand of the charging pile.
[0075] S220. Obtain the total load of the distribution transformer area at each moment according to the predicted power consumption load at each moment, the number of charging loads, and the power of a single charging load;
[0076] Specifically, the total load of the distribution transformer area at each moment includes the power consumption load of users and the connected charging load. Therefore, the total load of the distribution transformer area at each moment can be expressed as the sum of the predicted power consumption load at each moment and the connected charging load. Among them, the connected charging load can be expressed as the product of the number of charging loads and the power of a single charging load.
[0077] Therefore, the total load of the distribution substation area at each moment is expressed as:
[0078] P i = I i · P c + P load (i) (4)
[0079] Among them, I i is the number of charging loads at time Ti, P c is the power of a single charging load, P load (i) is the predicted electricity consumption load of the distribution substation area at time Ti;
[0080] S230. Establish an objective model based on the minimum load fluctuation between the total load and the average load of the distribution substation area at each moment.
[0081] Specifically, the minimum load fluctuation can reflect the change range of the distribution network load through the difference between the total load and the average load of the distribution network at each moment. Further, the standard deviation of the total load and the average load of the distribution network at each moment can be used. The value of the standard deviation is closer to the change range of the original data. The size of the standard deviation can directly reflect the degree of data dispersion, which helps to more intuitively evaluate the volatility of the data.
[0082] Therefore, the objective model is expressed as:
[0083]
[0084] Among them, P i is the total load of the distribution substation area at time Ti, P avg is the average load, and P MAX is the upper limit of the distribution substation area load. The total load of the distribution substation area being less than or equal to the upper limit of the distribution substation area load is used as a constraint condition for the objective model to avoid the total load of the distribution substation area exceeding the upper limit of the distribution substation area load, resulting in insufficient power of the distribution network or reduced stability of the distribution network.
[0085] S240. Set a triple including a state space S, an action space A, and a reward function R according to MDP. Among them, the state space S represents the total load of the distribution substation area at each moment, the action space A represents the action of scheduling the newly added charging load to a preset time period, and the reward function R represents the total number of charging loads after the action; obtain the Q-value function representing the action taken by the distribution substation area in the current system state according to the state space S and the action space A.
[0086] Specifically, the target model is converted into an MDP problem, and a Q-learning algorithm model is constructed. The Q-learning algorithm model evaluates the expected utility of taking a certain action in a given state by learning an action Q-value function, and then solves the MDP problem. According to the MDP setting, a triple including a state space, an action space, and a reward function can be expressed as {S, A, R}, where S is the state space, A is the action space, and R is the reward function.
[0087] The state space S represents the total load of the distribution transformer area at each moment. The system state at time slot t is set as:
[0088]
[0089] The action space A represents the actions that affect the system performance, and schedules the newly added charging load to a preset time period [T j , T j+L for charging. The action at time slot t is set as:
[0090]
[0091] That is to say, the newly added charging load is scheduled to the preset time period [T j , T j+L for charging, then the number of charging loads at the corresponding moment in this time period increases by one, and the number of charging loads at other moments remains unchanged. Therefore, if the preset time period is set as the load valley period of the distribution network, the load valley period can be fully utilized for charging to improve the power distribution utilization rate of the distribution network.
[0092] The reward function R sets the reward as choosing a certain action in the current system state S, obtaining the reward function for the next time slot t + 1, and entering the next system state S (t+1) . The reward function R can represent the total sum of the number of charging loads after the action:
[0093]
[0094] Taking the action space A in the state space S to obtain the Q-value function, that is, the Q-learning algorithm model, denoted as Q(s, a).
[0095] S250. Establish an equation relationship between the Q-value function and the target model; with the minimum value of the target model as the goal, according to the currently input state space S, solve the Q-value function to obtain the access strategy with the minimum value of the target model.
[0096] Specifically, taking the action space A in the state space S to obtain the Q-value function, let Q(s, a) = fp, establish a Q table, and set the Q value in the Q table to be initialized as P load(i), after performing an action in each time slot system state, the next system state and the obtained reward are observed. During the operation, the Q value corresponding to each pair of system state and action is updated. The above process is repeated to continuously optimize the Q value function until the performance standard is reached. At the same time, the updated Q table is also convenient for better selecting corresponding actions in the future.
[0097] By setting the goal of the MDP to obtain the policy with the minimum fp, a series of Q values corresponding to the minimum value of Q(s,a) are obtained. That is to say, a system state and an action correspond to a Q value. The Q value maps each state-action pair to a real number, which represents the expected minimum load fluctuation obtained after performing the action from this state. Therefore, according to the Q value that meets the minimum load fluctuation, that is, the target model is the minimum value, the corresponding state-action mapping is determined, and the access strategy of the charging load can be guided according to the current state.
[0098] Among them, the update of the Q value can be updated according to the update function, and the update function is:
[0099]
[0100] α is the learning rate, β is the discount rate; Q old (s (t) ,a (t) ) represents the Q value of taking the action in the current system state; Q new (s (t) ,a (t) ) represents the updated Q value in the current system state; R (t+1) is the reward function of the next system state after performing the action; s (t+1) represents the next system state entered after performing the current action; represents the minimum Q value among the actions taken in the next system state.
[0101] Optionally, after converting the target model into a Markov decision process and constructing a Q-learning algorithm model for orderly charging scheduling, it further includes:
[0102] Training the Q-learning algorithm model;
[0103] Among them, training the Q-learning algorithm model includes:
[0104] S1. Observe the current system state during the training period;
[0105] Specifically, perform initialization, initialize the learning rate α and the discount rate β, and the initial value of the Q table is 0, that is, the initial value of all state-action pairs is set to 0. In each time slot t within the training period Ts, observe the current system state s(t).
[0106] S2. Select an action according to the random policy to obtain the reward function R of the next system state;
[0107] Specifically, the system needs to randomly select an action based on the current state. The random action can be achieved by querying the Q-table or, when using deep learning, randomly generated by a neural network. Thus, the reward function R of the next system state can be obtained. (t+1) 。
[0108] S3. Obtain the next system state according to the action space A after the action;
[0109] Specifically, after the action, it means that the charging load is connected. Update the action space A after the action to obtain the quantity of the charging load at each moment. Combining Equation (4) and Equation (1), the next system state s can be obtained. (t+1) 。
[0110] S4. Update the Q-value function of the current system state according to the next system state and the reward function R of the next system state;
[0111] Specifically, substitute the next system state s (t+1) and the reward function R of the next system state (t+1) into the update function to obtain the updated Q-value function of the current system state.
[0112] S5. Repeat steps S2 to S4 until the training cycle ends.
[0113] Figure 3 The following is a schematic structural diagram of an orderly charging control device for a distribution network provided by an embodiment of the present invention. Refer to Figure 3 ,including:
[0114] The first model establishment module 110 is used to obtain data of the electricity consumption load and the charging load in the distribution substation area, and establish a target model for orderly charging in the distribution substation area with the minimum load fluctuation of the distribution substation area as the goal;
[0115] The second model establishment module 120 is used to convert the target model into a Markov decision process and construct a Q-learning algorithm model for orderly charging scheduling;
[0116] The calculation module 130 is used to obtain the access strategy of each charging load according to the Q-learning algorithm model and the target model.
[0117] Specifically, the distribution substation area supplies power to electricity users. Among them, electricity users include regular users and users composed of charging piles. The electricity consumption of regular users can be called the electricity load, and the electricity consumption connected to the charging piles can be called the charging load. Through the corresponding acquisition unit of the distribution substation area, historical data or real-time data related to the electricity load and the charging load can be acquired. Considering the volatility of various distributed resources and the load volatility caused by electric vehicles connecting to the charging piles, the first model establishment module 110 establishes a target model using the minimum load fluctuation of the distribution substation area, and reduces the change range between the load values of the distribution network by restricting the minimum load fluctuation, thereby improving the transmission efficiency of the power grid. Exemplarily, the minimum load fluctuation can be reflected by the difference between the total load and the average load of the distribution network at each moment, which reflects the change range of the distribution network load. Further, the standard deviation of the total load and the average load of the distribution network at each moment can be used. The value of the standard deviation is closer to the change range of the original data, and the size of the standard deviation can directly reflect the degree of data dispersion, which helps to more intuitively evaluate the volatility of the data.
[0118] The second model establishment module 120 converts the target model into an MDP problem and constructs a Q-learning algorithm model. The Q-learning algorithm model evaluates the expected utility of taking a certain action in a given state by learning an action Q-value function, and then solves the MDP problem. Exemplarily, according to the MDP setting, a triple including a state space, an action space, and a reward function can be represented as {S, A, R}, where S is the state space, A is the action space, and R is the reward function.
[0119] The state space S represents the total load of the distribution substation area at each moment. The system state at time slot t is set as:
[0120]
[0121] The action space A represents the actions that affect the system performance, and the action of scheduling the newly added charging load to a preset time period [T j , T j+L for charging. The action at time slot t is set as:
[0122]
[0123] That is to say, if the newly added charging load is scheduled to the preset time period [T j , T j+L for charging, the number of charging loads at the corresponding moment in this time period increases by one, and the number of charging loads at other moments remains unchanged. Therefore, if the preset time period is set as the load valley period of the distribution network, the load valley period can be fully utilized for charging in the action result, thereby improving the power distribution utilization rate of the distribution network.
[0124] The reward function R sets the reward for selecting an action in the current system state S, obtains the reward function for the next time slot t+1, and enters the next system state S (t+1) . The reward function R can represent the total number of charging loads after the action:
[0125]
[0126] Taking the action space A in the state space S obtains the Q-value function, denoted as Q(s,a).
[0127] The calculation module 130 establishes a relationship between the Q-value function and the target model. The target of the Q-value function can be set to the minimum value of the target model for solution to obtain the Q-value. The Q-value maps each state-action pair to a real number, which represents the minimum value of the expected load fluctuation obtained after executing the action from this state. In the Q-Learning algorithm, the Q-value continuously interacts with the environment, observes the rewards obtained by taking different actions in different states, and updates the Q-value accordingly to gradually obtain the optimal value of the target model. Therefore, according to the Q-value that conforms to the target model, the state-action mapping can be obtained, so as to guide how to select the access strategy of the charging load according to the current state.
[0128] Figure 4 FIG. is a schematic structural diagram of an orderly charging control system for a distribution network provided by an embodiment of the present invention. Refer to Figure 4 , an edge computing intelligent device and a communication module 420, wherein the edge computing intelligent device includes an electricity consumption information acquisition unit 411, an electricity load acquisition unit 412, a charging load acquisition unit 413, an orderly charging control device 414 for the distribution network, and a database 415;
[0129] Specifically, the edge computing intelligent device establishes a communication connection with the external smart meter 430 and the charging pile 440 through the communication module 420. The relevant data processed by the edge computing intelligent device can be output to the server or external control device through the communication interface 416, enabling remote data monitoring and device control. Among them, the communication module 420 can provide communication connections with the smart meter 430 and the charging pile 440 in the form of high-speed power line communication (HPLC) or wireless communication. The electricity consumption information collection unit 411 can collect the historical data and real-time data of the smart meters in the distribution area through the communication module 420 and store them in the database 415; the electricity load collection unit 412 predicts the daily electricity load of the distribution area based on the historical data and real-time data of the electricity load of the charging pile 440 to obtain the predicted electricity load at each moment; exemplarily, based on the real-time data and historical data of the electricity load in the distribution area, methods such as time series method and neural network prediction method are used to predict the daily electricity load, so as to obtain the predicted electricity load at each moment in the distribution area. It should be noted that the above electricity load prediction method can be related existing technologies and will not be elaborated here. The charging load collection unit 413 is used to collect the quantity of the charging load at each moment and the power of a single charging load of the charging pile 440 in the distribution area and store them in the database 415;
[0130] The orderly charging control device 414 of the distribution network can obtain the data of the electricity load and charging load of the distribution area from the database 415, establish an objective model for orderly charging in the distribution area with the minimum load fluctuation of the distribution area as the goal; convert the objective model into a Markov decision process, construct a Q-learning algorithm model for orderly charging scheduling, consider calculating with the minimum load fluctuation of the distribution network as the goal to reduce the complexity of the objective operation, use the Q-learning algorithm model to obtain the state-action mapping, and guide the access strategy of the charging load according to the current state, so as to control the access of the charging load to reduce the impact of the access charging load on the load fluctuation of the distribution network and achieve peak shaving and valley filling of the distribution network load to improve the operation stability of the distribution network.
[0131] Figure 5FIG. shows a schematic structural diagram of an electronic device 10 that can be used to implement an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0132] As Figure 5 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other via a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.
[0133] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0134] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the orderly charging control of the distribution network.
[0135] In some embodiments, the orderly charging control of the distribution network can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the orderly charging control of the distribution network described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the orderly charging control of the distribution network in any other suitable manner (e.g., by means of firmware).
[0136] The various embodiments of the systems and techniques described above can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems on a chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0137] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer programs are executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer programs can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0138] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0139] In order to provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0140] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0141] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The client-server relationship is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0142] It should be understood that various forms of processes shown above can be used, steps can be reordered, added or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.
[0143] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or equivalently replace some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An orderly charging control method for a distribution network, characterized in that, Including: Obtain the data of the power consumption load and charging load in the distribution substation area, and establish an objective model for orderly charging in the distribution substation area with the minimum load fluctuation of the distribution substation area as the goal. Convert the objective model into a Markov decision process, and construct a Q-learning algorithm model for orderly charging scheduling. Obtain the access strategies of each charging load according to the Q-learning algorithm model and the objective model. Among them, converting the objective model into a Markov decision process and constructing a Q-learning algorithm model for orderly charging scheduling includes: Set a triple including a state space, an action space, and a reward function according to the Markov decision process. Among them, the state space represents the total load of the distribution substation area at each moment, the action space represents the action of scheduling the newly added charging load to be charged during a preset time period, and the reward function represents the total number of charging loads after the action. According to the state space and the action space, obtain the Q-value function representing the action taken by the distribution substation area in the current system state.
2. The orderly charging control method for a distribution network according to claim 1, wherein Obtain the data of the power consumption load and charging load in the distribution substation area, and establish an objective model for orderly charging in the distribution substation area with the minimum load fluctuation of the distribution substation area as the goal, including: Obtain the predicted power consumption load at each moment in the distribution substation area, the number of charging loads at each moment, and the power of a single charging load. Obtain the total load of the distribution substation area at each moment according to the predicted power consumption load at each moment, the number of charging loads at each moment, and the power of a single charging load. Obtain the minimum load fluctuation according to the difference between the total load of the distribution substation area at each moment and the average load, thereby establishing the objective model.
3. The orderly charging control method for a distribution network according to claim 2, wherein, The total load of the distribution substation area at each moment is expressed as: I i is the number of charging loads at time Ti, P c is the power of a single charging load, P load (i) is the predicted electricity consumption load of the distribution substation area at time Ti; The objective model is expressed as: Among them, P i is the total load of the distribution transformer area at time Ti, P avg is the average load, and P MAX is the upper limit of the load of the distribution transformer area, and T is the total amount of time.
4. The orderly charging control method for a distribution network according to claim 1, characterized in that, The preset time period is the low load period of the total load of the distribution substation area.
5. The orderly charging control method for a distribution network according to claim 4, characterized in that, Obtain the access strategies of each charging load according to the Q-learning algorithm model and the objective model, including: Establish an equation relationship between the Q-value function and the objective model. With the minimum value of the objective model as the goal, solve the Q-value function according to the currently input state space, thereby obtaining the access strategy when the objective model is the minimum value.
6. The orderly charging control method for a distribution network according to claim 5, characterized in that, After converting the objective model into a Markov decision process and constructing a Q-learning algorithm model for orderly charging scheduling, it further includes: Train the Q-learning algorithm model. Among them, training the Q-learning algorithm model includes: S1. Observe the current system state during the training period. S2. Select an action according to a random policy and obtain the reward function of the next system state. S3. Obtain the next system state according to the action space of the current system state after the action. S4. Update the Q-value function of the current system state according to the next system state and the reward function of the next system state. S5. Repeat steps S2 to S4 until the training period ends.
7. The orderly charging control method for a distribution network according to claim 6, characterized in that, Update the Q-value function of the current system state according to the next system state and the reward function R of the next system state, including: Update the Q-value of the current system state according to the update function, and the update function is: Among them, α is the learning rate and β is the discount rate; represents the Q value of taking an action in the current system state; represents the updated Q value in the current system state; is the reward function of the next system state after executing the action; represents the next system state entered after executing the current action; represents the minimum Q value among the actions taken in the next system state.
8. An orderly charging control device for a distribution network, characterized in that, Including: The first model establishment module is used to obtain the data of the electricity consumption load and charging load in the distribution transformer area, and establish an objective model for orderly charging in the distribution transformer area with the minimum load fluctuation of the distribution transformer area as the goal; The second model establishment module is used to convert the objective model into a Markov decision process and construct a Q-learning algorithm model for orderly charging scheduling; Among them, converting the objective model into a Markov decision process and constructing a Q-learning algorithm model for orderly charging scheduling includes: Setting a triple including a state space, an action space, and a reward function according to the Markov decision process. Among them, the state space represents the total load of the distribution transformer area at each moment, the action space represents the action of scheduling the newly added charging load to a preset time period for charging, and the reward function represents the total number of charging loads after the action; Obtaining a Q-value function representing the action taken by the distribution transformer area in the current system state according to the state space and the action space; The calculation module is used to obtain the access strategy of each charging load according to the Q-learning algorithm model and the objective model.
9. An orderly charging control system for a distribution network, characterized in that, It includes an edge computing intelligent device and a communication module. Among them, the edge computing intelligent device includes an electricity consumption information collection unit, an electricity consumption load collection unit, a charging load collection unit, an orderly charging control device for the distribution network, and a database; The communication module is used to establish a communication connection between the edge computing intelligent device and external smart meters and charging piles; The electricity consumption information collection unit is used to collect the historical data and real-time data of the smart meters in the distribution transformer area and store them in the database; The electricity consumption load collection unit is used to predict the daily electricity consumption load in the distribution transformer area according to the historical data and real-time data of the electricity consumption load of the charging piles, and obtain the predicted electricity consumption load at each moment; The charging load collection unit is used to collect the number of charging loads and the power of a single charging load at each moment of the charging piles in the distribution transformer area and store them in the database; The orderly charging control device for the distribution network is used to obtain the data of the electricity consumption load and charging load in the distribution transformer area from the database, and establish an objective model for orderly charging in the distribution transformer area with the minimum load fluctuation of the distribution transformer area as the goal; convert the objective model into a Markov decision process and construct a Q-learning algorithm model for orderly charging scheduling; obtain the access strategy of each charging load according to the Q-learning algorithm model and the objective model; Among them, converting the objective model into a Markov decision process and constructing a Q-learning algorithm model for orderly charging scheduling includes: Setting a triple including a state space, an action space, and a reward function according to the Markov decision process. Among them, the state space represents the total load of the distribution transformer area at each moment, the action space represents the action of scheduling the newly added charging load to a preset time period for charging, and the reward function represents the total number of charging loads after the action; Obtaining a Q-value function representing the action taken by the distribution transformer area in the current system state according to the state space and the action space.
Citation Information
Patent Citations
Distributed electric vehicle real-time optimization scheduling method and system, terminal and medium
CN113515884A
Electric vehicle power distribution network regulation and control method based on V2G technology
CN115441484A