Electric vehicle charging and discharging scheduling method and device
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 国电投锦润新能源科技有限公司
- Filing Date
- 2023-12-28
- Publication Date
- 2026-08-07
AI Technical Summary
[0004]然而,上述方案存在无法满足车辆充电需求的问题
[0049]This application provides a method and apparatus for scheduling the charging and discharging of electric vehicles. It constructs a reward model to indicate the scheduling results at different observation times when different scheduling strategies are employed. The scheduling strategy indicates the charging and discharging strategy of at least one vehicle in the power distribution network. The scheduling result indicates the impact on the power distribution network and/or at least one vehicle after charging and discharging according to the scheduling strategy. Reinforcement learning is performed based on the reward model to determine a target scheduling strategy at at least one observation time. The charging and discharging of at least one vehicle in the power distribution network is then scheduled according to the target scheduling strategy at at least one observation time. This application can determine the target scheduling strategy when the impact meets the requirements through the reward model, thereby maximizing the performance stability of the power distribution network to meet vehicle charging needs. Furthermore, vehicles can be scheduled as energy storage to discharge to charging piles. In this way, discharging vehicles can provide power to the power distribution network when there are a large number of charging vehicles, enabling the power distribution network to meet vehicle charging needs.
Smart Images

Figure CN117841764B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of new energy technology, and in particular to a method and equipment for scheduling the charging and discharging of electric vehicles. Background Technology
[0002] With the rapid development of new energy technologies, new energy vehicles have emerged, which can be powered by electricity. These vehicles can also be charged from the power distribution network to obtain electrical energy.
[0003] In existing technologies, power distribution networks can be set up with multiple charging stations, each providing a preset number of charging piles. Each charging pile can provide charging functionality to one new energy vehicle at the same observation time. New energy vehicles can then find an available charging pile after entering the station.
[0004] However, the above solutions have the problem of failing to meet the charging needs of vehicles. Summary of the Invention
[0005] This application provides a method and equipment for scheduling the charging and discharging of electric vehicles, which can meet the needs of stable operation of the power distribution network and vehicle charging as much as possible.
[0006] In a first aspect, this application provides a method for scheduling the charging and discharging of an electric vehicle, the method comprising:
[0007] A reward model is constructed, which is used to indicate the scheduling results corresponding to different scheduling strategies at different observation times. The scheduling strategy is used to indicate the charging and discharging strategy of at least one vehicle in the power distribution network. The scheduling result is used to indicate the impact on the power distribution network and / or the at least one vehicle after the charging and discharging scheduling is carried out according to the scheduling strategy.
[0008] Reinforcement learning is performed based on the reward model to determine the target scheduling strategy at at least one observation time.
[0009] The charging and discharging of the at least one vehicle is scheduled in the power distribution network according to the target scheduling strategy at the at least one observation time.
[0010] Optionally, the scheduling result includes at least one of the following: electricity cost, discharge revenue, and the degree of state change of the distribution network.
[0011] Optionally, the electricity cost includes at least one of the following: the electricity cost of the power distribution network and the electricity cost of the at least one vehicle; the discharge revenue includes at least one of the following: the discharge revenue of the power distribution network and the discharge revenue of the at least one vehicle; and the degree of state change includes at least one of the following: the load power change of the power distribution network and the peak-valley difference of the active power of the power distribution network.
[0012] Optionally, the step of performing reinforcement learning based on the reward model to determine the target scheduling strategy at at least one observation time includes:
[0013] Under preset constraints, a target scheduling strategy is determined based on the reward model for at least one observation time. The preset constraints include vehicle constraints and power distribution network constraints. The vehicle constraints are used to constrain at least one of the following of the vehicle: battery level, power, and status. The power distribution network constraints are used to constrain the power of the power distribution network.
[0014] Optionally, the vehicle constraint conditions include at least one of the following: the vehicle's battery level is within a preset battery level range, the vehicle's charging power is within a preset charging power range, the vehicle's discharging power is within a preset discharging power range, and the vehicle's charging and discharging state at the same time is a preset state. The preset state includes one of the following: charging state, discharging state, and idle state. The idle state is a state other than the charging state and the discharging state.
[0015] The power distribution network constraints include at least one of the following: the sum of the squares of the active power and reactive power of the power distribution network is less than the square of the maximum apparent power; and the active power of the power distribution network is within a preset active power range.
[0016] Optionally, determining the target scheduling strategy for the at least one observation time based on the reward model under preset constraints includes:
[0017] Based on the scheduling results, multiple learning objectives are determined under the preset constraints. The multiple learning objectives include: a first learning objective of minimizing the degree of state change, a second learning objective of minimizing the electricity cost, and a third learning objective of maximizing the discharge revenue.
[0018] The priority of each learning objective is determined, and multiple learning tasks based on the multiple learning objectives are constructed according to the priority of each learning objective. The reward models corresponding to different learning tasks are different, and one learning task corresponds to one or more learning objectives.
[0019] The multiple learning tasks are performed to determine the target scheduling strategy.
[0020] Optionally, constructing multiple learning tasks based on the priority of each learning objective includes:
[0021] When the priority of the first learning objective is higher than the priority of the second learning objective, and the priority of the second learning objective is higher than the priority of the third learning objective, the following multiple learning tasks are constructed: a first learning task that minimizes the degree of state change under preset constraints, a second learning task that minimizes the electricity cost based on minimizing the degree of state change, and a third learning task that maximizes the discharge benefit based on minimizing the degree of state change and the electricity cost.
[0022] Optionally, performing the plurality of learning tasks to determine the target scheduling strategy includes:
[0023] In each learning task, an optimal action value model is constructed based on the reward model of the learning task. The optimal action value model is used to indicate the scheduling reward corresponding to the scheduling strategy and state within an observation time period. The scheduling reward includes the expected value of the scheduling result in the observation time period. The scheduling result corresponding to the first learning task is the degree of state change. The scheduling result corresponding to the second learning task is the degree of state change and the electricity cost. The scheduling result corresponding to the third learning task is the degree of state change, the electricity cost, and the discharge benefit.
[0024] Based on the optimal action value model, the scheduling strategy corresponding to the maximum scheduling reward is determined as the target scheduling strategy.
[0025] Secondly, this application provides an electric vehicle charging and discharging scheduling device, the device comprising:
[0026] The model building module is used to build a reward model, which is used to indicate the scheduling results corresponding to different scheduling strategies at different observation times. The scheduling strategy is used to indicate the charging and discharging strategy of at least one vehicle in the power distribution network. The scheduling result is used to indicate the impact on the power distribution network and / or the at least one vehicle after the charging and discharging scheduling is carried out according to the scheduling strategy.
[0027] The target scheduling strategy determination module is used to perform reinforcement learning based on the reward model to determine the target scheduling strategy at at least one observation time.
[0028] The scheduling module is used to schedule the charging and discharging of at least one vehicle in the power distribution network according to the target scheduling strategy at the at least one observation time.
[0029] Optionally, the scheduling result includes at least one of the following: electricity cost, discharge revenue, and the degree of state change of the distribution network.
[0030] Optionally, the electricity cost includes at least one of the following: the electricity cost of the power distribution network and the electricity cost of the at least one vehicle; the discharge revenue includes at least one of the following: the discharge revenue of the power distribution network and the discharge revenue of the at least one vehicle; and the degree of state change includes at least one of the following: the load power change of the power distribution network and the peak-valley difference of the active power of the power distribution network.
[0031] Optionally, the target scheduling strategy determination module is further configured to:
[0032] Under preset constraints, a target scheduling strategy is determined based on the reward model for at least one observation time. The preset constraints include vehicle constraints and power distribution network constraints. The vehicle constraints are used to constrain at least one of the following of the vehicle: battery level, power, and status. The power distribution network constraints are used to constrain the power of the power distribution network.
[0033] Optionally, the vehicle constraint conditions include at least one of the following: the vehicle's battery level is within a preset battery level range, the vehicle's charging power is within a preset charging power range, the vehicle's discharging power is within a preset discharging power range, and the vehicle's charging and discharging state at the same time is a preset state. The preset state includes one of the following: charging state, discharging state, and idle state. The idle state is a state other than the charging state and the discharging state.
[0034] The power distribution network constraints include at least one of the following: the sum of the squares of the active power and reactive power of the power distribution network is less than the square of the maximum apparent power; and the active power of the power distribution network is within a preset active power range.
[0035] Optionally, the target scheduling strategy determination module is further configured to:
[0036] Based on the scheduling results, multiple learning objectives are determined under the preset constraints. The multiple learning objectives include: a first learning objective of minimizing the degree of state change, a second learning objective of minimizing the electricity cost, and a third learning objective of maximizing the discharge revenue.
[0037] The priority of each learning objective is determined, and multiple learning tasks based on the multiple learning objectives are constructed according to the priority of each learning objective. The reward models corresponding to different learning tasks are different, and one learning task corresponds to one or more learning objectives.
[0038] The multiple learning tasks are performed to determine the target scheduling strategy.
[0039] Optionally, the target scheduling strategy determination module is further configured to:
[0040] When the priority of the first learning objective is higher than the priority of the second learning objective, and the priority of the second learning objective is higher than the priority of the third learning objective, the following multiple learning tasks are constructed: a first learning task that minimizes the degree of state change under preset constraints, a second learning task that minimizes the electricity cost based on minimizing the degree of state change, and a third learning task that maximizes the discharge benefit based on minimizing the degree of state change and the electricity cost.
[0041] Optionally, the target scheduling strategy determination module is further configured to:
[0042] In each learning task, an optimal action value model is constructed based on the reward model of the learning task. The optimal action value model is used to indicate the scheduling reward corresponding to the scheduling strategy and state within an observation time period. The scheduling reward includes the expected value of the scheduling result in the observation time period. The scheduling result corresponding to the first learning task is the degree of state change. The scheduling result corresponding to the second learning task is the degree of state change and the electricity cost. The scheduling result corresponding to the third learning task is the degree of state change, the electricity cost, and the discharge benefit.
[0043] Based on the optimal action value model, the scheduling strategy corresponding to the maximum scheduling reward is determined as the target scheduling strategy.
[0044] Thirdly, this application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;
[0045] The memory stores computer-executed instructions;
[0046] The processor executes computer execution instructions stored in the memory to implement the method as described in the first aspect.
[0047] Fourthly, this application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the method described in the first aspect.
[0048] Fifthly, this application provides a computer program product for implementing the method of the first aspect.
[0049] This application provides a method and apparatus for scheduling the charging and discharging of electric vehicles. It constructs a reward model to indicate the scheduling results at different observation times when different scheduling strategies are employed. The scheduling strategy indicates the charging and discharging strategy of at least one vehicle in the power distribution network. The scheduling result indicates the impact on the power distribution network and / or at least one vehicle after charging and discharging according to the scheduling strategy. Reinforcement learning is performed based on the reward model to determine a target scheduling strategy at at least one observation time. The charging and discharging of at least one vehicle in the power distribution network is then scheduled according to the target scheduling strategy at at least one observation time. This application can determine the target scheduling strategy when the impact meets the requirements through the reward model, thereby maximizing the performance stability of the power distribution network to meet vehicle charging needs. Furthermore, vehicles can be scheduled as energy storage to discharge to charging piles. In this way, discharging vehicles can provide power to the power distribution network when there are a large number of charging vehicles, enabling the power distribution network to meet vehicle charging needs. Attached Figure Description
[0050] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0051] Figure 1 This is a schematic diagram of an electric vehicle charging scenario provided in an embodiment of this application;
[0052] Figure 2 This is a flowchart illustrating the steps of a charging and discharging scheduling method for electric vehicles provided in an embodiment of this application.
[0053] Figure 3 This is a flowchart of another electric vehicle charging and discharging scheduling method provided in the embodiments of this application;
[0054] Figure 4 This is a structural block diagram of an electric vehicle charging and discharging scheduling device provided in an embodiment of this application;
[0055] Figure 5 This is a structural block diagram of an electronic device provided in an embodiment of this application.
[0056] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0057] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0058] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with relevant laws, regulations and standards, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0059] This application embodiment can be used to schedule the charging and discharging of one or more vehicles to determine whether each vehicle is charging or discharging, and the charging and discharging power. It should be noted that the vehicles in this application embodiment are powered by electrical energy and may include purely electric vehicles, hybrid vehicles, etc.
[0060] Figure 1 This is a schematic diagram of an electric vehicle charging scenario provided in an embodiment of this application. Referring to Figure 1, an exemplary charging station provides three charging piles A1, A2, and A3, and three vehicles B1, B2, and B3 are charging at charging piles A1, A2, and A3 respectively. At this time, vehicle B4 needs to wait.
[0061] In existing technologies, vehicles charge based on their power needs; for example, when a vehicle's battery is low, it searches for an available charging station. This method causes the performance of the power distribution network to vary with the number of vehicles charging, leading to network instability. When there are many vehicles charging, this may result in the inability to meet the charging needs of all vehicles.
[0062] To address the aforementioned technical problems, embodiments of this application construct a reward model to indicate the scheduling results corresponding to different scheduling strategies sampled at different observation times, thereby measuring the impact of the scheduling strategies on vehicles and the power distribution network. This reward model can then determine the target scheduling strategy when the impact meets the requirements for charging and discharging scheduling. In this way, the performance stability of the power distribution network can be improved as much as possible to meet the vehicle charging needs.
[0063] As can be seen, the embodiments of this application can not only schedule vehicles as loads to charge from charging piles, but also schedule vehicles as energy storage to discharge to charging piles. In this way, when the number of charging vehicles is large, the discharging vehicles can provide power to the power distribution network, so that the power distribution network can meet the charging needs of the vehicles.
[0064] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0065] Figure 2 This is a flowchart illustrating the steps of a charging and discharging scheduling method for an electric vehicle provided in an embodiment of this application. (Refer to...) Figure 2 As shown, the electric vehicle charging and discharging scheduling method of this application may include:
[0066] S201: Construct a reward model, which is used to indicate the corresponding scheduling results when different scheduling strategies are adopted at different observation times. The scheduling strategy is used to indicate the charging and discharging strategy of at least one vehicle in the power distribution network. The scheduling result is used to indicate the impact on the power distribution network and / or at least one vehicle after the charging and discharging scheduling is carried out according to the scheduling strategy.
[0067] The reward model can be understood as a reward function, and the scheduling result can be understood as the reward brought by the charging and discharging strategy. Therefore, a larger scheduling result indicates a better charging and discharging strategy; a smaller scheduling result indicates a worse charging and discharging strategy. The embodiments of this application aim to perform charging and discharging scheduling with a target scheduling strategy to maximize the scheduling result, that is, to maximize the positive impact of charging and discharging scheduling on the power distribution network and vehicles, and the greater the positive impact, the better.
[0068] The input to the above reward model can be the target scheduling strategy at each observation time, which indicates the charging and discharging information of each vehicle. The charging and discharging information of each vehicle includes: whether the vehicle is charging or discharging at that observation time, the charging power used when charging, and the discharging power used when discharging.
[0069] In this embodiment, whether the I-th vehicle is charging at observation time T can be represented by a variable CD(I, T). When CD(I, T) is 1, it means that the I-th vehicle is charging at observation time T; when CD(I, T) is 0, it means that the I-th vehicle is not charging at observation time T.
[0070] Whether the I-th vehicle discharges at observation time T can be represented by a variable FD(I,T). When FD(I,T) is 1, it means that the I-th vehicle discharges at observation time T; when FD(I,T) is 0, it means that the I-th vehicle does not discharge at observation time T.
[0071] For the I-th vehicle, its charging power at observation time T can be represented by the variable CGL(I,T), and its discharging power at observation time T can be represented by the variable FGL(I,T).
[0072] Charging and discharging scheduling according to the above-described scheduling strategy will impact the power distribution network and vehicles. This impact can be evaluated using indicators of any dimension. For example, commonly used indicators for evaluating this impact may include at least one of the following: electricity cost, discharging revenue, and the degree of change in the state of the power distribution network. In this way, the impact of the scheduling strategy can be evaluated from multiple dimensions, improving the comprehensiveness of the reward model's assessment of the impact and helping to ensure better scheduling results when charging and discharging is performed using the target scheduling strategy.
[0073] Electricity cost indicates the amount of electricity consumed or the corresponding economic cost. Discharge revenue indicates the amount of discharge or the corresponding economic benefit. The degree of state change can be assessed from one or more dimensions to evaluate the state changes of the distribution network.
[0074] As can be seen from the above, the embodiments of this application not only minimize electricity costs and maximize discharge benefits, but also minimize the degree of change in the state of the power distribution network. Thus, it can reduce costs, increase benefits, and ensure the stability of the power distribution network.
[0075] In some implementations, the aforementioned electricity cost includes at least one of the following: the electricity cost of the distribution network and the electricity cost of at least one vehicle. Discharge revenue includes at least one of the following: the discharge revenue of the distribution network and the discharge revenue of at least one vehicle. The degree of change in the state of the distribution network includes at least one of the following: the change in load power of the distribution network and the peak-to-valley difference in active power of the distribution network. Thus, electricity cost, discharge revenue, and the degree of change in the state of the distribution network can be determined from multiple perspectives, which helps to further improve the comprehensiveness of the dispatching results and thereby improve the effectiveness of charge and discharge dispatching.
[0076] The electricity cost of the aforementioned power distribution network used for electric vehicle charging can be calculated using the following formula:
[0077]
[0078] Wherein, PCB(T) is the electricity cost of the distribution network at time T. PGJ(T) is the unit electricity price paid by the distribution network to the main grid at time T. IMAX is the total number of vehicles that can be dispatched in this embodiment. CD(I,T) indicates whether the I-th vehicle is charging at time T, with 1 indicating charging. CGL(I,T) indicates the charging power of the I-th vehicle at time T. In this embodiment, the dispatch time segment can be divided into TNUM observation time segments, with the start time of each observation time segment being an observation time. TC is the ratio of the number of dispatch time segments to the number of observation time segments TNUM, called the observation duration. For example, if the dispatch time segment is 1 hour, then 1 hour can be divided into 60 observation time segments with an observation duration of 1 minute each.
[0079] It can be seen that the above formula (1) is used to determine the cost required for the distribution network to purchase electricity within the observation period of time T based on the scheduling strategy at time T, which is used as the electricity cost of the distribution network. Of course, in some implementations, the above formula (1) can also be used to determine the cost of electricity purchase within the observation period of time T. The electricity cost of the distribution network is defined as the amount of electricity that the distribution network needs to provide within the observation period of time T, based on the dispatch strategy at time T.
[0080] The electricity cost of the distribution network is positively correlated with the number of charging vehicles, the cost of purchasing electricity for the distribution network, the charging power of the vehicles, and the observation duration. The higher the electricity cost of the distribution network, the greater the number of charging vehicles, the greater the cost of purchasing electricity for the distribution network, the greater the charging power of the vehicles, and the longer the observation duration.
[0081] The electricity cost for at least one of the aforementioned vehicles can be calculated using the following formula:
[0082]
[0083] Where CCB(T) is the electricity cost of at least one vehicle at time T. CGJ(I,T) is the unit electricity price paid by the I-th vehicle to the distribution network when it is charging at time T.
[0084] It can be seen that the above formula (2) is used to determine the cost required for all vehicles to charge within the observation period at time T based on the scheduling strategy at time T, which is used as the vehicle's electricity cost. The vehicle's electricity cost is positively correlated with the number of charging vehicles, the unit electricity price paid by the vehicle to the distribution network when charging at time T, the vehicle's charging power, and the observation period. The higher the vehicle's electricity cost is, the greater the number of charging vehicles, the higher the unit electricity price paid by the vehicle to the distribution network when charging at time T, the higher the vehicle's charging power, and the longer the observation period.
[0085] The discharge revenue of the distribution network can be calculated using the following formula:
[0086]
[0087] Where PSY(T) is the discharge revenue of the distribution network at time T, CGJ(I,T) is the unit electricity price paid by the I vehicle to the distribution network when charging at time T, and PGJ(T) is the unit electricity price paid by the distribution network to the main grid when purchasing electricity at time T.
[0088] As can be seen from the above formula (3), the discharge revenue of the distribution network is used to represent the revenue obtained by the distribution network from the charging process of all dispatchable vehicles within the observation period TC corresponding to time T. The discharge revenue of the distribution network is positively correlated with whether each discharging vehicle is charging and the charging power, the observation period, and the unit electricity price difference corresponding to the distribution network. The unit electricity price difference corresponding to the distribution network is the difference between the unit electricity price paid by the vehicle to the distribution network when charging at time T and the unit electricity price paid by the distribution network to the main grid when purchasing electricity at time T. Therefore, the more charging vehicles, the greater the charging power, the longer the observation period, and the greater the unit electricity price difference of the distribution network, the greater the discharge revenue of the distribution network.
[0089] The discharge benefit of an electric vehicle can be calculated using the following formula:
[0090]
[0091] Where CSY(T) represents the discharge revenue of all dispatchable vehicles at time T, CFJ(I,T) is the unit electricity price paid by the distribution network to the vehicle I when it discharges to the distribution network at time T, CGJ(I,T') is the unit electricity price paid by the vehicle I when it charges to the distribution network at historical time T', FD(I,T) indicates whether the vehicle I discharges at time T, and FGL(I,T) is the discharge power of the vehicle I at time T.
[0092] As can be seen from the above formula (4), the discharge revenue of a vehicle is the revenue obtained by all schedulable vehicles discharging during the discharge process within the observation time period TC corresponding to time T. This discharge revenue is positively correlated with whether each vehicle discharges, its discharge power, the observation duration of the observation time period, and the unit electricity price difference corresponding to the vehicle. The unit electricity price difference corresponding to the vehicle is the difference between the unit electricity price paid by the distribution network to the vehicle when it discharges to the distribution network and the unit electricity price paid by the vehicle to the distribution network when it is charging at the historical time T'. Therefore, the more vehicles discharging, the greater the discharge power, the longer the observation duration, and the greater the unit electricity price difference of the vehicle, the greater the discharge revenue of the vehicle.
[0093] The change in load power in a power distribution network can be calculated using the following formula:
[0094]
[0095] Where FHBH(T) is the change in load power of the distribution network at time T, and TNUM is the preset number of time points for statistically analyzing the change. TNUM can be 24, and T' is a historical time point at or before time T. The TNUM time points statistically analyzed in formula (5) include time T and TNUM-1 time points before time T. FH(T') is the load power of the distribution network at time T', and PFH(T) is the average load power of the distribution network at time T.
[0096] As can be seen from the above formula (5), the change in load power of the distribution network at time T is the degree of deviation between the load power of the distribution network and the average load power in the historical time period before and after time T, which is the variance of the load power of the distribution network in that historical time period.
[0097] In some implementations, the above FH(T') is calculated using the following formula:
[0098]
[0099] Where JCFH(T') is the base load power of the distribution network at time T', which is the power required for the distribution network to operate when no charging vehicles are present. CD(I,T') indicates whether the I-th vehicle is charging at time T'. CGL(I,T') is the charging power of the I-th vehicle at time T'.
[0100] As can be seen from the above formula (6), the load power of the distribution network at any time is the sum of the basic load power of the distribution network at that time and the total charging power of all dispatchable vehicles when charging.
[0101] PFH(T) in the above formula (5) can be calculated using the following formula:
[0102]
[0103] It can be seen that the average load power of the distribution network at time T is the average of the load power at time T and the TNUM-1 times preceding time T.
[0104] The peak-to-valley difference of active power in a power distribution network at time T can be represented as the difference between the maximum and minimum active power of the network during a historical time period based on time T. Here, active power can be the power used for vehicle charging. Therefore, the peak-to-valley difference of active power in the power distribution network at time T can be calculated using the following formula:
[0105] YGBH(T) = MAX[YGL(T')] T'=T-TNUM+1~T ]-MIN[YGL(T') T'=T-TNUM+1~T (8)
[0106] Where YGBH(T) is the peak-to-valley difference of active power in the distribution network at time T. YGL(T') is used to indicate the active power of the distribution network at time T'. MAX[YGL(T')] T'=T-TNUM+1~T MIN[YGL(T')] represents the maximum active power of the distribution network at time T and within the preceding TNUM-1 time intervals. T'=T-TNUM+1~T [T] represents the minimum active power of the distribution network at time T and within the preceding TNUM-1 time points. Therefore, the peak-to-valley difference of the active power of the distribution network at time T is the difference between the maximum and minimum active power of the distribution network within the preceding historical time points.
[0107] The active power in formula (8) above can be calculated using the following formula:
[0108]
[0109] Where YGL(T') is the active power of the power distribution network at time T'. CD(I,T') is used to indicate whether the I-th vehicle is charging at that time, and CGL(I,T') is the charging power of the I-th vehicle at time T'.
[0110] After determining the electricity cost, discharge revenue, load power variation of the distribution network and peak-valley difference of the distribution network through the above formulas (1) to (9), the reward model can be determined based on the evaluation model of electricity cost, the evaluation model of discharge revenue, the evaluation model of load power variation of the distribution network and the evaluation model of peak-valley difference of the distribution network.
[0111] In some implementations, the scheduling result can be obtained by weighting the electricity cost, discharge revenue, load power variation of the distribution network, and peak-valley difference of active power in the distribution network. Therefore, a reward model can be constructed by combining the weights corresponding to the electricity cost, discharge revenue, load power variation of the distribution network, and peak-valley difference of active power in the distribution network, along with the evaluation model.
[0112] Specifically, the reward model can be referenced as follows:
[0113]
[0114] Here, JL(T) represents the scheduling result obtained based on the scheduling policy at time T. Based on the above reward model, reinforcement learning can be performed using S202 to determine the target scheduling policy.
[0115] In the above formula (10), W1 to W6 are the weights corresponding to the electricity cost of the distribution network, the electricity cost of the vehicle, the discharge revenue of the distribution network, the discharge revenue of the vehicle, the load power change, and the peak-valley difference of active power, respectively. These weights can be flexibly set according to actual needs, and the embodiments of this application do not impose any restrictions on them.
[0116] As can be seen from the above formula (10), the scheduling result takes into account the electricity cost, discharge revenue and the degree of change in the state of the distribution network, so as to minimize the electricity cost, increase the discharge revenue and maintain the stability of the state of the distribution network, so as to achieve a balance among these objectives.
[0117] S202: Reinforcement learning based on a reward model is used to determine the target scheduling strategy at at least one observation time.
[0118] The target scheduling strategy can be the scheduling strategy that maximizes the scheduling result output of the reward model. This target scheduling strategy can be obtained through reinforcement learning of the reward model. The learning objective can be to maximize the output of the reward model. Of course, for greater accuracy, an observation period can be set, with the maximization of the expected output of the reward model within that observation period as the learning objective to obtain the target scheduling strategy.
[0119] The aforementioned reinforcement learning needs to be conducted under pre-defined constraints. These constraints include vehicle constraints and power distribution network constraints. Vehicle constraints address at least one of the following: vehicle charge, power, and state. Power distribution network constraints address the power of the power distribution network. This ensures that the vehicle's charge, power, and state, as well as the power of the power distribution network, remain within controllable ranges during the learning process.
[0120] In some implementations, the vehicle constraints include at least one of the following: the vehicle's battery level is within a preset range, the vehicle's charging power is within a preset range, the vehicle's discharging power is within a preset range, and the vehicle's charging / discharging state at any given time is a preset state. The preset state includes one of the following: charging state, discharging state, or idle state, where the idle state is a state other than the charging state or the discharging state.
[0121] For the I-th vehicle, the battery level within the preset range can be expressed by the following formula:
[0122] SOC(MIN)<=SOC(I,T)<=SOC(MAX) (11)
[0123] Where SOC(I,T) is the battery level of vehicle I at time T, SOC(MIN) is the minimum battery level within the preset battery range, and SOC(MAX) is the maximum battery level within the preset battery range.
[0124] For charging vehicles, the above SOC(I,T) can be calculated using the vehicle's charging power and charging efficiency, as shown in the following formula:
[0125]
[0126] Where E0 is the charge level of the I-th vehicle at the start of charging (T0). CGL(I,t) is the charging power of the I-th vehicle at time t. η1 is the charging efficiency, indicating the ratio of the charge flowing into the vehicle's battery to the charge flowing out of the charging station. EMAX(I) is the maximum charge level of the I-th vehicle.
[0127] For vehicles undergoing discharge, the aforementioned SOC(I,T) can be calculated using the vehicle's discharge power and discharge efficiency, as detailed in the following formula:
[0128]
[0129] Where E0 is the charge of the I-th vehicle at the start of discharge time T0. FGL(I,t) is the discharge power of the I-th vehicle at time t. η2 is the discharge efficiency, used to indicate the ratio of the charge flowing into the charging station to the charge flowing out of the vehicle's battery.
[0130] In some implementations, the battery level can be expressed as a percentage of the maximum battery level; for example, 20% represents 20% of the maximum battery level. The minimum battery level can be flexibly set, for example, 40%, to ensure that there is still remaining battery power after the vehicle has discharged to meet its driving needs. The maximum battery level can be 1.
[0131] For the I-th vehicle, the charging power of the vehicle within the preset charging power range can be expressed by the following formula:
[0132] CGL(I,T,MIN)<=CGL(I,T)<=CGL(I,T,MAX) (14)
[0133] Where CGL(I,T) is the charging power of the I-th vehicle at time T. CGL(I,T,MIN) and CGL(I,T,MAX) are the minimum and maximum charging power of the I-th vehicle at time T, respectively.
[0134] For the I-th vehicle, the discharge power of the vehicle within the preset discharge power range can be expressed by the following formula:
[0135] FGL(I,T,MIN)<=FGL(I,T)<=FGL(I,T,MAX) (15)
[0136] Where FGL(I,T) is the discharge power of the I-th vehicle at time T. FGL(I,MAX) is the maximum discharge power of the I-th vehicle.
[0137] For the I-th vehicle, the preset charging and discharging state of the vehicle at any given time can be expressed by the following formula:
[0138] CD(I,T)+FD(I,T)<=1 (16)
[0139] Wherein, CD(I,T) indicates whether the I-th vehicle is charging at time T, and FD(I,T) indicates whether the I-th vehicle is discharging at time T. CD(I,T) = 1 indicates that the I-th vehicle is charging at time T, and CD(I,T) = 0 indicates that the I-th vehicle is not charging at time T. FD(I,T) = 1 indicates that the I-th vehicle is discharging at time T, and FD(I,T) = 0 indicates that the I-th vehicle is not discharging at time T.
[0140] By using the constraints of the above formula (16), the state of each vehicle can be controlled to be either charging, discharging, or idle, thus avoiding the situation where a vehicle is both charging and discharging.
[0141] Based on the above charging state variables CD(I,T), CGL(I,T,MIN) in formula (14) can be expressed as CD(I,T) multiplied by the minimum charging power CGL(I,MIN) of the I-th vehicle. CGL(I,T,MAX) in formula (14) can be expressed as CD(I,T) multiplied by the maximum charging power CGL(I,MAX) of the I-th vehicle.
[0142] Accordingly, based on the above discharge state variable FD(I,T), FGL(I,T,MIN) in formula (15) can be expressed as FD(I,T) multiplied by the minimum discharge power FGL(I,MIN) of the I-th vehicle. FGL(I,T,MAX) in formula (15) can be expressed as FD(I,T) multiplied by the maximum discharge power FGL(I,MAX) of the I-th vehicle.
[0143] In addition to the vehicle constraints mentioned above, the power distribution network constraints include at least one of the following: the sum of the squares of the active power and reactive power of the power distribution network is less than the square of the maximum apparent power; and the active power of the power distribution network is within the preset active power range.
[0144] The preset active power range is determined by the minimum and maximum active power of the distribution network at time T. The minimum active power can be referred to MIN[YGL(T') in the aforementioned formula (8). T'=T-TNUM+1~T The maximum active power can be referred to as MAX[YGL(T') in the aforementioned formula (8). T'=T-TNUM+1~T ].
[0145] When performing reinforcement learning on the reward model based on the above-mentioned preset constraints to obtain the target scheduling strategy, a well-known strong learning algorithm can be called for learning.
[0146] In some implementations, the target scheduling strategy can be determined through the following learning process: First, based on the scheduling results, multiple learning objectives are determined under preset constraints. These objectives include: a first learning objective that minimizes the degree of state change, a second learning objective that minimizes electricity costs, and a third learning objective that maximizes discharge revenue. Then, the priority of each learning objective is determined, and multiple learning tasks are constructed based on these objectives. Different learning tasks correspond to different reward models, and one learning task corresponds to one or more learning objectives. Finally, multiple learning tasks are executed to determine the target scheduling strategy. In this way, learning can be performed sequentially through multiple learning tasks, reducing the complexity of determining the target scheduling strategy and improving the efficiency of vehicle charging and discharging scheduling. Furthermore, the priority of the learning objectives can be flexibly adjusted to adjust the learning tasks, achieving a flexible learning process.
[0147] Understandably, each scheduling result can correspond to a learning objective. For example, when the scheduling result includes the degree of state change in the distribution network, the corresponding learning objective is to minimize the degree of state change; when the scheduling result includes electricity costs, the corresponding learning objective is to minimize electricity costs; and when the scheduling result includes discharge revenue, the corresponding learning objective is to maximize discharge revenue.
[0148] The priorities of the different learning objectives mentioned above can be flexibly set according to actual needs. Higher-priority learning objectives are executed first, and lower-priority learning objectives are executed later. Therefore, when multiple learning objectives have different priorities, different learning tasks are obtained. Specifically, a learning task is first constructed based on the highest-priority learning objective to achieve that objective. Then, for each subsequent lower-priority learning task, a learning task is constructed based on that task and other learning tasks with higher priorities to achieve those learning objectives. In this way, multiple learning tasks are generated from high-priority learning objectives to low-priority learning objectives. These multiple learning tasks are interdependent, so they need to be executed one by one according to the dependencies between them. This dependency relationship is determined by the learning objectives of each learning task; learning tasks with fewer objectives are executed first, and those with more objectives are executed later.
[0149] In some implementations, when the priority of the first learning objective is higher than that of the second learning objective, and the priority of the second learning objective is higher than that of the third learning objective, the following multiple learning tasks are constructed: a first learning task that minimizes the degree of state change under preset constraints; a second learning task that minimizes electricity costs based on minimizing the degree of state change; and a third learning task that maximizes discharge benefits based on minimizing both the degree of state change and electricity costs. It can be seen that the learning process in this embodiment can prioritize ensuring the stability of the distribution network's state, avoiding large state change intensities, and thus helping to improve the stability of the distribution network. Secondly, it ensures low electricity costs, and finally, it maximizes discharge benefits. In this way, while ensuring the stability of the distribution network's state, it achieves low electricity costs and, finally, maximizes discharge benefits as much as possible.
[0150] In each learning task, an optimal action value model is constructed based on the reward model of that learning task. The optimal action value model is used to indicate the scheduling reward corresponding to the scheduling policy and state within an observation time period. The scheduling reward includes the expected value of the scheduling result within the observation time period. Then, based on the optimal action value model, the scheduling policy corresponding to the maximum scheduling reward is determined as the target scheduling policy.
[0151] Different learning tasks correspond to different reward models, specifically in that the reward models output different scheduling results, while the input is always a scheduling strategy. The scheduling result for the first learning task is the degree of state change; the scheduling result for the second learning task is the degree of state change and electricity cost; and the scheduling result for the third learning task is the degree of state change, electricity cost, and discharge revenue.
[0152] Specifically, the scheduling result of the second learning task can be a comprehensive result of the degree of state change and the cost of electricity, while the scheduling result of the third learning task can be a comprehensive result of the degree of state change, the cost of electricity, and the discharge benefit. This comprehensive result can be understood as a weighted average or a simple summation.
[0153] Therefore, it is equivalent to splitting the reward model represented by formula (10) into multiple parts, resulting in the reward models corresponding to the following three learning tasks:
[0154] JL(1,T)=W5·FHBH(T)+W6·YGBH(T) (17)
[0155] JL(2,T)=JL(1,T)+W1·PCB(T)+W2·CCB(T) (18)
[0156] JL(3,T)=JL(2,T)+W3·PSY(T)+W4·CSY(T) (19)
[0157] In this context, formula (17) is the reward model for the first learning task, and JL(1,T) is the scheduling result output by the first learning task. Formula (18) is the reward model for the second learning task, and JL(2,T) is the scheduling result output by the second learning task. Formula (19) is the reward model for the third learning task, and JL(3,T) is the scheduling result output by the third learning task.
[0158] The optimal action value model described above represents the learning objective of reinforcement learning. This model can be obtained by recursively solving the Bellman equation. For the Nth learning task, the Bellman equation is as follows:
[0159] Q(π*(S(T),A(T)))=E[JL(N,T+1)+γ·Q(π*(S(T+1),A(T+1)))] (20)
[0160] Where Q(π*(S(T),A(T))) is the optimal action value model at time T. Q(π*(S(T+1),A(T+1))) is the optimal action value model at time T+1.
[0161] S(T) is the set of states corresponding to time T, including the energy level of each vehicle at time T, the energy level of each vehicle at the end of charging, the start and end times of charging for each vehicle, the net load power demand of the distribution network at time T, and the historical unit electricity price when the distribution network provides electricity at time T. The net load power demand of the distribution network at time T is the difference between the load power of the distribution network at time T and the power obtained by the distribution network from the main grid. The load power of the distribution network at time T can be calculated with reference to formula (6).
[0162] A(T) is the scheduling policy at time T, and JL(N,T) is the reward model for the Nth subtask at time T. JL(1,T), JL(2,T), and JL(3,T) are the reward models for the first to third learning tasks, respectively.
[0163] π*(S(T),A(T)) is used to indicate the scheduling policy A(T) to be used when the state set is S(T).
[0164] γ is the coefficient of the equation.
[0165] S203: Perform charging and discharging scheduling for at least one vehicle in the power distribution network according to the target scheduling strategy at at least one observation time.
[0166] The target scheduling strategy for each observation time is used to indicate the charging and discharging information of each vehicle. This information includes whether the vehicle is charging or discharging at that observation time, the charging power used during charging, and the discharging power used during discharging. Therefore, the vehicle can be scheduled for charging and discharging based on its charging and discharging information. For example, if a vehicle's charging and discharging information indicates that it is charging at that observation time, then the vehicle can be scheduled to charge at that charging power at that observation time. As another example, if a vehicle's charging and discharging information indicates that it is discharging at that observation time, then the vehicle can be scheduled to discharge to the charging station at that discharging power during that observation time.
[0167] In summary, the embodiments of this application can utilize the aforementioned target scheduling strategy to schedule the charging and discharging of any vehicle at any observation time, enabling discharging vehicles to compensate for the discharging capacity of the power distribution network and thus maximizing the satisfaction of charging demands. Furthermore, the target scheduling strategy determined in this application is based on a reward model; therefore, the target scheduling strategy corresponds to the optimal scheduling result, thereby compensating for the discharging capacity of the power distribution network while ensuring the best possible scheduling outcome.
[0168] Figure 3 This is a flowchart illustrating the steps of another electric vehicle charging and discharging scheduling method provided in this application embodiment. (Refer to...) Figure 3 As shown, the above-mentioned electric vehicle charging and discharging scheduling process may include:
[0169] S301: Construct a reward model, which is used to indicate the corresponding scheduling results when different scheduling strategies are adopted at different observation times. The scheduling strategy is used to indicate the charging and discharging strategy of at least one vehicle in the power distribution network. The scheduling result is used to indicate the impact on the power distribution network and / or at least one vehicle after the charging and discharging is scheduled according to the scheduling strategy. The scheduling result includes at least one of the following: electricity cost, discharge benefit, and degree of change in the state of the power distribution network.
[0170] S302: Based on the scheduling results, determine multiple learning objectives under preset constraints. The multiple learning objectives include: a first learning objective of minimizing the degree of state change, a second learning objective of minimizing electricity costs, and a third learning objective of maximizing discharge benefits.
[0171] S303: When the priority of the first learning objective is higher than the priority of the second learning objective, and the priority of the second learning objective is higher than the priority of the third learning objective, construct the following multiple learning tasks: the first learning task of minimizing the degree of state change under preset constraints, the second learning task of minimizing the electricity cost based on minimizing the degree of state change, and the third learning task of maximizing the discharge benefit based on minimizing the degree of state change and the electricity cost.
[0172] S304: Execute the first learning task, the second learning task, and the third learning task in sequence. In each learning task, construct the optimal action value model based on the reward model of the learning task. The optimal action value model is used to indicate the scheduling reward corresponding to the scheduling strategy and state within an observation time period. The scheduling reward includes the expected value of the scheduling result in the observation time period. The scheduling result corresponding to the first learning task is the degree of state change, the scheduling result corresponding to the second learning task is the degree of state change and the electricity cost, and the scheduling result corresponding to the third learning task is the degree of state change, the electricity cost, and the discharge benefit.
[0173] S305: Based on the optimal action value model, determine the scheduling strategy that maximizes the scheduling reward as the target scheduling strategy.
[0174] Understandably, each time step S305 is executed for a learning task, the target scheduling policy corresponding to that learning task can be determined. After the third learning task is completed, the obtained target scheduling policy is used in S306.
[0175] S306: Perform charging and discharging scheduling for at least one vehicle in the power distribution network according to the target scheduling strategy at at least one observation time.
[0176] It should be noted that the order of steps S301 to S306 can be flexibly adjusted on the basis of mutual independence. The corresponding descriptions of steps S301 to S306 can be referred to the corresponding descriptions in the method embodiments, and will not be repeated here.
[0177] Figure 4 This is a structural block diagram of an electric vehicle charging and discharging scheduling device provided in an embodiment of this application.
[0178] Reference Figure 4As shown, the above-mentioned electric vehicle charging and discharging scheduling device 400 includes:
[0179] The model building module 401 is used to build a reward model, which is used to indicate the scheduling results corresponding to different scheduling strategies at different observation times. The scheduling strategy is used to indicate the charging and discharging strategy of at least one vehicle in the power distribution network. The scheduling result is used to indicate the impact on the power distribution network and / or the at least one vehicle after the charging and discharging is scheduled according to the scheduling strategy.
[0180] The target scheduling strategy determination module 402 is used to perform reinforcement learning based on the reward model to determine the target scheduling strategy at at least one observation time.
[0181] The scheduling module 403 is used to schedule the charging and discharging of the at least one vehicle in the power distribution network according to the target scheduling strategy at the at least one observation time.
[0182] Optionally, the scheduling result includes at least one of the following: electricity cost, discharge revenue, and the degree of state change of the distribution network.
[0183] Optionally, the electricity cost includes at least one of the following: the electricity cost of the power distribution network and the electricity cost of the at least one vehicle; the discharge revenue includes at least one of the following: the discharge revenue of the power distribution network and the discharge revenue of the at least one vehicle; and the degree of state change includes at least one of the following: the load power change of the power distribution network and the peak-valley difference of the active power of the power distribution network.
[0184] Optionally, the target scheduling strategy determination module 402 is further configured to:
[0185] Under preset constraints, a target scheduling strategy is determined based on the reward model for at least one observation time. The preset constraints include vehicle constraints and power distribution network constraints. The vehicle constraints are used to constrain at least one of the following of the vehicle: battery level, power, and status. The power distribution network constraints are used to constrain the power of the power distribution network.
[0186] Optionally, the vehicle constraints include at least one of the following: the vehicle's battery level is within a preset range; the vehicle's charging power is within a preset range; the vehicle's discharging power is within a preset range; and the vehicle's charging / discharging state at any given time is a preset state, where the preset state includes one of the following: charging state, discharging state, or idle state, and the idle state is a state other than the charging state or the discharging state. The power distribution network constraints include at least one of the following: the sum of the squares of the active power and reactive power of the power distribution network is less than the square of the maximum apparent power; and the active power of the power distribution network is within a preset active power range.
[0187] Optionally, the target scheduling strategy determination module 402 is further configured to:
[0188] Based on the scheduling results, multiple learning objectives are determined under the preset constraints. These multiple learning objectives include: a first learning objective that minimizes the degree of state change, a second learning objective that minimizes the electricity cost, and a third learning objective that maximizes the discharge revenue. The priority of each learning objective is determined, and multiple learning tasks based on the multiple learning objectives are constructed according to the priority of each learning objective. Different learning tasks correspond to different reward models, and one learning task corresponds to one or more learning objectives. The multiple learning tasks are executed to determine the target scheduling strategy.
[0189] Optionally, the target scheduling strategy determination module 402 is further configured to:
[0190] When the priority of the first learning objective is higher than the priority of the second learning objective, and the priority of the second learning objective is higher than the priority of the third learning objective, the following multiple learning tasks are constructed: a first learning task that minimizes the degree of state change under preset constraints, a second learning task that minimizes the electricity cost based on minimizing the degree of state change, and a third learning task that maximizes the discharge benefit based on minimizing the degree of state change and the electricity cost.
[0191] Optionally, the target scheduling strategy determination module 402 is further configured to:
[0192] In each learning task, an optimal action value model is constructed based on the reward model of the learning task. The optimal action value model indicates the scheduling reward corresponding to the scheduling strategy and state within an observation time period. The scheduling reward includes the expected value of the scheduling result within the observation time period. The scheduling result corresponding to the first learning task is the degree of state change, the scheduling result corresponding to the second learning task is the degree of state change and the electricity cost, and the scheduling result corresponding to the third learning task is the degree of state change, the electricity cost, and the discharge benefit. Based on the optimal action value model, the scheduling strategy corresponding to the maximum scheduling reward is determined as the target scheduling strategy.
[0193] Figure 5 This is a structural block diagram of an electronic device provided in an embodiment of this application. The electronic device 600 includes a memory 602 and at least one processor 601, and the memory 602 and the processor 601 are communicatively connected.
[0194] Among them, memory 602 stores computer-executed instructions.
[0195] At least one processor 601 executes computer execution instructions stored in memory 602, causing electronic device 600 to perform the aforementioned functions. Figure 2 The method in the middle.
[0196] In addition, the electronic device 600 may also include a receiver 603 and a transmitter 604, wherein the receiver 603 is used to receive information from other devices or equipment and forward it to the processor 601, and the transmitter 604 is used to send information to other devices or equipment.
[0197] In an exemplary embodiment, a non-transitory computer-readable storage medium is also provided, which stores computer-executable instructions that, when executed by a processor, are used to implement the above-described electric vehicle charging and discharging scheduling method.
[0198] In an exemplary embodiment, a computer program product is also provided for implementing the aforementioned electric vehicle charging and discharging scheduling method.
[0199] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.
[0200] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A method for scheduling the charging and discharging of electric vehicles, characterized in that, include: A reward model is constructed to indicate the scheduling results corresponding to different scheduling strategies at different observation times. The scheduling strategy is used to indicate the charging and discharging strategy of at least one vehicle in the power distribution network. The scheduling result is used to indicate the impact on the power distribution network and / or the at least one vehicle after the charging and discharging scheduling is carried out according to the scheduling strategy. The scheduling result includes at least one of the following: electricity cost, discharging benefit, and the degree of state change of the power distribution network. Reinforcement learning is performed based on the reward model to determine the target scheduling strategy at at least one observation time. The charging and discharging scheduling of the at least one vehicle in the power distribution network is carried out according to the target scheduling strategy at the at least one observation time. The reinforcement learning based on the reward model to determine the target scheduling strategy at at least one observation time includes: Under preset constraints, a target scheduling strategy is determined based on the reward model for at least one observation time. The preset constraints include vehicle constraints and power distribution network constraints. The vehicle constraints are used to constrain at least one of the following of the vehicle: battery level, power, and status. The power distribution network constraints are used to constrain the power of the power distribution network. The step of determining the target scheduling strategy for at least one observation time based on the reward model under preset constraints includes: Based on the scheduling results, multiple learning objectives are determined under the preset constraints. The multiple learning objectives include: a first learning objective of minimizing the degree of state change, a second learning objective of minimizing the electricity cost, and a third learning objective of maximizing the discharge revenue. The priority of each learning objective is determined, and multiple learning tasks based on the multiple learning objectives are constructed according to the priority of each learning objective. The reward models corresponding to different learning tasks are different, and one learning task corresponds to one or more learning objectives. Execute the multiple learning tasks to determine the target scheduling strategy; The step of constructing multiple learning tasks based on the priority of each learning objective includes: When the priority of the first learning objective is higher than the priority of the second learning objective, and the priority of the second learning objective is higher than the priority of the third learning objective, the following multiple learning tasks are constructed: a first learning task that minimizes the degree of state change under preset constraints, a second learning task that minimizes the electricity cost based on minimizing the degree of state change, and a third learning task that maximizes the discharge benefit based on minimizing the degree of state change and the electricity cost.
2. The method according to claim 1, characterized in that, The electricity cost includes at least one of the following: the electricity cost of the power distribution network and the electricity cost of the at least one vehicle; the discharge revenue includes at least one of the following: the discharge revenue of the power distribution network and the discharge revenue of the at least one vehicle; and the degree of state change includes at least one of the following: the load power change of the power distribution network and the peak-valley difference of the active power of the power distribution network.
3. The method according to claim 1, characterized in that, The vehicle constraints include at least one of the following: the vehicle's battery level is within a preset battery level range, the vehicle's charging power is within a preset charging power range, the vehicle's discharging power is within a preset discharging power range, and the vehicle's charging and discharging state at the same time is a preset state. The preset state includes one of the following: charging state, discharging state, and idle state. The idle state is a state other than the charging state and the discharging state. The power distribution network constraints include at least one of the following: the sum of the squares of the active power and reactive power of the power distribution network is less than the square of the maximum apparent power; and the active power of the power distribution network is within a preset active power range.
4. The method according to claim 1, characterized in that, The execution of the plurality of learning tasks to determine the target scheduling strategy includes: In each learning task, an optimal action value model is constructed based on the reward model of the learning task. The optimal action value model is used to indicate the scheduling reward corresponding to the scheduling strategy and state within an observation time period. The scheduling reward includes the expected value of the scheduling result in the observation time period. The scheduling result corresponding to the first learning task is the degree of state change. The scheduling result corresponding to the second learning task is the degree of state change and the electricity cost. The scheduling result corresponding to the third learning task is the degree of state change, the electricity cost, and the discharge benefit. Based on the optimal action value model, the scheduling strategy corresponding to the maximum scheduling reward is determined as the target scheduling strategy.
5. A charging and discharging scheduling device for electric vehicles, characterized in that, include: The model building module is used to build a reward model, which is used to indicate the scheduling results corresponding to different scheduling strategies at different observation times. The scheduling strategy is used to indicate the charging and discharging strategy of at least one vehicle in the power distribution network. The scheduling result is used to indicate the impact on the power distribution network and / or the at least one vehicle after the charging and discharging scheduling is carried out according to the scheduling strategy. The scheduling result includes at least one of the following: electricity cost, discharging benefit, and the degree of state change of the power distribution network. The target scheduling strategy determination module is used to perform reinforcement learning based on the reward model to determine the target scheduling strategy at at least one observation time. The scheduling module is used to schedule the charging and discharging of the at least one vehicle in the power distribution network according to the target scheduling strategy at the at least one observation time. When performing reinforcement learning based on the reward model to determine the target scheduling policy at at least one observation time, the target scheduling policy determination module is configured to: Under preset constraints, a target scheduling strategy is determined based on the reward model for at least one observation time. The preset constraints include vehicle constraints and power distribution network constraints. The vehicle constraints are used to constrain at least one of the following of the vehicle: battery level, power, and status. The power distribution network constraints are used to constrain the power of the power distribution network. When determining the target scheduling strategy for at least one observation time based on the reward model under preset constraints, the target scheduling strategy determination module is used to: Based on the scheduling results, multiple learning objectives are determined under the preset constraints. The multiple learning objectives include: a first learning objective of minimizing the degree of state change, a second learning objective of minimizing the electricity cost, and a third learning objective of maximizing the discharge revenue. The priority of each learning objective is determined, and multiple learning tasks based on the multiple learning objectives are constructed according to the priority of each learning objective. The reward models corresponding to different learning tasks are different, and one learning task corresponds to one or more learning objectives. Execute the multiple learning tasks to determine the target scheduling strategy; When constructing multiple learning tasks based on the multiple learning objectives according to the priority of each learning objective, the objective scheduling strategy determination module is used to: When the priority of the first learning objective is higher than the priority of the second learning objective, and the priority of the second learning objective is higher than the priority of the third learning objective, the following multiple learning tasks are constructed: a first learning task that minimizes the degree of state change under preset constraints, a second learning task that minimizes the electricity cost based on minimizing the degree of state change, and a third learning task that maximizes the discharge benefit based on minimizing the degree of state change and the electricity cost.
6. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Electric vehicle charging optimization guiding strategy considering reward mechanism
CN112193116A
Electric vehicle multi-objective optimization charging scheduling method under hybrid demand response
CN115630796A