Real-time charging and discharging scheduling method and system for electric vehicles based on distributed ADP
By constructing an EV multi-cell model using the distributed ADP algorithm and consensus optimization method, the computational complexity and privacy leakage issues in real-time scheduling of electric vehicles are resolved, enabling efficient and rapid electric vehicle charging and discharging scheduling, and improving the economic efficiency and security of power grid operation.
Patent Information
- Application Number
- CN202510682980.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2045-05-26
AI Technical Summary
Existing real-time scheduling methods for electric vehicles suffer from high computational complexity, prolonged response time, privacy leaks, and poor flexibility, making them particularly difficult to deploy effectively in geographically dispersed and structurally complex power grids.
The distributed ADP algorithm is adopted to construct an EV multi-cell charging and discharging model through multi-cell modeling and consensus distributed optimization. Combined with Markov decision process, multi-stage sequential decision-making is carried out, and scheduling optimization is performed by utilizing local information to reduce computational load and protect user privacy.
It improves the executability and optimization of scheduling strategies, enhances the reliability and adaptability of the system, and supports rapid response and privacy protection in complex power grids.
Smart Images

Figure CN120542862B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of electric vehicle charging and discharging scheduling, in particular to an electric vehicle real-time charging and discharging scheduling method and system based on distributed ADP. BACKGROUND
[0002] The statements in this section merely provide background information related to the present application and do not necessarily constitute the prior art.
[0003] Electric vehicles (EV) are distributed energy storage resources with bidirectional energy flow capability. During the operation of the power system, if the charging and discharging behavior of EVs can be reasonably guided and coordinated, the operation economy and safety of the power grid can be effectively improved, and the absorption capacity of new energy generation (such as wind energy, solar energy, etc.) fluctuations can be enhanced. In the context of increasing new energy penetration, the rapid response capability and mobile distributed characteristics of EV groups can improve the regulation capability of the power system. By designing a reasonable EV participation type optimization scheduling mechanism, active charging and orderly discharging of EVs can be realized to respond to the load regulation demand of the power grid, relieve system peak load, reduce reserve capacity requirements, and improve overall energy utilization efficiency.
[0004] However, the current EV real-time scheduling still faces many technical bottlenecks: first, with the rapid expansion of the scale of EVs, the EV aggregation strategy is generally used in research to reduce the scheduling dimension. The current main methods include external approximation method and internal approximation method. The external approximation method (such as simple geometric envelope) has the advantage of computational efficiency, but cannot support the disaggregation operation, limiting its applicability in individual EV control. The internal approximation method (such as box, zonotope, and polytope) can more accurately approximate the feasible region, especially the polytope method provides higher fidelity, but its construction and solving process still has large computational overhead, making it difficult to meet the fast response demand of the actual system. Therefore, the traditional modeling method is difficult to adapt to the rapid growth of computational complexity, and the scheduling system faces the dual challenges of high load and response time delay. Second, the centralized optimization scheduling framework adopted by the mainstream scheduling method needs to collect real-time information from all EVs and other elements and uniformly calculate the scheduling strategy. While obtaining global information, it also exposes the problems of user privacy leakage and poor system robustness. Once the central controller is attacked or fails, the whole scheduling system will be paralyzed. Third, the demand for electric vehicle charging and the output of renewable energy sources are highly random. The representative method currently used to cope with uncertainty is the Approximate Dynamic Programming (ADP) algorithm, which uses historical data to train the value function in the day-ahead stage and uses it to approximate the optimal decision in the day-ahead stage. However, most of these methods rely on complete system global information to complete the training process, limiting their flexibility and adaptability in actual deployment. SUMMARY
[0005] To solve the above problems, the present application provides an electric vehicle real-time charging and discharging scheduling method and system based on distributed ADP, which realizes multi-period, distributed, and privacy-friendly optimization control.
[0006] To achieve the above purpose, the present application adopts the following technical solutions:
[0007] One or more embodiments provide an electric vehicle real-time charging and discharging scheduling method based on distributed ADP, comprising the following steps:
[0008] All EVs are divided into multiple cluster groups, and the charging and discharging power set of each EV is constructed as a polytope. The polytope approximation aggregation method is used for each cluster group to obtain the constructed EV polytope charging and discharging model;
[0009] The total operating cost of the power grid is minimized and the adjustable potential of the thermal power unit is maximized as the target, and the constraint conditions include the EV charging and discharging constraints represented based on the EV polytope charging and discharging model. An EV centralized scheduling model considering aggregation is constructed;
[0010] The EV centralized scheduling model is reconstructed as a Markov decision process, and the scheduling problem is converted into a multi-stage sequential decision problem;
[0011] The day-ahead power system historical state data and the EV charging and discharging scheduling data are acquired, day-ahead offline training is performed, the multi-stage sequential decision problem is iteratively trained and solved by combining the consistency distributed optimization method and the ADP algorithm, and the segmented linear value function slope of the trained ADP distributed value function is obtained.
[0012] The running state of the current period of the day-ahead power system is acquired, the multi-stage sequential decision problem is solved by using the consistency distributed optimization method according to the trained ADP distributed value function, the current period scheduling decision is obtained, and the charging and discharging strategy of each EV is obtained based on the EV multi-cell charging and discharging model.
[0013] One or more embodiments provide a distributed ADP-based real-time charging and discharging scheduling system for electric vehicles, comprising:
[0014] The multi-cell charging and discharging model construction module is configured to divide all EVs into a plurality of cluster groups, construct the charging and discharging power set of each EV into a multi-cell, and aggregate each cluster group by using the multi-cell approximate aggregation method to obtain the constructed EV multi-cell charging and discharging model.
[0015] The EV centralized scheduling model construction module is configured to minimize the total operation cost of the power grid and maximize the adjustable potential of the thermal power unit as the target, and the constraint conditions include the EV charging and discharging constraints represented based on the EV multi-cell charging and discharging model, and the EV centralized scheduling model considering aggregation is constructed.
[0016] The problem conversion module is configured to reconstruct the EV centralized scheduling model as a Markov decision process, and convert the scheduling problem into a multi-stage sequential decision problem.
[0017] The day-ahead training module is configured to acquire the day-ahead power system historical state data and the EV charging and discharging scheduling data, perform day-ahead offline training, iteratively train and solve the multi-stage sequential decision problem by combining the consistency distributed optimization method and the ADP algorithm, and obtain the segmented linear value function slope of the trained ADP distributed value function.
[0018] The day-ahead optimization module is configured to acquire the running state of the current period of the day-ahead power system, solve the multi-stage sequential decision problem by using the consistency distributed optimization method according to the trained ADP distributed value function, obtain the current period scheduling decision, and obtain the charging and discharging strategy of each EV based on the EV multi-cell charging and discharging model.
[0019] One or more embodiments provide a distributed ADP-based real-time charging and discharging scheduling system for electric vehicles, comprising:
[0020] a data acquisition device and a processor;
[0021] The data acquisition device is used to acquire the operating state of the power system and EV charging and discharging regulation data.
[0022] The processor is configured to perform the steps of the above-mentioned distributed ADP-based real-time charging and discharging scheduling method for electric vehicles.
[0023] Compared with the prior art, the present application has the following beneficial effects:
[0024] The present application improves the fitting accuracy of the EV aggregation model through the polytope modeling technology, significantly improves the executability and optimality of the scheduling strategy; introduces a consistent distributed training and reasoning mechanism to solve the problems of information concentration, heavy computing load and user privacy leakage in traditional centralized methods; uses the piecewise linear value function of approximate dynamic programming (ADP) for solving, so that the scheduling algorithm has good scalability and real-time response capability under the premise of ensuring performance; supports deployment in high geographical dispersion and complex structure power grid, significantly enhances the reliability and adaptability of the system.
[0025] The advantages of the present application and the advantages of the additional aspects will be described in detail in the following specific embodiments. BRIEF DESCRIPTION OF DRAWINGS
[0026] The drawings accompanying the specification of the present application form a part thereof and serve to provide further understanding of the present application, the illustrative embodiments of the present application and its description serve to explain the present application, and do not constitute limitations thereof.
[0027] Figure 1 is the first flowchart of the charging and discharging scheduling method of embodiment 1 of the present application;
[0028] Figure 2 is the relative error of the distributed value function updating method in the simulation example of embodiment 1 of the present application;
[0029] Figure 3 is the second flowchart of the charging and discharging scheduling method of embodiment 1 of the present application;
[0030] Figure 4 is the relative error comparison chart of two distributed algorithms and a centralized global optimal algorithm in the simulation example of embodiment 1 of the present application; DETAILED DESCRIPTION
[0031] The present application will be further described below in conjunction with the drawings and embodiments.
[0032] It should be noted that the following detailed description is exemplary in nature and is intended to provide further description of the application. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs.
[0033] It is to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of example embodiments in accordance with the present application. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, steps, operations, elements, components, and / or groups thereof, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups thereof. It will be noted that various embodiments of the present application and features of the embodiments can be combined with each other as appropriate, unless explicitly stated otherwise. Embodiments will now be described in detail with reference to the accompanying drawings.
[0034] Embodiment 1
[0035] In the technical solutions disclosed in one or more embodiments, as shown in the drawings, a real-time charging and discharging scheduling method for electric vehicles based on distributed ADP includes the following steps: Figures 1 to 4
[0036] Step 1, divide all EVs into multiple cluster groups, construct the charging and discharging power set of each EV as a polytope, and aggregate each cluster group using the polytope approximation aggregation method to obtain the constructed EV polytope charging and discharging model;
[0037] Step 2, taking minimizing the total operation cost of the power grid and maximizing the adjustable potential of the thermal power unit as the target, and the constraint conditions including the EV charging and discharging constraints represented based on the EV polytope charging and discharging model, construct an aggregated EV centralized scheduling model;
[0038] Step 3, reconstruct the EV centralized scheduling model as a Markov decision process, and convert the scheduling problem into a multi-stage sequential decision problem;
[0039] Step 4, obtain the historical state data of the power system and the EV charging and discharging scheduling data in advance, perform offline training in advance, and solve the multi-stage sequential decision problem through iteration by combining the consistency distributed optimization method and the ADP algorithm to obtain the segmented linear value function slope of the trained ADP distributed value function;
[0040] Step 5, obtain the operation state of the current period of the power system, solve the multi-stage sequential decision problem by using the consistency distributed optimization method according to the trained ADP distributed value function, and obtain the scheduling decision of the current period , and based on the EV polytope charging and discharging model, the charging and discharging strategy of each EV is obtained by de-aggregation.
[0041] In this embodiment, the plug-in time of EV is pattern recognized by a clustering algorithm, a plurality of EV clustering groups with similar behavior characteristics are formed, and a polytope approximation modeling method is used to convert the clustering groups into a unified aggregate model, so as to construct an EV polytope charging and discharging model, and realize modeling of the EV space after aggregation. The scheduling model expresses the charging and discharging behavior boundary of the aggregate through constraints, and constructs a centralized optimization problem with a predetermined objective function. The model is then reconstructed as a Markov decision process (MDP), so that the scheduling problem can be converted into a value function optimization problem. To avoid dependence on global information, a consistent distributed optimization algorithm is introduced, historical data and ADP algorithm are combined for training in the day-ahead stage, the optimal slope of the piecewise linear value function is obtained through local information and neighborhood communication, and then the scheduling strategy can be solved only with the current state information in the intra-day scheduling, so as to improve the responsiveness and flexibility of the system.
[0042] The embodiment improves the fitting accuracy of the EV aggregation model through the polytope modeling technology, and significantly improves the executability and optimality of the scheduling strategy; introduces a consistent distributed training and reasoning mechanism to solve the problems of information concentration, heavy computing load and user privacy leakage in traditional centralized methods; uses the piecewise linear value function of approximate dynamic programming (ADP) for solving, so that the scheduling algorithm has good scalability and real-time response ability under the premise of ensuring performance; supports deployment in high geographical dispersion and complex structure power grid, and significantly enhances the reliability and adaptability of the system.
[0043] In this embodiment, the ADP algorithm is combined with the consistency theory to propose a distributed piecewise linear function slope updating method, which can embed the experience information of system randomness factors into the value function to assist decision making only through adjacent element communication and local information on the EV side, so as to obtain an approximately globally optimal EV scheduling strategy in real-time optimization in a distributed manner.
[0044] In step 1, the EV polytope charging and discharging model is constructed, and the process is as follows:
[0045] In step 11, a charging and discharging model is constructed for each electric vehicle (EV), and based on the charging and discharging model constraints, the feasible charging and discharging power set of the electric vehicle is represented as a convex polytope defined in a multi-dimensional space.
[0046] The i-th electric vehicle The charging and discharging model after accessing the power system is shown in equations (1) to (5):
[0047] The charging power limit is: (1);
[0048] The discharging power limit is: (2);
[0049] The accumulated charge and discharge energy of the t period The calculation formula is:
[0050] (3);
[0051] The charge and discharge energy constraint is: (4);
[0052] The target energy constraint: (5);
[0053] Wherein, and respectively are The charging power and discharging power in the t period; and respectively are The maximum charging power and maximum discharging power of the t period; and respectively are The charging efficiency and discharging efficiency of the t period; and respectively are the charging plug-in and plug-out time; is The energy value in the t period; is The initial energy at the time of plugging in; is The target minimum energy of the t period; represents each scheduling period from the plugging-in time, is the scheduling step; and
[0054] respectively are The upper and lower limits of the energy in the t period, which are calculated as follows:
[0055] ;
[0056] In the formula, is The battery capacity of the t period.
[0057] As can be seen from formula (1) to formula (5), The model of the t period does not introduce 0-1 variables to represent its charging and discharging state, and the scheduling method optimized in the subsequent steps of the embodiment will only appear charging or discharging in each decision, that is, either charging or discharging, and will not appear charging and discharging at the same time, which is due to the effect of charging and discharging efficiency, so as to avoid the charging and discharging at the same time which will cause power loss, which does not comply with the optimal optimization principle.
[0058] To facilitate the aggregation of different EVs, the above charging and discharging model constraints (i.e., as shown in equations (1) to (5)) are expressed in matrix form:
[0059] (6);
[0060] In the formula, is a charging and discharging power column vector of dimension n; is a restriction matrix of dimension n; is the total charging time of n, and is the number of charging and discharging time periods; is a restriction column vector of dimension n; the calculation formula is as follows:
[0061] ;
[0062] (7);
[0063] In the formula, is a unit matrix of dimension n; is a column vector of dimension n, with all elements being 1; , , and are column vectors of dimension n, , , , ; and are diagonal matrices of dimension n, , ; is a lower triangular matrix of dimension n, and the calculation formula is as follows:
[0064] (8);
[0065] The set of feasible charging and discharging powers of an electric vehicle is represented as a convex polytope defined in a multi-dimensional space. Specifically, the charging and discharging powers (vectors) of an EV at different time periods are taken as variables, and the set of all feasible charging and discharging power combinations is constructed as a convex polyhedron defined by linear inequalities, i.e., a polytope.
[0066] For an electric vehicle , the charging and discharging power column vector The acceptable range is specifically represented as a... A multicellular entity in 3D space, which is represented as :
[0067] (9);
[0068] In the formula, the charge / discharge power column vector In multicellular organisms It can be satisfied within the range Required charge / discharge constraints.
[0069] Step 12, EV Cluster and Base Set Construction: Based on the insertion and removal times of each EV, a clustering algorithm is used to divide all EVs into K clusters; for each cluster k (k=1,2,3,…K), by statistically analyzing the average operating parameters of the EV individuals within cluster k, a base set representing the feasible domain characteristics of the EVs in each cluster is constructed, i.e., the base set.
[0070] In the above implementation, the modeling method can treat similar EVs as a typical multi-cell in scheduling optimization, thereby effectively reducing the number of constraints, improving solution efficiency, and retaining the physical feasibility and flexibility of EV aggregation scheduling.
[0071] Specifically, before constructing the base set and aggregation, EVs are clustered. All EVs are clustered using insertion and removal times as clustering metrics, and clustering algorithms such as K-means++ are used to obtain... One cluster;
[0072] Construct base set As an aggregate The basic set used to represent aggregates The multicellular entities of each EV, the base set The representation is shown in equation (10). In this embodiment, the polymer is obtained by calculating the polymer. All EV limit matrices and restricted column vectors The average value is used to determine the base set. parameters and The constructed base set is represented as follows:
[0073] (10);
[0074] In the formula, The column vector of charging and discharging power is the base set.
[0075] Step 13, EV aggregation and disaggregation based on polytope approximation: based on the polytope approximation method, the approximate polytope of each polytope is obtained by transforming the basic polytope set, and the polytope in the clustering group is aggregated and disaggregated to obtain the EV polytope charging and discharging model;
[0076] Step 131, construction of theoretical feasible region: the aggregate accurate feasible region is expressed in the form of Minkowski sum as follows:
[0077] (11);
[0078] In the formula, is the accurate feasible region of the aggregate ; is the original polytope of ; denotes the Minkowski sum.
[0079] In this step, the theoretically accurate EV total feasible region is constructed, but the calculation complexity is high and the solution is irreversible.
[0080] Step 132, polytope approximation aggregation is adopted, and the same base set is used to transform and express the polytope feasible region of all EVs to obtain the approximate polytope of each polytope ;
[0081] The original feasible region of each EV is of different dimensions and shapes, but through polytope approximation aggregation, a common “shape template” is used to express the feasible region of all EVs; by assigning a set of transformation parameters to each EV, the heterogeneity can be expressed without re-modeling, and multiple EVs can be expressed with a unified base set in the least cost manner, facilitating subsequent unified modeling;
[0082] Specifically, to simplify the calculation of , the polytope approximation aggregation method is adopted in this embodiment, which principle is to cover as much as possible the available space of the polytope by stretching and translating the base set while ensuring that the polytope is not exceeded, so that the base set can be used to approximate the range of all EV polytopes in the aggregate , facilitating the operation of polytopes with the same base set.
[0083] As shown in formula (12), the base set is stretched by times and translated by , and the polytope Approximation as an approximation polytope ; approximation polytope of the base set and The Minkowski sum between them can be simplified as formula (13):
[0084] (12);
[0085] (13);
[0086] In the formula, is a polytope Transform the approximation polytope represented by the base set satisfies the relationship ; and and are the scaling amount and the translation amount, respectively.
[0087] Further, by constructing a linear programming problem, it is determined how the base set is transformed to cover each EV, and the base set Approximation represents a polytope The scaling amount and the translation amount of the approximation polytope The linear programming problem is constructed as follows:
[0088] (14);
[0089] In the formula, is a decision matrix.
[0090] Step 133, based on the Minkowski sum form constructed in step 131, sum the approximation polytopes of all EVs to obtain the aggregated polytope; complete the aggregation to form a reversible, low-dimensional aggregated model; the aggregated polytope of the aggregated polytope is:
[0091] (15);
[0092] In the formula, and are the scaling amount and the translation amount, respectively, through representation transformation;
[0093] , (16);
[0094] Step 134, decouple the obtained aggregated polytope according to the time period, and decompose the high-dimensional aggregated polytope into a two-dimensional polytope , the charging and discharging power and energy constraints of each time period are obtained, i.e. the EV polytope charging and discharging model;
[0095] is a 2m-dimensional polytope, and each dimension represents the charging and discharging power of the aggregation body in the respective time period , the 2m-dimensional polytope is converted into a corresponding polytope of each time period , as shown in equation (17);
[0096] (17);
[0097] In the formula, is the charging and discharging power vector of the EV aggregation body in the time period , which is a two-dimensional column vector, ; and are the intercept matrices of and in the time period, respectively; is a 6-dimensional column vector, and the first 4 rows are the intercept values of the time period , while the last 2 rows need to consider the relationship with the energy of the time period , and the last two elements are and , and are the average values of and of all EVs in the aggregation body , respectively, are calculated as follows:
[0098] (18);
[0099] In the formula, and are the elements in .
[0100] In the above scheme, the polytope approximation aggregation result is to approximate the original feasible region of each EV as an affine transformation of a unified basis set, so as to realize the compression of the diversity of a group of EVs in power scheduling into a small number of controllable parameters (scaling and translation), reduce the calculation amount while ensuring sufficient optimization accuracy.
[0101] At the same time, the state of charge SOC of the battery is introduced to represent the battery state before and after the decision of the EV aggregation body , and the state of charge of the battery is represented as follows:
[0102] (19);
[0103] (20);
[0104] wherein, and are EV aggregates In SOC before and after the time period decision; is the initial total energy of EV aggregate , .
[0105] In this embodiment, the analytical structure of the Minkowski sum is used to construct a charging and discharging feasible region model with clear structure and geometric solvability, so that the scheduling modeling has good mathematical structure and computational feasibility; a unified basic polytope set transformation strategy is used to reduce the model complexity and improve the expansion ability and consistent expression of the model; the high-dimensional polytope processing difficulty is reduced through time period decoupling, so that the scheduling problem is converted into a decouplable low-dimensional optimization problem, thereby realizing efficient modeling and fast optimization of EV scheduling.
[0106] Further, after obtaining the current time period scheduling decision through the intra-day solution of step 5, the value of the single-dimensional polytope is obtained, and the charging and discharging power column vector of the aggregate scheduling result includes charging power and discharging power, and each time period is charged or discharged. During the scheduling optimization process, if charging and discharging occur at the same time, energy will be damaged, so the optimization result has an element of 0.
[0107] In this embodiment, the EV polytope charging and discharging model is constructed, the EV aggregation and disaggregation method based on polytope approximation is used to reduce the computational amount while ensuring sufficient optimization accuracy; by establishing a charging and discharging model for each electric vehicle, a multi-dimensional convex polytope is constructed as its feasible charging and discharging power set according to the upper and lower limits of its charging and discharging power, the time window and the charging demand. After clustering, in each cluster group, the average operating parameters are calculated to form a basic polytope set representing the characteristics of the EV class, which is used as the basis for subsequent approximation. By linearly transforming the basic polytope set (such as scaling and translation), the approximate polytope of each EV is obtained. On this basis, the polytope charging and discharging model of the EV class is constructed by performing aggregation operation on multiple EV approximate polytopes in each class, while supporting disaggregation operation to restore the control ability of individual EVs. This modeling method greatly reduces the computational resources required for individual modeling and improves the efficiency of EV charging and discharging strategy aggregation modeling; using polytope as the representation form enhances the geometric expression ability of the model to the EV feasible region, and can flexibly adjust the approximation accuracy of the model; through the clustering-based processing method, the structure of the model is simplified, while the statistical characteristics of the key operating parameters are retained, and the model accuracy and computational efficiency are considered.
[0108] In step 2, with the objectives of minimizing the total operating cost of the power grid and maximizing the adjustable potential of thermal power units, an EV centralized dispatch model considering aggregation is established, and the model is reconstructed into a Markov decision process:
[0109] The real-time scheduling model aims to minimize the total operating cost of the power grid and maximize the adjustable range of thermal power units. While utilizing the flexibility resources of EV to reduce system operating costs, it ensures that the adjustable range of thermal power units is maintained at an appropriate level, enhances the system's ability to cope with extreme scenarios, and improves the system's reliability.
[0110] (twenty one);
[0111] (twenty two);
[0112] (twenty three);
[0113] (twenty four);
[0114] (25);
[0115] In the formula, For power grid The objective function for the time period; For the power grid The total operating cost for the period includes the operating cost of thermal power units, the cost of curtailment of solar power, the cost of curtailment of wind power, the cost of purchasing and selling electricity from the external grid, and the cost of EV battery life depreciation. A collection of thermal power units; , and For thermal power units The power generation cost coefficient; and These are collections of photovoltaic power sources and wind turbine units, respectively. and These are the penalty cost coefficients for abandoning solar and wind power; and They are respectively Periodic photovoltaic power Wind turbine The maximum available power generation capacity; and They are respectively Periodic photovoltaic power Wind turbine The actual absorption capacity; For external power grid nodes; For the system Electricity purchase and sale cost coefficient for external power grid during the specified time period; for Period external grid node of the purchase and sale of electricity power; is the set of EV aggregates; is the total number of segments of the life loss piecewise linear function; is Period EV aggregate In Segmented discharge power; is the EV aggregate In Segmented battery life loss cost coefficient; is the EV battery installation cost coefficient; is Period thermal power unit Adjustable interval size; is the thermal power unit adjustable potential coefficient, determined according to the system reliability requirement, generally ; and are the upper and lower limits of the thermal power unit active power output; is the maximum climbing rate of the thermal power unit .
[0116] When the output of the thermal power unit is , it can achieve global regulation, at which time the reliability is the highest. The farther away from the interval, the lower the potential for up and down regulation, and the less conducive to system reliability.
[0117] EV centralized scheduling model constraint conditions, including: power balance constraint (26), thermal power unit constraints (27) and (28), photovoltaic and wind power output constraints (29) and (30), external grid purchase and sale of electricity power constraints (31) and EV charging and discharging constraints (32), each constraint formula is as follows:
[0118] (26);
[0119] (27);
[0120] (28);
[0121] (29);
[0122] (30);
[0123] (31);
[0124] (32);
[0125] In the formula, It is the collection of loads in the power grid; for Time-of-use load The load; and External power grid Upper and lower limits of power consumption for purchasing and selling electricity.
[0126] In this embodiment, an EV centralized scheduling model considering aggregation is established, and the model is reconstructed into a Markov decision process (MDP) to conform to the architecture of real-time scheduling.
[0127] In step 3, the EV centralized scheduling model is reconstructed into a Markov decision process, including the following:
[0128] Step 31: Output of conventional thermal power units The prediction errors of photovoltaic power output, wind power output, electricity price and load are used as state variables; the dispatch output decisions of each energy source are used as decision variables. The state of the current period is calculated from the state of the previous period, and the state transition function of the Markov decision is constructed.
[0129] The EV real-time scheduling model after MDP reconstruction mainly includes state variables. Decision variables External information factors and It also includes the state transition function (equation), as follows:
[0130] (33);
[0131] (34);
[0132] (35);
[0133] (36);
[0134] The state transition function is:
[0135] (37);
[0136] In the formula, and They are respectively Forecasting errors for photovoltaic power output, wind power output, electricity price, and load during different time periods; , , and They are respectively Day-ahead forecast values of multi-period photovoltaic output, wind power output, electricity price and load.
[0137] Step 32, in order to obtain the intra-day real-time scheduling and global optimal decision, based on the obtained state transition function and Bellman optimal principle, the multi-period EV charging and discharging scheduling process is abstracted as a multi-stage sequential decision problem, and the reconstructed objective function is:
[0138] (38);
[0139] In the formula, is the state variable after the multi-period decision, generally ; is the total value function of the multi-period to the end of the scheduling time ; is the decision variable of the multi-period; represents the influence of the multi-period state and decision on the future value function of the system, and represents the optimal value function of the state reached after executing the decision ;
[0140] The ADP-based EV real-time scheduling framework, however, is based on a centralized solving mode and relies on a central controller to make centralized decisions for scheduling, and there is a problem of lack of privacy protection mechanism. In the ADP value function training process, the global information is used in the objective function, which violates the principle of distributed optimization. To solve the above problems, the embodiment starts from the slope update principle, combines the slope calculation with the consistency algorithm, and proposes a completely distributed slope update method, which can update the piecewise linear function slope only by using the EV itself information.
[0141] In step 4, the day-ahead offline training method comprises the following steps:
[0142] Step 41, EV polytope aggregation modeling: obtaining historical state data of the power system and EV charging and discharging scheduling data, constructing an EV polytope charging and discharging model, and forming an EV aggregate by using a polytope approximation aggregation method;
[0143] Specifically, the K-means++ algorithm is used for the EV cluster, an EV matrix model (formula (7)) is constructed, a polytope approximation aggregation method is used to form an EV aggregate, and formulas (17) and (18) are used.
[0144] Step 42, generating a training scene: generating a plurality of training scenes according to the day-ahead forecast values for EV charging and discharging regulation;
[0145] According to the day-ahead prediction value of the EV charging and discharging regulation data (such as wind and light output, electricity price and the like), NA random scenarios are constructed, and each scenario corresponds to a set of power system state sequences; specifically, a plurality of training scenarios are randomly generated according to the day-ahead prediction value by Monte Carlo;
[0146] The day-ahead prediction value is a value predicted based on the day-ahead power system operation or dispatch data, and in the EV dispatching model involved in the embodiment, the day-ahead prediction value includes photovoltaic output prediction, wind power output prediction, load prediction, and electricity purchase price prediction.
[0147] Step 43, for each training scenario data, the constructed multi-stage sequential decision problem is solved by combining the consistent distributed optimization method and the ADP algorithm to obtain the optimal decision of each period; and the value function slope of the ADP is iteratively trained based on the optimal decision scheme of each period to obtain the final value function slope;
[0148] In step 43, the process of iteratively training the value function slope based on the optimal decision scheme of each period to obtain the final value function slope is as follows:
[0149] Step 431, the multi-stage sequential decision problem is converted into an ADP piecewise linear form by using an approximate dynamic optimization method to obtain an ADP piecewise linear value function;
[0150] Step 432, for the obtained ADP piecewise linear value function, each training scenario and each period under each training scenario are iteratively operated as follows:
[0151] (1) The optimization problem of the ADP piecewise linear function is solved by using the consistent distributed optimization method to obtain the current period dispatching decision;
[0152] (2) Distributed ADP value function training: based on the current period dispatching decision, disturbance analysis is performed by using the state of charge change as the disturbance quantity, and the ADP value function slope is updated based on the local information of the EV.
[0153] In step 431, the value function is initialized in a piecewise linear form, the state variable SOC (state of charge) of the EV aggregate is selected as the input dimension of the value function, and a concave piecewise linear function is used to fit the approximate value function to obtain a value function in a piecewise linear form, and the process is as follows:
[0154] The randomness and uncertainty in the system operation process make it very difficult to calculate in formula (38), and the SOC of the EV aggregate has a period-coupled constraint. In this embodiment, the SOC of the EV aggregate is used to represent the state variable of the system, the state variable SOC (state of charge) of the EV aggregate is selected as the input dimension of the value function, and a concave piecewise linear function is used for fitting:
[0155] (39);
[0156] In the formula, for After the decision-making process, the EV aggregate The relevant approximation function.
[0157] The EV centralized scheduling model proposed in this embodiment is a concave optimization model, which uses a concave piecewise linear function to fit an approximate function. The resulting ADP value function is:
[0158] (40);
[0159] In the formula, The total number of segments in a piecewise linear function; EV aggregate exist Time period The slope of the segment; EV aggregate exist Time period Segmented battery status.
[0160] The existing scheduling scheme updates the formulas as shown in equations (42) and (43). By embedding uncertainties into the value function through offline training, it can effectively cope with the randomness of the source, grid and load in the power system and ensure that an approximately global optimal decision is made during the real-time charging and discharging scheduling of EVs.
[0161] In the In this training iteration, a small perturbation ΔSOC is set, and its combined effect on the current operating cost and the next state value function is observed. Based on this, the marginal effect is estimated, as expressed below:
[0162] (42);
[0163] (43);
[0164] In the formula, Indicates the first The next iteration; For the first EV aggregate during the next iteration exist Sampling estimates of the slope of the piecewise linear function over a given time period; for The piecewise linear function is segmented into pieces; To update the step size, the SOC of the EV polymer is linearly decomposed into D segments;
[0165] Equation (42) and Equation (43) can only update The slope of the segment, updating the slope of other segments requires using a concave adaptive value estimation (CAVE) algorithm to adjust the slope, and through gradient constraint and slope adjustment, the slope sequence of the piecewise linear function is always decreasing, ensuring the concavity of the piecewise linear function.
[0166] Thus, Equation (23) to Equation (43) constitute an ADP-based EV real-time scheduling framework, but this framework is based on a centralized solution mode and relies on a central controller to make centralized decision scheduling, and there is a problem of lack of privacy protection mechanism, so a consistent distributed optimization method is adopted in this embodiment.
[0167] Further, in step 432, a consistent distributed optimization method is used to solve the current period scheduling decision, including the following steps:
[0168] Step 432-1, the multi-stage sequential decision problem of the Markov decision process of the reconstructed EV centralized scheduling model is decoupled into distributed sub-optimization problems related to each power generation or power consumption element through the Lagrange equation;
[0169] Specifically, the distributed sub-optimization problems include thermal power, photovoltaic, wind power, external grid, and EV sub-problems; and a power limit factor is added to the constraint conditions of the photovoltaic, wind power, external grid, and EV sub-problems;
[0170] In some cases, the marginal cost of the external grid or other elements will repeatedly oscillate around the critical value (i.e., the boundary point); if directly optimized, the behavior of this element can only jump between "maximum power purchase" and "maximum power sale", making it difficult for the system to converge; to avoid such "oscillation caused by non-continuous points", a power limit factor ( ) is introduced to limit the maximum output amplitude of the element and suppress the boundary jumping behavior of the optimization process.
[0171] Step 432-2, determine the distributed sub-optimization problem with the added power limit factor and the constraint of the added power limit factor;
[0172] For the method of determining the distributed sub-optimization problem with the added power limit factor, specifically:
[0173] First, all power generation or power consumption elements do not introduce a power limit factor for consistent iteration; if it can converge, it means that the period does not need a limit factor; otherwise, consistent iteration oscillates, and the marginal cost of each element will oscillate around the marginal cost Small-scale oscillations in the vicinity cause oscillations at the upper and lower limits of the element. At this point, the element that needs to be introduced with a constraint factor can be identified as the marginal element.
[0174] The constructed distributed sub-optimization problem is as follows:
[0175] 1) Optimization problem of thermal power unit:
[0176] (44);
[0177] 2) Photovoltaic sub-optimization problem:
[0178] (45);
[0179] 3) Wind power optimization issues:
[0180] (46);
[0181] 4) External network optimization problem:
[0182] (47);
[0183] 5) EV sub-optimization problem:
[0184] (48);
[0185] In the formula, , , , and The first Second iteration thermal power units Photovoltaic power supply Wind turbine External power grid and EV aggregate The corresponding marginal cost; , , and The first Second-generation photovoltaic power supply Wind turbine External power grid and EV aggregate The corresponding power limiting factor; It is a 6th-order diagonal matrix used to constrain EV power in the EV multicell model.
[0186] Step 432-3, each sub-optimization problem is respectively optimized, and the marginal cost, the power mismatch amount, the power limiting factor and the decision power of the distributed node element are updated through the local information of the adjacent distributed node element, and the iteration is continuously carried out until the convergence condition is met, and the current period scheduling decision is obtained ; the distributed node element includes a power generation element and a power consumption element, and the specific iteration process is as follows:
[0187] Step 432-3-1: constructing consistent optimization initial conditions, including initializing the marginal cost , the decision power , the power mismatch amount and the power limiting factor , so that the power system scheduling problem satisfies the power supply and demand balance constraint; specifically, the initialized parameters satisfy the power balance condition as formula (49):
[0188] Based on the consistency theory, the decision vector is set as:
[0189] ;
[0190] The marginal cost vector is set as
[0191] ;
[0192] The power mismatch vector is set as:
[0193] ;
[0194] Wherein, , , , , and are the power mismatch amounts of the thermal power unit , the external power grid , the wind power unit , the load , the photovoltaic power supply and the EV aggregate .
[0195] Optionally, the communication weight matrix is set, the adjacent node elements communicate with each other through the matrix , and only through these information, the sub-problems are continuously optimized, and the matrix is updated until convergence, that is, the consistent distributed optimization is realized; the adjacent node elements include power generation or power consumption elements, each power generation element can be a node, and each power consumption element can be a node;
[0196] The power balance condition (49) is:
[0197] (49);
[0198] wherein the superscript 0 in the formula denotes the initial state value, , , , and respectively denote the initial value of the power output of the thermal power unit, the power purchased / sold by the external grid node, the power consumption of the photovoltaic power source, the power consumption of the wind power unit, the power consumption of the load; and respectively denote the initial value of the charging / discharging power of the EV aggregation; and respectively denote the initial value of the charging / discharging power of the EV aggregation; , , , , and respectively denote the initial value of the power mismatch of the thermal power unit, the external grid, the wind power unit, the photovoltaic power source, the EV aggregation, and the load; denotes the initial value of the charging / discharging power vector of the EV aggregation. Step 432-3-1 is to construct an initial power solution, so that before the distributed solution of all elements, the overall power system is in a state of supply-demand balance or approximate balance; Step 432-3-2: through the information exchange between the adjacent node elements, and the feedback of the current node element power supply and demand state, the marginal cost and the power limiting factor
[0199] are updated adaptively, and the updating process is shown in formula (50), and the updated and
[0200] update the decision power and the power mismatch , as shown in formula (51); (50);
[0201] (51);
[0202] (51);
[0203] (52);
[0204] (53);
[0205] wherein, is a double stochastic matrix, and is a node incidence matrix; is a node incidence matrix; and are marginal cost and power limitation factor respectively, and is the update step, whose dimension is the same as the dimension of the variable to be updated. In formula (51), the update of the scheduling decision is completed through formulas (44) to (48);
[0206] The above update mechanism: when the power imbalance of the current node element is significant, the marginal cost estimation value of the node is automatically adjusted to force it to approach the regional average marginal value, so as to gradually realize the marginal cost consistency of the whole system; if a node element is frequently at the upper or lower limit of output in the scheduling, the power mismatching amount will oscillate in a small range near its marginal cost, at this time the power limitation factor will gradually shrink its adjustable interval, so as to prevent the system from oscillating due to boundary jumping. This mechanism does not require a central coordinator, but only relies on the local state and the communication of adjacent nodes to complete the parameter synchronization and adjustment between sub-elements, and has good distributed characteristics and scalability.
[0207] Step 432-3-3: judge whether the convergence criteria of marginal cost consistency and power balance are reached in the current iteration round; if yes, terminate the iteration and output the final scheduling result, and get the current period scheduling decision , otherwise continue the next iteration.
[0208] Convergence judgment is made based on formula (54):
[0209] and (54);
[0210] In formula (54), the first term is the marginal consistency check: it indicates that the marginal cost change between all elements has tended to be stable; the second term is the power balance check: the difference between the total power supply and demand of the current system is lower than the error limit, and the system has basically met the real-time balance. Among them, and are both set values;
[0211] If the second norm of the marginal cost and the power mismatching amount are both within the allowed range, the iteration converges successfully, which is the distributed optimization optimal solution based on the consistency theory; otherwise, return to step 432-3-2 and continue iteration.
[0212] In this embodiment, in order to realize distributed optimization scheduling, by introducing consistency theory and Lagrange factor, the centralized scheduling problem is converted into a plurality of distributed optimization problems of subsystems (such as thermal power, photovoltaic, wind power, external power grid and EV), each of which makes a local decision, and updates through communication coordination, updates the variables and parameters of the distributed unit through the local information of the adjacent distributed unit, and iterates constantly, and finally globally converges;
[0213] (2) Distributed ADP value function training in step 432: based on the current period scheduling decision, through disturbance analysis (such as formulas (55) to (57)), the slope is updated based on the local information of EV, including the following:
[0214] The value function slope updating method of the centralized ADP algorithm is shown in formulas (42) and (43), and the objective function in the value function training process is Using global information violates the principle of distributed optimization. To solve the above problems, the slope calculation is combined with the consistency algorithm based on the slope updating principle, and a completely distributed slope updating method is proposed, which can update the piecewise linear function slope only by using the information of EV itself.
[0215] Step 4321: based on the current period scheduling decision, initialize the local state information, and perform initialization operation for each EV, including: reading the power state at the current time ; read the prediction information in the planned scheduling period, which can include electricity price, PV output, load demand and the like; initialize the local value function slope estimation value;
[0216] Step 4322: generate a small state of charge disturbance for each EV , a small disturbance is applied to the SOC state of the EV itself , and the charging and discharging power strategy before and after the disturbance is calculated;
[0217] Step 4323: according to whether the disturbance causes different types of power changes, the sampling estimation value of the local slope is calculated, and the specific description is as follows:
[0218] As shown in formula (42), the physical meaning of the sampling estimation value is the marginal influence of SOC on the value function , which is actually a kind of special marginal cost. By increasing a small disturbance , the influence of the charging and discharging state and the charging and discharging efficiency of the EV on the system operating state is considered, the effect of the small disturbance on the value function is analyzed, and the marginal influence is further calculated distributedly by using the known information and local information to obtain the sampling estimation value , and the sampling estimation value of the slope is calculated in the following 3 cases:
[0219] 1) when a small disturbance without affecting the charging / discharging power of the EV aggregation , the operation of the system will not be affected, so the objective function of the time period will not change due to the disturbance, and the sample estimation value of the slope can be derived from equation (42) using only the EV side information, as shown in equation (55):
[0220] (55);
[0221] 2) When the charging power of the EV aggregation is affected, under the source-load power balance constraint (34), the change in charging power will inevitably cause the output power of other elements to change. According to the different power changing elements, it can be divided into the following 4 cases: i)
[0222] The change in charging power caused by the thermal power unit causes the power of the thermal power unit to change. This case can be identified by the change of the Lagrange factor in the consistency distributed optimization process, and the sample estimation value of the slope is calculated by equation (56) (a).
[0223] ii) The change in charging power caused by the thermal power unit causes the power of the thermal power unit to change. This case can be identified by the change of the Lagrange factor in the consistency distributed optimization process, and the sample estimation value of the slope is calculated by equation (56) (a).
[0224] iii) The change in charging power caused by the thermal power unit causes the power of the thermal power unit to change. This case can be identified by the change of the Lagrange factor in the consistency distributed optimization process, and the sample estimation value of the slope is calculated by equation (56) (a).
[0225] iv) The change in charging power caused by the thermal power unit causes the power of the thermal power unit to change. This case can be identified by the change of the Lagrange factor in the consistency distributed optimization process, and the sample estimation value of the slope is calculated by equation (56) (a).
[0226] (56);
[0227] 3) When the discharging power of the EV aggregation is affected, by the same reasoning, according to the different power changing elements, it can be divided into the following 4 cases:
[0228] i) When the change of the EV discharging power caused by the change of the power of the thermal power unit, the sample estimation value of the slope is solved and calculated by equation (57) (a);
[0229] ii) When the change of the EV discharging power caused by the change of the power purchase and sale of the external power grid, the sample estimation value of the slope is solved and calculated by equation (57) (b);
[0230] iii) When the change of the EV discharging power caused by the change of the output of the photovoltaic power source, the sample estimation value of the slope is solved and calculated by equation (57) (c);
[0231] iv) When the change of the EV discharging power is compensated by the change of the output of the wind power unit, the sample estimation value of the slope is solved and calculated by equation (57) (d);
[0232] (57);
[0233] As shown in equation (56) and equation (57), the parameters used in the calculation process are all EV side information or known information, the sample estimation value is obtained by equation (56) distribution, and on the basis of the slope increasing property of the CAVE algorithm, the convergence of the distributed value function training method is ensured, and then the distributed training of the value function is realized. The method cooperates with the consistency algorithm, and finally a completely distributed ADP real-time scheduling framework is constructed.
[0234] The distributed day-ahead offline training mainly includes EV aggregation, day-ahead training scene generation, consistent distributed optimization scheduling and distributed ADP value function training. The specific training iteration example steps are as shown in Figure 3 , which can be as follows:
[0235] 1) Obtain the historical state data of the day-ahead power system and the EV charging and discharging scheduling data, use the K-means++ algorithm for the EV cluster, construct the EV matrix model (7), use the polytope approximation aggregation method to form the EV aggregation, equations (14) and (15);
[0236] 2) Set the training number NA , generate the training scene, initialize the value function slope, and set the current training number ;
[0237] 3) For the first training, set the current time period ;
[0238] 4) Determine the first training Periodic power system operation state, based on the consistent distributed optimization method, step 432-3 iteration to solve the current period scheduling decision ;
[0239] 5) Based on the distributed ADP value function training method and CAVE algorithm, update the value function slope of the first training period by equations (55) to (57);
[0240] 6) Update the time, if it does not exceed the time boundary then return 4), otherwise, go to the next step;
[0241] 7) Update the training times, if it does not exceed the training times then return 3), otherwise, go to the next step;
[0242] 8) Output the trained piecewise linear value function slope.
[0243] In step 5, the consistent distributed optimization method is used to solve the ADP value function optimization problem, and the current period scheduling decision is obtained, the specific steps are as follows:
[0244] 1) Input the piecewise linear value function slope obtained by day-ahead training;
[0245] 2) Initialize the time, ;
[0246] 3) Get the actual operation state of the power system in the period, that is, the operation state, and use the consistent distributed optimization method in step 432 to iteratively solve the current period scheduling decision ;
[0247] 4) Execute the scheduling decision, wait for the next period, update the time, if it does not exceed the time boundary then return 3), otherwise, the process is ended.
[0248] The distributed ADP algorithm proposed in this embodiment includes two cooperative stages: in the day-ahead offline training stage, the distributed value function training is realized by fusing the consistency algorithm and the value function slope updating mechanism; in the real-time scheduling stage, based on the trained value function representing the state-decision coupling relationship across periods, the distributed cooperative optimization is completed with the help of the consistency algorithm. Through the two-stage consistency interaction architecture, the algorithm effectively avoids the grid structure limitation and the problem that day-ahead training still needs global information of the existing method.
[0249] To verify the effectiveness of the model and method proposed in this paper, the improved IEEE6 node system is selected as the test system, including thermal power unit nodes, external grid AC nodes, wind turbine nodes, load nodes, photovoltaic power supply nodes and EV charging device nodes. The maximum and minimum output of the thermal power unit in the system is 40 MW and 10 MW respectively, and the maximum ramp rate of the unit is 25 MW / h. The cost coefficients of the thermal power unit are , and . The maximum purchase and sale power of the external grid is set to 15 MW. The parameters of the selected EV are shown in Table 1, the day-ahead forecast values of the load, electricity price, wind power output and photovoltaic output and the training scene are considered, the EV charging continuity is considered, the 8:00 is taken as the starting point of the real-time scheduling within a day, the forecast value also starts from 8:00, and the step is 1 h, which is divided into 24 periods. The existing load and electricity price prediction method is relatively mature, and under the confidence level of 90%, the confidence interval of the load and electricity price prediction error is set to , the confidence interval of the wind power and photovoltaic output prediction error is ±10%, and 500 training scenes are generated by using Monte Carlo simulation for the training of the value function of the distributed ADP algorithm. The penalty cost coefficient of abandoned wind and light is set to .
[0250] Table 1 EV parameter setting;
[0251]
[0252] 1.1 Verification of effectiveness of distributed ADP value function training method
[0253] To verify the effectiveness of the distributed ADP value function training method used, the day-ahead training method of the embodiment is used for 50 times of training, and the slope obtained by the centralized value function training method is compared, and the result is shown in Figure 2 .
[0254] As can be seen from Figure 2 , the relative error of the distributed value function training method used in the embodiment and the true value is very small, the maximum relative error is only , and the average relative error is only , and the order of magnitude is very small, therefore, the value function training effect is equivalent to the centralized training effect, and has effectiveness.
[0255] 1.2 Performance analysis of distributed ADP algorithm under conventional scene
[0256] To verify the superiority of the overall performance of the proposed EV real-time scheduling method based on distributed ADP, based on the parameter settings in Table 1, 100 random scenes are generated again using the Monte Carlo simulation method, and the minimum operation cost and the adjustable potential of thermal power units of each scene are obtained by using the centralized global optimal algorithm, the distributed greedy algorithm and the distributed ADP algorithm proposed in this embodiment to optimize the generated scenes respectively, and the results of the two distributed algorithms are compared with the centralized global optimal algorithm respectively, and the relative error is as shown in Figure 4 The average error and the maximum error are shown in Table 2.
[0257] Table 2 shows the average error and the maximum error of the two distributed algorithms;
[0258]
[0259] From Figure 4 and Table 2, it can be seen that the relative error of the distributed greedy algorithm is basically maintained at 6%-8%, and the relative error of the distributed ADP algorithm proposed in this paper is below 1%. The average relative error of the distributed greedy algorithm and the distributed ADP algorithm is 7.05% and 0.43% respectively. In the distributed random scene, the use of the distributed ADP algorithm proposed in this paper can reduce the operation cost by about 5%-8%. It can be seen that the distributed ADP algorithm proposed in this paper is obviously superior to the distributed greedy algorithm, and it has excellent performance for most random scenes and good applicability for possible random scenes within a day.
[0260] Embodiment 2
[0261] Based on Embodiment 1, the electric vehicle real-time charging and discharging scheduling system based on distributed ADP is provided in this embodiment, and the system comprises:
[0262] The multi-cell charging and discharging model construction module is configured to divide all EVs into a plurality of cluster groups, construct the charging and discharging power set of each EV into a multi-cell, and aggregate each cluster group using the multi-cell approximate aggregation method to obtain the constructed EV multi-cell charging and discharging model;
[0263] The EV centralized scheduling model construction module is configured to minimize the total operation cost of the power grid and maximize the adjustable potential of the thermal power unit as the target, and the constraint conditions include the EV charging and discharging constraints represented based on the EV multi-cell charging and discharging model, and the aggregated EV centralized scheduling model is constructed;
[0264] The problem conversion module is configured to reconstruct the EV centralized scheduling model into a Markov decision process, and convert the scheduling problem into a multi-stage sequential decision problem;
[0265] The day-ahead training module is configured to acquire day-ahead power system historical state data and EV charging and discharging scheduling data, perform day-ahead offline training, solve a multi-stage sequential decision problem through iteration by combining a consistent distributed optimization method and an ADP algorithm, and obtain a segmented linear value function slope of the trained ADP distributed value function.
[0266] The intra-day optimization module is configured to acquire a current period operation state of the power system, solve the multi-stage sequential decision problem by using the consistent distributed optimization method according to the trained ADP distributed value function, obtain a current period scheduling decision, and obtain a charging and discharging strategy of each EV based on the EV multi-cell charging and discharging model.
[0267] It should be noted that each module in the embodiment corresponds to each step in Embodiment 1 one by one, and the specific implementation process is the same, which will not be repeated here.
[0268] Embodiment 3
[0269] Based on Embodiment 1, the embodiment provides a distributed ADP-based real-time charging and discharging scheduling system for electric vehicles, which comprises:
[0270] A data acquisition device and a processor;
[0271] The data acquisition device is configured to acquire an operation state of the power system and EV charging and discharging regulation data.
[0272] The processor is configured to perform the steps of the distributed ADP-based real-time charging and discharging scheduling method for electric vehicles in Embodiment 1.
[0273] The above only describes the preferred embodiments of the present application and is not used to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A real-time charging and discharging scheduling method for electric vehicles based on distributed ADP, characterized in that, The method comprises the following steps: All EVs are divided into multiple cluster groups, and the charging and discharging power set of each EV is constructed as a polytope. The polytope approximation aggregation method is used for aggregation of each cluster group to obtain the constructed EV polytope charging and discharging model; A centralized scheduling model of EVs is constructed by taking minimization of total operation cost of the power grid and maximization of adjustable potential of thermal power units as targets, and by taking EV charging and discharging constraints represented based on the EV polytope charging and discharging model as constraint conditions; The centralized scheduling model of EVs is reconstructed as a Markov decision process, and the scheduling problem is converted into a multi-stage sequential decision problem; Day-ahead power system historical state data and EV charging and discharging scheduling data are acquired for day-ahead offline training. The multi-stage sequential decision problem is iteratively trained and solved by combining a consistent distributed optimization method and an ADP algorithm to obtain a segmented linear value function slope of the trained ADP distributed value function; The operation state of the power system in the current period is acquired. The multi-stage sequential decision problem is solved by using the consistent distributed optimization method based on the trained ADP distributed value function to obtain a scheduling decision in the current period. The charging and discharging strategy of each EV is obtained based on the de-aggregation of the EV polytope charging and discharging model; The EV polytope charging and discharging model is constructed as follows: A charging and discharging model is constructed for each electric vehicle. The feasible charging and discharging power set of the electric vehicle is represented as a convex polytope defined in a multi-dimensional space based on the charging and discharging model constraints; Based on the plug-in time and the pull-out time of each EV, a clustering algorithm is used to divide all EVs into K cluster groups. For each cluster k, the average operating parameters of the EV individuals in the cluster are counted to construct a basic polytope set representing the characteristics of the feasible region of the polytope of the EVs in the cluster. Based on the polytope approximation method, the approximate polytope of each polytope is transformed based on the basic polytope set. The polytope in the cluster group is aggregated and de-aggregated to obtain the EV polytope charging and discharging model. Based on the polytope approximation method, the approximate polytope of each polytope is transformed based on the basic polytope set. The polytope in the cluster group is aggregated and de-aggregated, including the following steps: The accurate feasible region of the aggregate is represented as a Minkowski sum form; The polytope approximation aggregation is used to transform the same basic polytope set to represent the polytope feasible region of all EVs to obtain the approximate polytope of each polytope. The approximate polytopes of all EVs are summed based on the Minkowski sum form to obtain the aggregate polytope. The obtained aggregate polytope is decoupled by time period, and the high-dimensional aggregate polytope is disassembled into a single-dimensional polytope corresponding to each time to obtain the EV polytope charging and discharging model. The centralized scheduling model of EVs is reconstructed as a Markov decision process, including the following: The prediction errors of the output of the traditional thermal power unit, the output of the photovoltaic power, the output of the wind power, the electricity price and the load are taken as state variables. The scheduling output decision of each energy is taken as a decision variable. The state of the current period is calculated based on the state of the previous period to construct a state transition function of the Markov decision. Based on the obtained state transition function and Bellman optimal principle, the objective function is reconstructed to obtain a multi-stage sequential decision problem.
2. The distributed ADP-based real-time charging and discharging scheduling method for electric vehicles according to claim 1, characterized in that: The day-ahead offline training method comprises the following steps: Obtain historical state data of the day-ahead power system and EV charging and discharging regulation data, construct an EV polytope charging and discharging model, and form an EV aggregate by using a polytope approximation aggregation method; Generate a plurality of training scenarios according to day-ahead predicted values of the EV charging and discharging regulation data; For each training scenario data, the constructed value function optimization problem is iteratively solved by using a consistent distributed optimization method to obtain an optimal decision for each time period; and the value function slope is iteratively trained based on the optimal decision scheme of each time period to obtain a final value function slope.
3. The distributed ADP-based real-time charging and discharging scheduling method for electric vehicles according to claim 2, characterized in that: The process of iteratively training the value function slope based on the optimal decision scheme of each time period to obtain a final value function slope is as follows: The multi-stage sequential decision problem is converted into an ADP piecewise linear form by using an approximate dynamic optimization method to obtain an ADP piecewise linear value function; For the obtained ADP piecewise linear value function, the following operations are performed for each training scenario and each time period under each training scenario: The ADP piecewise linear function optimization problem is solved by using a consistent distributed optimization method to obtain a current time period scheduling decision; Based on the current time period scheduling decision, distributed ADP value function training is performed, disturbance analysis is performed based on state of charge change as a disturbance quantity, and the ADP value function slope is updated based on local information of the EV.
4. The distributed ADP-based real-time charging and discharging scheduling method for electric vehicles according to claim 1 or 3, characterized in that, The ADP value function optimization problem is solved by using a consistent distributed optimization method, including the following steps: The multi-stage sequential decision problem of the Markov decision process of the reconstructed EV centralized scheduling model is decoupled into distributed sub-optimization problems related to each power generation or power consumption element by using a Lagrange equation; The distributed sub-optimization problem of determining the power limit factor is determined, and the constraint of the power limit factor is added; Each sub-optimization problem is optimized respectively, the marginal cost, power mismatch amount, power limit factor and decision power of the distributed node element are updated based on local information of adjacent distributed node elements, and the iteration is continuously performed until the convergence condition is met to obtain a current time period scheduling decision; the iteration process is as follows: Consistent optimization initial conditions are constructed, including initialization of the marginal cost, decision power, power mismatch amount and power limit factor, so that the power system scheduling problem satisfies the power supply and demand balance constraint; The marginal cost and power limit factor are adaptively updated through information exchange between adjacent node elements and feedback of the power supply and demand state of the current node element; It is judged whether the convergence standards of marginal cost consistency and power balance are reached in the current iteration round; if yes, the iteration is terminated and the final scheduling result is output to obtain a current time period scheduling decision.
5. The distributed ADP-based real-time charging and discharging scheduling method for electric vehicles according to claim 3, characterized in that: Based on the current time period scheduling decision, distributed ADP value function training is performed, disturbance analysis is performed based on state of charge change as a disturbance quantity, and the ADP value function slope is updated based on local information of the EV, including the following steps: Based on the current period scheduling decision, initialize the local state information, and perform initialization operation for each EV; Generate a small charge state disturbance for each EV, apply disturbance to the SOC state of the EV itself, and calculate the charge and discharge power strategy before and after the disturbance; According to whether the disturbance causes different types of power changes, calculate the sample estimation value of the local slope.
6. The distributed ADP-based real-time charging and discharging scheduling system for electric vehicles according to claim 1, characterized in that, It includes: A polytope charge and discharge model construction module configured to divide all EVs into multiple cluster groups, construct the charge and discharge power set of each EV as a polytope, and aggregate each cluster group using the polytope approximate aggregation method to obtain the constructed EV polytope charge and discharge model; An EV centralized scheduling model construction module configured to minimize the total operating cost of the power grid and maximize the adjustable potential of the thermal power unit as the target, and the constraint conditions include the EV charge and discharge constraints represented based on the EV polytope charge and discharge model, and construct the aggregated EV centralized scheduling model; A problem conversion module configured to reconstruct the EV centralized scheduling model as a Markov decision process and convert the scheduling problem into a multi-stage sequential decision problem; A day-ahead training module configured to obtain day-ahead power system historical state data and EV charge and discharge scheduling data, perform day-ahead offline training, and solve the multi-stage sequential decision problem through the combination of consistent distributed optimization method and ADP algorithm to obtain the segmented linear value function slope of the trained ADP distributed value function; An intra-day optimization module configured to obtain the operating state of the intra-day power system in the current period, solve the multi-stage sequential decision problem using the consistent distributed optimization method based on the trained ADP distributed value function, obtain the current period scheduling decision, and perform de-aggregation based on the EV polytope charge and discharge model to obtain the charge and discharge strategy of each EV.
7. The real-time charging and discharging scheduling system for electric vehicles based on distributed ADP, characterized in that, It includes: A data acquisition device and a processor; The data acquisition device is used to obtain the operating state of the power system and the EV charge and discharge regulation data; The processor is configured to perform the steps of the distributed ADP-based real-time charge and discharge scheduling method for electric vehicles according to any one of claims 1-5.
Citation Information
Patent Citations
Electric vehicle load dynamic distribution prediction method based on vehicle network development inherent parameters
CN117010088A
Online scheduling method for industrial park comprehensive energy system containing electric vehicle load
CN118134228A