Micro-grid optimization scheduling method of multi-target particle swarm combining SARSA and electricity price regulation and control
By combining SARSA reinforcement learning and multi-objective particle swarm optimization algorithm, the inertia weight is dynamically adjusted to generate the optimal solution set. Through a two-stage sub-Bruker optimization model, the problem of coordinated optimization of economic cost, environmental emissions and voltage stability of microgrids in low-inertia environments with high proportion of renewable energy and electric vehicles is solved, realizing a flexible and efficient scheduling scheme.
Patent Information
- Application Number
- CN202511448181.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-11
- Publication Date
- 2026-01-23
AI Technical Summary
Traditional microgrid optimization and scheduling methods are difficult to coordinate and optimize economic costs, environmental emissions, and voltage stability in low-inertia microgrids with a high proportion of renewable energy and random access to electric vehicles. Furthermore, they do not fully utilize the potential of electric vehicle V2G and are prone to getting trapped in local optima and facing the risk of instability.
By combining SARSA reinforcement learning and multi-objective particle swarm optimization algorithm, the inertia weight is dynamically adjusted to generate the optimal solution set. Then, through a two-stage sub-Bruker optimization model, a robust optimal scheduling scheme is generated. Taking into account the power output error of new energy sources and the net energy error of electric vehicle clusters, the electricity price regulation is optimized to guide EV charging and discharging.
It achieves efficient synergistic optimization of the economy, environmental protection and voltage stability of microgrids in complex environments, improves voltage stability, adapts to the randomness of EV charging and the dynamic changes of microgrids, generates flexible and efficient scheduling schemes, and supports the 'dual carbon' target.
Smart Images

Figure CN121390408A_ABST
Abstract
Description
[0001] Microgrid optimal scheduling method combining SARSA and electricity price regulation multi-objective particle swarm optimization TECHNICAL FIELD
[0002] The application belongs to the technical field of smart grid, more specifically, relates to a microgrid optimal scheduling method combining SARSA and electricity price regulation multi-objective particle swarm optimization. BACKGROUND
[0003] With the rapid development and popularization of renewable energy represented by photovoltaic and wind energy and electric vehicles, microgrid as a key system that can effectively integrate these distributed energy sources plays a crucial role in promoting green energy transformation. A typical microgrid usually contains photovoltaic power generation units, energy storage systems, gas turbines and other devices. In order to achieve the best balance between economic cost, environmental emissions and grid voltage stability, these devices need to be finely optimized and scheduled. With the increase of new energy into the grid, its intermittency and volatility greatly increase the frequency and difficulty of grid scheduling. Under the background of the continuous rise of new energy vehicle ownership and charging power in China, the influence of its charging behavior on the grid is increasingly significant, and the voltage stability problem is particularly prominent. Although the traditional microgrid scheduling method attempts to manage the charging behavior of electric vehicles through demand response (DR) mechanism, it mainly focuses on peak clipping and valley filling, often ignoring the inherent randomness of electric vehicle charging behavior, and not fully realizing that electric vehicles can be used as mobile energy storage devices to participate in grid scheduling.
[0004] In this field, the existing methods still have the following technical problems: the traditional MOPSO algorithm is easy to fall into local optimum in multi-objective optimization, and it is difficult to balance economic cost, environmental emissions and voltage stability, especially in low inertia (note: refers to high proportion of power electronic devices and few rotating devices, resulting in weak inertial response ability of the system to power fluctuation) microgrid. Insufficient dynamic adaptability: Q-learning relies on predicted optimal actions and ignores actual execution feedback, making it difficult to cope with the randomness of EV charging (such as access time, location, battery state) and the dynamic changes of distribution network (such as operation mode switching). EV charging management is limited: existing electricity price regulation only focuses on demand response, ignores the impact of EV charging on distribution network power flow distribution and voltage stability, and does not fully utilize the discharge potential of V2G mode. Multi-objective conflict: economic cost, emissions and voltage stability targets conflict, existing methods lack comprehensive optimization mechanism, making it difficult to generate diversified solution set. Voltage stability is insufficient: insufficient consideration of the impact of microgrid low inertia, high proportion of power electronic devices, low line impedance ratio (X / R), load concentration and operation mode switching, resulting in voltage fluctuation or instability risk.
[0005] In summary, existing microgrid optimization and scheduling methods are insufficient to synergistically optimize economic efficiency, environmental friendliness, and voltage stability in low-inertia microgrids with a high proportion of renewable energy and random access to electric vehicles. Furthermore, they lack effective utilization of the V2G potential of electric vehicles, which makes the system prone to local optima and faces the risk of instability.
[0006] Therefore, improving the efficiency and robustness of microgrid optimal scheduling is a technical problem that urgently needs to be solved. Summary of the Invention
[0007] To address the shortcomings of existing technologies, this application aims to provide a microgrid optimization scheduling method that combines SARSA and electricity price regulation with multi-objective particle swarm optimization. This method is intended to solve the problem that traditional multi-objective optimization algorithms are prone to getting trapped in local optima when dealing with multi-objective optimization problems, making it difficult to achieve efficient and robust scheduling.
[0008] To achieve the above objectives, in a first aspect, this application provides a microgrid optimal scheduling method combining SARSA and electricity price regulation with multi-objective particle swarm optimization, comprising: A microgrid EV model is established and the optimization objective function of the model is defined. The operating constraints of the microgrid system are set. The optimization objective function includes an economic cost objective function, an environmental emission objective function, and a voltage stability objective function. An improved multi-objective particle swarm optimization algorithm is constructed based on the SARSA reinforcement learning algorithm. The optimization objective function is solved based on the multi-objective particle swarm optimization algorithm under the running constraints to generate the optimal solution set. Based on the optimal solution set, a two-stage sub-Blu-rod optimization model considering the power output error of new energy sources and the net energy error of EV clusters is constructed. The optimal scheduling scheme is obtained by solving the two-stage sub-Blu-rod optimization model.
[0009] Optionally, the process of generating the optimal solution set includes: The SARSA update strategy for inertia weights is determined by constructing a state vector by calculating the rate of change of the objective function and updating the value evaluation table based on the reward function and temporal difference learning. The inertial weights are determined based on the SARSA update strategy. The particle velocity and position are updated by combining the particle's own optimal position and the group's optimal position. The optimal solution set is obtained by filtering using an adaptive grid method.
[0010] Optionally, the SARSA algorithm implementation process includes: Calculate the rate of change of the optimization objective function under the target number of iterations, and construct the state vector for reinforcement learning based on the rate of change; Define an action space, which includes increasing the inertia weight, decreasing the inertia weight, and keeping the inertia weight unchanged; constructing a reward function, and obtaining calculation reward data based on an update of the solution set according to the reward function; determining a current value estimation under a current state vector and a new value estimation of a new state, updating a value evaluation table of the reinforcement learning by using a time difference learning algorithm according to the current value estimation and the new value estimation, a learning rate, a discount factor and the reward data, to determine a SARS A update strategy.
[0011] Optionally, the particle optimization algorithm implementation process comprises: determining an inertia weight based on the SARS A, performing linear superposition on a historical optimal position of the particle and a group optimal position based on the inertia weight, updating a velocity vector of the particle, and obtaining a flight speed of a new generation of particles; performing particle position updating and constraint processing based on the flight speed of the new generation of particles and a position vector of the particle, and obtaining a new position of the particle as a new scheduling scheme; adopting an adaptive grid method to uniformly divide a grid for the multi-objective space, and screening solutions corresponding to the new scheduling scheme to obtain the optimal solution set.
[0012] Optionally, based on the optimal solution set, a two-stage distribution robust optimization model considering new energy output error and EV cluster net energy error is constructed, and the two-stage distribution robust optimization model is solved to obtain an optimal scheduling scheme, comprising: selecting a target scheduling scheme from the optimal solution set; constructing an error set based on the new energy output error and the EV cluster net energy error, assigning the error set with uncertainty, and converting the error set to an uncertainty set by a Wasserstein distance method; establishing a two-stage distribution robust optimization model based on the microgrid EV model and the uncertainty set; inputting the target scheduling scheme into the two-stage distribution robust optimization model, determining a scheduling plan through a first stage, and adjusting the scheduling plan according to actual uncertainty through a second stage; converting the two-stage distribution robust optimization model into a mixed integer linear programming problem, calling a mathematical optimization solver to solve the mixed integer linear programming problem, and outputting an optimal robust optimal scheduling scheme.
[0013] Optionally, the operation constraint conditions comprise power balance constraints, energy storage system capacity constraints, EV charging cluster power upper and lower limit constraints, EV charging cluster net energy constancy constraints in a scheduling period, power ramping constraints, reactive power margin constraints, equivalent short-circuit ratio constraints, and voltage lower limit constraints.
[0014] Optionally, the economic cost objective function is as shown in the following formula:
[0015] where, is the total economic cost, is the set of number of distributed generators such as gas turbines, is the set of number of microgrid nodes, is the set of number of all devices, is the unit generation cost of gas turbine, is the power of gas turbine, is the time step, is the grid purchase price of node , is the grid purchase power of node , is the unit power operation and maintenance cost of device , is the power of device at time , is the sale price of node , is the sale power to grid of node , is the dynamic price of the th EV charging cluster,
[0016] Optionally, the environmental emission objective function is as shown in the following formula:
[0017] where, is the total emission, is the set of number of all devices, is the pollutant type, is the emission coefficient of device producing pollutant , is the generation power of device , is the time step.
[0018] Optionally, the voltage stability objective function is as shown in the following formula:
[0019] where, is the voltage stability objective, is the voltage of node at time , is the reference voltage, is the active power and reactive power flowing through the line connected with node , respectively, are the resistance and reactance of the line connected to the node are the reactive loads of the node are the weight factors determined by the entropy weight method, are the operation mode weights, is the lower limit of the voltage.
[0020] In a second aspect, the present application provides an electronic device, comprising: at least one memory for storing a program; and at least one processor for executing the program stored in the memory, wherein the processor is configured to execute the method described in the first aspect or any possible implementation manner of the first aspect.
[0021] In a third aspect, the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is run on a processor, the processor is caused to execute the method described in the first aspect or any possible implementation manner of the first aspect.
[0022] In a fourth aspect, the present application provides a computer program product, and when the computer program product is run on a processor, the processor is caused to execute the method described in the first aspect or any possible implementation manner of the first aspect.
[0023] It can be understood that the beneficial effects of the second aspect to the fourth aspect described above can be referred to the related description in the first aspect, and will not be repeated here.
[0024] Overall, compared with the prior art, the above technical solutions conceived by the present application have the following beneficial effects: (1) By introducing the SARS A reinforcement learning mechanism, the present application can perceive the optimization process of each target in real time, increase the inertia weight by dynamically adjusting the search strategy to expand the exploration range or reduce the weight to develop finely, and through the adaptive learning ability based on feedback, it ensures that the algorithm can effectively jump out of the local optimal trap, systematically search the entire feasible region, and generate an optimal solution set with wide distribution and good convergence, and then obtain the optimal scheduling scheme. By fusing SARS A reinforcement learning and multi-objective particle swarm optimization, and introducing two-stage distribution robust optimization, the efficient collaborative optimization and robust scheduling of microgrid economy, environmental protection and voltage stability in a complex uncertain environment are realized.
[0025] (2) The application monitors the optimization state of the three objective functions of economy, environmental protection and stability in real time through the SARSA algorithm, and then intelligently selects to increase, decrease or maintain the inertia weight, a key parameter, the dynamic adjustment mechanism of the application has the ability of autonomous optimization between global exploration and local development, can effectively avoid the search process converging to a suboptimal solution area too early, so as to ensure that the final generated solution set not only comprehensively covers various optimal trade-off schemes between the three objectives. And the application is suitable for the randomness of EV charging and the dynamic changes of microgrid such as photovoltaic output fluctuation and operation mode switching, and provides flexible and efficient scheduling scheme.
[0026] (3) The application can use dynamic price regulation to guide EVs to charge at low valley and discharge at peak (V2G), optimize power distribution network flow distribution, and improve voltage stability; comprehensively consider the influence of microgrid low inertia, low line impedance ratio, load reactive power demand and operation mode switching, and enhance voltage stability.
[0027] (4) The application can realize multi-objective optimization of economy, environmental protection and grid stability, and support the "double carbon" target. BRIEF DESCRIPTION OF DRAWINGS
[0028] Figure 1 is one of the flow schematic diagrams of the microgrid optimization scheduling method of the multi-objective particle swarm combined with SARSA and price regulation provided by the embodiment of the application; Figure 2 is a system architecture diagram of the microgrid optimization scheduling method of the multi-objective particle swarm combined with SARSA and price regulation provided by the embodiment of the application; Figure 3 is the second flow schematic diagram of the microgrid optimization scheduling method of the multi-objective particle swarm combined with SARSA and price regulation provided by the embodiment of the application; Figure 4 is a price regulation effect diagram provided by the embodiment of the application; Figure 5 is a constant net energy of EV cluster diagram provided by the embodiment of the application; Figure 6 is a structural schematic diagram of an electronic device provided by the embodiment of the application. DETAILED DESCRIPTION
[0029] In order to make the purpose, technical scheme and advantages of the application more clear and explicit, the application will be further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the application and not to limit the application.
[0030] The term "and / or", used in the present document, is used to describe the association relationship of associated objects, which means that there can be three relationships, for example, A and / or B can represent the following three cases: A exists alone, A and B exist together, and B exists alone. The symbol " / " in the present document represents an or relationship of associated objects, for example, A / B represents A or B.
[0031] The terms "first" and "second" and the like in the description and claims of the present document are used to distinguish different objects, rather than to describe a specific order of the objects. For example, the first response message and the second response message are used to distinguish different response messages, rather than to describe a specific order of the response messages.
[0032] In the embodiments of the present application, the words "exemplary" or "for example" are used to mean serving as an example, instance, or illustration, and not necessarily to imply any preference or superiority. In fact, an embodiment or design scheme described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or superior than other embodiments or design schemes.
[0033] In the description of the embodiments of the present application, unless otherwise specified, "a plurality of" means two or more, for example, a plurality of processing units means two or more processing units, and the like; a plurality of elements means two or more elements, and the like.
[0034] The embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application.
[0035] Referring to Figure 1 The present application provides a micro-grid optimization scheduling method combining SARSA and electricity price regulated multi-objective particle swarm, comprising: S101. Establishing a micro-grid EV model and defining an optimization objective function of the model, setting operation constraint conditions of the micro-grid system, and the optimization objective function including an economic cost objective function, an environmental emission objective function, and a voltage stability objective function; S102. Constructing an improved multi-objective particle swarm optimization algorithm according to the SARSA reinforcement learning algorithm, solving the optimization objective function under the operation constraint conditions based on the multi-objective particle swarm optimization algorithm, and generating an optimal solution set; S103. Based on the optimal solution set, constructing a two-stage distribution robust optimization model considering new energy output error and EV cluster net energy error, and solving the two-stage distribution robust optimization model to obtain an optimal scheduling scheme.
[0036] Specifically, first, a system model that can accurately reflect the physical characteristics of the micro-grid needs to be constructed. The model core includes photovoltaic power generation units, energy storage systems, gas turbines, and other traditional devices, and innovatively aggregates a large number of electric vehicles (EVs) in the region into one or more charge clusters that can be uniformly dispatched.
[0037] Referring to Figure 2 , Figure 2 is a system architecture diagram of an embodiment of the present application, showing the micro-grid structure: photovoltaic, energy storage, gas turbine, EV charging cluster, power distribution network node. The EV charging cluster interacts with the power grid through V2G, and the electricity price regulation module adjusts , optimizes the power flow distribution and voltage stability.
[0038] On this basis, three key optimization objective functions are clearly defined: the economic cost objective function aims to minimize the total operating cost of the system, including power generation fuel cost, operation and maintenance cost, power grid power purchase cost minus power sales revenue, and EV cluster charging and discharging electricity cost; the environmental emission objective function aims to quantify and minimize the total emissions of pollutants such as carbon dioxide and nitrogen oxides generated during operation; the voltage stability objective function is a comprehensive index for evaluating and improving the stability of the grid voltage, which takes into account factors such as voltage deviation, line voltage drop, and reactive load impact.
[0039] At the same time, to ensure the feasibility and safety of the dispatching scheme, strict system operation constraints need to be set, including power balance constraints, energy storage system capacity constraints, EV charging cluster power upper and lower limit constraints, EV charging cluster net energy constancy constraints within the dispatching period, power ramping constraints, reactive power margin constraints, equivalent short-circuit ratio constraints, and voltage lower limit constraints.
[0040] Secondly, an improved multi-objective particle swarm optimization (MOPSO) algorithm is used, the key improvement of which is the deep integration of the SARSA reinforcement learning algorithm. In the optimization process, the SARSA algorithm is used to monitor the optimization progress of the three objectives of economy, environment, and voltage stability in real time, dynamically decide and adjust the key inertia weight parameter in the MOPSO algorithm, to balance the global exploration and local development capabilities of the algorithm, and effectively avoid the defect that the traditional algorithm is easy to fall into local optimum.
[0041] Under the guidance of SARSA, the MOPSO algorithm searches in the feasible region that meets all the operation constraints, constantly updates the candidate dispatching scheme, and finally outputs a Pareto optimal solution set.
[0042] Finally, step S103 is the decision and robustness improvement stage of the scheme, which aims to select a final scheme from multiple excellent schemes that can best resist the uncertainty of actual operation. First, select an optimal scheduling scheme as a benchmark from the Pareto optimal solution set generated in step S102.
[0043] A two-stage distribution robust optimization model is constructed. The first stage of the model is to develop a day-ahead scheduling plan based on the selected optimal scheme. The second stage considers the error between the actual output and the predicted value of new energy, as well as the error between the actual net energy demand and the predicted value of the EV cluster, and constructs a fuzzy set to describe the probability distribution range of these uncertainties based on historical error data. The optimization objective of the model is to find a robust scheme with the minimum adjustment cost considering the worst probability distribution.
[0044] Finally, the two-stage DRO model is converted into a standard mixed integer linear programming problem, and a professional mathematical optimization solver is called to solve it. Finally, an optimal scheduling scheme that inherits the superiority of Pareto solution and has strong robustness to prediction error is output to guide the actual safe, economic and green operation of the microgrid.
[0045] Optionally, the economic cost objective function is as shown in the following formula:
[0046] wherein, is the total economic cost, is a set of the number of distributed generators such as gas turbines, is a set of the number of microgrid nodes, is a set of the number of all devices, is the unit power generation cost of the gas turbine, is the power of the gas turbine, is the time step, is the node grid purchase power price, grid purchase power, is the unit power operation and maintenance cost of the device is the power of the device at time is the node selling price, selling power to the grid, is the dynamic electricity price of the EV charging cluster is the charging / discharging power of the EV charging cluster , positive for charging and negative for discharging, represented as:
[0047] The economic cost objective function is defined as follows: calculating the total operating cost of a microgrid, including gas turbine power generation costs, grid purchase costs, and equipment operation and maintenance costs, minus revenue from electricity sales (such as V2G discharge or sales of surplus photovoltaic power), and adding the cost of EV charging clusters. The cluster approach utilizes dynamic electricity pricing. Simplify scheduling and enhance adaptability to EV stochasticity.
[0048] Optionally, the environmental emission objective function is shown in the following formula:
[0049] in, Total emissions, The set of all device counts. As for the type of pollutant, For equipment Pollutants are generated The emission coefficient, For equipment Power generation capacity, For time step.
[0050] The environmental emission objective function is used to measure the emissions generated by different equipment (gas turbines, photovoltaics, etc.). and Emissions, taking into account multiple pollutants.
[0051] Optionally, the voltage stability objective function is shown in the following formula:
[0052] in, For voltage stability objectives, For nodes In time voltage, For reference voltage, They flow through and nodes respectively The active and reactive power of the connected lines, They are respectively with nodes The resistance and reactance of the connected circuits, For nodes reactive load, The weighting factors are determined using the entropy weighting method. Weights for operating modes This is the lower limit of voltage.
[0053] It should be noted that the weight factor is determined by the entropy weight method (Note: the entropy weight method is an objective weighting method, which calculates the weight based on the information entropy size of each target component (voltage deviation rate, impedance voltage drop term, reactive load term) in the solution set. The smaller the information entropy (the greater the data variation), the greater the weight, reflecting the greater influence of the component on decision-making, for example ).
[0054] The formula of the operation mode weight is as follows:
[0055] The significance of voltage stability is as follows: the first term: voltage deviation rate, which measures the deviation of voltage from the standard value. The second term: reflects the voltage drop caused by line impedance (R ), based on the approximate relationship . The third term: represents the influence of node reactive load on voltage. Constant power reactive load will draw more current when the voltage drops , thereby further exacerbating voltage drop and affecting stability. The is used to reflect the sensitivity of the load to voltage and avoid too small denominator. The operation mode weight further considers the sensitivity of voltage stability in different operation modes (island / black start mode is more vulnerable).
[0056] Further, the constraint conditions in the embodiments of the present application are as follows: Power balance: (define positive for charging, negative for discharging)
[0057] wherein, is the index of the distributed generator, is the power generation of the kth gas turbine (or other distributed generator) at time t, is the power purchased from the main grid, is the photovoltaic power generation, is the discharging power of the energy storage battery, is the charging power of the energy storage battery, is the charging and discharging power of the mth EV charging cluster, is the total power of the local load, is the power sold to the main grid.
[0058] Significance: the total power of all power sources (distributed generators, grid purchase, photovoltaic, battery discharge, EV discharge) in the microgrid is equal to the total power of all loads (local load, grid sale, battery charge, EV charge), ensuring real-time power balance of the system. The EV power is represented by the sum of the clusters , simplifying the scheduling.
[0059] Battery capacity:
[0060] where, is the battery self-discharge rate, is the stored energy, is the charging efficiency, is the discharging efficiency, is the time length of the scheduling period.
[0061] Meaning: Track battery power changes, considering self-discharge loss rate , charging and discharging status and efficiency .
[0062] EV charging cluster power:
[0063]
[0064] where, is a function representing the inverse relationship between EV power and electricity price, for example , or by optimizing its tendency, is the lower power limit, is the upper power limit.
[0065] Meaning: The power of each EV charging cluster is limited, and tends to charge more when the electricity price is low and to discharge when the electricity price is high. The cluster power is the sum of the power of individual vehicles within the cluster, which is uniformly regulated by the electricity price .
[0066] EV cluster net energy constancy (weekly): (Note: Net energy refers to the total charging amount minus the total discharging amount of the cluster from the grid within the scheduling period, reflecting the actual driving energy demand of the EVs in the cluster).
[0067]
[0068] where, is the net energy demand (kWh) of the th cluster within the scheduling period (one week, ).
[0069] Meaning: Considering the energy consumption of electric vehicles during driving, the total charging amount minus the total discharging amount (net energy) of each EV charging cluster from the grid within the scheduling period (one week) can be considered as a constant value . This reflects the actual driving demand of electric vehicle users within a week, while allowing electricity price regulation to flexibly allocate charging and discharging power in each period.
[0070] Power ramping:
[0071] where, Pi(t) is the power of device i at time t, Pi(t-1) is the power of device i at time t-1, Pi,max is the maximum power ramp-down rate of device i, Pi,max is the maximum power ramp-up rate of device i.
[0072] Meaning: the power change of a device cannot be too fast, to protect the device.
[0073] Reactive margin:
[0074] where, Qi(t) is the reactive power that the inverter at node i is required to provide or absorb at time t. Qi,max is the apparent power capacity (rated power) of the inverter connected to node i, Qi,other is the reactive power that has been consumed or generated by other devices at node i, except the target reactive power Qi,target. Qi,max is the apparent power capacity (rated power) of the inverter connected to node i,
[0075] Meaning: to ensure that the inverter connected to node i has enough reactive margin to prevent voltage instability. Equivalent short circuit ratio:
[0076]
[0077] where, SCR is the equivalent circuit ratio, SCi is the short circuit capacity at node i, Pk is the active power of the kth load, V(t) is the voltage at time t, Zth is the Thevenin equivalent impedance of the system seen from node i.
[0078] Meaning: to ensure that the microgrid has enough short circuit capacity to reduce the risk of voltage instability.
[0079] Voltage lower limit:
[0080] Meaning: to ensure that the voltage at node i is not lower than (e.g. 0.9 p.u.), to guarantee the voltage quality. Optionally, the generating process of the optimal solution set comprises:
[0081] The state vector is constructed by calculating the change rate of the objective function, and the value evaluation table is updated based on the reward function and the time difference learning to determine the SARSA update strategy of the inertia weight; The inertia weight is determined based on the SARSA update strategy, the particle speed and position are updated based on the optimal position of the particle itself and the optimal position of the group, and the optimal solution set is obtained by screening through the adaptive grid method.
[0082] Specifically, the reinforcement learning SARSA calculates the change rate of the economic, environmental, and voltage stability three objective functions in the iteration process in real time, discretizes it into rising, stable or falling states, and thus constructs a state vector that can represent the current multi-objective optimization process; then, based on a reward function that comprehensively considers the dominance relationship of new solutions and the diversity improvement of solution set, and using the time difference learning algorithm, the internal value evaluation table is dynamically updated, and finally the adjustment strategy (increase, decrease or keep) for the inertia weight parameter is output.
[0083] Then, the optimization algorithm module (MOPSO) executes the above strategy, linearly combines the inertia weight dynamically determined by SARSA with the historical optimal position of each particle itself and the global optimal position of the entire population, and introduces bounded random disturbance to jointly update the flying speed and direction of all particles. After the particles move according to the new speed, the scheduling scheme represented by the new position will be processed to ensure feasibility. Finally, the adaptive grid method is used to manage all newly generated non-dominated solutions, by dividing the target space into grids and removing redundant solutions in dense areas, a Pareto optimal solution set with good convergence and wide distribution is finally screened and output.
[0084] Reference Figure 3 The SARSA module receives the microgrid state (load , voltage , EV charging cluster data ), updates the Q-table through and , and guides MOPSO to generate Pareto solution set.
[0085] Initialization: particle swarm, Q-table, Pareto solution set; S1: Establish microgrid / EV model, define F1 (economic), F2 (emission), F3 (voltage stability); S2: Set constraints: power balance, equipment safety, EV cluster; SARSA dynamic optimization module: state s=(ΔF1, ΔF2, ΔF3); action a (adjust inertia weight w); Improved MOPSO optimization, speed / position update (with w), Pareto solution set maintenance; Judgment: No → Continue iteration; Calculate the reward: r = new solution dominance rate + diversity;
[0086] Determine: Has the maximum number of iterations been reached? Output: Pareto optimal scheduling scheme.
[0087] Furthermore, the SARSA algorithm implementation process includes: Calculate the rate of change of the optimization objective function under the target number of iterations, and construct the state vector for reinforcement learning based on the rate of change; Define an action space, which includes increasing the inertia weight, decreasing the inertia weight, and keeping the inertia weight unchanged. Construct a reward function, and calculate reward data based on the update status of the statistical solution set of the reward function; Determine the current value estimate under the current state vector and the new value estimate under the new state. Based on the current value estimate and the new value estimate, combined with the learning rate, discount factor and reward data, update the reinforcement learning value evaluation table using the temporal difference learning algorithm to determine the SARSA update strategy.
[0088] This application details the operational mechanism of the SARSA algorithm. First, it employs state awareness, periodically (e.g., every 5 iterations) sampling and calculating the average rate of change of each objective function value. Based on a preset threshold, this rate of change is quantified into three basic states: decreasing, stable, or increasing. These three objective states are then combined into a comprehensive state vector, providing a basis for decision-making. Subsequently, the algorithm selects a strategy within a predefined action space containing three basic actions: increasing inertia weight, decreasing inertia weight, and keeping inertia weight constant.
[0089] After executing a selected action, the algorithm immediately calculates a reward signal. This signal is composed of the proportion of the new solution dominating the old solution minus the proportion of the old solution dominating the new solution, plus the improvement in solution set diversity, and is used to evaluate the quality of the action. The most crucial step is to employ a temporal difference learning algorithm, which incrementally updates the value evaluation table by combining the value estimate of the current state-action pair, the immediate reward obtained, and the value estimate of the new state-action pair after transitioning to the new state, with a fixed learning rate and discount factor. Through continuous iteration of this process, the algorithm eventually learns which action to take in which optimization state to obtain the best long-term reward.
[0090] Specifically, the SARSA algorithm implementation process includes the following steps: State space construction: Calculate the rate of change of the objective function (sampled every 5 iterations):
[0091] in, Refers to the number of iterations. The economic cost of the current iteration ( ),emission( Voltage stability The target value serves to quantify the optimization progress and determine whether the objective function is improving, stagnating, or deteriorating.
[0092] Discretized state:
[0093] Final state vector: There are 27 possible combinations; function: to construct the state space of reinforcement learning and provide input for decision-making.
[0094] Action space definition: : Increase inertia weight (Strengthen exploration); Reduce inertial weight (Enhance development); :Keep Unchanged (equilibrium state).
[0095] Formula meaning: Defines three actions that the algorithm can perform.
[0096] Reward function calculation: Statistics on Pareto solution set updates:
[0097] Formula meaning: Evaluates the quality of the search action. Purpose: Guides the algorithm to balance convergence and diversity.
[0098] in, The number of solutions in the old Pareto solution set dominated by the newly generated non-dominated solution. The number of newly generated non-dominated solutions that are dominated by solutions in the old Pareto solution set. : The total number of Pareto solutions (or the total number of particles, depending on the definition). Improvement in solution set diversity (based on Euclidean distance standard deviation); : Diversity weighting coefficient.
[0099] Q-table update (temporal difference learning): Updated formula:
[0100] where, : learning rate, : discount factor (measure the importance of future rewards); : new state, : based on greedy policy selection ( ); : state value estimation of performing action ; action: update Q value based on time difference learning.
[0101] Optionally, the particle optimization algorithm implementation process comprises: based on the inertia weight determined by the SARS A, the historical optimal position of the particle and the group optimal position are combined and linearly superimposed based on the inertia weight, the speed vector of the particle is updated, and the flight speed of the new generation of particles is obtained; based on the flight speed of the new generation of particles and the position vector of the particle, the particle position is updated and constrained, and the new position of the particle is obtained as a new scheduling scheme; the adaptive grid method is used for the new scheduling scheme, the multi-objective space is uniformly divided into grids, and the solution corresponding to the new scheduling scheme is screened to obtain the optimal solution set.
[0102] In this embodiment, the search speed of the particle is first updated, which is linearly superimposed by a plurality of key parts: the product of the dynamic inertia weight decided by the SARS A algorithm in real time and the speed of the last generation of particles, which determines the inheritance and trend of the search, and the pursuit of the particle to its own historical optimal experience and the pursuit of the current global optimal experience of the group, and the randomness is injected through the update of the search direction to avoid premature convergence.
[0103] Each particle moves to a new position according to the current position and new speed, and the position encodes a new potential scheduling scheme. Subsequently, through constraint processing techniques such as projection method, it is ensured that all decision variables (such as generator output, EV charging and discharging power) in the new scheme meet the upper and lower limit constraints of the physical equipment operation. Finally, all newly generated feasible schemes and their historical schemes are sent into the adaptive grid filter together, by dividing the grid in the multi-objective space and limiting the number of solutions in each grid, those solutions that are not dominated and evenly distributed are retained, so as to maintain and update the high-quality Pareto optimal solution set.
[0104] Specifically, the embodiment is an MOPSO optimization process, which is as follows: particle speed update: introduce SARSA guided inertia weight :
[0105] Effect: Introduce inertia weight to update particle search direction, introduce bounded randomness to avoid premature convergence.
[0106] The symbols in the above formulas are defined as follows: : Position vector of particle ; : Inertia weight (adjusted dynamically by SARS A); : Acceleration coefficient; : Random number; : Historical best position of particle ; : Swarm best position; : Random disturbance angle, provides bounded randomness (independent for each particle).
[0107] Position update and constraint handling: Position update:
[0108] Effect: Generate new scheduling scheme.
[0109] Boundary repair (projection method):
[0110] Effect: Ensure that the scheme meets the physical limits of the device.
[0111] The symbols in the above formulas are defined as follows: : Position vector of particle (representing the scheduling scheme, such as decision variables , etc.).
[0112] : Lower and upper limits of decision variables (such as ).
[0113] Pareto solution set maintenance: Use adaptive grid method (Adaptive Grid): Divide the target space into grid (economic × emission × voltage three-dimensional).
[0114] Keep non-dominated solutions, and remove redundant solutions in dense areas (density threshold ).
[0115] Formula meaning: solution set maintenance rule; Symbol definition: : Number of target space grids (economic × emission × voltage three-dimensional); : single grid maximum solution number; : function: reserve high-quality non-dominated solutions and maintain solution set diversity.
[0116] Optionally, based on the optimal solution set, a two-stage distribution robust optimization model considering new energy output error and EV cluster net energy error is constructed, and the two-stage distribution robust optimization model is solved to obtain an optimal scheduling scheme, including: constructing an uncertainty set based on the new energy output error and the EV cluster net energy error, establishing a two-stage distribution robust optimization model based on the microgrid EV model and the uncertainty set; transforming the two-stage distribution robust optimization model into a mixed integer linear programming problem; calling a mathematical optimization solver to solve the mixed integer linear programming problem, and outputting an optimal robust optimal scheduling scheme.
[0117] It should be noted that in the embodiments of the present application, a target scheduling scheme is selected from the Pareto optimal solution set according to a specific preference as the basis for robust optimization. At the same time, error samples of new energy output prediction and EV cluster net energy demand in historical data are collected, and a fuzzy set describing the range of uncertainty probability distribution, i.e. an uncertainty set, is constructed around these empirical distributions using Wasserstein distance measurement.
[0118] Further, a two-stage distribution robust optimization model is established, the first stage of which formulates a day-ahead scheduling plan based on the selected target scheme, and the second stage considers the worst probability distribution in the uncertainty set to reserve adjustment space.
[0119] Subsequently, the process enters the model solving and scheme output stage, and through the application of duality theory and linearization technology, the above-mentioned complex distribution robust optimization model is transformed into a mixed integer linear programming problem that can be efficiently solved. The transformed model completely retains the two-stage decision structure and robustness requirements of the original model. Finally, the mixed integer linear programming model is input into a professional mathematical optimization solver for calculation, and the result output by the solver is the optimal robust scheduling scheme sought. The embodiments in this embodiment have the strongest adaptability to prediction errors while following the basic plan.
[0120] Specifically, after generating a diversified Pareto solution set, the embodiments further determine the final optimal scheduling scheme therefrom and consider the uncertainty in system operation: 1. new energy output error (actual photovoltaic / wind power generation ≠ predicted value); 2. Net energy error of EV cluster (actual charging and discharging demand ≠ predicted value). The final scheme is required to not only have good performance, but also be robust to prediction error. The scheduling problem is finally converted into a mixed integer linear programming (MILP) problem for solution in the application: A two-stage distributionally robust optimization scheduling model is established: Based on the defined collaborative optimization scheduling model, i.e., the microgrid EV model, and the uncertain set based on the Wasserstein distance, a two-stage distributionally robust optimization (Two-Stage Distributionally Robust Optimization, DRO) scheduling model is proposed.
[0121] The decision is divided into two stages: the first stage (pre-operation / prediction stage) determines the scheduling plan; the second stage (operation / adjustment stage) adjusts according to the actual uncertainty. The core expression is as follows:
[0122]
[0123]
[0124] Wherein, : First stage decision variable (scheduling plan, such as unit start-stop, planned power generation, EV cluster planned charging and discharging power, and power selling plan), determined before the uncertainty (reconciliation error) is revealed, is its feasible region; : First stage objective function (collaborative optimization model objective); : Random vector representing uncertainty (new energy output error, EV net energy error); : Uncertain distribution set (fuzzy set) constructed based on the Wasserstein distance; : Find the worst probability distribution in the fuzzy set that maximizes the expected adjustment cost : Second stage decision variable (adjustment amount, such as power generation adjustment, power purchase and sale adjustment, and EV charging and discharging adjustment), determined after is observed, is its feasible region, which depends on and ; : Second stage objective function (adjustment cost); : First stage constraints (device power limits, ramping, EV cluster power constraints, net energy constraints, lower voltage limits, etc.) : Second stage constraints to ensure that the adjusted solution satisfies all physical and security constraints (power balance, device limits, voltage constraints, etc.) relying on .
[0125] Further, the two-stage distribution robust optimization scheduling model is converted into a mixed integer programming problem. The inner optimization problem is converted into its dual form and the coupling between the uncertain set and the two-stage decision variables are handled by applying duality theory and linearization techniques (e.g. KKT conditions, affine policy). The converted mixed integer linear programming (MILP) problem expression is as follows:
[0126]
[0127] where, denotes the total number of historical or generated samples used to construct the fuzzy set ; denotes the sample index, .
[0128]
[0129]
[0130]
[0131]
[0132]
[0133]
[0134]
[0135] where, : First stage decision variable; : Second stage decision variable corresponding to the th sample ; : Second stage decision variable expression that has linear / affine relationship with the uncertain quantity after affine policy (or other linearization methods) handling; : First and second stage objective function coefficient vectors transpose; : total number of historical or generated samples for constructing Wasserstein fuzzy set; : sample index ( ); : the th uncertainty sample (new energy output error, EV net energy error); : upper and lower bounds of uncertainty ; : auxiliary variable; : dual variable (related to Wasserstein ball radius); : Wasserstein ball radius (controls the size of the fuzzy set, reflects the degree of conservatism to the distribution uncertainty); : coefficient matrix and vector of the first stage constraints; : coefficient matrix / vector of the second stage constraints, dependent on ; It should be noted that this transformation is based on a specific linearization technique (such as linear decision rule LDR or limited adaptability strategy), and the specific form depends on the transformation method.
[0136] Finally, the transformed MILP problem is input into a professional optimization solver (such as Gurobi, CPLEX, MOSEK, etc.) for solving. The solver will find the values of the decision variables that satisfy all the constraints and minimize the objective function, thus determining the optimal robust scheduling scheme of the microgrid under consideration of uncertainty . This scheme maximizes the benefits of both the distribution network operator and the electric vehicle cluster, and ensures the robustness to prediction errors.
[0137] Specifically, referring to Figure 4 , Figure 4 is the effect diagram of the electricity price regulation of the embodiments of the present application. The red in the diagram represents the actual electricity price, the blue represents the actual EV power, and the green represents the actual voltage deviation; the diagram shows the power distribution during peak (high electricity price, high, EV charging cluster discharging) and valley (low electricity price, low, EV charging cluster charging), and the voltage deviation is significantly reduced.
[0138] Referring to Figure 5 , Figure 5is a schematic diagram of the net energy constancy of the EV cluster of the embodiment of the present application, which shows the charging and discharging power change (regulated by the electricity price) of each period in a week, but the total sum of the net energy remains constant.
[0139] With reference to Figure 6 Based on the method in the above embodiment, an electronic device is provided in the embodiment of the present application, which can include a processor 610, a communications interface 620, a memory 630 and a communications bus 640, wherein the processor 610, the communications interface 620 and the memory 630 complete mutual communication through the communications bus 640. The processor 610 can invoke the logical instructions in the memory 630 to execute the method in the above embodiment.
[0140] In addition, the logical instructions in the memory 630 described above can be implemented in the form of a software functional unit and sold or used as an independent product, which can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product or part of the technical solutions of the present application, and the computer software product is stored in a storage medium, including a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the embodiments of the present application.
[0141] Based on the method in the above embodiment, a computer readable storage medium is provided in the embodiment of the present application, and the computer readable storage medium stores a computer program, which causes the processor to execute the method in the above embodiment when the computer program runs on the processor.
[0142] Based on the method in the above embodiment, a computer program product is provided in the embodiment of the present application, which causes the processor to execute the method in the above embodiment when the computer program product runs on the processor.
[0143] It can be understood that the processor in the embodiments of the present application can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. The general-purpose processor can be a microprocessor or any conventional processor.
[0144] The method steps in the embodiments of the present application can be implemented in the form of hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, and the software modules can be stored in a random access memory (RAM), a flash memory, a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically EPROM (EEPROM), a register, a hard disk, a mobile hard disk, a CD-ROM or any other form of storage medium well known in the art. An exemplary storage medium is coupled to the processor, so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and the storage medium can be located in an ASIC.
[0145] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are generated. The computer can be a general purpose computer, a special purpose computer, a computer network, or other programmable apparatus. The computer instructions can be stored in or transmitted by a computer readable storage medium. The computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through a wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) manner. The computer readable storage medium can be any available medium accessible by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media sets. The available media can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0146] It can be understood that various numerical numbers involved in the embodiments of the present application are only distinguished for convenience of description, and are not used to limit the scope of the embodiments of the present application.
[0147] Those skilled in the art easily understand that the above only describes the preferred embodiments of the present application and is not used to limit the present application. Any modification, equivalent replacement and improvement made within the spirit and principle of the present application should be included in the protection scope of the present application.
Claims
1. A micro-grid optimal scheduling method of multi-objective particle swarm combined with SARS A and electricity price regulation, characterized in that, include: A microgrid EV model is established and its optimization objective function is defined. The operating constraints of the microgrid EV model are set. The optimization objective function includes an economic cost objective function, an environmental emission objective function, and a voltage stability objective function. An improved multi-objective particle swarm optimization algorithm is constructed based on the SARSA reinforcement learning algorithm. The optimization objective function is solved based on the multi-objective particle swarm optimization algorithm under the running constraints to generate the optimal solution set. Based on the optimal solution set, a two-stage sub-Blu-rod optimization model considering the power output error of new energy sources and the net energy error of EV clusters is constructed. The optimal scheduling scheme is obtained by solving the two-stage sub-Blu-rod optimization model. 2.The microgrid optimal scheduling method of multi-objective particle swarm combined with SARSA and electricity price regulation according to claim 1, wherein, The process of generating the optimal solution set includes: The SARSA update strategy for inertia weights is determined by constructing a state vector through calculating the rate of change of the objective function and updating the value evaluation table based on the reward function and temporal difference learning. The inertial weights are determined based on the SARSA update strategy. The particle velocity and position are updated by combining the particle's own optimal position and the group's optimal position. The optimal solution set is obtained by filtering using an adaptive grid method. 3.The microgrid optimal scheduling method of multi-objective particle swarm combined with SARSA and electricity price regulation according to claim 2, wherein, The SARSA algorithm implementation process includes: Calculate the rate of change of the optimization objective function under the target number of iterations, and construct the state vector for reinforcement learning based on the rate of change; Define an action space, which includes increasing the inertia weight, decreasing the inertia weight, and keeping the inertia weight unchanged; Construct a reward function, and calculate reward data based on the update status of the statistical solution set of the reward function; Determine the current value estimate under the current state vector and the new value estimate under the new state. Based on the current value estimate and the new value estimate, combined with the learning rate, discount factor and reward data, update the reinforcement learning value evaluation table using the temporal difference learning algorithm to determine the SARSA update strategy.
4. The method of claim 2, wherein the multi-objective particle swarm optimization scheduling of microgrid with SARSA and electricity price regulation is characterized by, The implementation process of the particle optimization algorithm includes: The inertial weight is determined by dynamic adjustment based on SARSA. The inertial weight is then combined with the historical best position of the particle and the group's best position to linearly superimpose and update the particle's velocity vector, thus obtaining the next generation of particle flight speed. Based on the new generation of particle flight velocity and particle position vector, particle position update and constraint processing are performed to obtain the new position of the particle as a new scheduling scheme. The new scheduling scheme employs an adaptive grid method, uniformly dividing the multi-objective space into grids, and then filtering the solutions corresponding to the new scheduling scheme to obtain the optimal solution set.
5. The method of claim 1, wherein the multi-objective particle swarm optimization scheduling of microgrid with SARSA and electricity price regulation is characterized by, Based on the optimal solution set, a two-stage sub-Brussels bar optimization model considering both new energy output error and EV cluster net energy error is constructed. Solving the two-stage sub-Brussels bar optimization model yields the optimal scheduling scheme, including: Select a target scheduling scheme from the optimal solution set; An error set is constructed based on the new energy output error and the net energy error of the EV cluster. The error set is then given uncertainty and transformed using the Wasserstein distance method to obtain an uncertainty set. A two-stage sub-Bruker optimization model is established based on the microgrid EV model and the uncertainty set. inputting the target scheduling scheme into the two-stage distribution robust optimization model, determining a scheduling plan through a first stage, and adjusting the scheduling plan according to actual uncertainty through a second stage; transforming the two-stage distribution robust optimization model into a mixed integer linear programming problem, calling a mathematical optimization solver to solve the mixed integer linear programming problem, and outputting an optimal robust optimal scheduling scheme.
6. The method of claim 1, wherein the multi-objective particle swarm optimization scheduling of microgrid with SARSA and electricity price regulation is characterized by, The operation constraints include power balance constraints, energy storage system capacity constraints, EV charging cluster power upper and lower limit constraints, EV charging cluster net energy constancy constraints in a scheduling period, power ramping constraints, reactive power margin constraints, equivalent short-circuit ratio constraints, and voltage lower limit constraints.
7. The method of claim 1, wherein the multi-objective particle swarm optimization scheduling of microgrid with SARSA and electricity price regulation is characterized by, The economic cost objective function is as shown in the following formula: in, The total economic cost, The set of microgrid nodes The total number of all devices Unit power generation cost of gas turbines Gas turbine power, Time step node The grid purchase price of electricity, Power purchased by the power grid equipment The unit power operation and maintenance cost, equipment In time power, node Electricity sales price, Electricity sold to the grid No. Dynamic electricity price for individual EV charging clusters No. The charging / discharging power of each EV charging cluster.
8. The method of claim 1, wherein the multi-objective particle swarm optimization scheduling of microgrid with SARSA and electricity price regulation is characterized by, The environmental emission objective function is as shown in the following formula: wherein, is the total emissions, is the set of numbers of all devices, is the pollutant type, is the device produces the pollutant with the emission factor, is the power production of the device at the time step, is the time step. 9.The microgrid optimal scheduling method of multi-objective particle swarm combined with SARSA and electricity price regulation of claim 1, wherein, The voltage stability objective function is as shown in the following formula: wherein, is the voltage stability target, is the node at time is the voltage, is the reference voltage, are the active and reactive power flowing through the line connected to the node , respectively, are the resistance and reactance of the line connected to the node , respectively, is the reactive load of the node , is the weight factor determined by the entropy weight method, is the operating mode weight.
10. A computer-readable storage medium storing a computer program, the computer-readable storage medium comprising: When the computer program runs on the processor, the processor is caused to perform the method according to any one of claims 1-9.