Micro-grid electric vehicle charging and discharging control method
By constructing a multi-dimensional state perception system and an improved Markov decision model, combined with the DDPG algorithm and the Actor-Critic network, the problem of insufficient power for electric vehicles in emergency scenarios is solved, and the coordinated scheduling of electric vehicles and multi-energy systems is realized to ensure power supply to important loads and system stability.
Patent Information
- Application Number
- CN202511862795.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-11
- Publication Date
- 2026-02-24
AI Technical Summary
Existing vehicle-to-grid (V2G) technologies suffer from insufficient battery power in electric vehicles during emergencies such as extreme weather or equipment failures, failing to guarantee power supply to critical loads. Furthermore, they lack the ability to coordinate and schedule multiple energy systems, making it difficult to cope with fluctuations in electricity prices and the randomness of user behavior, leading to deviations in scheduling strategies.
A multi-dimensional state perception system is constructed, and an improved Markov decision model and DDPG algorithm are adopted. Combined with the Actor-Critic dual network structure and experience playback mechanism, the optimal charging and discharging control strategy is generated. The model parameters are corrected through edge computing and real-time feedback to realize the coordinated scheduling of electric vehicles and multi-energy systems.
It achieves synergistic optimization of economic efficiency and resilience, ensures continuous power supply to critical loads, improves the capacity for renewable energy absorption, reduces strategy deviation, and enhances system stability and user participation.
Smart Images

Figure CN121566569A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of microgrid energy dispatching technology, and more specifically, to a microgrid electric vehicle charging and discharging control method. Background Technology
[0002] With the rapid growth of electric vehicle ownership, vehicle-to-grid (V2G) technology has become a key link connecting transportation and energy systems, transforming electric vehicles from simple transportation tools into flexibly deployable "mobile energy storage units." As a comprehensive energy system integrating distributed energy resources, energy storage devices, and diverse loads, microgrids can switch to independent operation modes in the event of external grid failures, making them a core carrier for enhancing the resilience of regional power supply.
[0003] Existing research on vehicle-to-grid (V2G) interaction technologies often focuses on single objectives such as reducing operating costs and improving economic efficiency, resulting in the following shortcomings: First, it neglects the resilience requirements of the power distribution system under sudden scenarios such as extreme weather and equipment failures, often leading to insufficient power for electric vehicles in emergencies due to excessive daily discharge, thus failing to guarantee power supply for important loads such as hospitals and communication base stations; Second, it lacks a coordinated mechanism for the scheduling of electric vehicle charging and discharging in multi-energy systems such as combined heat and power units and gas boiler energy conversion equipment, failing to fully leverage the advantages of multi-energy complementarity; Third, it lacks sufficient ability to cope with uncertainties such as electricity price fluctuations and the randomness of user charging and discharging behavior, resulting in significant deviations in scheduling strategies in practical applications.
[0004] Therefore, a microgrid electric vehicle charging and discharging control method is proposed to address the above problems. Summary of the Invention
[0005] The purpose of this application is to provide a method for controlling the charging and discharging of electric vehicles in a microgrid.
[0006] The microgrid electric vehicle charging and discharging control method provided in this application adopts the following technical solution:
[0007] A method for controlling the charging and discharging of a microgrid electric vehicle, the method comprising the following steps:
[0008] S1: Construct a multi-dimensional state perception system, collect microgrid energy supply data, load demand data and electric vehicle operation data, and complete data standardization after preprocessing;
[0009] S2: Establish an improved Markov decision-making model, with the reward function employing a composite mechanism of economic rewards, resilience rewards, and penalty terms;
[0010] S3: The improved DDPG algorithm is used to solve the Markov decision model. The optimal charging and discharging control strategy is generated through the Actor-Critic dual network structure, experience replay mechanism and adaptive learning rate adjustment.
[0011] S4: Transform the optimal control strategy into execution instructions and send them to the terminal device, and correct the model parameters based on real-time feedback data.
[0012] Furthermore, the multi-dimensional state perception system in S1 deploys distributed sensing nodes to collect real-time operation data of the microgrid, including three types of core information, specifically:
[0013] Energy supply data: real-time status of solar and wind power, operating status and output power of combined heat and power units and gas boiler energy conversion equipment;
[0014] Load demand data: Real-time power consumption of ordinary loads and important loads are collected in categories, as well as historical fluctuation characteristics of loads of cogeneration units and gas boiler energy conversion equipment;
[0015] Electric vehicle data: number of electric vehicles connected to the microgrid, state of charge (SoC), estimated dwell time, user charging and discharging preferences, and battery health status.
[0016] Furthermore, in S1, the collected data is preprocessed through edge computing nodes to remove outliers and complete the time synchronization and format standardization of multi-source data.
[0017] Furthermore, the state space of the improved Markov decision model in S2 includes the microgrid's multi-energy supply and demand balance, renewable energy output coefficient, electric vehicle cluster state of charge, real-time electricity price level, resilience reserve coefficient, and user energy consumption preference coefficient.
[0018] Furthermore, the action space of the improved Markov decision model includes electric vehicle charging and discharging control, cogeneration unit output regulation, heat pump operation mode switching, and energy storage device charging and discharging actions.
[0019] Furthermore, the resilience reserve coefficient in S2 is calculated by combining the power supply guarantee duration for critical loads with the emergency response speed; the charging and discharging actions include charging power level, discharging power level, and standby state.
[0020] Furthermore, in the S3 Actor-Critic dual network structure, the Actor network is responsible for outputting continuous charge and discharge control actions, while the Critic network is responsible for evaluating the value of the actions and guiding network updates.
[0021] Furthermore, in S3, an experience replay mechanism is introduced to store system operation data. By randomly sampling, the correlation between data is broken, thereby improving the convergence stability of the algorithm.
[0022] Furthermore, in S3, an adaptive learning rate is set. When the system is stable, a smaller learning rate is used to ensure the stability of the strategy, and in case of emergencies, the learning rate is increased to achieve a rapid response.
[0023] Furthermore, the emergency dispatch mode in S4 ensures the power supply needs of medical equipment and communication base station loads by increasing the weight of resilience rewards in the reward function.
[0024] The technical effects and advantages of this application are as follows:
[0025] Compared with existing technologies, this microgrid electric vehicle charging and discharging control method achieves synergistic optimization of economy and resilience. By dynamically adjusting the charging and discharging strategy, it effectively optimizes the energy configuration structure in daily operation to control the overall operating cost of the microgrid. At the same time, it relies on the energy storage potential of the electric vehicle cluster to reserve sufficient emergency power resources for sudden scenarios, ensuring the continuous power supply of important loads such as medical equipment and communication base stations under extreme conditions, breaking the limitation of traditional technologies that "emphasize economy but neglect resilience".
[0026] Construct a comprehensive multi-energy coordinated dispatch mechanism for cogeneration units and gas-fired boiler energy conversion equipment, deeply integrate the charging and discharging process of electric vehicles with the multi-energy network, give full play to the complementary characteristics of different energy forms, significantly improve the absorption capacity of renewable energy such as photovoltaic and wind power, effectively mitigate the impact of renewable energy output fluctuations on the stable operation of microgrids, and enhance the overall efficiency of system energy utilization.
[0027] Leveraging the strong learning and adaptive capabilities of the DDPG algorithm, the system accurately captures the patterns of system state changes through an Actor-Critic dual network structure. Combined with an experience replay mechanism, it reduces interference from uncertainties. For complex scenarios such as electricity price fluctuations and random user charging and discharging behaviors, it can output scheduling strategies with stronger stability and higher adaptability, significantly reducing the risk of strategy deviation in practical applications.
[0028] Establish an interactive incentive system centered on user needs, dynamically optimize service strategies through a real-time user satisfaction feedback mechanism, and combine it with a scientific dynamic pricing mechanism and incentive measures such as discharge points redemption to fully mobilize the enthusiasm of electric vehicle users to participate in vehicle-to-grid interaction. Combined with mature models verified in multiple pilot areas, significantly improve the overall vehicle-to-grid interaction participation rate and form a virtuous cycle of interaction between the microgrid and users. Attached Figure Description
[0029] Figure 1 This is a flowchart illustrating a microgrid electric vehicle charging and discharging control method according to this application. Detailed Implementation
[0030] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0031] like Figure 1 The method for controlling the charging and discharging of a microgrid electric vehicle, as shown, includes the following steps:
[0032] S1: Construct a multi-dimensional state perception system, collect microgrid energy supply data, load demand data and electric vehicle operation data, and complete data standardization after preprocessing;
[0033] The multi-dimensional state perception system in S1 deploys distributed sensing nodes to collect real-time operation data of the microgrid, including three core types of information, specifically:
[0034] Energy supply data: real-time status of renewable energy sources such as solar and wind power, and operating status and output power of energy conversion equipment such as combined heat and power units and gas boilers;
[0035] Load demand data: Real-time power consumption of ordinary loads (residential electricity consumption) and important loads (medical equipment, communication base stations, etc.) are collected in categories, as well as the historical fluctuation characteristics of the load of cogeneration units and gas boiler energy conversion equipment;
[0036] Electric vehicle data: number of electric vehicles connected to the microgrid, state of charge (SoC), estimated dwell time, user charging and discharging preferences, and battery health status.
[0037] In S1, edge computing nodes are used to preprocess the collected data, remove outliers, and complete the time synchronization and format standardization of multi-source data.
[0038] S2: Establish an improved Markov decision-making model, with the reward function employing a composite mechanism of economic rewards, resilience rewards, and penalty terms;
[0039] The state space of the improved Markov decision model in S2 includes the microgrid's multi-energy supply and demand balance, renewable energy output coefficient, electric vehicle cluster state of charge, real-time electricity price level, resilience reserve coefficient, and user energy consumption preference coefficient. The microgrid's multi-energy supply and demand balance includes the supply and demand balance of electricity, heat, and gas. The resilience reserve coefficient is calculated by comprehensively considering the power supply guarantee duration for important loads and the emergency response speed.
[0040] The action space of the improved Markov decision model includes electric vehicle charging and discharging control, cogeneration unit output regulation, heat pump operation mode switching, and energy storage device charging and discharging actions. The charging and discharging actions include charging power level, discharging power level, and standby state.
[0041] The improved Markov decision model, also known as the improved MDP model, focuses on "economic efficiency and resilience dual-objective optimization." The specific strategy is as follows: Let the optimization period of the improved MDP model be... The basic objective is to minimize the total operating cost of microgrid operators (MMOs) while maximizing the system's resilience reserve capacity. The mathematical expression is: ,in, For action space set; for The total operating cost of the time-tracking system, which includes energy procurement costs, equipment maintenance costs, and electric vehicle battery depreciation costs; for The system resilience gain at any given moment (dimensionless, normalized to [0,1]) represents the ability to guarantee critical loads; for System status at all times; for Constantly control your actions;
[0042] state space The specific construction strategy is as follows:
[0043] A1: Determine the state dimensions: Select 6 core dimensions to represent the system state, forming a state vector. ;
[0044] A2: Quantifying parameters across various dimensions: The balance between supply and demand of multiple energy sources (electricity, heat, and gas) at any given time (dimensionless). The calculation formula is: ( ),in for time Total power supply of energy sources; for time Total power demand for energy types; (Exceeding the range is considered a supply-demand imbalance);
[0045] for The power output factor of renewable energy (photovoltaic + wind power) at any given time is calculated using the following formula: ,in, , They are respectively Real-time output of solar and wind power; , These are the rated outputs of photovoltaic and wind power, respectively. ;
[0046] for The average state of charge (SoC, dimensionless) of the electric vehicle cluster at any given time is calculated as follows: ,in for The number of electric vehicles connected to the microgrid at all times; For the first The electric vehicle's state of charge at time t ( (To avoid overcharging and over-discharging)
[0047] for Real-time electricity price levels (dimensionless, discrete values), quantified according to local electricity pricing policies. (These correspond to off-peak, flat, peak, and high-peak electricity prices, respectively).
[0048] for The time-resilient reserve coefficient (dimensionless) is calculated as follows: ,in for Emergency power supply guarantee duration for critical loads at all times; The minimum required guarantee duration (usually 4 hours) is specified in the regulations. Emergency response time (h, the smaller the better); ;
[0049] for The user energy consumption preference coefficient at any given time (dimensionless) is obtained based on historical data clustering. (The higher the value, the more price-sensitive the user is.)
[0050] Action space The specific construction strategy is as follows:
[0051] B1: Classification of Action Types: Control actions are divided into two categories: "electric vehicle charging and discharging control" and "multi-energy coordinated regulation";
[0052] B2: Quantify the parameters of each action:
[0053] Specifically, this refers to the charging and discharging operations of electric vehicles. Discrete charge / discharge power ratings The sign indicates discharging, the sign indicates charging, and 0 indicates standby.
[0054] Combined heat and power unit output regulation action Continuously adjust the output ratio. (Dimensionless, corresponding to 50%-100% of rated output);
[0055] Heat pump operation mode switching action Discrete mode selection (0 = Normal heating, 1 = Waste heat recovery, 2 = Shutdown);
[0056] Energy storage device charging and discharging operation Continuous power regulation The negative sign indicates discharging, and the positive sign indicates charging. Rated power of energy storage;
[0057] B3: Combined Action Vector: The final action vector is ;
[0058] Satisfy action constraints: ,in The scheduling step size is typically set to 0 or 25 hours. For the first Vehicle battery rated capacity;
[0059] Energy supply and demand balance constraints: (Real-time balance of electrical energy), among which for The total power supply of electrical energy in the microgrid at any given time; for The total number of electric vehicles that are connected to the microgrid and participate in dispatch at any given time (dimensionless, a positive integer). for The real-time power consumption of a microgrid (unit: kW) includes the power consumption of various loads such as residential, industrial, medical, and communication loads; the same applies to the balance of heat and gas energy.
[0060] reward function A composite reward mechanism is adopted to balance economic efficiency, resilience, and user satisfaction;
[0061] The specific strategy is as follows: ,in, For the weighting coefficients, satisfying It can be adjusted according to the scenario;
[0062] The calculation of each item is as follows: Economic incentive It is positively correlated with the amount of operating cost savings, and the calculation formula is: ,in This represents the lowest operating cost in history. This represents the highest operating cost in history. ;
[0063] Resilience Rewards The calculation is based on the increase in the resilience reserve coefficient, specifically as follows: ,in To perform the action The resilience reserve coefficient at the next moment. ;
[0064] Penalty items This covers user dissatisfaction, battery degradation, and supply-demand imbalance, specifically: ,in, The preset state of charge when the electric vehicle leaves; For the first When the car left The actual state of charge; Battery loss coefficient (dimensionless) The greater the charging and discharging power, The more extreme the situation, the greater the loss. ;
[0065] State transition function The construction strategy is as follows: Based on historical data and real-time parameters, a state transition probability matrix is established to quantify the mapping relationship between "current state - action - next state". The steps are as follows:
[0066] C1. Collect microgrid operation data (including energy output, load demand, user behavior, etc.) for the past 12 months to form a sample set. , The number of samples, and ;
[0067] C2: The transition probability is fitted using the kernel density estimation (KDE) method, with the following formula: ,in, For the state space dimension, This is the bandwidth parameter (determined through cross-validation). For Gaussian kernel function: ;
[0068] C3: Introduction of Uncertainty Correction Term: Adding correction factors to address fluctuations in renewable energy output and random user behavior. ( ,normal distribution), .
[0069] S3: The improved DDPG algorithm is used to solve the Markov decision model. The optimal charging and discharging control strategy is generated through the Actor-Critic dual network structure, experience replay mechanism and adaptive learning rate adjustment.
[0070] The S3 uses an Actor-Critic dual-network structure. The Actor network is responsible for outputting continuous charge and discharge control actions, while the Critic network is responsible for evaluating the value of the actions and guiding network updates.
[0071] The specific initialization steps of the improved DDPG algorithm are as follows:
[0072] D1: Construct 4 neural networks, namely:
[0073] Actor's Current Network Input status Output continuous motion The output layer activation function adopts (Constrain the action within a preset range);
[0074] Actor Target Network :parameter Initialize to Used for stable training;
[0075] Critic Current Network :enter Output action value ;
[0076] Critic target network :parameter Initialize to , used to calculate target value;
[0077] D2: Initialize the experience replay pool The capacity is set to ;
[0078] D3: Setting Hyperparameters: Learning Rate (Actor Network) (Critic network), target network soft update coefficients Discount factor Batch sampling size .
[0079] S3 introduces an experience replay mechanism to store system operation data, and breaks data correlation through random sampling to improve the algorithm's convergence stability;
[0080] The specific steps for implementing the experience replay mechanism are as follows:
[0081] E1: Real-time acquisition of system status Generate actions through the current network of the Actor. ( (Add exploration noise);
[0082] E2: Execute action Get rewards and a comfortable state ;
[0083] E3: Experience Store in the experience replay pool ;
[0084] E4: When At that time, from Random sampling An empirical sample is used for network updates.
[0085] The steps for updating the Actor network are as follows:
[0086] F1: Calculate the policy gradient: the gradient of the Critic network output with respect to the Actor network output, using the following formula: ;
[0087] F2: Update the current network parameters of the Actor using the Adam optimizer: ;
[0088] The specific steps for updating the Critic network are as follows:
[0089] G1: Calculate the target action value (based on the Critic target network): ,in, For the first Rewards for each sample The next state;
[0090] G2: Calculate the Critic network loss (mean squared error loss):
[0091] ;
[0092] G3: Update the current Critic network parameters using the Adam optimizer: ;
[0093] The specific steps for a target network soft update are as follows:
[0094] Update the parameters of the Actor and Critic target networks to avoid training oscillations. The formula is: , .
[0095] In S3, an adaptive learning rate is set. When the system is stable, a smaller learning rate is used to ensure the stability of the strategy, and the learning rate is increased to achieve a rapid response in sudden scenarios.
[0096] The specific strategy for adaptive learning rate adjustment is as follows: A dynamic learning rate is designed for unexpected scenarios, and the steps are as follows:
[0097] Calculate the state fluctuation coefficient (Quantifying system stability);
[0098] when If the system is stable, then maintain the initial learning rate. ;
[0099] when In case of sudden fluctuations, such as extreme weather or equipment failure, the learning rate will be increased to [a higher value]. , This accelerates policy updates and also increases the resilience weight in the reward function. Increase it to 0.5 to prioritize critical loads.
[0100] S4: The optimal control strategy is translated into execution instructions and sent to the terminal device. Based on real-time feedback data, the model parameters are corrected. In case of emergencies, an emergency dispatch mode is triggered, and the resilience reward weight is increased to prioritize power supply to critical loads. The emergency dispatch mode increases the weight of the resilience reward in the reward function. Specifically, the strategy is to use the optimal action output by the DDPG algorithm. ( To train the converged Actor network, control commands are converted into instructions and sent to charging piles and energy equipment via the OCPP1.6J protocol (vehicle-to-grid interaction standard protocol); every 15 minutes ( Re-collection Substitute these values into the model to update the state transition probability and reward value, thereby achieving dynamic correction.
Claims
1. A method for controlling the charging and discharging of a microgrid electric vehicle, characterized in that: The control method includes the following steps: S1: Construct a multi-dimensional state perception system, collect microgrid energy supply data, load demand data and electric vehicle operation data, and complete data standardization after preprocessing; S2: Establish an improved Markov decision-making model, with the reward function employing a composite mechanism of economic rewards, resilience rewards, and penalty terms; S3: The improved DDPG algorithm is used to solve the Markov decision model. The optimal charging and discharging control strategy is generated through the Actor-Critic dual network structure, experience replay mechanism and adaptive learning rate adjustment. S4: Transform the optimal control strategy into execution instructions and send them to the terminal device, and correct the model parameters based on real-time feedback data.
2. The microgrid electric vehicle charging and discharging control method according to claim 1, characterized in that: The multi-dimensional state perception system in S1 deploys distributed sensing nodes to collect real-time operation data of the microgrid, including three types of core information, specifically: Energy supply data: real-time status of solar and wind power, operating status and output power of combined heat and power units and gas boiler energy conversion equipment; Load demand data: Real-time power consumption of ordinary loads and important loads are collected in categories, as well as historical fluctuation characteristics of loads of cogeneration units and gas boiler energy conversion equipment; Electric vehicle data: number of electric vehicles connected to the microgrid, state of charge (SoC), estimated dwell time, user charging and discharging preferences, and battery health status.
3. The microgrid electric vehicle charging and discharging control method according to claim 1, characterized in that: In step S1, edge computing nodes are used to preprocess the collected data, remove outliers, and complete the time synchronization and format standardization of multi-source data.
4. The microgrid electric vehicle charging and discharging control method according to claim 1, characterized in that: The state space of the improved Markov decision model in S2 includes the microgrid's multi-energy supply and demand balance, renewable energy output coefficient, electric vehicle cluster state of charge, real-time electricity price level, resilience reserve coefficient, and user energy consumption preference coefficient.
5. The microgrid electric vehicle charging and discharging control method according to claim 4, characterized in that: The action space of the improved Markov decision model includes electric vehicle charging and discharging control, cogeneration unit output regulation, heat pump operation mode switching, and energy storage device charging and discharging actions.
6. The microgrid electric vehicle charging and discharging control method according to claim 5, characterized in that: The resilience reserve coefficient in S2 is calculated by combining the power supply guarantee duration for critical loads with the emergency response speed; the charging and discharging actions include charging power level, discharging power level, and standby state.
7. The microgrid electric vehicle charging and discharging control method according to claim 1, characterized in that: The Actor-Critic dual-network structure in S3 includes an Actor network and a Critic network. The Actor network is responsible for outputting continuous charge and discharge control actions, while the Critic network is responsible for evaluating the value of the actions and guiding network updates.
8. The microgrid electric vehicle charging and discharging control method according to claim 1, characterized in that: The S3 introduces an experience replay mechanism to store system operation data, and breaks data correlation through random sampling to improve the convergence stability of the improved DDPG algorithm.
9. The microgrid electric vehicle charging and discharging control method according to claim 1, characterized in that: The adaptive learning rate is set in S3 to ensure strategy stability when the system is in a stable state, and to increase the learning rate to achieve rapid response in sudden scenarios.
10. The microgrid electric vehicle charging and discharging control method according to claim 1, characterized in that: The emergency dispatch mode in S4 ensures the power supply needs of medical equipment and communication base station loads by increasing the weight of resilience rewards in the reward function.