A charging and discharging strategy scheduling method and system for an electric vehicle
By employing personalized federated reinforcement learning and dynamic bidding mechanisms, the system addresses issues such as privacy leaks, unquantified user intentions, and static rules in the vehicle-to-grid (V2G) interactive charging scheduling system. This enables efficient and fair electric vehicle charging and discharging strategy scheduling, improving system adaptability and user satisfaction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- UNIV OF SCI & TECH BEIJING
- Filing Date
- 2026-02-06
- Publication Date
- 2026-05-01
AI Technical Summary
The existing vehicle-grid interactive charging dispatch system has problems such as high risk of user privacy leakage, lack of quantification of user participation willingness, difficulty in adapting static rules to dynamic changes in electricity prices and user behavior, and insufficient utilization of heterogeneous user characteristics, resulting in large deviations between dispatch results and execution and low acceptability.
A personalized federated reinforcement learning approach is adopted to construct a global collaborative scheduling model by aggregating the parameters of the first two layers of the Critic network. Combined with a user willingness model and a dynamic bidding mechanism, cross-vehicle collaborative optimization is achieved, which can dynamically respond to changes in grid load and user behavior.
While protecting user privacy, the feasibility of the scheduling scheme and user participation rate have been improved, the system's environmental adaptability and fairness have been enhanced, and communication latency and computational load have been reduced.
Smart Images

Figure CN121663604B_ABST
Abstract
Description
A method and system for scheduling charging and discharging strategies for electric vehicles Technical Field
[0001] This application relates to the fields of vehicle-to-grid interaction and smart grid scheduling technology, and in particular to a method and system for scheduling the charging and discharging strategies of electric vehicles. Background Technology
[0002] With increasing global focus on environmental protection and sustainable development, electric vehicles, as a clean energy mode of transportation, are experiencing rapid growth in their ownership. The widespread adoption of electric vehicles has not only changed the energy consumption structure in the transportation sector but also brought new opportunities and challenges to the operation and management of the power system. Vehicle-to-grid (V2G) technology, acting as a bridge connecting electric vehicles and the power grid, enables bidirectional energy flow and information exchange between the two systems. This is of great significance for improving grid stability, promoting the integration of renewable energy, and enhancing the economic benefits for electric vehicle users.
[0003] In vehicle-to-grid (V2G) charging dispatching systems, distributed dispatching has become a commonly used solution. This approach typically employs a three-tier architecture of "grid dispatch center—aggregator—client" to alleviate data transmission pressure and privacy risks associated with centralized systems. However, under this architecture, the aggregator or dispatch center still needs to collect detailed raw data on user vehicles for optimization calculations, essentially failing to fundamentally avoid the risk of user behavior data leakage. Furthermore, the following shortcomings remain: related technologies often assume complete user compliance with dispatching, failing to quantify their willingness to participate; optimization rules are static, making it difficult to adapt to real-time fluctuations in electricity prices, renewable energy output, and user behavior; raw data still needs to be collected centrally within the region, posing a privacy risk; and there is a lack of mechanisms for collaboratively utilizing heterogeneous user characteristics, resulting in large deviations between dispatching results and execution, and low acceptability.
[0004] Therefore, there is an urgent need for a charging and discharging strategy scheduling method for electric vehicles that can dynamically assess and respond to users' willingness to participate while fully protecting user privacy, and adapt to complex and ever-changing environmental conditions, thereby improving the feasibility of the scheduling scheme, user participation rate, and privacy security level. Summary of the Invention
[0005] The purpose of this application is to provide a method and system for scheduling charging and discharging strategies for electric vehicles. This method can achieve cross-vehicle collaborative scheduling optimization by aggregating model parameters rather than raw data, while fully protecting user privacy. By introducing a user willingness model and a dynamic bidding mechanism, it can effectively improve the acceptability and execution rate of scheduling schemes. Furthermore, by utilizing the adaptive capabilities of federated reinforcement learning, it can achieve more flexible, fair, and efficient vehicle-grid interaction scheduling in environments with fluctuating electricity prices, changes in grid load, and uncertain user behavior.
[0006] To achieve the above objectives, this application provides the following solution:
[0007] Firstly, this application provides a method for scheduling charging and discharging strategies for electric vehicles, comprising: acquiring scheduling-related data for the current microgrid area within the current time period; the scheduling-related data includes electricity price, microgrid load demand, and charging and discharging behavior data of each electric vehicle within the current microgrid area; the charging and discharging behavior data includes arrival time, departure time, state of charge, driving demand, battery capacity, and user participation willingness; aggregating the parameters of the first two layers of the Critic network in the pre-trained initial scheduling model corresponding to all electric vehicles within the current microgrid area to obtain the global parameters of the first two layers of the Critic network; the initial scheduling model is based on scheduling-related data within the historical time period of the current microgrid area, aiming to maximize the scheduling objective function, and adjusting the Actor network. The deep reinforcement learning model of the Critic network framework is obtained through iterative training using a near-end policy optimization algorithm. Based on the global parameters, the parameters of the first two layers of the Critic network are updated to obtain the collaborative scheduling model for each electric vehicle in the current microgrid area, corresponding to the initial scheduling model. Based on the scheduling-related data within the current time period, the charging and discharging strategies for each electric vehicle are obtained through the collaborative scheduling model corresponding to each electric vehicle. The charging and discharging strategies include charging power and discharging power. Based on the user participation willingness of each electric vehicle within the current time period, the charging and discharging strategies are optimized using a willingness-based bidding method, resulting in the optimized charging and discharging strategies for each electric vehicle in the current time period within the current microgrid area.
[0008] Secondly, this application provides a charging and discharging strategy scheduling system for electric vehicles, comprising: a scheduling-related data acquisition module, used to acquire scheduling-related data for the current microgrid area within the current time period; the scheduling-related data includes electricity price, microgrid load demand, and charging and discharging behavior data of each electric vehicle within the current microgrid area; the charging and discharging behavior data includes arrival time, departure time, state of charge, driving demand, battery capacity, and user participation willingness; and a global parameter aggregation module, used to aggregate the parameters of the first two layers of the Critic network in the trained initial scheduling model corresponding to all electric vehicles within the current microgrid area to obtain the global parameters of the first two layers of the Critic network; the initial scheduling model is based on scheduling-related data within the historical time period of the current microgrid area, aiming to maximize the scheduling objective function, and applies the Actor network-Critic network to the target microgrid area. The deep reinforcement learning model of the itic network framework is obtained through iterative training using a near-end policy optimization algorithm. The collaborative model generation module updates the parameters of the first two layers of the Critic network of the initial scheduling model for all electric vehicles in the current microgrid area based on the global parameters, thus obtaining the collaborative scheduling model for each electric vehicle in the current microgrid area. The policy generation module, based on scheduling-related data within the current time period, obtains the charging and discharging policy for each electric vehicle within the current time period using the collaborative scheduling model for each electric vehicle. The charging and discharging policy includes charging power and discharging power. The policy optimization module optimizes the charging and discharging policy based on the user participation willingness of each electric vehicle within the current time period using a willingness-based bidding method, thus obtaining the optimized charging and discharging policy for each electric vehicle in the current time period within the current microgrid area.
[0009] According to the specific embodiments provided in this application, this application has the following technical effects:
[0010] This application provides a method and system for scheduling charging and discharging strategies for electric vehicles. By acquiring scheduling-related data for the current microgrid area within the current time period, it solves the problems of existing scheduling systems, such as the lack of quantitative modeling of users' dynamic participation intentions and the high risk of privacy leakage due to reliance on aggregators to centrally collect raw data. It achieves localized collection and privacy-secure input of multi-dimensional, fine-grained scheduling data that includes users' subjective intentions. By aggregating the parameters of the first two layers of the Critic network in the initial scheduling model corresponding to all electric vehicles in the current microgrid area, the global parameters of the first two layers of the Critic network are obtained. This solves the problem in traditional distributed optimization that still requires the exchange of sensitive user data and cannot achieve cross-vehicle collaborative learning while protecting privacy. It enables the extraction of regional common scheduling features and the construction of a global collaborative representation simply through model parameter interaction. By updating the parameters of the first two layers of the Critic network of the initial scheduling model corresponding to all electric vehicles in the current microgrid area based on the global parameters, the collaborative scheduling model corresponding to each electric vehicle in the current microgrid area is obtained. This model addresses the challenge of balancing personalization and generalization in heterogeneous vehicle scheduling models. It allows each vehicle to share common feature extraction capabilities while retaining its local personalized decision-making layer, enhancing the model's adaptability to diverse user behaviors. By leveraging scheduling-related data within the current time period and employing the collaborative scheduling model corresponding to each electric vehicle, it obtains the charging and discharging strategy for each electric vehicle within that time period. This solves the problem of traditional static rules or centralized optimization failing to respond in real-time to dynamic electricity prices, load, and user status changes, achieving low-latency, adaptive charging and discharging decision generation based on local models and real-time data. Furthermore, by optimizing the charging and discharging strategy based on user participation willingness for each electric vehicle within the current time period using a willingness-based bidding method, it obtains the optimized charging and discharging strategy for each electric vehicle in the current microgrid area within the current time period. This addresses the issues of mismatch between scheduling instructions and actual user acceptance, insufficient fairness, and large execution deviations, achieving a dynamic trade-off between grid control objectives and user willingness. The bidding mechanism ensures priority response for users with high willingness, improving the executability of the scheduling scheme and the fairness of user participation. This application constructs a complete scheduling closed loop from intention perception, privacy protection, collaborative learning to dynamic bidding, which significantly improves the executability, environmental adaptability and system fairness of vehicle-to-grid interactive scheduling while ensuring user data security. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 is a flowchart illustrating a charging and discharging strategy scheduling method for an electric vehicle according to an embodiment of this application.
[0013] Figure 2 is a schematic diagram of the distributed scheduling mode execution flow in an existing vehicle-to-grid interactive charging scheduling system provided in an embodiment of this application.
[0014] Figure 3 is a schematic diagram of the deployment phase of a charging and discharging strategy scheduling system for electric vehicles provided in an embodiment of this application.
[0015] Figure 4 is a schematic diagram of the functional modules of a charging and discharging strategy scheduling system for an electric vehicle provided in an embodiment of this application.
[0016] Figure 5 is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0017] First, some technical terms involved in the embodiments of this application will be introduced.
[0018] As shown in Figure 2, distributed scheduling has become a commonly used technical solution in existing vehicle-grid interactive charging scheduling systems. This solution typically adopts a three-tier architecture of "grid dispatch center—aggregator—client". Specifically, the client (i.e., the electric vehicle terminal or its accompanying user APP) collects and stores real-time vehicle status information locally, including battery state of charge (SOC), charging demand, and estimated departure time. For privacy protection and communication efficiency considerations, this raw data is not directly uploaded to the grid dispatch center, but is first transmitted to the corresponding regional aggregator.
[0019] As a regional control node, the aggregator collects local status data from all vehicles within its region and performs local optimization scheduling based on this data, such as rationally allocating charging power and charging / discharging time windows. After completing local optimization, the aggregator reports summary information such as adjustable capacity and load response capability to the power grid dispatch center. The power grid dispatch center, based on the status summaries from different aggregators and combined with real-time grid load, electricity prices, and renewable energy output information, performs global optimization calculations and generates an overall dispatch strategy or regional adjustment instructions.
[0020] During the execution phase, the instructions from the dispatch center are refined by each aggregator into charging and discharging schedules for individual vehicles and then sent to the client. The vehicles then actually perform the charging or discharging operations, thereby achieving peak shaving and valley filling or providing ancillary services. This distributed solution alleviates, to some extent, the data transmission pressure and privacy leakage risks of the centralized model, and can improve the real-time performance of local responses.
[0021] However, this approach still has the following shortcomings: First, most existing optimization models assume that users will completely comply with scheduling, without modeling or dynamically evaluating users' willingness to participate, leading to a discrepancy between scheduling results and actual execution. Second, traditional distributed optimization is usually based on preset rules or static optimization methods, making it difficult to adaptively adjust strategies in environments with fluctuating electricity prices, random output from renewable energy sources, and uncertain user behavior. Third, although the distributed model reduces the centralized transmission of raw data, user information still needs to be collected within the region, and in the process of joint optimization across regions and multiple aggregators, some sensitive feature information may still need to be exchanged, posing privacy risks. Fourth, different users have different travel patterns and load characteristics, and existing methods lack effective mechanisms for utilizing heterogeneous data in multi-participant collaborative optimization.
[0022] In summary, this application addresses the issue of willingness to discharge by introducing a willingness-based bidding mechanism during both the local reinforcement learning training and deployment phases. When the total discharge demand of the vehicle group exceeds the actual demand of the power grid, the bidding mechanism is triggered: the system sorts users according to their willingness to discharge from highest to lowest; priority is given to satisfying the discharge demand of users with higher willingness to discharge until the power grid demand is met; users with lower willingness to discharge continue to charge or remain idle. Through this mechanism, the power grid demand and user interests can be dynamically balanced under uncertain environments, improving the acceptability of scheduling execution.
[0023] To address the issue of user heterogeneity and ensure adaptability of heterogeneous users in vehicle-to-grid (V2G) charging and discharging scheduling, this application employs a personalized federated reinforcement learning method. During the aggregation process, only the first two layers of the Critic network are shared and updated to extract common features of different vehicles' charging and discharging behaviors. In this way, different vehicles can learn universally applicable decision-making patterns at the global level, while retaining personalized parameters locally to adapt to the differentiated scheduling needs of vehicles with different characteristics. This mechanism not only ensures the model's adaptability to differences in vehicle usage characteristics but also improves the overall generalization effect of scheduling.
[0024] To address privacy risks, this application employs a federated learning algorithm, ensuring that users' raw charging and discharging data is always stored locally. During training, only necessary model parameters are uploaded for aggregation, fundamentally avoiding the risk of user behavior data leakage. Simultaneously, the distributed training process further enhances the system's security and scalability.
[0025] To address the issues of power grid and price fluctuations, this application employs an online heuristic update mechanism based on reinforcement learning, enabling the model to dynamically schedule operations during the deployment phase based on real-time electricity prices, grid load, and vehicle status. This process enhances the policy's adaptability to uncertain environments, thereby ensuring optimal user economics and grid coordination even under fluctuating conditions.
[0026] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0027] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0028] In an exemplary embodiment, as shown in FIG1, a method for scheduling the charging and discharging strategy of an electric vehicle is provided, including the following steps 201 to 205. Wherein:
[0029] Step 201: Obtain scheduling-related data for the current microgrid area within the current time period; the scheduling-related data includes electricity price, microgrid load demand, and charging / discharging behavior data for each electric vehicle within the current microgrid area; the charging / discharging behavior data includes arrival time, departure time, state of charge, driving demand, battery capacity, and user participation willingness.
[0030] Step 202: Aggregate the parameters of the first two layers of the Critic network in the trained initial scheduling model corresponding to all electric vehicles in the current microgrid area to obtain the global parameters of the first two layers of the Critic network; the initial scheduling model is obtained by iteratively training the deep reinforcement learning model of the Actor network-Critic network framework using a near-end policy optimization algorithm based on scheduling-related data in the current microgrid area over a historical period, with the goal of maximizing the scheduling objective function.
[0031] Step 203: Based on the global parameters, update the parameters of the first two layers of the Critic network of the initial scheduling model corresponding to all electric vehicles in the current microgrid area to obtain the collaborative scheduling model corresponding to each electric vehicle in the current microgrid area.
[0032] Step 204: Based on the scheduling-related data within the current time period, obtain the charging and discharging strategy for each electric vehicle within the current time period through the collaborative scheduling model corresponding to each electric vehicle; the charging and discharging strategy includes charging power and discharging power.
[0033] Step 205: Based on the user participation willingness of each electric vehicle in the current time period, the charging and discharging strategy is optimized using the willingness bidding method to obtain the optimized charging and discharging strategy for each electric vehicle in the current microgrid area in the current time period.
[0034] By implementing steps 201 to 205 above, this application can achieve cross-vehicle collaboration by aggregating the parameters of the first two layers of the Critic network without transmitting the original vehicle data, thereby significantly improving the level of privacy protection; by utilizing a bidding mechanism driven by real-time electricity prices, microgrid load, and individual willingness, the scheduling strategy can dynamically adapt to changes in electricity price fluctuations and grid demand, effectively improving peak shaving and valley filling accuracy; by quantifying user participation willingness and incorporating it into discharge priority ranking, the application solves the execution deviation caused by the traditional "complete obedience" assumption, improving user satisfaction and participation rate; and by using parallel computing on the vehicle side and lightweight parameter interaction, the application reduces communication latency and computational load, achieving a win-win situation of maximizing economic operation on the microgrid side and maximizing benefits on the vehicle owner side.
[0035] In another exemplary embodiment of this application, after obtaining the optimized charging and discharging strategy for each electric vehicle in the current microgrid area during the current time period, the method further includes:
[0036] The actual operating status data of each electric vehicle in the current microgrid area during the current time period is obtained after charging and discharging according to the corresponding optimized charging and discharging strategy; the actual operating status data includes the actual state of charge and charging and discharging power curve of each electric vehicle at the end of the scheduling period.
[0037] An experience buffer is established to store the state-action-reward-new state quadruple data; the state is the scheduling-related data, the action is the optimization of the charging and discharging strategy, the reward is the reward value determined by the scheduling objective function based on the actual operating state data, and the new state is the scheduling-related data updated according to the actual operating state data.
[0038] If the number of quadruplet data in the experience buffer does not reach the preset threshold, the next time period will be used as the current time period, and the process will return to the step "obtain the actual operating status data of each electric vehicle in the current microgrid area after charging and discharging according to the corresponding optimized charging and discharging strategy in the current time period".
[0039] If the number of quadruplet data in the experience buffer reaches a preset threshold, then based on the quadruplet data in the experience buffer, with the goal of maximizing the scheduling objective function, the cooperative scheduling model of the corresponding electric vehicle is updated using a near-end strategy optimization algorithm, and the buffer is cleared; the updated cooperative scheduling model is used as the initial scheduling model, and the next time period is used as the current time period, returning to the step "aggregate the parameters of the first two layers of the Critic network of all initial scheduling models to obtain the global parameters of the first two layers of the Critic network".
[0040] This application further enhances the system's adaptability in actual operation by introducing an online heuristic learning mechanism. Specifically, after executing the scheduling strategy, the system continuously collects actual operating data and builds an experience buffer. When the accumulated data reaches a preset threshold, it triggers online fine-tuning of the model based on a near-end strategy optimization algorithm. This process enables the scheduling model to perform self-iterative optimization using real-time feedback data after deployment, forming a closed loop of "decision-execution-learning". Through the online heuristic learning mechanism, the scheduling system no longer relies on a fixed historical training model, but can dynamically adjust strategy parameters according to actual operating results, effectively solving the problem of model performance degradation caused by dynamic factors such as changes in user behavior patterns and fluctuations in power grid conditions. Through continuous small-step online optimization, the system avoids the computational overhead of large-scale retraining and ensures that the scheduling strategy always adapts to the current environmental state, thereby significantly improving the stability, adaptability, and continuous optimization capability of the overall scheduling performance in long-term operation.
[0041] In another exemplary embodiment of this application, the willingness-based bidding method specifically includes:
[0042] The total planned discharge power of the current microgrid area during the current time period is obtained by summing up the discharge power of all electric vehicles in the current microgrid area during the current time period.
[0043] Based on the planned total discharge power, the optimized charging and discharging strategies for each electric vehicle in the current microgrid area during the current time period are obtained through the following steps:
[0044] If the planned total discharge power is greater than the microgrid load demand in the current microgrid area during the current time period, then all electric vehicles with non-zero discharge power in the current microgrid area during the current time period are sorted according to user participation willingness from high to low to obtain the set of discharge electric vehicles.
[0045] The electric vehicles in the discharge electric vehicle set are discharged sequentially, and the discharge power is accumulated to obtain the cumulative discharge power. When the cumulative discharge power equals the microgrid load demand in the current microgrid area for the current time period, the discharge power of the undischarged electric vehicles in the discharge electric vehicle set is set to zero.
[0046] In another exemplary embodiment of this application, the scheduling objective function is:
[0047] .
[0048] .
[0049] in, Represent the scheduling objective function; This indicates a microgrid regulation reward; Indicates a reward for user satisfaction; and Let represent the microgrid reward weighting coefficient at time t and the user satisfaction reward weighting coefficient at time t, respectively.
[0050] In another exemplary embodiment of this application, the microgrid reward weighting coefficient is:
[0051] .
[0052] .
[0053] .
[0054] in, This represents the competitive regulation weight of the microgrid at time t; This represents the competitive adjustment weight of user satisfaction at time t; This represents the basic weight of the microgrid at time t; The microgrid regulation competition factor at time t; This represents the initial weights of the microgrid at time t; This indicates the urgency of microgrid regulation at time t.
[0055] The weighting coefficient for user satisfaction rewards is:
[0056] .
[0057] .
[0058] .
[0059] in, This represents the basic weight of user satisfaction at time t; This represents the user satisfaction competition factor at time t; This represents the initial weight of user satisfaction at time t; This represents the urgency of user satisfaction at time t.
[0060] In another exemplary embodiment of this application, the microgrid regulation urgency is:
[0061] .
[0062] .
[0063] .
[0064] in, This indicates the urgency of the load at time t; This indicates the urgency of electricity price retardation at time t; This represents the microgrid load at time t; Indicates safe load; Indicates the maximum load that can be withstood; This indicates the average electricity price over the reference period. This represents the electricity price at time t; This represents the standard deviation of electricity prices within the reference time period.
[0065] The urgency level for user satisfaction is:
[0066] .
[0067] in, This indicates the current battery level of the electric vehicle; Indicates the target battery level of the electric vehicle at the end of the dispatch process; This indicates the total parking time of the electric vehicle; This indicates the remaining parking time for the electric vehicle.
[0068] In another exemplary embodiment of this application, the microgrid regulation competition factor is:
[0069] .
[0070] .
[0071] .
[0072] in, This represents the peak-shaving competition adjustment factor at time t; This represents the electricity price competition adjustment factor at time t; This represents the first fusion weight coefficient; This indicates the average electricity price over the reference period. This represents the electricity price at time t; This indicates the lowest and highest electricity prices within the reference time period; This represents the standard deviation of electricity prices within the reference period. The load indication function at time t; This represents the electricity price adjustment coefficient; This indicates the urgency of electricity price retardation at time t; This represents the set of peak electricity consumption periods; This represents the set of periods with low electricity consumption.
[0073] The competitive factors for user satisfaction are:
[0074] .
[0075] .
[0076] in, This represents the conflict severity adjustment factor; This indicates the severity of the conflict between the microgrid and the user's objective at time t; This indicates the urgency of microgrid regulation at time t; This represents the urgency of user satisfaction at time t.
[0077] In another exemplary embodiment of this application, the microgrid regulation reward is:
[0078] .
[0079] .
[0080] .
[0081] in, and These represent the second and third fusion weight coefficients, respectively. This indicates a peak-shaving and valley-filling reward based on peak and valley periods; This indicates an optimization reward based on real-time electricity prices; This represents the normalized charging power; This represents the set of peak electricity consumption periods; This represents the set of periods with low electricity consumption. This represents the normalized deviation of the electricity price at time t from the reference electricity price. This indicates the average electricity price over the reference period. This represents the electricity price at time t; and These represent the lowest and highest electricity prices within the reference time period, respectively. This represents the standard deviation of electricity prices within the reference time period.
[0082] In another exemplary embodiment of this application, the user satisfaction reward is:
[0083] .
[0084] .
[0085] .
[0086] .
[0087] in, Indicates the target battery level of the electric vehicle at the end of the dispatch process; This indicates the actual battery level of the electric vehicle at the end of the dispatch process; This indicates the initial electricity demand of electric vehicle users; Indicates the rated capacity of the electric vehicle's battery; This represents the excess margin based on user engagement levels; and These represent the minimum excess margin and the maximum excess margin, respectively.
[0088] The following uses a specific example of the charging and discharging strategy scheduling process of an electric vehicle to illustrate this application.
[0089] This application addresses the scheduling process between regional aggregators and clients, aiming to provide a personalized federated reinforcement learning-based electric vehicle intention-driven scheduling method. This method fully considers the participation intentions and privacy protection needs of heterogeneous users, enabling substantial scheduling results by retaining only the original data at the vehicle end. Furthermore, it introduces an intention-based bidding mechanism to maximize user-side economic benefits, thereby improving the fairness and feasibility of scheduling. Simultaneously, leveraging the dynamic decision-making characteristics of reinforcement learning, it achieves adaptive scheduling capabilities to fluctuations in grid load and user behavior. The power grid comprises several microgrid regions, each with a microgrid control center. The grid dispatch center generates and distributes grid-level scheduling reference information, including regional time-of-use / real-time electricity prices and microgrid load demand signals, to each microgrid control center based on the overall operating status.
[0090] The method proposed in this application is applicable to the vehicle-side local training mode in heterogeneous vehicle scenarios, as shown in Figure 3. Its technical solution includes the following steps:
[0091] 1. Training phase.
[0092] Client-side: ① Local data collection and willingness modeling: Each electric vehicle terminal collects charging behavior-related data locally, including vehicle arrival and departure times, state of charge (SOC), driving demand, battery capacity and electricity price, grid demand, and electric vehicle user participation willingness.
[0093] ② Local Reinforcement Learning Training: A scheduling objective function with user charging economy as the core goal is established at the vehicle end. A charging time-aligned reinforcement learning algorithm (Proximal Policy Optimization (PPO) algorithm) is used to perform local iterative training on the charging and discharging strategies (charging and discharging power) to obtain optimized local policy parameters (weights, biases). The deep reinforcement learning model is trained using the Proximal Policy Optimization algorithm within an Actor-Critic framework. The Actor network acts as the policy function, responsible for generating charging and discharging actions based on the environmental state; the Critic network acts as the value function, responsible for evaluating the long-term expected reward of the state and state-action pairs. The PPO algorithm, through an importance sampling ratio pruning mechanism, constrains the update magnitude of the Actor network each time, guided by the advantage function calculated by the Critic network, thereby improving policy performance while ensuring the stability of the training process. During training, the parameters of the Actor network and the Critic network are alternately optimized until the model converges.
[0094] Federation server side: ③ Federation parameter aggregation: Each vehicle periodically uploads the trained model parameters to the federated server. The federated server aggregates the parameters without touching the original data to generate a global model and achieve cross-user collaborative optimization.
[0095] The client receives the global model from the federated server to perform local training and updates.
[0096] ④ Global model update and distribution: The aggregated global model is distributed to each vehicle terminal to guide subsequent local training and policy updates.
[0097] Repeat steps ②, ③, and ④ until convergence.
[0098] Aggregator: 2. Deployment phase.
[0099] The system employs an online heuristic learning mechanism to dynamically update dispatch strategies based on real-time electricity prices, grid demand, and vehicle status. This process does not require access to raw user data but rather adapts quickly to environmental changes through lightweight strategy adjustments.
[0100] The entire approach is divided into two phases: a training phase and a deployment phase. The training phase involves training the model based on historical data, simulating real-world deployment scenarios as closely as possible, such as temporal alignment. However, to reduce training costs, the model update frequency and federated aggregation frequency are not very frequent. The deployment phase utilizes online heuristic learning to quickly adapt and optimize the model in a dynamic environment.
[0101] The entire process involves local reinforcement learning training with time-series alignment across clients. A bidding mechanism based on willingness is required; the total discharge requests from all clients in any given time period must not exceed a certain value. Priority is given to users with high willingness to discharge, and the discharge requests from each client are aggregated and bid on by the aggregation server. After a certain number of training rounds, the model is uploaded to the federated server for model aggregation, and then distributed to each client for a new round of training.
[0102] The process of coordination among the above three parties:
[0103] Step 1, Aggregation End: Implement data collection, which includes electricity demand and electricity price.
[0104] Step 2, Client (Electric Vehicle):
[0105] 2.1 Based on locally collected charging behavior data, including vehicle arrival and departure times, state of charge (SOC), driving demand, battery capacity and electricity price, grid demand, and electric vehicle user participation willingness, etc.
[0106] 2.2 Establish a reward function with user charging economy as the core objective at the vehicle end, and use a reinforcement learning algorithm with charging time alignment (proximal policy optimization algorithm, PPO) to perform local iterative training on the charging and discharging strategy to obtain the optimized local policy parameters.
[0107] The overall formula for the reward function (scheduling objective function) is as follows:
[0108] The total reward function for vehicle-to-grid (V2G) charging scheduling consists of two core components, using a standardized weighted summation method:
[0109] .
[0110] .
[0111] in, Represent the scheduling objective function; This indicates a microgrid regulation reward; Indicates a reward for user satisfaction; and Let represent the microgrid reward weighting coefficient at time t and the user satisfaction reward weighting coefficient at time t, respectively.
[0112] The calculation formulas for each of the above reward components are as follows:
[0113] The microgrid regulation reward is:
[0114] .
[0115] .
[0116] .
[0117] in, and These represent the second and third fusion weight coefficients, respectively. This indicates a peak-shaving and valley-filling reward based on peak and valley periods; This indicates an optimization reward based on real-time electricity prices; This represents the normalized charging power (-1 to 1, negative values indicate discharging power). This indicates the collection of peak electricity consumption periods (e.g., 5 PM - 9 PM). This indicates the set of off-peak electricity consumption periods (e.g., 0:00-7:00). This represents the normalized deviation of the electricity price at time t from the reference electricity price. This indicates the average electricity price over a reference period (the next 24 hours). During the deployment phase, the electricity price is based on the past 24 hours. This represents the electricity price at time t; and These represent the lowest and highest electricity prices within the reference time period, respectively. This represents the standard deviation of electricity prices within the reference time period.
[0118] The optimized reward based on real-time electricity prices involves the agent learning charging and discharging timings based on the load conditions reflected in the real-time electricity prices, serving as a reward for... This is a supplement.
[0119] The user satisfaction reward component involves imposing a secondary penalty on the agent for failing to meet the user's final energy requirement, encouraging it to achieve results based on user wishes. Personalized target energy .
[0120] .
[0121] .
[0122] .
[0123] .
[0124] in, Indicates the target battery level of the electric vehicle at the end of the dispatch process; This indicates the actual battery level of the electric vehicle at the end of the dispatch process; This indicates the initial electricity demand of electric vehicle users; Indicates the rated capacity of the electric vehicle's battery; Indicates based on user willingness to participate (0 represents the minimum willingness, 1 represents the maximum willingness) excess margin; and These represent the minimum excess margin and the maximum excess margin, respectively.
[0125] As an optional implementation method, the microgrid reward weighting coefficient is:
[0126] .
[0127] .
[0128] .
[0129] in, This represents the competitive regulation weight of the microgrid at time t; This represents the competitive adjustment weight of user satisfaction at time t; This represents the basic weight of the microgrid at time t; The microgrid regulation competition factor at time t; This represents the initial weights of the microgrid at time t; This indicates the urgency of microgrid regulation at time t.
[0130] The weighting coefficient for user satisfaction rewards is:
[0131] .
[0132] .
[0133] .
[0134] in, This represents the basic weight of user satisfaction at time t; This represents the user satisfaction competition factor at time t; This represents the initial weight of user satisfaction at time t; This represents the urgency of user satisfaction at time t.
[0135] In implementing this method, the basic urgency calculation includes the following calculations: microgrid regulation urgency and user satisfaction urgency.
[0136] The urgency of microgrid regulation is:
[0137] .
[0138] .
[0139] .
[0140] in, This indicates the load urgency at time t (based on time period division); This indicates the urgency of electricity price at time t (based on real-time electricity price). This represents the microgrid load at time t; Indicates safe load; This indicates the maximum tolerable load. When the grid load approaches the safe threshold, the peak shaving weight is significantly increased. This indicates the average electricity price over the reference period. This represents the electricity price at time t; This represents the standard deviation of electricity prices over the reference period. The greater the deviation of electricity prices from the mean, the more urgent the opportunity.
[0141] The urgency level for user satisfaction is:
[0142] .
[0143] in, This indicates the current battery level of the electric vehicle; Indicates the target battery level of the electric vehicle at the end of the dispatch process; This indicates the total parking time of the electric vehicle; This indicates the remaining parking time for the electric vehicle. The larger the battery deficit and the shorter the remaining time, the higher the weight of user satisfaction.
[0144] Among the above, the competitive regulation factor in the application of competitive regulation parameters... The calculations include the following processes for calculating microgrid regulation competition factors and user satisfaction competition factors:
[0145] Microgrid regulation competition factor: Peak shaving competition regulation (signal consistency enhancement). If the electricity price is also high during peak hours, the weight should be appropriately increased.
[0146] .
[0147] Electricity price competition regulation (dominated by flat-hour periods):
[0148] .
[0149] Unified microgrid regulation competition factor:
[0150] .
[0151] in, This represents the peak-shaving competition adjustment factor at time t; This represents the electricity price competition adjustment factor at time t; This represents the first fusion weight coefficient; This indicates the average electricity price over the reference period. This represents the electricity price at time t; This indicates the lowest and highest electricity prices within the reference time period; This represents the standard deviation of electricity prices within the reference period. The load indication function at time t; This represents the electricity price adjustment coefficient; This indicates the urgency of electricity price retardation at time t; This represents the set of peak electricity consumption periods; This represents the set of periods with low electricity consumption.
[0152] Competitive factors for user satisfaction: Protecting users' basic needs when both power grid and economic objectives are urgent.
[0153] Conflict severity calculation:
[0154] .
[0155] Competitive factors in user satisfaction:
[0156] .
[0157] in, This represents the conflict severity adjustment factor; This indicates the severity of the conflict between the microgrid and the user's objective at time t; This indicates the urgency of microgrid regulation at time t; This represents the urgency of user satisfaction at time t.
[0158] 2.3 Parameter update.
[0159] This application employs the Proximal Policy Optimization (PPO) algorithm to iteratively update the policy.
[0160] 2.3.1. Dominance function estimation.
[0161] Time step Advantage function Calculated using Generalized Advantage Estimation (GAE).
[0162] .
[0163] in, For time step The dominance function represents the advantage function at time step [0, 1]. Take action The additional benefit relative to the average strategy. Positive values. A negative value indicates that the action is better than the average strategy, while a negative value indicates that it is worse. Indicates a future time step ; This is a discount factor, ranging from 0 to 1, used to measure the current value of future rewards. The closer it is to 1, the more importance is placed on future rewards. The smoothing factor is used in generalized advantage estimation (GAE) to control the tradeoff between bias and variance. Between 0 and 1 This indicates that only TD (Temporal Difference) residuals are used. This indicates that Monte Carlo estimation was used. Indicates TD residual, Indicates time step The reward is calculated from the scheduling objective function. Let be the state value function, representing the state value at time t. The expected cumulative return that can be obtained by following the current strategy.
[0164] Function: To measure the quality of a particular action relative to the average policy, and to provide direction for subsequent policy gradient updates.
[0165] 2.3.2. Strategy update mechanism.
[0166] Define probability ratio Measure the new strategy Compared to the old strategy In the same state Select the same action below The possibility of change. When When, it indicates that the new strategy is more inclined to choose that action; when When this is the case, it indicates that the new strategy is less inclined to choose that action.
[0167] .
[0168] in, This represents the probability ratio between the old and new strategies; Represents the parameters of the policy network; Indicates the current strategy. This indicates the old strategy.
[0169] PPO introduces a pruning mechanism on top of the standard policy gradient to limit the magnitude of a single-step update:
[0170] .
[0171] in, Let be the loss function for the pruning strategy. This represents all time steps in the current batch of data. Expectations based on experience; This is the clipping threshold.
[0172] Function: While increasing the probability of advantageous actions, it limits the differences between new and old strategies and ensures the stability of updates.
[0173] In addition, an entropy regularization term was added to encourage exploration:
[0174] .
[0175] in, This represents the entropy regularization term corresponding to the parameters of the policy network; The entropy coefficient controls the strength of entropy regularization. The larger the value, the more exploration is encouraged. Indicates the policy in the state Entropy measures the randomness of a strategy. Higher entropy indicates a more random strategy.
[0176] 2.3.3. Value function update.
[0177] To ensure the stability of the value network, a mean squared error loss for pruning is introduced:
[0178] .
[0179] in, Representing value network parameters The corresponding mean squared error loss function for cropping; Value estimation for value network computation is based on value network parameters. Parameterize the value network and output the state. Value estimate. Value estimation calculated for the old value network, old value network parameters Parameterize the value network and output its state. The value estimate is based on a fixed value network before the update, used to calculate the target value. , representing the reward at time step t. The pruning threshold is also used for value function updates to prevent the value function from being updated too large.
[0180] 2.3.4. Synthetic objective function and parameter update.
[0181] The final optimization objective function of the model is:
[0182] .
[0183] in, Parameters representing the policy network The corresponding optimization objective function; The weights are the values for the loss function. These are the parameters of the policy network. These are the parameters of the value network.
[0184] The model parameters are updated following the gradient descent principle:
[0185] .
[0186] .
[0187] in, These represent the learning rates for the Actor network and the Critic network, respectively. The Actor network and the Critic network use different learning rates. The optimizer chosen is Adam.
[0188] To enhance training stability, a gradient clipping mechanism is introduced:
[0189] .
[0190] in, Represents the gradient; This is the preset gradient clipping threshold.
[0191] In addition, there is a KL divergence early termination mechanism, which immediately terminates the current update when the average KL divergence between the old and new strategies exceeds the preset KL divergence threshold.
[0192] .
[0193] in, Indicates KL divergence; This is the preset KL divergence threshold.
[0194] 2.4 Each vehicle periodically uploads the trained model parameters to the federated server.
[0195] Step 3, Federation Server: The federated server performs personalized parameter aggregation without touching the original data, that is, it only aggregates the first two layers of the critic network to generate a global model, realizes cross-user collaborative optimization, and distributes the global model to each vehicle terminal.
[0196] The central server collects data from... Parameters of the first two layers of the Critic network for each client Then, the averaging method is used for aggregation to generate the parameters of the first two layers of the next-generation Critic network. The aggregation formula is as follows:
[0197] .
[0198] in, It represents the number of clients.
[0199] Step 4, Client: Receive the global model from the federated server: The aggregated global model is distributed to each vehicle terminal to guide subsequent local training and policy updates.
[0200] The local model also uses an actor-critic architecture, where the first two layers of the critic network are directly replaced with the parameters aggregated from the global model, and reinforcement learning training and optimization are performed locally. The client output is the charging and discharging power.
[0201] Step 5, Aggregator: Based on the client's update results, the scheduling strategy is dynamically updated according to real-time electricity prices, grid demand, and vehicle status through an online heuristic learning mechanism.
[0202] Specifically, online heuristic mechanisms: online means acquiring data in real time and dynamically changing strategies; heuristic means relying on pre-trained heuristic strategies, which can help the model make decisions within limited time and resources.
[0203] Based on the same inventive concept, this application also provides a charging and discharging strategy scheduling system for electric vehicles to implement the charging and discharging strategy scheduling method for electric vehicles described above. The solution provided by this system is similar to the implementation scheme described in the above method. Therefore, the specific limitations of one or more embodiments of the charging and discharging strategy scheduling system for electric vehicles provided below can be found in the limitations of the charging and discharging strategy scheduling system for electric vehicles described above, and will not be repeated here.
[0204] In an exemplary embodiment, as shown in FIG4, a charging and discharging strategy scheduling system for an electric vehicle is provided, comprising:
[0205] The scheduling-related data acquisition module 301 is used to acquire scheduling-related data for the current microgrid area within the current time period. The scheduling-related data includes electricity price, microgrid load demand, and charging and discharging behavior data of each electric vehicle in the current microgrid area. The charging and discharging behavior data includes arrival time, departure time, state of charge, driving demand, battery capacity, and user participation willingness.
[0206] The global parameter aggregation module 302 is used to aggregate the parameters of the first two layers of the Critic network in the trained initial scheduling model corresponding to all electric vehicles in the current microgrid area to obtain the global parameters of the first two layers of the Critic network. The initial scheduling model is obtained by iteratively training the deep reinforcement learning model of the Actor network-Critic network framework using a near-end policy optimization algorithm based on scheduling-related data in the current microgrid area over a historical period, with the goal of maximizing the scheduling objective function.
[0207] The collaborative model generation module 303 is used to update the parameters of the first two layers of the Critic network of the initial scheduling model corresponding to all electric vehicles in the current microgrid area based on the global parameters, so as to obtain the collaborative scheduling model corresponding to each electric vehicle in the current microgrid area.
[0208] The strategy generation module 304 is used to obtain the charging and discharging strategy of each electric vehicle in the current time period based on the scheduling-related data in the current time period and through the cooperative scheduling model corresponding to each electric vehicle; the charging and discharging strategy includes charging power and discharging power.
[0209] The strategy optimization module 305 is used to optimize the charging and discharging strategy based on the user participation willingness of each electric vehicle in the current time period using the willingness bidding method, so as to obtain the optimized charging and discharging strategy for each electric vehicle in the current microgrid area in the current time period.
[0210] In summary, this application has the following technical effects:
[0211] 1. Improved Feasibility and Participation Rate: This application explicitly introduces user willingness into the local data collection and willingness modeling stages, and combines it with a willingness-based auction bidding mechanism, which effectively avoids the distortion caused by the traditional method's reliance on the "users completely obey the assumption" and thus improves the feasibility of the scheduling scheme and the actual user participation rate.
[0212] 2. Enhanced Adaptability to Fluctuations: Relying on the dynamic decision-making characteristics of reinforcement learning, this application can achieve real-time adjustment of strategies under uncertain conditions such as electricity price fluctuations and load fluctuations. During the deployment phase, the strategy is updated in a lightweight manner through online heuristic learning, thereby ensuring adaptability to fluctuations in complex environments.
[0213] 3. Privacy Protection and Compliance: This application adopts a federated learning algorithm to ensure that the original data is kept only on the vehicle side, and the scheduling server only receives model parameters and aggregates them, thus avoiding the risk of original data leakage and meeting the dual requirements of privacy protection and data compliance.
[0214] 4. Heterogeneous Adaptation and Generalization Ability: This application uses personalized federated learning to share only the first two layers of the Critic network during the aggregation process to learn a general representation across users, while the remaining layers are updated independently locally. This not only maintains the personalized adaptation of different user characteristics but also achieves cross-user generalization ability.
[0215] In an exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram is shown in Figure 5. The computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is connected to the system bus via the I / O interfaces. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database of the computer device is used for processing data related to the charging and discharging strategy of electric vehicles. The I / O interfaces of the computer device are used for exchanging information between the processor and external devices. The communication interface of the computer device is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements a charging and discharging strategy scheduling method for electric vehicles.
[0216] Those skilled in the art will understand that the structure shown in Figure 5 is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. A specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0217] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0218] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0219] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0220] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).
[0221] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0222] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0223] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A method for scheduling the charging and discharging strategy of an electric vehicle, characterized in that, include: Obtain scheduling-related data for the current microgrid area within the current time period; The scheduling-related data includes electricity price, microgrid load demand, and charging / discharging behavior data for each electric vehicle within the current microgrid area. The charging / discharging behavior data includes arrival time, departure time, state of charge, driving demand, battery capacity, and user participation willingness. The parameters of the first two layers of the Critic network in the pre-trained initial scheduling model corresponding to all electric vehicles within the current microgrid area are aggregated to obtain the global parameters of the first two layers of the Critic network. The initial scheduling model is obtained by iteratively training a deep reinforcement learning model of the Actor network-Critic network framework using a near-end policy optimization algorithm, based on scheduling-related data from historical time periods within the current microgrid area, with the goal of maximizing the scheduling objective function. During the aggregation process, only the parameters of the first two layers of the Critic network are shared and updated to extract common features of different vehicles in charging and discharging behavior. Based on the global parameters, the parameters of the first two layers of the Critic network of the initial scheduling model corresponding to all electric vehicles in the current microgrid area are updated to obtain the collaborative scheduling model corresponding to each electric vehicle in the current microgrid area. Based on the scheduling-related data in the current time period, the charging and discharging strategy of each electric vehicle in the current time period is obtained through the collaborative scheduling model corresponding to each electric vehicle. The charging and discharging strategy includes charging power and discharging power. Based on the user participation willingness of each electric vehicle in the current time period, the charging and discharging strategy is optimized using the willingness bidding method to obtain the optimized charging and discharging strategy of each electric vehicle in the current time period in the current microgrid area.
2. The electric vehicle charging and discharging strategy scheduling method according to claim 1, characterized in that, After obtaining the optimized charging and discharging strategy for each electric vehicle in the current microgrid area within the current time period, the method further includes: acquiring the actual operating status data of each electric vehicle in the current microgrid area within the current time period after charging and discharging according to the corresponding optimized charging and discharging strategy; the actual operating status data includes the actual state of charge and charging / discharging power curve of each electric vehicle at the end of the scheduling period; establishing an experience buffer to store the state-action-reward-new state quadruple data; the state is scheduling-related data, the action is the optimized charging and discharging strategy, the reward is the reward value determined by the scheduling objective function based on the actual operating status data, and the new state is the scheduling-related data updated according to the actual operating status data; if the number of quadruple data in the experience buffer does not reach a preset threshold... If the next time period is taken as the current time period, the process returns to the step "obtain the actual operating status data of each electric vehicle in the current microgrid area after charging and discharging according to the corresponding optimized charging and discharging strategy in the current time period"; if the number of quadruple data in the experience buffer reaches the preset number threshold, then based on the quadruple data in the experience buffer, with the goal of maximizing the scheduling objective function, the cooperative scheduling model of the corresponding electric vehicle is updated using the near-end strategy optimization algorithm, and the buffer is cleared; the updated cooperative scheduling model is taken as the initial scheduling model, and the next time period is taken as the current time period, and the process returns to the step "aggregate the parameters of the first two layers of the Critic network of all initial scheduling models to obtain the global parameters of the first two layers of the Critic network".
3. The electric vehicle charging and discharging strategy scheduling method according to claim 1, characterized in that, The willingness-based bidding method specifically includes: accumulating the discharge power of all electric vehicles in the current microgrid area during the current time period to obtain the planned total discharge power of the current microgrid area during the current time period; based on the planned total discharge power, obtaining the optimized charging and discharging strategy for each electric vehicle in the current microgrid area during the current time period through the following steps: if the planned total discharge power is greater than the microgrid load demand in the current microgrid area during the current time period, then sorting all electric vehicles with non-zero discharge power in the current microgrid area during the current time period according to user participation willingness from high to low, to obtain a set of electric vehicles to be discharged; discharging the electric vehicles in the set of electric vehicles to be discharged in sequence, and accumulating the discharge power to obtain the cumulative discharge power, until the cumulative discharge power equals the microgrid load demand in the current microgrid area during the current time period, then setting the discharge power of the undischarged electric vehicles in the set of electric vehicles to zero.
4. The electric vehicle charging and discharging strategy scheduling method according to claim 1, characterized in that, The scheduling objective function is: ; ;in, Represent the scheduling objective function; This indicates a microgrid regulation reward; Indicates a reward for user satisfaction; and Let represent the microgrid reward weighting coefficient at time t and the user satisfaction reward weighting coefficient at time t, respectively.
5. The electric vehicle charging and discharging strategy scheduling method according to claim 4, characterized in that, The microgrid reward weighting coefficient is: ; ; ;in, This represents the competitive regulation weight of the microgrid at time t; This represents the competitive adjustment weight of user satisfaction at time t; This represents the basic weight of the microgrid at time t; The microgrid regulation competition factor at time t; This represents the initial weights of the microgrid at time t; This represents the urgency of microgrid regulation at time t; the user satisfaction reward weighting coefficient is: ; ; ;in, This represents the basic weight of user satisfaction at time t; This represents the user satisfaction competition factor at time t; This represents the initial weight of user satisfaction at time t; This represents the urgency of user satisfaction at time t.
6. The electric vehicle charging and discharging strategy scheduling method according to claim 5, characterized in that, The urgency of microgrid regulation is: ; ; ;in, This indicates the urgency of the load at time t; This indicates the urgency of electricity price retardation at time t; This represents the microgrid load at time t; Indicates safe load; Indicates the maximum load that can be withstood; This indicates the average electricity price over the reference period. This represents the electricity price at time t; This represents the standard deviation of electricity prices within the reference time period; the urgency of user satisfaction is: ;in, This indicates the current battery level of the electric vehicle; Indicates the target battery level of the electric vehicle at the end of the dispatch process; This indicates the total parking time of the electric vehicle; This indicates the remaining parking time for the electric vehicle.
7. The electric vehicle charging and discharging strategy scheduling method according to claim 5, characterized in that, The microgrid regulation competition factor is: ; ; ;in, This represents the peak-shaving competition adjustment factor at time t; This represents the electricity price competition adjustment factor at time t; This represents the first fusion weight coefficient; This indicates the average electricity price over the reference period. This represents the electricity price at time t; This indicates the lowest and highest electricity prices within the reference time period; This represents the standard deviation of electricity prices within the reference period. The load indication function at time t; This represents the electricity price adjustment coefficient; This indicates the urgency of electricity price retardation at time t; This represents the set of peak electricity consumption periods; This represents the set of off-peak electricity consumption periods; the user satisfaction competition factor is: ; ;in, This represents the conflict severity adjustment factor; This indicates the severity of the conflict between the microgrid and the user's objective at time t; This indicates the urgency of microgrid regulation at time t; This represents the urgency of user satisfaction at time t.
8. The electric vehicle charging and discharging strategy scheduling method according to claim 4, characterized in that, The microgrid regulation reward is: ; ; ;in, and These represent the second and third fusion weight coefficients, respectively. This indicates a peak-shaving and valley-filling reward based on peak and valley periods; This indicates an optimization reward based on real-time electricity prices; This represents the normalized charging power; This represents the set of peak electricity consumption periods; This represents the set of periods with low electricity consumption. This represents the normalized deviation of the electricity price at time t from the reference electricity price. This indicates the average electricity price over the reference period. This represents the electricity price at time t; and These represent the lowest and highest electricity prices within the reference time period, respectively. This represents the standard deviation of electricity prices within the reference time period.
9. The electric vehicle charging and discharging strategy scheduling method according to claim 4, characterized in that, User satisfaction rewards are: ; ; ; ;in, Indicates the target battery level of the electric vehicle at the end of the dispatch process; This indicates the actual battery level of the electric vehicle at the end of the dispatch process; This indicates the initial electricity demand of electric vehicle users; Indicates the rated capacity of the electric vehicle's battery; This represents the excess margin based on user engagement levels; and These represent the minimum excess margin and the maximum excess margin, respectively.
10. A charging and discharging strategy scheduling system for electric vehicles, characterized in that, The electric vehicle charging and discharging strategy scheduling system uses the electric vehicle charging and discharging strategy scheduling method according to any one of claims 1-9. The electric vehicle charging and discharging strategy scheduling system includes: a scheduling-related data acquisition module, used to acquire scheduling-related data for the current microgrid area within the current time period; the scheduling-related data includes electricity price, microgrid load demand, and charging and discharging behavior data of each electric vehicle within the current microgrid area; the charging and discharging behavior data includes arrival time, departure time, state of charge, driving demand, battery capacity, and user participation willingness; a global parameter aggregation module, used to aggregate the parameters of the first two layers of the Critic network in the trained initial scheduling model corresponding to all electric vehicles within the current microgrid area to obtain the global parameters of the first two layers of the Critic network; the initial scheduling model is based on scheduling-related data within the historical time period of the current microgrid area, aiming to maximize the scheduling objective function, and uses an Actor network-Critic network framework. The deep reinforcement learning model is obtained through iterative training using a near-end policy optimization algorithm. During the aggregation process, only the first two layers of the Critic network are shared and updated to extract common features of different vehicles' charging and discharging behaviors. A collaborative model generation module updates the parameters of the first two layers of the Critic network of the initial scheduling model for all electric vehicles within the current microgrid area based on the global parameters, obtaining a collaborative scheduling model for each electric vehicle within the current microgrid area. A policy generation module, based on scheduling-related data within the current time period, obtains the charging and discharging strategy for each electric vehicle within the current time period using the collaborative scheduling model for each electric vehicle. The charging and discharging strategy includes charging power and discharging power. A policy optimization module optimizes the charging and discharging strategy based on the user participation willingness of each electric vehicle within the current time period using a willingness-based bidding method, obtaining an optimized charging and discharging strategy for each electric vehicle within the current time period in the current microgrid area.
Citation Information
Patent Citations
Federal reinforcement learning-based data enhancement method and related equipment
CN119740027A
Internet of vehicles communication resource allocation method based on federal multi-agent deep reinforcement learning
CN120390236A