An electric vehicle charging and discharging intelligent scheduling method based on user behavior modeling and TD3

CN122823552APending Publication Date: 2026-09-25STATE GRID ZHEJIANG ELECTRIC POWER CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610877539.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-17
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

然而,DDPG 算法固有地存在Q值高估问题——由于 Critic 网络使用同一个目标值进行更新,对动作价值的估计容易产生正向偏差并随训练传播,导致策略陷入次优或训练发散,这一问题在高维连续动作空间调度中尤为严重,并且缺乏对用户实际充电行为价格弹性的量化建模,使得调度策略与经济激励脱节

Benefits of technology

[0016]本发明的有益效果是:通过将公共充电场景中确定的用户行为价格弹性参数,嵌入电动汽车充放电调度的马尔可夫决策过程中的当前时刻环境内,并设置融合电网侧削峰填谷目标与用户侧经济成本的加权复合奖励函数,引导用户从“消费者”转变为“产销者”;通过TD3模型的双Q网络最小值截断机制、目标策略平滑机制和延迟更新机制,有效抑制了传统DDPG算法固有的Q值高估问题,实现了在高维连续动作空间中的快速收敛与稳定决策;设计的负荷联动型动态价格公式将实时电网负荷偏差率直接映射为车网互动回购电价,实现了电网物理状态到用户经济信号的分钟级实时传导,克服了静态分时电价无法响应实时电网状态的缺陷;同时,本发明形成的“用户签约与分群、策略下发与调度、过程计量与响应核验、收益结算与反馈优化”四环节实施路径及用户收益分配,为车网互动提供了可工程落地的商业闭环。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122823552A_ABST
    Figure CN122823552A_ABST
Patent Text Reader

Abstract

The application discloses an electric vehicle charging and discharging intelligent scheduling method based on user behavior modeling and TD3, which comprises the following steps: step one, a three-party cooperation framework of aggregators, charging station operators and vehicle owners is constructed; step two, user behavior price elasticity parameters are determined based on public charging scene operation data, and a Markov decision process of electric vehicle charging and discharging scheduling is established; step three, a double-delay deep deterministic policy gradient algorithm model is constructed to obtain a pre-trained Actor network model; step four, a load linkage type dynamic price incentive mechanism is designed, and a real-time power grid load deviation rate of a power grid is mapped to a vehicle-grid interaction repurchase price; and step five, the above method is sequentially divided into a day-ahead declaration, an intra-day scheduling and a closed-loop execution process of after-the-fact settlement according to time and business logic. Through the application, the power grid can realize peak shaving and valley filling, the user's power consumption cost can be reduced, line congestion can be relieved, and an engineering-landed commercial closed-loop solution for vehicle-grid interaction is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of power system operation optimization and intelligent charging and discharging technology for electric vehicles, specifically involving an intelligent scheduling method for electric vehicle charging and discharging based on user behavior modeling and TD3. Background Technology

[0002] With the explosive growth of electric vehicle (EV) ownership, large-scale disorderly charging exacerbates the peak-valley difference in the power grid, increasing the pressure on distribution network expansion and safe operation. Traditional scheduling methods rely heavily on precise mathematical models and deterministic optimization algorithms. However, these methods often struggle to guarantee real-time performance and optimality when faced with the complex travel demands of massive EV users, fluctuating market prices, and random charging behavior. Some studies have attempted to apply reinforcement learning algorithms such as Deep Deterministic Policy Gradient (DDPG) to EV scheduling, addressing the key bottleneck of discrete methods by employing an Actor-Critic architecture to output continuous actions. However, the DDPG algorithm inherently suffers from Q-value overestimation—because the Critic network uses the same target value for updates, the estimation of action value is prone to positive bias that propagates during training, leading to suboptimal strategies or training divergence. This problem is particularly severe in high-dimensional continuous action space scheduling and lacks quantitative modeling of the price elasticity of actual user charging behavior, causing the scheduling strategy to become disconnected from economic incentives.

[0003] Existing business models, largely based on static time-of-use pricing, lack the ability to dynamically adjust according to real-time grid load conditions, making it difficult to fully leverage the flexibility of user-side resources. Furthermore, there is a lack of solutions that simultaneously consider the distribution of benefits and incentive transmission among aggregators, charging station operators, and end-user vehicle owners, often neglecting the price elasticity and heterogeneity of user behavior. This leads to a disconnect between incentive strategies and actual user responses, hindering the achievement of a win-win situation for the grid, aggregators, and users. Therefore, constructing a dynamic adjustment mechanism that can respond to grid conditions in real time while also considering user interests is crucial to resolving this supply-demand imbalance. Summary of the Invention

[0004] The purpose of this invention is to overcome the shortcomings of the prior art and provide an intelligent scheduling method for electric vehicle charging and discharging based on user behavior modeling and TD3. This method constructs a collaborative framework that takes into account grid security, user travel, and aggregator profitability, enabling load-linked dynamic pricing, precise incentives, and an executable business closed loop, thereby improving the efficiency and feasibility of vehicle-grid interactive scheduling.

[0005] To achieve the above effects, this invention proposes an intelligent scheduling method for electric vehicle charging and discharging based on user behavior modeling and TD3, which includes the following steps: Step 1: Build a three-party collaborative framework among aggregators, charging station operators, and car owners, and clarify the core responsibilities of each party and the path for aggregators to participate in the market; Step 2: Collect historical operation datasets of public charging scenarios and clean the data to obtain a clean operation dataset; based on the clean operation dataset, determine the price elasticity parameter of charging quantity, the price elasticity parameter of charging duration, and the marginal effect of early termination of charging; based on the tripartite collaboration framework and the price elasticity parameters of charging quantity, charging duration, and the marginal effect coefficient of early termination of charging, construct a Markov decision process by defining a quintuple. Step 3: Map the Markov decision process to a trainable neural network architecture; based on this neural network architecture, construct the TD3 model using the dual-delay deep deterministic policy gradient algorithm; and obtain the pre-trained Actor network model based on the TD3 model. Step 4: Design a dynamic price incentive mechanism based on real-time grid load status, generate charging and discharging instructions using a pre-trained Actor network model, and establish a feedback optimization mechanism to iteratively update the model and parameters. Step five involves dividing steps one through four into three stages based on time and business logic: daily reporting, intraday scheduling, and post-event settlement. This results in a closed-loop execution process, clearly defining the boundaries of responsibilities and settlement rules for each stage, thereby enabling intelligent scheduling of electric vehicle charging and discharging.

[0006] Furthermore, the tripartite collaborative framework described in step one includes: the vehicle owner, as a flexible resource provider, responds to price signals and provides adjustable load; the charging station operator is responsible for the investment, operation and maintenance of physical facilities and the execution of local strategies; and the aggregator, as a bridge connecting dispersed users and the power grid, is responsible for receiving grid dispatch instructions, formulating internal incentive prices, and aggregating adjustable capacity on the user side.

[0007] Furthermore, step two includes the following sub-steps: (2.1) Collect historical operation datasets of public charging scenarios and clean the data to obtain clean operation datasets; the historical operation datasets include data fields and several charging session records, the data fields include transaction time, electricity, cost, location, termination reason and abnormal marker; (2.2) Based on the clean operation dataset, a charging quantity response model and a charging duration response model are constructed using a double log-linear regression model to determine the charging quantity price elasticity parameter and the charging duration price elasticity parameter; based on the clean operation dataset, an early termination of charging behavior model is constructed using a binary choice model to determine the early termination of charging marginal effect coefficient. (2.3) Based on the aforementioned tripartite collaborative framework and the price elasticity parameters of charging quantity, charging duration, and early termination of charging marginal effect coefficient, a Markov decision process is constructed by defining a quintuple and instantiating its elements into a programmable simulation environment interface; the quintuple includes a multidimensional state space S, a continuous action space A, a state transition probability P, a weighted composite reward function R, and a discount factor. The Markov decision process models the aggregator as a reinforcement learning agent and describes its interactive decision-making process with the environment model through a quintuple.

[0008] Furthermore, the expressions for the power response model and charging time response model mentioned in step (2.2) are as follows: in, For price variables, Control variables include time period, week, and grouping characteristics. The amount of charge for the i-th charging session. For the intercept term, This is a parameter representing the price elasticity of charge volume. For control variables The corresponding coefficient vector, This is the random error term; Let i be the duration of the i-th charging session. For the intercept term, This is a parameter related to the price elasticity of charging time. For control variables The corresponding coefficient vector, This is the random error term; The expression for the early termination of charging behavior model is as follows: in, The probability of premature termination of charging. Define an indicator variable to indicate the early termination of charging. This indicates that charging has ended early; otherwise, the value is 0. For the distribution function of the binary choice model, For the intercept term, The marginal effect coefficient for early termination of charging represents the direction and strength of the influence of the price variable on the probability of early termination of charging. For price variables, Control variables include time period, week, and grouping characteristics. For control variables The corresponding coefficient vector.

[0009] Furthermore, the definition of the quintuple in step (2.3) specifically includes the following sub-steps: (2.3.1) Define a multidimensional state space S, which reflects the state characteristics of the environment and vehicles at the current moment. To support the scheduling decision of the aggregator, output the observable state vector at time t: in, Let be the observable state vector at time t. This indicates the state of charge of the electric vehicle battery at time t, and also sets the physical safety boundary of the battery to prevent overcharging and over-discharging. ; and Encode time using trigonometric functions to eliminate abrupt changes in periodic boundaries. The total number of time steps in a scheduling cycle; The base electricity price at time t is determined according to the time-of-use pricing rules of the region. Represents the base load deviation rate of the power grid at time t. The time in the daily load curve The base load, This represents the daily average load. (2.3.2) Define a continuous action space A, where the continuous action space A is the charging and discharging power command applied by the aggregator agent to the electric vehicle charging pile, and the observable state vector is used as the basis for the action space A. Input is normalized continuous motion, output is normalized continuous motion. ,in Indicates charging. Indicates discharge; (2.3.3) Define the state transition probability P, which is based on the current observable state vector. and the normalized continuous action As input, based on the price elasticity parameters of charging quantity, charging duration, and early termination of charging marginal effect coefficient, the price signal is converted into the user's response behavior, and the state transition probability is output. It quantifies the next state after the current state and the execution of the scheduling instruction. The evolutionary pattern; (2.3.4) Define a weighted composite reward function R, which includes a weighted sum of grid-side rewards, user-side rewards, and constraint penalties. The weighted composite reward function at time t... The expression is as follows: As a reward weight; User-side rewards The aim is to minimize user costs: in, For real-time electricity prices, This is the battery loss cost coefficient. For actual power, when At that time, the user receives revenue from electricity sales, and the reward is positive; Grid-side rewards The aim is to smooth out peak and valley peaks. A reward will be given when the power grid is at its peak and vehicles are discharging electricity; Restraint and Punishment To ensure users' travel needs are met: in, The penalty coefficient is... The battery state of charge as desired by the user when offline. The actual state of charge (SOC) of the off-grid battery is used as the basis for determining the actual SOC. If the SOC does not meet the user's expected value at the time of off-grid connection, a constraint penalty will be imposed. (2.3.5) Define the discount factor The discount factor The weights used to balance current and future rewards are determined based on the time series length and decision-making process of the scheduling problem.

[0010] Furthermore, step three includes the following sub-steps: (3.1) Map the multidimensional state space S and continuous action space A in the Markov decision process to the input-output structure of a trainable neural network, and obtain the neural network architecture; (3.2) Based on the neural network architecture, a TD3 model is constructed using the double-delay deep deterministic policy gradient algorithm; the TD3 model includes an Actor network, two Critic networks, a double Q network minimum truncation mechanism, a target policy smoothing mechanism, and a delayed update mechanism; (3.3) Experience samples generated by the interaction between the aggregator agent and the environment Stored in the experience replay pool ; (3.4) Based on the experience replay pool, obtain training samples, iteratively update the TD3 model until the convergence condition is met, and extract the pre-trained Actor network model.

[0011] Furthermore, in step (3.2), based on the neural network architecture, the TD3 model is constructed using a dual-delay deep deterministic strategy gradient algorithm, including the following components: a. Construct an Actor network and two Critic networks: This results in an Actor network containing the main network. and the corresponding target network The Actor main network is used for decision-making and updates, while the Actor target network is used to provide stable objectives; the two Critic networks contain the main network. and the corresponding target network , The Critic main network is used to estimate the Q-value of the current state-action, and the Critic target network is used to calculate the Q-value in the Bellman objective.

[0012] b. Constructing a dual-Q network minimum truncation mechanism: The dual-Q network minimum truncation mechanism refers to using the smaller value in the output of two independent Critic target networks to suppress the overestimation of Q value in the DDPG algorithm when calculating the Bellman objective value; c. Constructing a target policy smoothing mechanism: The target policy smoothing mechanism refers to adding truncated normal noise to the smoothed target action when calculating the Q value, so that the learning target of the Critic target network becomes... This is equivalent to performing a local averaging of the Q-function of the Critic target network in the action space; d. Constructing a delayed update mechanism: The delayed update mechanism refers to making the update frequency of the Actor network lower than that of the Critic network, so that the Critic network has more time to converge value estimation under a fixed policy, thereby improving the accuracy of policy updates.

[0013] Furthermore, step four includes the following sub-steps: (4.1) Based on the real-time grid load status, the physical constraints of the grid are transformed into economic signals that users can perceive through price levers. A load-linked dynamic pricing mechanism is designed, and a modified aggregated load curve is obtained. (4.2) Based on the pre-trained Actor network model, the actions output by the model are converted into actual hardware charging and discharging power instructions and fund allocation schemes, and the execution data is recorded; (4.3) Based on the execution data, establish a feedback optimization mechanism to iteratively update the incentive parameters and model parameters.

[0014] Furthermore, step (4.1) includes the following sub-steps: (4.1.1) Based on the data provided by the power grid side in the aforementioned tripartite collaborative framework, calculate the real-time aggregated power grid load; the calculation formula for the real-time aggregated power grid load is as follows: in, The aggregate load for time period t, For time period The total number of charging sessions currently in a charging state. The instantaneous charging power of the i-th charging session during time period t; (4.1.2) Based on the real-time aggregated load of the power grid, calculate the real-time power grid load deviation rate. The calculation formula is as follows: in, This represents the real-time grid load deviation rate. For real-time grid load aggregation, This represents the average load for the day. (4.1.3) When At that time, the vehicle-to-grid repurchase price, which reflects the real-time congestion level of the power grid, is obtained: in, The electricity price will be repurchased for vehicle-to-everything (V2X) interaction. As a benchmark for time-of-use electricity pricing, For time period coefficients; This is the incentive step coefficient; (4.1.4) Based on the aforementioned charging quantity price elasticity parameter And charging time price elasticity parameter The charging amount for a single charging session is adjusted based on the vehicle-to-grid (V2G) repurchase electricity price incentive. and correct charging time By correcting the elasticity formula, the corrected aggregate load curve for time period t under excitation conditions is obtained.

[0015] Furthermore, the three stages in step five are as follows: the day-ahead declaration stage is to declare adjustable capacity based on the user-set off-grid time, target electricity volume, and response floor price; the day-ahead scheduling stage is to generate dynamic electricity price and power instructions based on real-time load deviation rate and pre-trained Actor network model and execute them; the post-event settlement stage is to settle fees and distribute revenue based on the actual response electricity volume measured on the charging pile side, and to feed back the execution effect to the feedback optimization mechanism in step four.

[0016] The beneficial effects of this invention are as follows: By embedding the price elasticity parameters of user behavior determined in public charging scenarios into the current moment environment of the Markov decision process for electric vehicle charging and discharging scheduling, and setting a weighted composite reward function that integrates the grid-side peak shaving and valley filling objectives with the user-side economic costs, users are guided to transform from "consumers" to "producers and sellers." Through the dual-Q network minimum truncation mechanism, target policy smoothing mechanism, and delayed update mechanism of the TD3 model, the inherent Q-value overestimation problem of the traditional DDPG algorithm is effectively suppressed, achieving rapid convergence and stable decision-making in a high-dimensional continuous action space. The designed load-linked dynamic price formula directly maps the real-time grid load deviation rate to the vehicle-grid interaction buyback price, realizing minute-level real-time transmission from the grid physical state to the user's economic signal, overcoming the defect that static time-of-use pricing cannot respond to the real-time grid state. At the same time, the four-stage implementation path and user revenue distribution formed by this invention—"user contracting and grouping, strategy issuance and scheduling, process measurement and response verification, revenue settlement and feedback optimization"—provides a commercially viable closed loop for vehicle-grid interaction. Attached Figure Description

[0017] Figure 1 This is a flowchart of the present invention; Figure 2 This is a network architecture diagram of the system of the present invention; Figure 3 This is a schematic diagram of the simulation system composition of the present invention; Figure 4 This describes the system load peak shaving and valley filling effect in a specific embodiment of the present invention; Figure 5 The charging and discharging power of the EV cluster and the dynamic buyback price are described in a specific embodiment of the present invention. Figure 6 This is the trajectory of the average state of charge (SOC) of the on-grid EV changing over time in a specific embodiment of the present invention; Figure 7 This is a specific embodiment of the invention showing the distribution of daily electricity costs for users under disordered charging and coordinated scheduling. Figure 8 This invention provides a comparison of node-based line power flow congestion prevention and collaborative scheduling in specific embodiments. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are only for explaining the invention and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without inventive effort are within the scope of protection of this invention.

[0019] The following is combined with Figure 1The overall flowchart shown below details the implementation process of the intelligent scheduling method for electric vehicle charging and discharging based on user behavior modeling and TD3 provided by this invention, including the following steps: This embodiment will focus on Figure 1 The process framework is explained in detail, with each step elaborated on in conjunction with real-world data scenarios.

[0020] Step one: Build a three-party collaborative framework among aggregators, charging station operators and car owners, and clarify the core responsibilities of each party and the path for aggregators to participate in the market.

[0021] Specifically, the tripartite collaborative framework includes: the vehicle owner as a flexible resource provider, responding to price signals and providing adjustable load; the charging station operator responsible for the investment, operation and maintenance of physical facilities and the execution of local strategies; and the aggregator as a bridge connecting dispersed users and the power grid, responsible for receiving grid dispatch instructions, setting internal incentive prices, and aggregating adjustable capacity on the user side.

[0022] The aggregator's participation in the market includes demand response, ancillary services, and vehicle-to-everything (V2X) interaction-related transaction mechanisms.

[0023] In addition, the hardware connection and communication links between the aggregator side, the charging station operator side, the power grid side, and the vehicle owner side, such as Figure 2 As shown; the power grid sends LMP and congestion warnings to the aggregator's cloud platform via dedicated line / VPN, the aggregator's cloud platform sends power commands to the charging station via 4G / 5G, the car owner sets travel needs and receives settlement bills through a mobile Internet App, and the in-vehicle T-BOX can optionally upload BMS data.

[0024] Step 2: Collect historical operational datasets of public charging scenarios and clean the data to obtain a clean operational dataset; based on the clean operational dataset, determine the price elasticity parameters of charging volume, charging duration, and the marginal effect of early termination of charging; based on the tripartite collaboration framework and the price elasticity parameters of charging volume, charging duration, and the marginal effect coefficient of early termination of charging, construct a Markov decision process by defining a quintuple.

[0025] Specifically, step two includes the following sub-steps: (2.1) Collect historical operation datasets of public charging scenarios and clean the data to obtain clean operation datasets; the historical operation datasets include data fields and several charging session records, the data fields include transaction time, electricity, cost, location, termination reason and abnormal marker, etc. Specifically, the process of collecting historical operation datasets for public charging scenarios is as follows: In this embodiment, real load operation data of all public charging stations in a certain region for 6 consecutive months are collected as historical operation datasets; based on the transaction time and electricity in the data fields, a daily load curve containing 96 points / 24 points is obtained.

[0026] Specifically, the process of cleaning the historical operation dataset includes cleaning the charging volume in the several charging session records. Charging time Charging sessions with missing fields or marked as 1 for abnormal interruption were removed, resulting in 359,433 valid sessions.

[0027] (2.2) Based on the clean operation dataset, a charging quantity response model and a charging duration response model are constructed using a double log-linear regression model to determine the charging quantity price elasticity parameter and the charging duration price elasticity parameter; based on the clean operation dataset, an early termination of charging behavior model is constructed using a binary choice model to determine the early termination of charging marginal effect coefficient. Specifically, based on the aforementioned cleaning operations dataset, a charging quantity response model and a charging duration response model are constructed using a double log-linear regression model; the expressions for the charging quantity response model and the charging duration response model are as follows: in, For price variables, Control variables include time period, week, and grouping characteristics. The amount of charge for the i-th charging session. For the intercept term, This is a parameter representing the price elasticity of charge volume. For control variables The corresponding coefficient vector, This is the random error term; Let i be the duration of the i-th charging session. For the intercept term, This is a parameter related to the price elasticity of charging time. For control variables The corresponding coefficient vector, This is the random error term.

[0028] The process of determining the price elasticity parameters of charging quantity and charging time based on the aforementioned charging quantity response model and charging duration response model includes: The pricing elasticity parameter for charging volume is determined using a combination of "group fixed effect + week fixed effect + robust standard error of group_id clustering". The group fixed effect introduces control variables based on group characteristics to control for the inherent heterogeneity between different charging station / user groups that does not change over time. The week fixed effect uses control variables based on week to control for macro-time trends such as seasons and holidays. The robust standard error of group_id clustering allows for correlation between samples within the same charging station / user, and the correction method of group_id clustering eliminates the influence of intra-group sample correlation on parameter estimation, ensuring the effectiveness of the coefficient significance test.

[0029] a. Determine the price elasticity parameter for charging volume: based on the unit price paid by the user. Substituting the price variable into the aforementioned charging response model, and based on the clean operations dataset, a regression estimate is performed on the charging response model employing "group fixed effects + week fixed effects + robust standard error clustering by group_id". Since the model uses a log-log form, the charging price elasticity parameter can be directly obtained. .

[0030] For example, when the charging quantity price elasticity parameter is -0.3748, its economic meaning is that, controlling for other factors to remain unchanged, when the user's unit payment price increases by 1%, the average charging quantity per session decreases by approximately 0.3748%. b. Determine the price elasticity parameter for charging time: based on the unit price paid by the user. Substituting the price variable into the aforementioned charging time response model, and based on the clean operations dataset, regression estimation is performed on the charging time response model employing "group fixed effects + week fixed effects + robust standard error clustering by group_id". Since the model uses a log-log form, the charging time price elasticity parameter in the charging time response model can be directly determined. .

[0031] For example, when the charging duration price elasticity parameter is -0.2956, its economic meaning is that, under the condition that other factors remain unchanged, for every 1% increase in the unit price paid by the user, the average charging duration of a single session is shortened by about 0.2956%.

[0032] Specifically, based on the cleaning operation dataset, a model for early termination of charging behavior is constructed using a binary selection model; the expression for the early termination of charging behavior model is as follows: in, The probability of premature termination of charging. Define an indicator variable to indicate the early termination of charging. This indicates that charging has ended early; otherwise, the value is 0. For the distribution function of the binary choice model, For the intercept term, The marginal effect coefficient for early termination of charging represents the direction and strength of the influence of the price variable on the probability of early termination of charging. For price variables, Control variables include time period, week, and grouping characteristics. For control variables The corresponding coefficient vector.

[0033] Based on the aforementioned early charging termination behavior model, the process of determining the marginal effect coefficient of early charging termination includes: The above-mentioned "group fixed effect + week fixed effect + robust standard error of clustering by group_id" is used to perform regression estimation on the early termination of charging behavior model.

[0034] c. Determine the marginal effect coefficient for early termination of charging: based on the unit price paid by the user. Substituting the price variable into the early charging behavior model, to obtain an interpretable single value, the partial derivative of the price with respect to the probability of early termination is calculated at each sample point. Then, the arithmetic mean is taken over all samples to obtain the marginal effect coefficient of early charging termination. The final estimated value.

[0035] For example, when the marginal effect coefficient of early termination of charging is +0.1921, the marginal effect coefficient of early termination of charging is positive, indicating that the increase in the unit price paid by the user will increase the probability of the user terminating the charging behavior early.

[0036] Specifically, in this embodiment, the intercept term and the random error term are both defined by the regression model standard.

[0037] (2.3) Based on the aforementioned tripartite collaborative framework and the price elasticity parameters of charging quantity, charging duration, and early termination of charging marginal effect coefficient, a Markov decision process is constructed by defining a quintuple and instantiating its elements into a programmable simulation environment interface; the quintuple includes a multidimensional state space S, a continuous action space A, a state transition probability P, a weighted composite reward function R, and a discount factor. The Markov decision process models the aggregator as a reinforcement learning agent and describes its interactive decision-making process with the environment model through a quintuple.

[0038] Specifically, defining the quintuple in step (2.3) includes the following sub-steps: (2.3.1) Define a multidimensional state space S, which reflects the environment at the current moment ( ) and vehicles ( The state characteristics of the aggregater are used to support the scheduling decisions of the aggregater, and the observable state vector at time t is output. ,in, This indicates the state of charge of the electric vehicle battery at time t, and also sets the physical safety boundary of the battery to prevent overcharging and over-discharging. ; and Encode time using trigonometric functions to eliminate abrupt changes in periodic boundaries. This represents the total number of time steps in a scheduling cycle. In this embodiment, we take 24 hours, which is a total of 96 time steps. The base electricity price at time t is determined according to the time-of-use pricing rules of the region. This represents the base load deviation rate of the power grid at time t. The time in the daily load curve The base load, This represents the daily average load. To improve the training efficiency of the neural network, all the above state variables are normalized.

[0039] (2.3.2) Define a continuous action space A, where the continuous action space A is the charging and discharging power command applied by the aggregator agent to the electric vehicle charging pile, and the observable state vector is used as the basis for the action space A. Input is normalized continuous motion, output is normalized continuous motion. ,in Indicates charging. This indicates a discharge.

[0040] (2.3.3) Define the state transition probability P, which is based on the current observable state vector. and the normalized continuous action As input, based on the price elasticity parameters of charging quantity, charging duration, and early termination of charging marginal effect coefficient, the price signal is converted into the user's response behavior, and the state transition probability is output. It quantifies the next state after the current state and the execution of the scheduling instruction. The evolutionary pattern.

[0041] (2.3.4) Define a weighted composite reward function R, which includes a weighted sum of grid-side rewards, user-side rewards, and constraint penalties. The weighted composite reward function at time t... The expression is as follows: In this embodiment, the reward weight .

[0042] User-side rewards The aim is to minimize user costs: in, For real-time electricity prices, The battery loss cost coefficient is set at 0.1 yuan / kWh in this embodiment, constituting the economic floor for users to participate in vehicle-to-grid interaction. For actual power, when When discharging electricity, users receive revenue from selling electricity, and the reward is positive. Grid-side rewards The aim is to smooth out peak and valley peaks. When the power grid is at peak And the vehicle discharges ( A reward will be given when this occurs; Restraint and Punishment To ensure users' travel needs are met: in, The penalty coefficient is... The battery state of charge as desired by the user when offline. This refers to the actual state of charge (SOC) of the off-grid battery. If the SOC does not reach the user's expected value at the time of off-grid connection ( If ), then a constraint penalty will be imposed.

[0043] (2.3.5) Define the discount factor The discount factor The weights used to balance current and future rewards are determined based on the time series length and decision-making process of the scheduling problem.

[0044] Step 3: Map the Markov decision process to a trainable neural network architecture; based on the neural network architecture, construct the TD3 model using the dual-delay deep deterministic policy gradient algorithm; and obtain the pre-trained Actor network model based on the TD3 model.

[0045] Step three includes the following sub-steps: (3.1) Map the multidimensional state space S and continuous action space A in the Markov decision process to the input-output structure of a trainable neural network, and obtain the neural network architecture; Specifically, in this embodiment, both the Actor network and the Critic network adopt a three-layer fully connected network structure, with 256 hidden layer nodes and ReLU activation function. The input to the Actor network is the predicted state vector. The output layer uses the Tanh function to normalize continuous actions. Constraints Interval output. The input to the Critic network is a state vector. and normalized continuous motion The output layer uses linear neurons to output continuous state-action value values. .

[0046] (3.2) Based on the neural network architecture, the TD3 model is constructed using the double-delay deep deterministic policy gradient algorithm; the TD3 model includes an Actor network, two Critic networks, a double Q network minimum truncation mechanism, a target policy smoothing mechanism, and a delayed update mechanism.

[0047] Specifically, based on the neural network architecture, the TD3 model is constructed using a dual-delay deep deterministic strategy gradient algorithm, comprising the following parts: a. Construct an Actor network and two Critic networks: This results in an Actor network containing the main network. and the corresponding target network The Actor main network is used for decision-making and updates, while the Actor target network is used to provide stable objectives; the two Critic networks contain the main network. and the corresponding target network , The Critic main network is used to estimate the Q-value of the current state-action, and the Critic target network is used to calculate the Q-value in the Bellman objective.

[0048] b. Constructing a dual-Q network minimum truncation mechanism: The dual-Q network minimum truncation mechanism refers to introducing a dual-Q network minimum truncation mechanism when calculating the Bellman objective value. This utilizes the smaller value from the outputs of two independent Critic objective networks to suppress the overestimation of Q-values ​​in the traditional DDPG algorithm. The formula for calculating the Bellman objective value is as follows: in, For Bellman's target value, The reward obtained at time t. As a discount factor, For the i-th Critic target network, Let be the state vector at time t+1. The target action has been smoothed.

[0049] c. Constructing a target policy smoothing mechanism: The target policy smoothing mechanism refers to smoothing the target action during Q-value calculation. Truncated normal noise is added to enhance the algorithm's robustness to value assessment; the expression for the target policy smoothing mechanism is: in, For the Actor target network, To conform to a mean of 0 and a standard deviation of The normal distribution is truncated at Random noise within the interval, Let be the truncation function. After adding truncated normal noise, the learning objective of the Critic target network becomes... This is equivalent to locally averaging the Q-function of the Critic target network in the action space, suppressing the "false spikes" that appear in narrow regions.

[0050] d. Constructing a delayed update mechanism: This delayed update mechanism reduces the update frequency of the Actor network to that of the Critic network, allowing the Critic network more time to converge value estimates under a fixed policy. This prevents unstable value estimates from affecting policy updates, thereby improving the accuracy of policy updates. In this embodiment, the Critic main network updates once for every two updates.

[0051] (3.3) Experience samples generated by the interaction between the aggregator agent and the environment Stored in the experience replay pool .

[0052] Specifically, the aggregator intelligent agent, based on the state vector Output continuous actions through the Actor network After the environment performs an action, it returns the state vector for the next time step. and rewards , to use empirical samples Store in the experience replay pool.

[0053] (3.4) Based on the above training samples, the TD3 model is iteratively updated until the convergence condition is met. Preferably, the training parameters in this embodiment are set as follows: 1500 training rounds, each round simulating 24 hours (96 time steps, 15 minutes per step); Actor network learning rate... Critic network learning rate The optimizer used in all cases is Adam; discount factor Soft update coefficient ;Target strategy smoothing noise parameters Cut-off range The delayed update frequency is 2 updates for the Critic network and 1 update for the Actor network. The vehicle battery capacity is set in the environment. Maximum charging and discharging power Charge and discharge efficiency The initial SOC follows a Gaussian distribution. Generate random user travel scenarios. User travel behavior: , , .

[0054] Specifically, before training, the experience replay pool is used. Randomly sample a small batch , and are used as training samples for the Critic main network, Actor main network and TD3 model target network in the TD3 model; Preferably, in this embodiment, the experience playback pool Capacity set to 10 5 The batch size of random sampling during each training session ; The random sampling method can break the temporal correlation between samples and improve sample utilization efficiency.

[0055] Specifically, the update process of the Critic network includes: based on the Bellman target value, defining the loss function of the two Critic main networks as the mean square Bellman error between the predicted Q value and the target Q value, wherein the loss function is expressed as: in, Let be the loss function of the i-th Critic main network. For small batches The number of samples in For Bellman's target value, For the i-th Critic main network at the input The output is the predicted state-action value.

[0056] Based on the loss functions of the two Critic main networks, the gradient descent method is used to update the parameters of the Critic main network. The gradient calculation formula and the update rule of the Critic main network parameters are as follows: in, loss function Regarding parameters gradient vector, The output of the i-th Critic master network For parameters gradient vector, is the learning rate of the Critic main network.

[0057] Specifically, the update process of the Actor main network includes: the optimization objective of the Actor main network is to maximize the expected cumulative reward of the current policy, and its policy gradient calculation formula and Actor main network parameter update rules are as follows: in, For the policy objective function Regarding Actor main network parameters gradient vector, In order to be in The output of the i-th Critic master network is calculated at point 1. The vector of partial derivatives with respect to action a, The output of the Actor main network is used to set parameters. The gradient matrix, is the learning rate of the Actor main network.

[0058] Specifically, the update process of the target network includes: the Actor network and the Critic network do not directly copy the parameters of the main network, but adopt a soft update method to slowly track the network parameters and improve training stability. The target network update rule is as follows: in, This is the soft update coefficient. These are the parameters of the i-th Critic target network; The parameters of the i-th Critic main network, For the parameters of the Actor target network, These are the parameters of the Actor's main network.

[0059] In the initial training phase, the aggregator agent adopts a random strategy, resulting in large fluctuations and low rewards. As the number of training rounds increases (>500 rounds), the dual-critic mechanism of the TD3 model begins to take effect, Q-value estimation becomes more accurate, and the Actor network gradually masters a composite strategy of "low-cost charging, high-cost discharging" and "cooperating with grid peak shaving." Around 1000 rounds later, the average reward curve stabilizes, indicating that the intelligence has converged to the optimal strategy. Finally, a pre-trained Actor network model is output. This network model can receive the power grid state vector in real time. It can directly output the optimal charging and discharging power within 50 milliseconds. It meets the real-time response window requirements of power grid dispatch instructions.

[0060] Step four: Design a load-linked dynamic price incentive mechanism based on real-time grid load status; map the real-time grid load deviation rate to the vehicle-grid interactive buyback price, so that the vehicle-grid interactive buyback price is positively correlated with the grid load deviation rate; at the same time, based on the pre-trained Actor network model, convert the actions output by the model into actual hardware charging and discharging power commands, and establish a feedback optimization mechanism for model iteration.

[0061] Step four includes the following sub-steps: (4.1) Based on the real-time aggregated load status of the power grid, the physical constraints of the power grid are transformed into economic signals that users can perceive through price levers. A load-linked dynamic price mechanism is designed, and a modified aggregated load curve is obtained.

[0062] Specifically, step (4.1) includes the following sub-steps: (4.1.1) Based on the data provided by the power grid side in the aforementioned tripartite collaborative framework, calculate the real-time aggregated power grid load; the calculation formula for the real-time aggregated power grid load is as follows: in, For time period The polymerization load, For time period The total number of charging sessions currently in a charging state. The instantaneous charging power of the i-th charging session during time period t.

[0063] (4.1.2) Based on the real-time aggregated load of the power grid, calculate the real-time power grid load deviation rate. The calculation formula is as follows: in, This represents the real-time grid load deviation rate. For real-time grid load aggregation, This represents the average load for the day.

[0064] (4.1.3) When At that time, the vehicle-to-grid repurchase price, which reflects the real-time congestion level of the power grid, is obtained: in, The electricity price will be repurchased for vehicle-to-everything (V2X) interaction. As a benchmark for time-of-use electricity pricing, The time period coefficients are as follows: 0.38 for the valley period (22:00-08:00), 1.4 for the peak period, and 1.7 for the apex period (specific peak times in summer and winter); To incentivize the step coefficient, when hour ; The adjustment factor is set to 1.5.

[0065] (4.1.4) Based on the aforementioned charging quantity price elasticity parameter And charging time price elasticity parameter The charging amount for a single charging session is adjusted based on the vehicle-to-grid (V2G) repurchase electricity price incentive. and correct charging time The corrected aggregate load curve for time period t under excitation conditions is obtained by correcting using the elastic formula; the elastic formula is: The modified polymerization load curve is as follows: in, For charging sessions at any time The instantaneous power, whose cumulative time is equal to the total charging amount of the session. And charging time It is the length of the time interval for its integration.

[0066] (4.2) Based on the pre-trained Actor network model, the actions output by the model are converted into actual hardware charging and discharging power instructions and fund allocation schemes, and the actual execution data is recorded.

[0067] Specifically, at each time t, the system collects the current state. Inputting these into a pre-trained Actor network model yields normalized actions. In normalized actions When translating this into actual physical commands, the physical boundary constraints of the power battery must be considered. Define the upper limit of the maximum available charging power at time t. and maximum discharge power limit : in, The rated power of the charging pile This is the battery's rated power. , These are charging efficiency and discharging efficiency, respectively. , These represent the maximum and minimum values ​​of the state of charge (SOC) of the electric vehicle battery, respectively. This is the scheduling time step.

[0068] Based on the maximum charging power limit and maximum discharge power limit The final actual power command is obtained. The mapping function is: This instruction is sent to the corresponding charging pile via the 4G / 5G communication network after the end of each settlement period (for details on the communication link, see...). Figure 2 In the L2 and L3 models, the aggregator distributes revenue to participating users according to the agreed-upon revenue sharing ratio and the agreed-upon electricity price calculation rules, and records the execution data.

[0069] (4.3) Based on the execution data, establish a feedback optimization mechanism to iteratively update the incentive parameters and model parameters.

[0070] Specifically, the system iteratively adjusts the incentive parameters weekly based on the execution results in the execution data: comparing the corrected aggregate load curve with the actual load curve and calculating the deviation rate; if the deviation exceeds a threshold, the parameters of the TD3 model are fine-tuned using the interaction data from the most recent week, and the step coefficient in the dynamic price formula is adjusted. Regulatory factors The parameters and model will be updated to adapt to changes in grid load characteristics and user behavior. The updated parameters and model will be used for scheduling in the next cycle.

[0071] Step five involves dividing steps one through four into three stages based on time and business logic: daily reporting, intraday scheduling, and post-event settlement. This results in a closed-loop execution process, clearly defining the boundaries of responsibilities and settlement rules for each stage, thereby enabling intelligent scheduling of electric vehicle charging and discharging.

[0072] The closed-loop execution process is a transferable, executable, and settlement-enabled closed-loop execution process from user signing, policy issuance, dynamic scheduling, process measurement, response verification to revenue settlement and feedback optimization.

[0073] Specifically, the three stages in step five are as follows: During the recent application phase, the power company issued congestion warnings and time-of-use pricing information for each node the following day. Users could then set their expected off-grid time for the next day via their terminals. Target power In addition to the response floor price for participating in dispatch, the aggregator summarizes the adjustable capacity of users in the area according to the division of responsibilities in the tripartite framework, and submits the aggregated response capacity curve to the electricity market; during the intraday dispatch phase, the aggregator receives the load deviation rate in real time. The dynamic pricing mechanism from step four is invoked to generate the vehicle-to-grid (V2G) electricity buyback price. And push to the user, while simultaneously sending the current state vector Input the pre-trained Actor network model from step three, output normalized actions, and convert them into actual power commands according to the power mapping rules from step four. The data is transmitted to the charging piles via the communication network for execution. In the post-event settlement phase, the power company retrieves the metering data from the charging piles to verify the actual response electricity volume and settles the relevant adjustment fees with the aggregator. The aggregator, through a "price difference sharing mechanism," covers charging costs and battery losses with its electricity sales revenue and then distributes the net profit to the participating vehicle owners using the settlement system. Simultaneously, the execution results are fed back to the feedback optimization mechanism in step four to update price incentive parameters and model parameters, forming a sustainable iterative business closed loop.

[0074] This embodiment is based on the aforementioned real load data and randomly generated EV travel scenarios. The simulation system model used is as follows: Figure 3 As shown. Figure 3 The core components and their data / instruction flows at four levels—grid side, aggregator side, charging station side, and user side—are illustrated in a block diagram. Each module in the diagram directly corresponds to the technical solutions described in steps one through five of this invention. Finally, a comparison is made between unordered charging (Baseline) and a collaborative scheduling strategy based on the TD3 algorithm, with the following results.

[0075] like Figure 4 As shown, this illustrates the comparison of the total system load curves before and after coordinated scheduling. In the disordered charging mode (black dashed line), user charging behavior exhibits a clear "synchronicity," meaning that a large number of users immediately engage in high-power charging after returning home during the evening peak (18:00-22:00). This behavior highly overlaps with the grid's original baseline load peak, causing a sharp increase in the system's peak load, reaching approximately 1.05 × 10⁻⁶ during the 22:00-23:00 period. 6 kW, which can easily lead to transformer overload risk. After introducing the TD3-based coordinated dispatch of this invention (solid red line), the system load curve has undergone a significant shape reshaping: during the period of 22:00-23:00 when the original load was highest, the red curve did not surge along with the black dashed line, but was suppressed to 0.75×10 6 With a peak load reduction of approximately 30%, the capacity pressure on the distribution network was effectively alleviated. Meanwhile, during the off-peak hours in the early morning (02:00-05:00), the red curve was significantly higher than the black dashed line, indicating that the EV cluster was utilizing off-peak electricity prices for concentrated energy replenishment, achieving a "valley filling" effect. Calculations of the load variance revealed a significant decrease in load curve volatility after coordinated scheduling, demonstrating the physical potential of large-scale EV clusters as flexible loads in mitigating grid fluctuations.

[0076] like Figure 5As shown, the graph illustrates the strong coupling between the charging and discharging power of the EV cluster (left axis, bar chart) and the dynamic buyback price (right axis, red dotted line). During the period from 02:00 to 06:00, the electricity price is at its lowest point of the day (approximately 0.3 yuan / kWh), at which time the green bar chart (charging power) reaches its peak. The cluster absorbs electricity from the grid at near full power, indicating that the agent has learned to "stockpile" energy when costs are lowest. The most critical response occurs between 22:00 and 23:00. With a surge in grid load, the dynamic pricing mechanism triggers a high incentive, causing the electricity price to soar to over 1.2 yuan / kWh. At this time, the blue bar chart (discharging power) increases significantly, and the EV cluster instantly transforms into a "virtual power plant," feeding electricity back into the grid. It is worth noting that during the 11:00-12:00 peak period, although there was some discharge behavior, the agent chose a more conservative discharge strategy because the electricity price incentive (about 0.8 yuan / kWh) was not as significant as that during the evening peak. This strategy preserved more electricity to cope with the higher price peaks later. This cross-time period global planning capability is the core manifestation of reinforcement learning's superiority over traditional greedy algorithms.

[0077] like Figure 6 As shown, the trajectory of the average SOC of the EV on the network changes over time. From 01:00 to 05:00, the SOC curve shows a linear upward trend, rapidly recovering from approximately 10% to 95% (near the green line of the Max SOC). This is consistent with... Figure 4 Corresponding to the concentrated charging behavior, this ensures that the vehicle remains in a "fully charged standby" state during the daytime despite prolonged parking. During the daytime (06:00-18:00), the State of Charge (SOC) remains at a high level, fluctuating only with travel consumption or minor two-way interaction responses, never reaching the minimum SOC threshold. This indicates that the intelligent agent (aggregator) has a strong sense of security, prioritizing sufficient battery power. However, from 20:00-23:00, the SOC curve experiences a sharp drop, from over 90% to around 40%, corresponding to the concentrated discharge during the evening rush hour. Despite the significant discharge depth, the final SOC is still significantly higher than the 10% threshold and exceeds the daily commuting needs of most users (typically 20%-30%), demonstrating that the strategy, while tapping into the potential of two-way interaction regulation, ensures the minimum battery safety and meets the user's travel needs.

[0078] like Figure 7As shown, the histogram visually illustrates the frequency distribution of daily electricity costs for users under two modes: disordered charging and coordinated scheduling. In the disordered charging mode (gray bars), the cost distribution approximates a normal distribution, centered in the 20-30 yuan range. The vast majority of users incur higher charging costs because their charging activity primarily occurs during periods of high electricity prices. After adopting TD3-based coordinated scheduling (yellow bars), the cost distribution shifts significantly to the left, with the central range dropping to 0-10 yuan, directly quantifying the economic benefits brought to users by vehicle-to-grid interaction technology. Most notably, the yellow bars show a significant distribution in the negative range on the horizontal axis (-20 yuan to 0 yuan). This means that some users, by participating in vehicle-to-grid interactive scheduling, achieve a discharge revenue (electricity sales revenue) exceeding the charging cost. In other words, by actively responding to reverse discharge commands during periods of high prices, they achieve electricity sales revenue covering charging costs and battery losses. This result strongly supports the business model proposed in this study: covering battery depreciation costs and incentivizing user participation through huge arbitrage opportunities, thereby transforming the user from a "consumer" to a "producer and seller," is the core economic driving force that enables the vehicle-to-everything (V2X) interactive project to move from theory to large-scale implementation.

[0079] like Figure 8 The diagram illustrates the evolution of power flow at different nodes (industrial parks and residential areas) under safe boundaries. In unordered charging scenarios, the power flow in residential areas during the evening rush hour far exceeds the maximum transmission capacity of the lines (red congestion zone), posing a very high risk of distribution network collapse. The TD3 strategy, relying on a physical safety shield, forcibly truncates the power flow curve and smoothly extends it close to the capacity limit, effectively mitigating the severe line over-limit crisis. Simultaneously, the limited EV charging demand during the evening rush hour is not simply abandoned, but is intelligently scheduled to be released during the grid's idle period in the early morning. This mechanism smoothly fills a "trapezoidal plateau" in the load trough, maximizing the utilization of the remaining capacity margin of the distribution network. It should be noted that when approaching the capacity limit, the power flow curve after coordinated scheduling is limited by the battery BMS ramp rate, resulting in a short-term, minor over-limit phenomenon (i.e., "physical hysteresis"). This is not an algorithm malfunction, but rather an optimal engineering game that balances the grid safety boundary and the physical characteristics of the power battery. Actual industrial standards allow core power equipment such as transformers to operate under short-term overload at 105% to 110% load rates. This slight deviation is well within the allowable thermal stability margin. If a zero-tolerance absolute red line is enforced in dispatching, it would necessitate exceeding the physical ramp-up limit of the battery and causing a sudden and drastic power cut-off. This would not only lead to a sudden change in cell temperature and accelerated decay of internal chemical activity, but also shift the cost reduction of the grid to users through high battery depreciation costs, violating the original intention of mutually beneficial vehicle-grid interaction. Therefore, the allowable slight deviations in this invention actually preserve necessary engineering flexibility for commercial implementation.

[0080] This embodiment comprehensively verifies the effectiveness of the proposed strategy in resolving supply and demand imbalances and achieving a business closed loop through quantitative analysis of four dimensions: grid load characteristics, cluster response behavior, battery state evolution, and user economic costs. Furthermore, this invention is not limited to the above-described embodiments. Various changes and modifications can be made without departing from the scope of this invention, and all such changes and modifications fall within the scope of protection claimed by this invention.

Claims

1. A method for intelligent scheduling of electric vehicle charging and discharging based on user behavior modeling and TD3, characterized in that, The method includes the following steps: Step 1: Build a three-party collaborative framework among aggregators, charging station operators, and car owners, and clarify the core responsibilities of each party and the path for aggregators to participate in the market; Step 2: Collect historical operation datasets of public charging scenarios and clean the data to obtain a clean operation dataset; based on the clean operation dataset, determine the price elasticity parameter of charging volume, the price elasticity parameter of charging duration, and the marginal effect of ending charging early. Based on the aforementioned tripartite collaborative framework, as well as the price elasticity parameters of charging quantity, charging duration, and early termination of charging marginal effect coefficient, a Markov decision process is constructed by defining a quintuple. Step 3: Map the Markov decision process to a trainable neural network architecture; based on this neural network architecture, construct the TD3 model using the dual-delay deep deterministic policy gradient algorithm; and obtain the pre-trained Actor network model based on the TD3 model. Step 4: Design a dynamic price incentive mechanism based on real-time grid load status, generate charging and discharging instructions using a pre-trained Actor network model, and establish a feedback optimization mechanism to iteratively update the model and parameters. Step five involves dividing steps one through four into three stages based on time and business logic: daily reporting, intraday scheduling, and post-event settlement. This results in a closed-loop execution process, clearly defining the boundaries of responsibilities and settlement rules for each stage, thereby enabling intelligent scheduling of electric vehicle charging and discharging.

2. The intelligent scheduling method for electric vehicle charging and discharging based on user behavior modeling and TD3 as described in claim 1, characterized in that, The tripartite collaborative framework described in step one includes: the vehicle owner, as a flexible resource provider, responds to price signals and provides adjustable load; the charging station operator is responsible for the investment, operation and maintenance of physical facilities and the execution of local strategies; and the aggregator, as a bridge connecting dispersed users and the power grid, is responsible for receiving grid dispatch instructions, setting internal incentive prices, and aggregating adjustable capacity on the user side.

3. The intelligent scheduling method for electric vehicle charging and discharging based on user behavior modeling and TD3 as described in claim 1, characterized in that, Step two includes the following sub-steps: (2.1) Collect historical operation datasets of public charging scenarios and clean the data to obtain clean operation datasets; the historical operation datasets include data fields and several charging session records, the data fields include transaction time, electricity, cost, location, termination reason and abnormal marker; (2.2) Based on the clean operation dataset, a charging quantity response model and a charging duration response model are constructed using a double log-linear regression model to determine the charging quantity price elasticity parameter and the charging duration price elasticity parameter; based on the clean operation dataset, an early termination of charging behavior model is constructed using a binary choice model to determine the early termination of charging marginal effect coefficient. (2.3) Based on the aforementioned tripartite collaborative framework and the price elasticity parameters of charging quantity, charging duration, and early termination of charging marginal effect coefficient, a Markov decision process is constructed by defining a quintuple and instantiating its elements into a programmable simulation environment interface; the quintuple includes a multidimensional state space S, a continuous action space A, a state transition probability P, a weighted composite reward function R, and a discount factor. The Markov decision process models the aggregator as a reinforcement learning agent and describes its interactive decision-making process with the environment model through a quintuple.

4. The method for intelligent scheduling of electric vehicle charging and discharging based on user behavior modeling and TD3 according to claim 3, characterized in that, The expressions for the charging quantity response model and the charging duration response model mentioned in step (2.2) are as follows: ; ; in, For price variables, Control variables include time period, week, and grouping characteristics. The amount of charge for the i-th charging session. For the intercept term, This is a parameter representing the price elasticity of charge volume. For control variables The corresponding coefficient vector, This is the random error term; Let i be the duration of the i-th charging session. For the intercept term, This is a parameter related to the price elasticity of charging time. For control variables The corresponding coefficient vector, This is the random error term; The expression for the early termination of charging behavior model is as follows: ; in, The probability of premature termination of charging. Define an indicator variable to indicate the early termination of charging behavior. This indicates that charging has ended early; otherwise, the value is 0. For the distribution function of the binary choice model, For the intercept term, The marginal effect coefficient for early termination of charging represents the direction and strength of the influence of the price variable on the probability of early termination of charging. For price variables, Control variables include time period, week, and grouping characteristics. For control variables The corresponding coefficient vector.

5. The method for intelligent scheduling of electric vehicle charging and discharging based on user behavior modeling and TD3 according to claim 3, characterized in that, The definition of the quintuple in step (2.3) specifically includes the following sub-steps: (2.3.1) Define a multidimensional state space S, which reflects the state characteristics of the environment and vehicles at the current moment. To support the scheduling decision of the aggregator, output the observable state vector at time t: ; in, Let be the observable state vector at time t. This indicates the state of charge of the electric vehicle battery at time t, and also sets the physical safety boundary of the battery to prevent overcharging and over-discharging. ; and Encode time using trigonometric functions to eliminate abrupt changes in periodic boundaries. The total number of time steps in a scheduling cycle; The base electricity price at time t is determined according to the time-of-use pricing rules of the region. This represents the base load deviation rate of the power grid at time t. Based on the load, This represents the daily average load. (2.3.2) Define a continuous action space A, where the continuous action space A is the charging and discharging power command applied by the aggregator agent to the electric vehicle charging pile, and the observable state vector. Input is normalized continuous motion, output is normalized continuous motion. ,in Indicates charging. Indicates discharge; (2.3.3) Define the state transition probability P, which is based on the current observable state vector. and the normalized continuous action As input, based on the price elasticity parameters of charging quantity, charging duration, and early termination of charging marginal effect coefficient, the price signal is converted into the user's response behavior, and the state transition probability is output. It quantifies the next state after the current state and the execution of the scheduling instruction. The evolutionary pattern; (2.3.4) Define a weighted composite reward function R, which includes a weighted sum of grid-side rewards, user-side rewards, and constraint penalties. The weighted composite reward function at time t... The expression is as follows: ; As a reward weight; User-side rewards The aim is to minimize user costs: ; in, For real-time electricity prices, This is the battery loss cost coefficient. For actual power, when At that time, the user receives revenue from electricity sales, and the reward is positive; Grid-side rewards The aim is to smooth out peak and valley peaks. ; A reward will be given when the power grid is at its peak and vehicles are discharging electricity; Restraint and Punishment To ensure users' travel needs are met: ; in, The penalty coefficient is... The battery state of charge as desired by the user when offline. The actual state of charge (SOC) of the off-grid battery is used as the basis for determining the actual SOC. If the SOC does not meet the user's expected value at the time of off-grid connection, a constraint penalty will be imposed. (2.3.5) Define the discount factor The discount factor The weights used to balance current and future rewards are determined based on the time series length and decision-making process of the scheduling problem.

6. The intelligent scheduling method for electric vehicle charging and discharging based on user behavior modeling and TD3 as described in claim 1, characterized in that, Step three includes the following sub-steps: (3.1) Map the multidimensional state space S and continuous action space A in the Markov decision process to the input-output structure of a trainable neural network, and obtain the neural network architecture; (3.2) Based on the neural network architecture, a TD3 model is constructed using the double-delay deep deterministic policy gradient algorithm; the TD3 model includes an Actor network, two Critic networks, a double Q network minimum truncation mechanism, a target policy smoothing mechanism, and a delayed update mechanism; (3.3) Experience samples generated by the interaction between the aggregator agent and the environment Stored in the experience replay pool ; (3.4) Based on the experience replay pool, obtain training samples, iteratively update the TD3 model until the convergence condition is met, and extract the pre-trained Actor network model.

7. The method for intelligent scheduling of electric vehicle charging and discharging based on user behavior modeling and TD3 according to claim 6, characterized in that, In step (3.2), based on the neural network architecture, the TD3 model is constructed using a dual-delay deep deterministic strategy gradient algorithm, including the following components: a. Construct an Actor network and two Critic networks: This results in an Actor network containing the main network. and the corresponding target network The Actor main network is used for decision-making and updates, while the Actor target network is used to provide stable objectives; the two Critic networks contain the main network. and the corresponding target network , The Critic main network is used to estimate the Q-value of the current state-action, and the Critic target network is used to calculate the Q-value in the Bellman objective. b. Constructing a dual-Q network minimum truncation mechanism: The dual-Q network minimum truncation mechanism refers to using the smaller value in the output of two independent Critic target networks to suppress the overestimation of Q value in the DDPG algorithm when calculating the Bellman objective value; c. Constructing a target policy smoothing mechanism: The target policy smoothing mechanism refers to adding truncated normal noise to the smoothed target action when calculating the Q value, so that the learning target of the Critic target network becomes... This is equivalent to performing a local averaging of the Q-function of the Critic target network in the action space; d. Constructing a delayed update mechanism: The delayed update mechanism refers to making the update frequency of the Actor network lower than that of the Critic network, so that the Critic network has more time to converge value estimation under a fixed policy, thereby improving the accuracy of policy updates.

8. The intelligent scheduling method for electric vehicle charging and discharging based on user behavior modeling and TD3 as described in claim 1, characterized in that, Step four includes the following sub-steps: (4.1) Based on the real-time load status, the physical constraints of the power grid are transformed into economic signals that users can perceive through price levers. A load-linked dynamic pricing mechanism is designed, and a modified aggregated load curve is obtained. (4.2) Based on the pre-trained Actor network model, the actions output by the model are converted into actual hardware charging and discharging power instructions and fund allocation schemes, and the execution data is recorded; (4.3) Based on the execution data, establish a feedback optimization mechanism to iteratively update the incentive parameters and model parameters.

9. The method for intelligent scheduling of electric vehicle charging and discharging based on user behavior modeling and TD3 according to claim 8, characterized in that, Step (4.1) includes the following sub-steps: (4.1.1) Based on the data provided by the power grid side in the aforementioned tripartite collaborative framework, calculate the real-time aggregated power grid load; the calculation formula for the real-time aggregated power grid load is as follows: ; in, The aggregate load for time period t, For time period The total number of charging sessions currently in a charging state. The instantaneous charging power of the i-th charging session during time period t; (4.1.2) Based on the real-time aggregated load of the power grid, calculate the real-time power grid load deviation rate. The calculation formula is as follows: ; in, This represents the real-time grid load deviation rate. For real-time grid load aggregation, This represents the average load for the day. (4.1.3) When At that time, the vehicle-to-grid repurchase price, which reflects the real-time congestion level of the power grid, is obtained: ; in, The electricity price will be repurchased for vehicle-to-everything (V2X) interaction. As a benchmark for time-of-use electricity pricing, For time period coefficients; This is the incentive step coefficient; (4.1.4) Based on the aforementioned charging quantity price elasticity parameter And charging time price elasticity parameter The charging amount for a single charging session is adjusted based on the vehicle-to-grid (V2G) repurchase electricity price incentive. And correct charging time By correcting the elasticity formula, the corrected aggregate load curve for time period t under excitation conditions is obtained.

10. The method for intelligent scheduling of electric vehicle charging and discharging based on user behavior modeling and TD3 according to claim 1, characterized in that, The three stages in step five are as follows: the day-ahead declaration stage is to declare adjustable capacity based on the user-set off-grid time, target power, and response base price; the day-ahead scheduling stage is to generate dynamic electricity price and power instructions based on real-time load deviation rate and pre-trained Actor network model and execute them; the post-event settlement stage is to settle fees and distribute revenue based on the actual response power measured on the charging pile side, and to feed back the execution effect to the feedback optimization mechanism in step four.