Electric vehicle charging and discharging scheduling method and system based on multi-agent mixed game, electronic device and medium
Patent Information
- Application Number
- CN202610526549.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-21
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2046-04-21
AI Technical Summary
[0004]本发明针对上述不足或缺点,提供了一种基于多主体混合博弈的电动汽车充放电调度方法、系统、电子设备及介质,解决了配电网运营商、电动汽车聚合商与电动汽车用户之间难以实现实时协同优化调度的技术问题,提升了规模化电动汽车参与电网互动的整体效能与调度可靠性
[0015] The present invention provides a method for electric vehicle charging and discharging scheduling based on multi-agent hybrid game theory. The method first acquires charging and discharging behavior data of multiple target electric vehicles; based on this data, it simultaneously constructs charging and discharging strategy optimization models for electric vehicle users, strategy optimization models for multiple electric vehicle aggregators, and power allocation optimization models for distribution network operators; subsequently, it inputs these models into a multi-agent non-cooperative and master-slave hybrid game framework. This framework is a three-layer dynamic interactive architecture constructed by hierarchically coupling the non-cooperative game relationship between distribution network operators and electric vehicle aggregators, and the master-slave game relationship between electric vehicle aggregators and electric vehicle users; finally, it executes the charging and discharging scheduling of electric vehicles according to the charging and discharging scheduling strategy generated by this framework.
Smart Images

Figure CN122203368B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to electric vehicle charging and discharging technology, and in particular to an electric vehicle charging and discharging scheduling method, system, electronic device, and medium based on multi-agent hybrid game theory. Background Technology
[0002] In recent years, with the rapid development of the electric vehicle industry, the number of electric vehicles in my country has continued to rise. While large-scale grid connection for charging has alleviated the energy crisis, it has also led to a significant increase in peak-valley load differences, raising the risk to grid operation. However, electric vehicles possess mobile energy storage characteristics. By rationally scheduling their charging and discharging loads, it is possible not only to effectively smooth the grid load curve and reduce peak-valley differences, but also to participate in diversified markets such as peak shaving and frequency regulation, and electricity trading, achieving a win-win situation of minimizing user charging costs and ensuring stable grid operation.
[0003] While research on electric vehicle (EV) charging and discharging scheduling has made some progress, it still faces several limitations. First, regarding incentive mechanism design, existing research lacks systematic modeling of EV aggregator incentive strategies and fails to fully consider the coupling relationship between electricity price compensation mechanisms and battery life degradation, leading to insufficient user participation. Second, in terms of game theory framework construction, most studies are limited to simple models combining a single game type with fixed contracts, or only employ basic two-layer game structures, making it difficult to characterize the dynamic interactions and interest games among distribution network operators, EV aggregators, and EV users. For example, existing technology (application publication number CN114662759A) proposes an EV scheduling method based on multi-agent game theory, but it focuses on static contracts and hierarchical optimization, failing to address the coupling problem of multi-dimensional constraints in real-time collaborative scheduling, and cannot adapt to the real-time and uncertain nature of multi-agent strategy interactions in the electricity market environment. In summary, existing technologies struggle to achieve real-time collaborative optimization scheduling among distribution network operators, EV aggregators, and EV users under complex market mechanisms, limiting the overall effectiveness of large-scale EV participation in grid interaction. Summary of the Invention
[0004] To address the aforementioned shortcomings or deficiencies, this invention provides a method, system, electronic device, and medium for electric vehicle charging and discharging scheduling based on multi-agent hybrid game theory. This solves the technical problem of difficulty in achieving real-time collaborative optimization scheduling among power grid operators, electric vehicle aggregators, and electric vehicle users, and improves the overall efficiency and scheduling reliability of large-scale electric vehicle participation in power grid interaction.
[0005] This invention provides a method for scheduling the charging and discharging of electric vehicles based on multi-agent hybrid game theory, comprising: Acquire charging and discharging behavior data of multiple target electric vehicles.
[0006] Based on charging and discharging behavior data, we construct a charging and discharging strategy optimization model for electric vehicle users, a strategy optimization model for multiple electric vehicle aggregators, and a power allocation optimization model for power distribution network operators.
[0007] The charging and discharging strategy optimization model, strategy optimization model, and power allocation optimization model are input into a multi-agent non-cooperative and master-slave hybrid game framework to generate a charging and discharging scheduling strategy. The multi-agent non-cooperative and master-slave hybrid game framework is a three-layer dynamic interaction architecture model constructed through a layered coupling method based on the non-cooperative game relationship parameters between the distribution network operator and each electric vehicle aggregator, and the master-slave game relationship parameters between each electric vehicle aggregator and electric vehicle users.
[0008] The charging and discharging scheduling of electric vehicles is executed according to the charging and discharging scheduling strategy.
[0009] According to a second aspect, this invention provides an electric vehicle charging and discharging scheduling system based on multi-agent hybrid game theory, comprising: The charging and discharging behavior data acquisition module is used to acquire charging and discharging behavior data of multiple target electric vehicles.
[0010] The multi-agent optimization model construction module is used to build optimization models for charging and discharging strategies of electric vehicle users, strategy optimization models for multiple electric vehicle aggregators, and power allocation optimization models for distribution network operators based on charging and discharging behavior data.
[0011] The charging and discharging scheduling strategy generation module is used to input the charging and discharging strategy optimization model, strategy optimization model, and power allocation optimization model into a multi-agent non-cooperative and master-slave hybrid game framework to generate a charging and discharging scheduling strategy. The multi-agent non-cooperative and master-slave hybrid game framework is a three-layer dynamic interactive architecture model constructed through a layered coupling method based on the non-cooperative game relationship parameters between the distribution network operator and each electric vehicle aggregator, and the master-slave game relationship parameters between each electric vehicle aggregator and electric vehicle users.
[0012] The charge / discharge scheduling strategy execution module is used to execute the charge / discharge scheduling of electric vehicles according to the charge / discharge scheduling strategy.
[0013] According to a third aspect, the present invention provides an electronic device comprising: At least one processor; and The memory that is communicatively connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor, which are executed by the at least one processor to enable the at least one processor to execute any of the electric vehicle charging and discharging scheduling methods based on multi-agent hybrid game in the embodiments of the present invention.
[0014] According to another aspect of the present invention, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute any of the electric vehicle charging and discharging scheduling methods based on multi-agent hybrid game in the embodiments of the present invention.
[0015] The present invention provides a method for electric vehicle charging and discharging scheduling based on multi-agent hybrid game theory. The method first acquires charging and discharging behavior data of multiple target electric vehicles; based on this data, it simultaneously constructs charging and discharging strategy optimization models for electric vehicle users, strategy optimization models for multiple electric vehicle aggregators, and power allocation optimization models for distribution network operators; subsequently, it inputs these models into a multi-agent non-cooperative and master-slave hybrid game framework. This framework is a three-layer dynamic interactive architecture constructed by hierarchically coupling the non-cooperative game relationship between distribution network operators and electric vehicle aggregators, and the master-slave game relationship between electric vehicle aggregators and electric vehicle users; finally, it executes the charging and discharging scheduling of electric vehicles according to the charging and discharging scheduling strategy generated by this framework.
[0016] In this technical solution, the present invention addresses the problems mentioned in the background art, namely, the lack of systematic modeling of incentive mechanisms for electric vehicle aggregators and the fact that "existing game theory-based research often considers single game types or simple two-layer games, failing to achieve real-time collaboration among the three parties." By constructing a hybrid game framework that includes optimization models for three main entities—electric vehicle users, electric vehicle aggregators, and distribution network operators—this invention achieves collaborative optimization of the interaction and balance of interests among the three parties under a unified architecture by layering non-cooperative game theory with master-slave game theory. Therefore, the technical solution of this invention effectively overcomes the shortcomings of existing technologies, such as a single game structure and insufficient model coupling, and solves the technical problem of difficulty in achieving real-time collaborative optimization scheduling among distribution network operators, electric vehicle aggregators, and electric vehicle users, thereby improving the overall efficiency and scheduling reliability of large-scale electric vehicle participation in grid interaction. Attached Figure Description
[0017] Figure 1 This is a flowchart of an embodiment of the electric vehicle charging and discharging scheduling method based on multi-agent hybrid game theory according to the present invention; Figure 2 This is a flowchart of a hierarchical iterative optimization algorithm for solving a multi-agent non-cooperative and master-slave hybrid game framework according to an embodiment of the present invention. Figure 3 This is a structural block diagram of an electric vehicle charging and discharging scheduling system based on multi-agent hybrid game theory according to an embodiment of the present invention; Figure 4 This is a block diagram of an electronic device used to implement embodiments of the present invention. Detailed Implementation
[0018] The following description, in conjunction with the accompanying drawings, illustrates exemplary embodiments of the present invention, including various details to aid understanding. These details should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the invention. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0019] This invention provides a method for electric vehicle charging and discharging scheduling based on multi-agent hybrid game theory, according to a first aspect. This method can be applied to an electric vehicle collaborative scheduling system (hereinafter referred to as the "system"). The distributed computing devices in this system should be able to run multi-agent game theory algorithms through local deployment or containerization, completing real-time processing tasks such as charging and discharging behavior data acquisition, optimization model construction, and hybrid game equilibrium solution, providing a decision support basis for subsequent dynamic scheduling instruction generation. Specifically, the distributed computing devices include, but are not limited to, edge computing gateways, cloud server clusters, and embedded controllers for smart charging piles. The edge computing gateways are responsible for collecting electric vehicle charging and discharging behavior data, the cloud server clusters undertake the collaborative computation of the multi-agent game model, and the embedded controllers for smart charging piles execute the final generated charging and discharging scheduling strategy.
[0020] like Figure 1 As shown, the method may include: Step S110: Obtain charging and discharging behavior data of multiple target electric vehicles.
[0021] Charging and discharging behavior data refers to a set of characteristic parameters collected from electric vehicles and their charging facilities that reflect the user's charging and discharging patterns. This includes the charging start time, charging end time, initial state of charge (SOC), and related probability distribution parameters (such as mean and variance). Charging and discharging behavior data is used to quantify the schedulable characteristics and behavioral uncertainties of electric vehicles.
[0022] Specifically, the system collects raw data through the charging pile sensor network and on-board communication modules (such as 4G, 5G, or CAN bus), and then performs data cleaning and calibration to ensure data quality. CAN stands for Controller Area Network.
[0023] For example, in a test cluster of 100 electric vehicles, the system collects data every 15 minutes: charging starts may be concentrated between 6 PM and 10 PM (following a normal distribution with a mean of 20 hours and a standard deviation of 1.5 hours), and charging ends may be concentrated between 6 AM and 8 AM the following day (following a normal distribution with a mean of 7 hours and a standard deviation of 0.8 hours). The initial State of Charge (SOC) fluctuates between 20% and 50% (following a uniform distribution). Specifically, data preprocessing includes normalizing the time series (e.g., mapping time to values from 0 to 24 hours) and standardizing the SOC percentage to provide a unified input format for subsequent modeling.
[0024] In some embodiments, the system can calculate the probability density value at the start of charging using the following formula (1): (1) Formula (1) represents the probability density function at the start of charging and is used to describe the statistical distribution characteristics of the time point when electric vehicle users start charging. The variable representing the start time of charging is in hours, and its value ranges from 0 to 24 hours. The mean parameter characterizing the start time of charging, in hours, is used to indicate the concentrated time trend of user charging behavior; The standard deviation parameter characterizing the start time of charging, in hours, is used to quantify the dispersion of charging time; probability density value. Characterized at a specific start time The probability density at a given point is used for weight calculation in subsequent Monte Carlo sampling.
[0025] Next, the system can also calculate the probability density value at the end of charging using the formula (2) shown below: (2) Formula (2) represents the probability density function at the end of charging and is used to describe the statistical distribution characteristics of the time point when electric vehicle users finish charging. The variable representing the time when charging ends, in hours, with a value range of 0 to 24 hours; The mean parameter representing the time when charging ends, in hours, is used to indicate the central tendency of the user's charging completion time. The standard deviation parameter characterizing the end time of charging, in hours, is used to quantify the dispersion of the end time; probability density value. Characterized at a specific end time The probability density at a given point is used for weight calculation in Monte Carlo sampling. Through the above formulas (1) to (2), the system can quantify the probability of occurrence at different charging end times, and together with the probability model at the charging start time, form a complete temporal distribution representation of charging behavior, providing an accurate probabilistic basis for the optimization of charging and discharging strategies.
[0026] Step S120: Based on charging and discharging behavior data, construct a charging and discharging strategy optimization model for electric vehicle users, a strategy optimization model for multiple electric vehicle aggregators, and a power distribution optimization model for power distribution network operators.
[0027] The charging and discharging strategy optimization model can be a mathematical programming model aimed at minimizing the total charging cost for users, which includes charging and discharging fees, dispatch subsidies, and battery wear costs. The strategy optimization model can be a decision-making model aimed at maximizing the revenue of electric vehicle aggregators from electricity purchase and sale, with the revenue function encompassing electricity trading revenue and dispatch deviation penalties. The power allocation optimization model can be a resource allocation model aimed at maximizing the revenue of distribution network operators, involving electricity trading benefits and efficiency loss compensation. Specifically, the system can train a probability distribution model (such as a normal distribution describing charging and discharging time) based on historical data, generate 1000 sets of behavioral samples through Monte Carlo sampling, and then use linear regression or a neural network to fit the charging and discharging power curves of individual electric vehicles. For example, the decision variables for the user model are the charging and discharging power at different times, with constraints including the battery's safe SOC range (e.g., 20% to 90%) and the maximum charging and discharging power (e.g., 7 kW).
[0028] In some embodiments, to quantify user costs, the system can calculate the minimum total charging cost for electric vehicle users using the following formula (3): (3) Formula (3) represents an optimization model aimed at minimizing the total charging cost for electric vehicle users. This represents the total charging cost for electric vehicle (EV) users during the scheduling period, expressed in yuan. Its value is minimized through optimization calculations. The index number representing the time period is a dimensionless integer; T represents the total number of time periods within the scheduling cycle, which is also a dimensionless integer. Characterizing the first The charging power of electric vehicles during the time period, expressed in kilowatts, is a decision variable. Characterizing the first The charging and discharging subsidy coefficient for a given time period, expressed in yuan per hour; Characterizing the first The fixed cost of battery degradation during different time periods, expressed in yuan, is calculated using a battery aging model. This formula allows the system to accurately quantify the economic cost of charging at different times and to solve for the optimal charging power using an optimization algorithm. This achieves the goal of minimizing total cost.
[0029] In some embodiments, the system can calculate the first step using the formula (4) shown below. Battery degradation cost over time: (4) Formula (4) represents the single-period loss cost model calculated based on battery aging characteristics. Characterizing the first t Net charge / discharge capacity during the period, in kilowatt-hours, during discharge. It is a negative value when charging and a positive value when charging. Cycle life, a dimensionless integer, is the number of complete charge-discharge cycles that a battery can complete before its capacity decays to 80% of its rated value at a specific depth of discharge. Depth of Discharge (DSD) is a percentage measure of the amount of discharge relative to the rated capacity in a single cycle. Characterizes the rated capacity of a battery, measured in kilowatt-hours; Characterized by the initial purchase cost of the battery, in yuan; This formula represents the residual value of a battery after retirement, expressed in yuan. Using this formula, the system can accurately quantify the economic impact of a single charge-discharge cycle on battery life.
[0030] In some embodiments, the system can calculate the maximum evaluation utility value of the j-th electric vehicle aggregator using the following formula (5): (5) Formula (5) represents an optimization model aimed at maximizing the evaluation utility of electric vehicle aggregators (hereinafter referred to as aggregators). The comprehensive evaluation utility value of the j-th aggregator during the scheduling cycle is represented by a dimensionless economic indicator, which is obtained by optimization calculation to obtain the maximum value. This represents the electricity price charged by the j-th aggregator to the user in time period t, expressed in yuan per kilowatt-hour. The total charging power of the j-th aggregator in time period t is represented by kilowatt-hours, and its value is a decision variable. The price at which the j-th aggregator purchases electricity from the distribution network operator during time period t is expressed in yuan per kilowatt-hour. The quantity represents the amount of electricity purchased by the j-th aggregator from the distribution network operator during time period t, in kilowatt-hours. The fixed operating costs of the j-th aggregator during time period t are represented in yuan, including communication, metering and other expenses. The weighting coefficient for power deviation penalties, expressed in yuan per kilowatt-hour, is used to quantify the severity of penalties for deviations between actual and reported charging power. This formula allows the system to comprehensively consider the aggregator's revenue from the electricity purchase and sale price difference, operating costs, and dispatch deviation penalties.
[0031] In some embodiments, the system can calculate the total charging capacity of the j-th aggregator in time period t using the following formula (6): (6) Formula (6) represents the mathematical model for calculating the total charging capacity using the power-time integration method. i represents the index number of the charging stage and is a dimensionless integer. The total number of charging stages is represented by a dimensionless integer. The power value of the i-th charging stage within the t-th time period is represented in kilowatts and is obtained by the battery management system in real time. This represents the duration of each time interval, expressed in hours. Using this formula, the system can accurately aggregate discrete power sample values to obtain the total power consumption for a given period.
[0032] In some embodiments, the system can calculate the response matching degree of the j-th aggregator in time period t using the following formula (7): (7) Formula (7) represents the calculation model for the matching degree between the actual response power and the declared power. The ratio representing the response matching degree of the j-th aggregator in time period t is a dimensionless ratio used to quantify the accuracy of aggregator scheduling and execution. Using this formula, the system can evaluate the scheduling and execution accuracy of aggregators in real time.
[0033] In some embodiments, the system can calculate the maximum utility value of the distribution network operator using the following formula (8): (8) Wherein, formula (8) represents the optimization model with the goal of maximizing the comprehensive utility of the distribution network operator; j represents the index number of the aggregator, which is a dimensionless integer; J represents the total number of aggregators, which is a dimensionless integer; denoted by , represents the unit electricity service fee rate charged by the distribution network operator to the j-th aggregator, in yuan per kilowatt-hour; denoted by , represents the electricity deviation penalty coefficient, in yuan per kilowatt-hour squared, used to quantify the penalty intensity of the deviation between the declared electricity and the actual electricity. The meanings of the remaining parameters in formula (8) are consistent with those in formulas (3) to (7). Through this formula, the system can comprehensively consider the electricity sales revenue, service fee income, and dispatch deviation penalty of the distribution network operator.
[0034] In some embodiments, the system can calculate the maximized utility value of the j-th aggregator using the following formula (9): (9) Formula (9) represents an optimization model that aims to maximize the overall utility of the j-th aggregator. The power deviation penalty coefficient, expressed in yuan per kilowatt-hour, is used to quantify the economic penalty intensity for the deviation between the declared power volume and the actual power sales volume. The lower limit threshold representing the operating cost in time period t, in yuan, is determined by equipment depreciation and minimum maintenance cost; The upper limit threshold representing the operating cost in time period t, in yuan, is determined by the highest historical operating cost and inflation factors. The meanings of the remaining parameters in formula (8) are consistent with those in formulas (3) to (8). Through this formula, the system can quantify the balance between the aggregator's revenue from the "purchase and sale price difference", operating costs, and dispatch deviation penalties.
[0035] Step S130: Input the charging and discharging strategy optimization model, the strategy optimization model, and the power allocation optimization model into the multi-agent non-cooperative and master-slave hybrid game framework to generate the charging and discharging scheduling strategy.
[0036] The multi-agent non-cooperative and master-slave hybrid game framework is a three-layer dynamic interaction architecture model constructed through a hierarchical coupling approach, based on the non-cooperative game relationship parameters between distribution network operators and aggregators, and the master-slave game relationship parameters between aggregators and electric vehicle users. This framework integrates the optimization problems of the upper layer (distribution network operators), the middle layer (aggregators), and the lower layer (electric vehicle users) into a unified game process, achieving equilibrium through iterative solution.
[0037] Specifically, the system can use distributed optimization algorithms (such as Alternating Direction Method of Multipliers (ADMM) or gradient descent) to solve the game equilibrium: First, initialize the strategies of each agent (such as electricity price and power allocation), then alternately update the lower-level user model, the middle-level aggregator model, and the upper-level operator model until the strategy change is less than a threshold (such as 0.01). The master-follower game relationship is realized through a leader-follower model: the aggregator acts as the leader, issuing electricity price signals, and electric vehicle users act as followers, adjusting their charging and discharging plans; the non-cooperative game relationship is realized through a competitive bidding mechanism, where each aggregator independently optimizes its own bid to maximize its profits. ADMM stands for Alternating Direction Method of Multipliers.
[0038] Step S140: Execute the charging and discharging scheduling of electric vehicles according to the charging and discharging scheduling strategy.
[0039] Among them, scheduling execution can refer to sending the strategies generated by the game framework (such as time-sharing charging and discharging power instructions) to the electric vehicle or charging pile controller to achieve real-time power control.
[0040] For example, the system can operate through a cloud-edge collaborative architecture: after the cloud server calculates the scheduling strategy, it sends instructions to the designated charging piles via the edge gateway. For instance, during off-peak hours in the evening (such as from 11 PM to 5 AM the next day), an electric vehicle is instructed to charge at 5 kW for 2 hours; during peak hours at noon (such as from 12 PM to 2 PM), it is instructed to discharge at 3 kW for 1 hour.
[0041] Therefore, according to the above implementation method, the system first acquires charging and discharging behavior data of multiple target electric vehicles; based on the charging and discharging behavior data, it simultaneously constructs a charging and discharging strategy optimization model for electric vehicle users, a strategy optimization model for multiple aggregators, and a power allocation optimization model for distribution network operators; subsequently, it inputs the charging and discharging strategy optimization model, the strategy optimization model, and the power allocation optimization model into a multi-agent non-cooperative and master-slave hybrid game framework. This framework is a three-layer dynamic interactive architecture constructed by layering the non-cooperative game relationship between distribution network operators and aggregators, and the master-slave game relationship between aggregators and electric vehicle users; finally, it executes the charging and discharging scheduling of electric vehicles according to the charging and discharging scheduling strategy generated by this framework.
[0042] Throughout the process, the above embodiments address the issues mentioned in the background art, namely, the lack of systematic modeling of aggregator incentive mechanisms and the fact that "existing game theory-based research often considers single game types or simple two-layer games, failing to achieve real-time collaboration among the three parties." By constructing a hybrid game framework that includes optimization models for three main entities—electric vehicle users, aggregators, and distribution network operators—this framework layers and couples non-cooperative game theory with master-slave game theory, achieving collaborative optimization of the interaction and balance of interests among the three parties under a unified architecture. Therefore, the technical solution of the above embodiments effectively overcomes the shortcomings of existing technologies, such as a single game structure and insufficient model coupling. It solves the technical problem of difficulty in achieving real-time collaborative optimization scheduling among distribution network operators, aggregators, and electric vehicle users, improving the overall efficiency and scheduling reliability of large-scale electric vehicle participation in grid interaction.
[0043] In some embodiments, the charging and discharging behavior data includes the charging start time, the charging end time, and the initial state of charge. The charging start time can refer to the specific time when the electric vehicle user begins charging, typically following a normal distribution with a mean of 18:00 and a standard deviation of 1 hour. The charging end time can refer to the time when charging is completed, and depending on the charging duration and start time, it may follow a normal distribution with a mean of 6:00 the next day and a standard deviation of 2 hours. The initial state of charge can refer to the percentage of battery charge at the start of charging, for example, ranging from 20% to 80%.
[0044] The steps for constructing the charging and discharging strategy optimization model include: Behavioral modeling is performed based on a pre-defined probability distribution model for the start time of charging, the end time of charging, and the initial state of charge.
[0045] Among them, the probability distribution model can refer to a mathematical model used to describe the statistical characteristics of random variables. For example, the normal distribution model is used at the start of charging. (This indicates a mean of 18 and a standard deviation of 1 hour). The charging end time is calculated using a normal distribution model. (This indicates a mean of 30, i.e., 6:00 AM the next day, with a standard deviation of 2 hours). The initial state of charge adopts a uniform distribution model. (This indicates a uniform distribution between 20% and 80%).
[0046] Behavioral samples are extracted from the probability distribution model using the Monte Carlo sampling method, and charging and discharging behavior features are extracted from the behavioral samples.
[0047] For example, the system performs 10,000 Monte Carlo samplings, each generating a set of random samples with charging start times, end times, and initial state of charge. Charging and discharging behavior characteristics can refer to statistics calculated from the samples to describe user behavior patterns. For example, extracted features include average charging time (in hours), number of daily charging events, charging power distribution (in kilowatts), and the rate of change of state of charge. Specifically, the average charging time can be calculated as the difference between the end time and the start time in the sample, for example, with a mean of 12 hours; the number of daily charging events can be calculated by counting the number of charging events within 24 hours, for example, with a mean of 1 time.
[0048] Based on the characteristics of charging and discharging behavior, a charging and discharging strategy optimization model is constructed with the goal of minimizing the total charging cost for users. The total charging cost for users includes the charging and discharging cost of the individual electric vehicle, charging and discharging subsidies, and battery loss costs.
[0049] The total charging cost for users can refer to the total economic burden incurred by electric vehicle users during the charging and discharging process. The individual charging and discharging cost for an electric vehicle can refer to the electricity fee calculated based on electricity prices. For example, using a time-of-use pricing model, the peak electricity price (e.g., 8:00 AM to 10:00 PM) is 1.2 yuan per kilowatt-hour, and the off-peak electricity price (e.g., 10:00 PM to 8:00 AM the next day) is 0.5 yuan per kilowatt-hour. Charging and discharging subsidies can refer to incentive fees provided by the government or operators to encourage users to participate in grid dispatch; for example, a subsidy of 0.1 yuan per kilowatt-hour discharged. Battery depreciation cost can refer to the cost of battery aging caused by charge-discharge cycles, estimated through a battery life model; for example, the depreciation cost per charge-discharge cycle is 0.05 yuan per kilowatt-hour. The system can construct the optimization model as a mathematical programming problem, such as linear programming or dynamic programming, with the objective function being to minimize the total cost, and constraints including battery capacity limitations and charging power limitations.
[0050] In some embodiments, the system can calculate the minimum total charging cost for electric vehicle users using the following formula (10): (10) Formula (10) represents an optimization model aimed at minimizing the total charging cost for electric vehicle users. The total charging cost for electric vehicle users during the scheduling period is represented in yuan, and its minimum value is obtained through optimization calculation. The cyclic index variable representing the time period is a dimensionless integer used for counting in the summation operation. The meanings of the remaining parameters in formula (10) are consistent with those in formulas (3) to (9). Through this formula, the system can accurately quantify the economic cost of charging for users at different times.
[0051] In some embodiments, the system can determine the feasible region constraint of the charging power using the following formula (11): (11) Formula (11) represents the upper and lower limit constraints of the charging power, which is used to ensure that the charging power is within the safe operating range of the equipment. The minimum operating power allowed for a charging device is expressed in kilowatts and is determined by the hardware performance of the charger. This represents the maximum permissible operating power of the charging device, measured in kilowatts, and is limited by circuit capacity and the battery management system. This formula is embedded as a constraint in the model solution process for the optimization problem.
[0052] In some embodiments, the system can calculate the battery state of charge at the next moment using the following formula (12): (12) Formula (12) represents the recursive calculation model of battery state of charge based on power direction and efficiency coefficient. The state of charge of the battery at time t+1 is represented as a percentage, which is the ratio of the remaining charge to the rated capacity at that time. The state of charge of the battery at time t is represented by a known initial state value; It represents the time step, in hours, which is the time interval between adjacent moments; The discharge efficiency coefficient is a dimensionless decimal that takes into account the energy loss during the discharge process. The charging efficiency coefficient is a dimensionless decimal that takes into account energy loss during the charging process. The rated capacity of a battery is indicated by its unit of kilowatt-hour.
[0053] Next, the system can calculate the safe operating range of the battery state of charge and output the validity verification results using the following formulas (13) and (14): (13) (14) Formula (13) represents the safety operating range constraint conditions of the battery's state of charge; Characterizes the minimum permissible state of charge threshold of a battery, used to prevent over-discharge; The highest permissible state of charge threshold for the battery is defined to prevent overcharging. Formula (14) defines the validity verification condition of the actual output state of charge relative to the required value. The minimum required state of charge of a battery for a specific application scenario, expressed in decimal form; The measured state of charge (SOC) value, representing the actual output of the battery, is expressed as a decimal and is acquired in real time by the battery management system. These two formulas together constitute a dual verification mechanism for battery energy management, providing a mathematical basis for the safe execution of charging and discharging strategies.
[0054] In some embodiments, the system can construct the Lagrangian function for electric vehicle charging and discharging optimization using the following formula (15): (15) Formula (15) represents the augmented objective function constructed using the Lagrange multiplier method (a mathematical method for handling constrained optimization problems). L represents the Lagrange function, which is a dimensionless comprehensive performance index. The optimal charging and discharging strategy can be obtained by solving the extreme points of this function. The charging price for time period t is expressed in yuan per kilowatt-hour. The charging power decision variable represents the charging power in time period t, in kilowatts; The discharge compensation price for time period t is expressed in yuan per kilowatt-hour. The discharge power decision variable represents the discharge power in time period t, in kilowatts; k represents the battery loss cost coefficient, in yuan per kilowatt-hour. Characterizes the rated capacity of the battery, in kilowatt-hours; , , , , The Lagrange multipliers, representing various constraints (auxiliary variables used to introduce constraints into the objective function), are all non-negative real numbers. Auxiliary variables characterizing the upper limit constraint on the state of charge; Characterized by the maximum allowable charging power in time period t, in kilowatts; Characterized by the maximum permissible discharge power in time period t, in kilowatts; The charging efficiency coefficient is a dimensionless decimal. The discharge efficiency coefficient is a dimensionless decimal. The meanings of the other parameters in formula (15) are the same as those in formulas (3) to (14).
[0055] Next, the system can calculate the optimal charging power and optimal discharging power that make the Lagrange function (15) reach its extreme value by solving the optimality conditions formed by the following formulas (16), (17) and (18): (16) (17) (18) (19) ; ; (20) Formula (16) represents the Lagrangian function L with respect to charging power. The optimality condition of the first-order partial derivative is that the gradient of the function with respect to the charging power is zero at the extreme point. Characterizing the Lagrangian function L with respect to the charging power in time period t The partial derivative of L is a dimensionless mathematical quantity, and its equality with zero is a necessary condition for the existence of an extremum; the meanings of other parameters are consistent with those in formula (15). Formula (17) characterizes the Lagrangian function L's effect on discharge power. The optimality condition of the first-order partial derivative. Characterizing the Lagrangian function L with respect to the discharge power in time period t The partial derivatives of . Formula (18) characterizes the boundary constraints of charging and discharging power. The charging power decision variable represents the charging power in time period t, in kilowatts; Characterizes the maximum allowable charging power in time period t, in kilowatts; The decision variable characterizing the discharge power in time period t is... Abbreviation for kilowatt, unit is kilowatt. The maximum allowable discharge power in time period t is represented by kilowatts. The system calculates the optimal power by simultaneously solving the system of equations (16) to (18). and Therefore, the system can verify the constraint satisfaction of the optimization model using the above formulas (19) and (20); formula (19) characterizes the hard constraint of the safe operating range of the battery's state of charge. Formula (20) characterizes the nonnegativity constraint condition of the Lagrange multipliers (an auxiliary variable used for constraint optimization). The Lagrange multiplier characterizing the upper limit constraint of the charged state in time period t; The Lagrange multiplier characterizes the bounded constraint of the charged state in time period t.
[0056] Furthermore, the system can be characterized by mixed-integer linear programming modeling of the charge / discharge state and charge state boundary constraints using the following formula (21): ;(twenty one) Formula (21) represents the linearized constraint set constructed using the Big-M method (a mathematical modeling technique for handling coupled constraints of discrete and continuous variables). The charging status flag bit, representing the charging status in time period t, is a binary decision variable (taking values of 0 or 1). The time period indicates that charging is allowed during that period; M represents a sufficiently large positive real constant (Big-M coefficient) used for linearizing the logic constraints; The flag representing the discharge state at time t is a binary decision variable. The time indicates that discharge is permitted during that period; The activation flag representing the upper limit of the state of charge in time period t is a binary decision variable. This indicates that the upper limit constraint has been triggered; The activation flag representing the charge state at time t is a binary decision variable. The time indicates that the lower bound constraint is triggered; by solving this mixed integer programming model, the system can simultaneously obtain the optimal solutions for discrete state decision-making and continuous power allocation.
[0057] In some embodiments, the system can calculate the angle parameter using the formula (22) shown below. : ;(twenty two) in, An angle parameter representing time t, in degrees, used to describe the angle value of the system control variables; The minimum reference value for characterizing angular parameters, in degrees; The basic incremental unit for representing angles is degrees, used to define angular resolution; N represents the resolution parameter of the angle encoding, a dimensionless integer; k represents the index number of the binary bits, a dimensionless integer ranging from 0 to... .
[0058] In some embodiments, the system can calculate the power-angle integrated control quantity using the formula (23) shown below. This parameter is used to quantify the synergistic effect of active power and phase angle to optimize grid stability. ;(twenty three) Formula (23) represents the power-angle coupling parameter calculation model based on binary weight encoding. The power-angle integrated control quantity at time t is represented by units of This parameter is used to quantify the synergistic effect of power and phase angle; The baseline value of active power in time period t, in kilowatts; The basic resolution unit representing the phase angle is degrees; N represents the total number of angle encoding segments, which is an integer power of 2; k represents the index number of the binary weight bits, which is a dimensionless integer with a value range of... .
[0059] Next, the system can construct the constraints in the optimization model using the formula (24) shown below, which is used to handle the logical relationship between binary decision variables and continuous variables: ;(twenty four) Formula (24) represents the constraint set using the Big M method, which is commonly found in mixed-integer linear programming problems. M represents a sufficiently large positive real constant (called the Big M constant) to ensure the logical completeness of the constraints; t and k represent subscript indices, which are dimensionless integers used to distinguish different time periods and decision dimensions. These constraints are often applied to problems such as electric vehicle charging and discharging strategy optimization and resource scheduling, to transform "if-then" logical conditions into linear inequalities, making it easier for the solver to calculate.
[0060] In some embodiments, the system can calculate the scheduling deviation penalty cost using the following formulas (25) and (26): (25) (26) Formulas (25) and (26) together constitute a scheduling deviation cost calculation model based on a linear penalty mechanism. Characterizes the total cost of scheduling deviation penalty, in yuan, used to quantify the economic penalty caused by the deviation between actual electricity generation and planned electricity generation; The penalty unit price representing the amount of unit deviation, expressed in yuan per kilowatt-hour; The auxiliary variable representing the positive deviation in time period t, in kilowatt-hours, represents the portion of actual electricity consumption that exceeds planned electricity consumption; The negative deviation auxiliary variable, represented by kilowatt-hours, indicates the portion of actual electricity consumption less than planned electricity consumption during time period t. The constraints of formula (26) ensure the complementarity of the positive and negative deviation variables: when hour, and ;when hour, and When both are equal, they are both zero.
[0061] In some embodiments, the system can calculate the maximized utility value of the aggregator using the formula (27) shown below: (27) Formula (27) represents an optimization model aimed at maximizing the overall utility of the aggregator. The value representing the aggregater's overall utility during the scheduling period is expressed in yuan, and its maximum value is obtained through optimization calculation. The unit represents the electricity price charged by the j-th aggregator to electric vehicle users in time period t, expressed in yuan per kilowatt-hour. The wholesale price at which the j-th aggregator purchases electricity from the distribution network operator during time period t is represented in yuan per kilowatt-hour. The total cost of the scheduling deviation penalty is represented by yuan, and its value is calculated using formula (25). Through this formula, the system can comprehensively consider the aggregator's electricity sales revenue, electricity purchase cost, operating expenses, and deviation penalty, thereby maximizing utility.
[0062] Therefore, according to the above implementation method, the system can generate the optimal charging and discharging strategy, dynamically adjust the charging period and power, thereby reducing the total cost to users, improving grid stability, and extending battery life.
[0063] In some embodiments, the charging and discharging behavior data further includes distribution parameters for the charging start time, distribution parameters for the charging end time, and schedulable energy parameters. The distribution parameters for the charging start time can refer to characteristic values describing the probability distribution of the charging start time, such as the mean (e.g., 18:00) and standard deviation (e.g., 1 hour) of a normal distribution; the distribution parameters for the charging end time can refer to characteristic values describing the probability distribution of the charging end time, such as the mean (e.g., 30:00) and standard deviation (e.g., 2 hours) of a normal distribution; the schedulable energy parameters can refer to the upper limit of the energy that a single electric vehicle can participate in grid dispatch (discharging or adjusting charging) within a specific time period, for example, ranging from 10 kWh to 30 kWh, typically related to battery capacity and user travel demand. The steps for constructing the aggregator's strategy optimization model include: The total schedulable power is calculated based on the distribution parameters at the start and end of charging, as well as the schedulable power parameters.
[0064] The total dispatchable power capacity refers to the sum of the dispatchable power capacity of electric vehicles managed by each aggregator during the demand response period. Specifically, the total dispatchable power capacity can be calculated by weighting or probabilistic expectation based on the number of electric vehicles and the dispatchable power capacity parameters of each vehicle.
[0065] Based on the total dispatchable electricity and the electricity pricing strategies for each time period, a strategy optimization model is constructed with the goal of maximizing the revenue from electricity purchase and sale.
[0066] In this model, the electricity pricing strategy for each time period constitutes the decision variables for the strategy optimization model. The electricity pricing strategy can refer to the sequence of electricity sales prices or purchase prices declared by aggregators for different time periods in the electricity market. For example, a day can be divided into 24 time periods (each time period is 1 hour), and a price can be declared for each time period (unit: yuan per kilowatt-hour). Electricity sales revenue can refer to the net income obtained by aggregators through market transactions, i.e., the difference between electricity sales revenue and electricity purchase costs.
[0067] Therefore, according to the above implementation method, the system can generate the optimal electricity pricing strategy for aggregators, obtain the maximum economic benefits in the electricity market by coordinating and scheduling the flexible resources of a large number of electric vehicles, and at the same time improve the renewable energy consumption capacity.
[0068] In some embodiments, the charging and discharging behavior data also includes historical benefit coefficient parameters and efficiency loss coefficient parameters. The historical benefit coefficient parameter can refer to an evaluation index used to quantify the aggregator's performance in historical demand response events. This coefficient is positively correlated with response speed and power completion rate; for example, its value ranges from 0 to 1, with higher values indicating stronger historical performance reliability. The efficiency loss coefficient parameter refers to the proportion of power loss caused by factors such as line losses and conversion efficiency during the transmission of power from the aggregator to the grid node; for example, its value ranges from 2% to 5%, and it is typically calculated based on grid topology and electrical distance. The steps of the power distribution optimization model for distribution network operators include: The contract response power is determined based on the total dispatchable power, historical benefit coefficient parameters, and efficiency loss coefficient parameters.
[0069] The allocation scheme for contract response power includes power allocation weighting parameters between the distribution network operator and each aggregator. Contract response power can refer to the total target power volume that the distribution network operator plans to dispatch to each aggregator during a demand response event. The power allocation weighting parameters can refer to coefficients that characterize the proportion of the allocated power volume to each aggregator relative to the total contract power volume, with a weight sum of 100%.
[0070] Based on the power allocation scheme of contract response, an optimization model for power allocation is constructed with the goal of maximizing grid revenue.
[0071] In the power allocation optimization model, the decision variables are determined by the power allocation weights of each aggregator. Grid revenue refers to the net economic benefits obtained by distribution network operators through demand response, primarily including reduced grid loss costs, delayed equipment investment gains, and potential ancillary service market revenue.
[0072] Therefore, according to the above implementation method, the system can generate the optimal power distribution scheme, and maximize the benefits on the distribution network side by comprehensively considering the reliability of aggregators and transmission losses, thereby improving the overall economy and execution efficiency of demand response projects.
[0073] In some embodiments, after constructing the power distribution optimization model, the above method further includes: Establish the objective function for optimizing power allocation for distribution network operators.
[0074] The objective function includes an electricity trading benefit term and an efficiency loss compensation term. The activation condition for the efficiency loss compensation term includes the contracted electricity volume completion rate failing to reach a preset threshold. The electricity trading benefit term refers to the direct economic benefits obtained by the distribution network operator through allocating and executing contracted electricity volumes to aggregators. Specifically, the efficiency loss compensation term refers to the additional compensation cost incurred by the operator to ensure grid balance when the actual delivered electricity volume is lower than the contract value due to efficiency losses by the aggregator. This cost is typically charged to the fulfilling party or borne by the defaulting party. The contracted electricity volume completion rate refers to the ratio of the actual delivered electricity volume to the contracted electricity volume during the settlement period, expressed as a percentage. The preset threshold refers to the threshold value that triggers the compensation mechanism, set by the operator according to reliability requirements. For example, setting the threshold to 85% means that the compensation term is activated when the contracted electricity volume completion rate is lower than 85%.
[0075] Based on the preset dynamic power allocation mechanism, each aggregator is prioritized according to the historical benefit coefficient parameter, and the efficiency loss of each aggregator that meets the activation conditions is quantified according to the efficiency loss coefficient parameter.
[0076] The dynamic power allocation mechanism can refer to an adaptive algorithm that adjusts power allocation weights based on the real-time reliability of aggregators and the grid status. Priority ranking can refer to sorting aggregators from highest to lowest according to their historical benefit coefficients, with aggregators having higher coefficients receiving priority in power allocation. For example, aggregator A with a historical benefit coefficient of 0.9 is ranked before aggregator B with a coefficient of 0.7. Efficiency loss quantification can refer to calculating the power deviation caused by efficiency losses for aggregators that have activated compensation conditions.
[0077] Therefore, according to the above implementation method, the system can automatically identify high-reliability aggregators and quantify potential performance losses during the optimization allocation process, thereby dynamically adjusting the allocation strategy and effectively improving the economic benefits and dispatch reliability of the power grid side.
[0078] In some embodiments, the steps of constructing a multi-agent non-cooperative and master-slave hybrid game framework include: Establish the characteristics of the non-cooperative game relationship between power distribution network operators and various electric vehicle aggregators.
[0079] Among them, the basic parameters of strategy interaction in non-cooperative game relations include power allocation and electricity pricing. Non-cooperative game relations can refer to a strategy interaction mode in which multiple decision-making entities (distribution network operators and aggregators) make independent decisions and compete with each other under the drive of their own interests. The strategy choice of any entity will affect the gains of other entities, but there is no mandatory agreement among the parties. The basic parameters of strategy interaction can refer to the core variables upon which each party relies to make decisions during the game. For example, the power allocation parameter refers to the proportion of contracted power allocated by the operator to each aggregator (e.g., aggregator A receives 30%), and the electricity pricing parameter refers to the electricity price declared by the aggregator in the electricity market (e.g., 0.65 yuan per kilowatt-hour).
[0080] Establish the master-slave game relationship characteristics between various electric vehicle aggregators and electric vehicle users.
[0081] In this master-slave game, aggregators act as leaders, and electric vehicle users as followers. The master-slave game characteristic can be described as a hierarchical decision-making structure where leaders (aggregators) first formulate strategies (such as announcing charging and discharging prices to users), and followers (users) respond according to the leaders' strategies with the goal of maximizing their own utility (such as adjusting charging times). Leaders can refer to the decision-making entity that has an initial advantage in the game and can anticipate the followers' response strategies; for example, aggregators guide user behavior by setting subsidized prices. Followers can refer to the passive party that makes optimal decisions based on the leader's announced strategies; for example, users decide their charging plans based on electricity price signals.
[0082] A three-layer dynamic interaction architecture model is constructed by layering and coupling the characteristics of non-cooperative game relations with the characteristics of master-slave game relations.
[0083] The three-layer dynamic interaction architecture model includes upper-layer distribution network operators, middle-layer aggregators, and lower-layer electric vehicle users. Layered coupling refers to linking the game relationships at different levels through decision variables and utility functions. For example, the operator's electricity allocation decisions influence the aggregator's pricing strategy, which in turn affects the user's charging and discharging behavior. The three-layer dynamic interaction architecture model can also be seen as a mathematical model describing the bidirectional strategy transmission and feedback between upper-layer operators, middle-layer aggregators, and lower-layer users. This model seeks an equilibrium state through iterative calculations. For example, upper-layer operators optimize electricity allocation with the goal of maximizing grid revenue; middle-layer aggregators engage in non-cooperative game theory to determine electricity prices with the goal of maximizing electricity purchase and sale revenue; lower-layer users respond to pricing signals with the goal of minimizing charging costs, and their aggregation behavior feeds back to the middle layer, influencing the game outcome.
[0084] Therefore, according to the above implementation method, the system can simulate complex interaction processes among multiple parties, and by solving the equilibrium solution of the hybrid game model, it can achieve coordinated optimization of the interests of the distribution network side, the aggregator side, and the user side, thereby improving the economy and stability of the overall power system.
[0085] In some embodiments, the step of inputting the charging / discharging strategy optimization model, the strategy optimization model, and the power allocation optimization model into a multi-agent non-cooperative and master-slave hybrid game framework to generate a charging / discharging scheduling strategy includes: The power distribution optimization model, strategy optimization model, and charging / discharging strategy optimization model are respectively input into the upper, middle, and lower layers of the three-layer dynamic interaction architecture model.
[0086] The corresponding input can refer to integrating each optimization model as the objective function and constraint condition of the corresponding level decision-making entity into the architecture. For example, the upper-level distribution network operator uses a power allocation optimization model, the middle-level aggregators use a strategy optimization model, and the lower-level electric vehicle user group uses a charging and discharging strategy optimization model. In the lower level, the constraints of the charging and discharging strategy optimization model are introduced into the equilibrium solution process to establish a master-slave game equilibrium constraint between the aggregators and electric vehicle users.
[0087] Specifically, the equilibrium solution process can refer to the computational flow of finding a stable state through mathematical methods that leaves neither the leader nor the followers with any incentive to unilaterally change their strategies. The equilibrium constraints in a leader-follower game refer to the mathematical conditions that must be satisfied in the equilibrium state, typically manifested as the optimal reaction function of the followers (users) being embedded in the leader's (aggregator's) optimization problem. For example, establishing equilibrium constraints involves using the user's cost minimization model (charging / discharging strategy optimization model) as the reaction function of the lower-level followers, and embedding its optimal solution as a parameter into the middle-level aggregator's profit maximization model (strategy optimization model), thus forming the equilibrium conditions. At the upper level, based on the strategy optimization model and the power distribution optimization model, non-cooperative game equilibrium constraints are constructed between the distribution network operator and each aggregator.
[0088] In non-cooperative game equilibrium constraints, the conditions that all participants' strategy combinations must satisfy to reach Nash equilibrium in a non-cooperative game are met, meaning that no single participant can obtain higher payoffs by changing their strategy alone. For example, the constructed constraints can be expressed as follows: given the strategies of other aggregators and the operator's allocation scheme, each aggregator's electricity bidding strategy is the optimal solution to its payoff maximization problem; simultaneously, given all aggregator bidding strategies, the operator's electricity allocation scheme is the optimal solution to its grid payoff maximization problem.
[0089] A charging and discharging scheduling strategy is generated by collaboratively solving the equilibrium constraints of master-slave games and non-cooperative games, and by performing steps such as strategy initialization, iterative update and convergence determination.
[0090] Collaborative solution can refer to using distributed algorithms to simultaneously handle the equilibrium constraints of multi-level games, such as using the alternating direction multiplier method or the optimal response dynamic algorithm.
[0091] Specifically, when performing the strategy initialization step, the system can set initial values for the decision variables of each game player. For example, it can initialize the electricity price of each aggregator to the average grid connection price (e.g., 0.6 yuan per kilowatt-hour) and initialize the power allocation weight of the operator to uniform distribution. When performing the iterative update step, the system can cyclically adjust the strategy variables of each player according to preset rules. For example, in each iteration, lower-level users update their charging plans based on the current electricity price, middle-level aggregators update their prices based on user feedback, and upper-level operators update their allocation weights based on the prices. When performing the convergence determination step, the system can determine whether the iterative process has reached an equilibrium state based on certain criteria. For example, convergence is determined when the maximum value of the change in all decision variables in two consecutive iterations is less than a preset threshold (e.g., 0.001). The generated charging and discharging scheduling strategy can refer to the stable strategy combination obtained after convergence, including the optimal charging and discharging time and power commands for each user, the optimal electricity price sequence for each aggregator, and the optimal power allocation scheme for the operator.
[0092] In some embodiments, the system can calculate the reported electricity volume of the j-th aggregator in the t-th time period using the following formula (28): (28) Formula (28) represents the power allocation calculation model based on the proportional factor. The power allocation ratio factor representing the j-th aggregator in time period t is a dimensionless decimal whose value is determined by the aggregator's resource capacity or contract. This represents the total allocable electricity provided by the distribution network operator in time period t, expressed in kilowatt-hours. Using this formula, the system can allocate total electricity resources proportionally and fairly.
[0093] Next, the system can calculate the rate of change of the electricity allocation ratio factor of the j-th aggregator at time t using the formula (29) shown below: (29) Formula (29) represents the differential equation model based on the dynamic rate of change of voltage to charge ratio. The power allocation ratio factor representing the j-th aggregator at time t The rate of change with time t, in units of hour, is obtained by solving this differential equation; The power distribution ratio factor representing the j-th aggregator at time t is a dimensionless decimal. The adjustment coefficient characterizing the j-th aggregator is a dimensionless real number used to control the sensitivity of the rate of change. The voltage measurement representing the j-th aggregator at time t, in volts; The average voltage reference value, in volts, represents the j-th aggregator at time t, and its value is calculated based on historical data. This formula allows the system to quantify the scaling factor's response speed to voltage fluctuations in real time.
[0094] Then, the system can calculate the equivalent voltage on the distribution network side at time t using the following formula (30): (30) Formula (30) represents the equivalent voltage calculation model for the distribution network side based on the weighted average algorithm. Through this formula, the system can comprehensively consider the voltage contribution of each access point and calculate a more representative equivalent voltage.
[0095] Furthermore, the system can calculate the updated power allocation ratio factor of the j-th aggregator in the (z+1)-th iteration using the formula (31) shown below: (31) Formula (31) represents the dynamic iterative model of the proportional factor based on the voltage-electric density deviation. The power allocation ratio factor representing the j-th aggregator at time t and in the (z+1)-th iteration is a dimensionless decimal, representing the latest share of the aggregator in the total power allocation; The power allocation ratio factor at the z-th iteration is a dimensionless decimal; z represents the index number of the iteration number, which is a dimensionless integer. Using this formula, the system can dynamically adjust the power allocation ratio of each aggregator based on the deviation between its real-time voltage-power density and the average value.
[0096] Next, the system can determine whether the iterative process has converged using the formula (32) shown below: (32) Formula (32) represents the convergence judgment condition based on the voltage-electric density deviation. When the inequality holds, it means that the deviation between the voltage-electric density of the j-th aggregator and the average value is within the tolerance range, and the iteration process converges; otherwise, iteration needs to continue. ε represents the convergence tolerance threshold, with the unit being volts per kilowatt-hour. It is a small positive real number used to control the convergence accuracy. The meanings of other parameters in the formula are consistent with those in formula (31). Through this formula, the system can automatically determine whether the iterative algorithm has reached the convergence state.
[0097] Then, the system can calculate the electricity price of the j-th aggregator in the (r+1)th iteration using the formula (33) shown below: (33) Formula (33) represents the iterative update model of electricity price based on target deviation, which uses gradient descent (an optimization algorithm) to gradually adjust the price to approach the ideal value. The unit represents the electricity price quoted by the j-th aggregator in time period t and the (r+1)-th iteration, in yuan per kilowatt-hour. λ represents the electricity price quoted by the j-th aggregator in time period t and at the r-th iteration, in yuan per kilowatt-hour; λ represents the learning rate coefficient, a dimensionless decimal used to control the update step size in each iteration; The target electricity price quoted by the j-th aggregator in time period t is expressed in yuan per kilowatt-hour, and its value is determined by market equilibrium or cost models; r represents the index number of the iteration number, which is a dimensionless integer. Using this formula, the system can dynamically optimize the pricing strategy based on the target deviation.
[0098] Finally, the system can determine whether the electricity price bidding iteration has converged using the formula (34) shown below: (34) Formula (34) represents the convergence criterion of whether the absolute deviation between the target electricity price and the current electricity price is less than a preset threshold. When the inequality holds, it indicates that the iterative process has converged; otherwise, iteration needs to continue. The target electricity price quoted by the j-th aggregator in time period t is expressed in yuan per kilowatt-hour, and its value is determined by the market equilibrium or cost optimization model. The threshold parameter for convergence judgment, expressed in yuan per kilowatt-hour, is a small positive real number used to control convergence accuracy. Using this formula, the system can automatically determine whether the bidding iteration has reached a stable state.
[0099] In some embodiments, the system can be as follows: Figure 2The flowchart of the hierarchical iterative optimization algorithm shown above performs multi-subject collaborative optimization calculation, specifically including the following steps: First, the system performs an initialization step, which can refer to the process of setting initial values for all variables. Next, the system performs two optimization update processes in parallel: the distribution network operator optimizes and updates the power allocation scheme, and the aggregator optimizes and updates the price and strategy. Among them, the distribution network operator optimization update can refer to calculating the optimal power allocation weight of each aggregator based on formula (8) with the goal of maximizing the grid revenue; the aggregator optimization update can refer to adjusting its power price and charging and discharging strategy based on formula (9) with the goal of maximizing the utility of the aggregator. Then, the system enters the convergence judgment step, which can refer to the process of evaluating whether the current solution meets the preset accuracy requirements. The judgment condition of this step is based on formula (34), that is, checking whether the absolute deviation between the target value and the current value is less than the threshold. If the judgment result is "No" (not converged), the system executes the iterative update step, updating the iteration counter from k to k+1 (e.g., from k=0 to k=1), and returns to the parallel optimization step to continue the loop; if the judgment result is "Yes" (converged), the system outputs the equilibrium solution, which can refer to the stable state that makes all agents have no incentive to unilaterally change their strategies, such as the final power allocation scheme. The electricity price quotation sequence is as follows The cost is per kilowatt-hour. This flowchart enables the system to perform hierarchical iterative optimization, ensuring policy coordination among distribution network operators, aggregators, and users, ultimately generating a globally optimal scheduling scheme.
[0100] Therefore, according to the above implementation method, the system can automatically generate a charging and discharging scheduling strategy that coordinates and optimizes the interests of the power grid, aggregators and users by solving the equilibrium solution of the complex game framework, which can effectively improve the system's operating efficiency and economic benefits.
[0101] Figure 3 This is a structural block diagram of an electric vehicle charging and discharging scheduling system based on multi-agent hybrid game theory according to an embodiment of the present invention.
[0102] like Figure 3As shown, the electric vehicle charging and discharging scheduling system based on multi-agent hybrid game theory includes: a charging and discharging behavior data acquisition module 210, used to acquire charging and discharging behavior data of multiple target electric vehicles; a multi-agent optimization model construction module 220, used to construct charging and discharging strategy optimization models for electric vehicle users, strategy optimization models for multiple electric vehicle aggregators, and power allocation optimization models for distribution network operators based on the charging and discharging behavior data; a charging and discharging scheduling strategy generation module 230, used to input the charging and discharging strategy optimization models, strategy optimization models, and power allocation optimization models into a multi-agent non-cooperative and master-slave hybrid game framework to generate a charging and discharging scheduling strategy. The multi-agent non-cooperative and master-slave hybrid game framework is a three-layer dynamic interactive architecture model constructed through a layered coupling method based on the non-cooperative game relationship parameters between the distribution network operator and each electric vehicle aggregator, and the master-slave game relationship parameters between each electric vehicle aggregator and electric vehicle users; and a charging and discharging scheduling strategy execution module 240, used to execute the charging and discharging scheduling of electric vehicles according to the charging and discharging scheduling strategy.
[0103] The specific functions and examples of each module and submodule of the device in this embodiment of the invention can be found in the relevant descriptions of the corresponding steps in the above method embodiments, and will not be repeated here.
[0104] According to embodiments of the present invention, the above-described method of the present invention can be applied to an electronic device and a readable storage medium. Figure 4 A schematic block diagram of an electronic device 600 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein. Figure 4As shown, the electronic device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. The RAM 603 may also store various programs and data required for the operation of the electronic device 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604. Multiple components in the electronic device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a disk, optical disk, etc.; and a communication unit 609, such as a network card, modem, wireless transceiver, etc. The communication unit 609 allows the electronic device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks. The computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as a multi-agent hybrid game-based electric vehicle charging and discharging scheduling method. For example, in some embodiments, a multi-agent hybrid game-based electric vehicle charging and discharging scheduling method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of the multi-agent hybrid game-based electric vehicle charging and discharging scheduling method described above can be performed. Alternatively, in other embodiments, the computing unit 601 may be configured, by any other suitable means (e.g., by means of firmware), to perform an electric vehicle charging and discharging scheduling method based on multi-agent hybrid game theory.
[0105] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip systems, payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0106] The program code used to implement the methods of the present invention can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a standalone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0107] In the context of this invention, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0108] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT or LCD monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual, auditory, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input). The systems and techniques described herein can be implemented in computing systems that include back-end components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include front-end components (e.g., a user computer with a graphical user interface or web browser through which the user interacts with embodiments of the systems and techniques described herein), or computing systems that include any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected via digital data communication (e.g., a communication network) of any form or medium. Examples of communication networks include: Local Area Networks (LANs), Wide Area Networks (WANs), and the Internet.
[0109] Computer systems may include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. Client-server relationships are created by computer programs running on respective computers and having a client-server relationship with each other. Servers may be cloud servers, distributed system servers, or servers incorporating blockchain technology. It should be understood that steps can be reordered, added, or deleted using the various forms of processes shown above. For example, the steps described in this invention may be executed in parallel, sequentially, or in different orders, as long as the desired results of the disclosed technical solutions are achieved; no limitation is imposed herein.
[0110] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for scheduling the charging and discharging of electric vehicles based on multi-agent hybrid game theory, characterized in that, include: Acquire charging and discharging behavior data of multiple target electric vehicles; Based on the charging and discharging behavior data, a charging and discharging strategy optimization model for electric vehicle users, a strategy optimization model for multiple electric vehicle aggregators, and a power distribution optimization model for power distribution network operators are constructed. The charging and discharging strategy optimization model, the strategy optimization model, and the power allocation optimization model are input into a multi-agent non-cooperative and master-slave hybrid game framework to generate a charging and discharging scheduling strategy. The multi-agent non-cooperative and master-slave hybrid game framework is a three-layer dynamic interaction architecture model constructed through a layered coupling method based on the non-cooperative game relationship parameters between the distribution network operator and each of the electric vehicle aggregators, and the master-slave game relationship parameters between each of the electric vehicle aggregators and the electric vehicle users. The power allocation optimization model, the strategy optimization model, and the charging / discharging strategy optimization model are respectively input into the upper, middle, and lower layers of the three-layer dynamic interactive architecture model. In the lower layer, the user's cost minimization model is used as the reaction function of the lower-layer followers, and its optimal solution is embedded as a parameter into the revenue maximization model of the middle-layer aggregator, establishing a master-slave game equilibrium constraint between each electric vehicle aggregator and the electric vehicle user. In the upper layer, the revenue maximization conditions of each aggregator's power pricing strategy under the given strategies of other aggregators and the operator's allocation scheme, and the grid revenue maximization conditions of the operator's power allocation scheme under the given pricing strategies of all aggregators are constructed, forming a non-cooperative game equilibrium constraint. By collaboratively solving the master-slave game equilibrium constraint and the non-cooperative game equilibrium constraint, and performing the steps of strategy initialization, iterative update, and convergence determination, the charging / discharging scheduling strategy is generated. The iterative update steps include: in each iteration, lower-level users update their charging plans based on the current electricity price, middle-level aggregators update their prices based on user feedback, and upper-level operators update their allocation weights based on the prices; the upper-level operator updating the allocation weights specifically includes: obtaining the voltage measurement value of the j-th aggregator at time t, calculating the equivalent voltage on the distribution network side at time t, and dynamically adjusting the power allocation ratio factor of the j-th aggregator based on the voltage-power density deviation; the convergence determination steps include: determining whether the change in the decision variable in two consecutive iterations is less than a preset threshold, and confirming iterative convergence based on the dual criteria of voltage-power density deviation and target power price deviation; The charging and discharging scheduling of electric vehicles is executed according to the charging and discharging scheduling strategy.
2. The method according to claim 1, characterized in that, The charging and discharging behavior data includes the charging start time, charging end time, and initial state of charge; the construction steps of the charging and discharging strategy optimization model include: The behavior of the charging start time, charging end time and initial state of charge is modeled based on a preset probability distribution model. Behavioral samples are extracted from the probability distribution model using the Monte Carlo sampling method, and charging and discharging behavior features are extracted from the behavioral samples. Based on the charging and discharging behavior characteristics, a charging and discharging strategy optimization model is constructed with the goal of minimizing the user's total charging cost. The user's total charging cost includes the charging and discharging cost of the individual electric vehicle, charging and discharging subsidies, and battery loss costs.
3. The method according to claim 2, characterized in that, The charging and discharging behavior data also includes distribution parameters at the start of charging, distribution parameters at the end of charging, and schedulable power parameters; the steps for constructing the strategy optimization model for the electric vehicle aggregator include: The total schedulable power is calculated based on the distribution parameters at the start of charging, the distribution parameters at the end of charging, and the schedulable power parameters. The total schedulable power refers to the sum of the schedulable power of the electric vehicles managed by each of the electric vehicle aggregators during the demand response period. Based on the total dispatchable electricity and the electricity pricing strategy for each time period, a strategy optimization model is constructed with the goal of maximizing the revenue from electricity purchase and sale. The electricity pricing strategies for each time period constitute the decision variables of the strategy optimization model.
4. The method according to claim 3, characterized in that, The charging and discharging behavior data also includes historical benefit coefficient parameters and efficiency loss coefficient parameters; the construction steps of the power distribution optimization model of the distribution network operator include: The contract response power is determined based on the total dispatchable power, the historical benefit coefficient parameter, and the efficiency loss coefficient parameter. The allocation scheme of the contract response power includes the power allocation weight relationship parameter between the distribution network operator and each of the electric vehicle aggregators. Based on the power allocation scheme of the contract response, the power allocation optimization model is constructed with the goal of maximizing grid revenue; The decision variables of the power allocation optimization model are determined by the power allocation weights of each electric vehicle aggregator.
5. The method according to claim 4, characterized in that, After constructing the power distribution optimization model, the method further includes: Establish an objective function for optimizing power allocation for the power distribution network operator. The objective function includes a power transaction benefit term and an efficiency loss compensation term. The activation condition for the efficiency loss compensation term includes that the contract power completion rate has not reached a preset threshold. Based on the preset dynamic power allocation mechanism, the electric vehicle aggregators are prioritized according to the historical benefit coefficient parameter, and the efficiency loss of each electric vehicle aggregator that meets the activation condition is quantified according to the efficiency loss coefficient parameter.
6. The method according to claim 4, characterized in that, The steps for constructing the multi-agent non-cooperative and master-slave hybrid game framework include: Establish the non-cooperative game relationship characteristics between the power distribution network operator and each of the electric vehicle aggregators, wherein the basic parameters of the strategy interaction of the non-cooperative game relationship characteristics include power allocation and power price. Establish the master-slave game relationship characteristics between each of the electric vehicle aggregators and the electric vehicle users. In the master-slave game relationship, each of the electric vehicle aggregators acts as the leader and the electric vehicle users act as the followers. By layering and coupling the non-cooperative game relationship features with the master-slave game relationship features, a three-layer dynamic interaction architecture model is constructed. The three-layer dynamic interaction architecture model includes an upper-layer power distribution network operator, a middle-layer electric vehicle aggregator, and a lower-layer electric vehicle user.
7. An electric vehicle charging and discharging scheduling system based on multi-agent hybrid game theory, characterized in that, include: The charging and discharging behavior data acquisition module is used to acquire charging and discharging behavior data of multiple target electric vehicles. The multi-agent optimization model construction module is used to construct, based on the charging and discharging behavior data, a charging and discharging strategy optimization model for electric vehicle users, a strategy optimization model for multiple electric vehicle aggregators, and a power distribution optimization model for power distribution network operators. The charging and discharging scheduling strategy generation module is used to input the charging and discharging strategy optimization model, the strategy optimization model, and the power distribution optimization model into a multi-agent non-cooperative and master-slave hybrid game framework to generate a charging and discharging scheduling strategy. The multi-agent non-cooperative and master-slave hybrid game framework is a three-layer dynamic interaction architecture model constructed through a layered coupling method based on the non-cooperative game relationship parameters between the distribution network operator and each of the electric vehicle aggregators, and the master-slave game relationship parameters between each of the electric vehicle aggregators and electric vehicle users. A charge / discharge scheduling strategy execution module is used to execute the charge / discharge scheduling of electric vehicles according to the charge / discharge scheduling strategy. The charging and discharging scheduling strategy generation module is further used to input the power allocation optimization model, the strategy optimization model, and the charging and discharging strategy optimization model into the upper, middle, and lower layers of the three-layer dynamic interactive architecture model, respectively. In the lower layer, the user's cost minimization model is used as the reaction function of the lower-layer follower, and its optimal solution is embedded as a parameter into the revenue maximization model of the middle-layer aggregator, establishing a master-slave game equilibrium constraint between each electric vehicle aggregator and the electric vehicle user. In the upper layer, the revenue maximization conditions of each aggregator's power bidding strategy under the given strategies of other aggregators and the operator's allocation scheme, and the grid revenue maximization conditions of the operator's power allocation scheme under the given bidding strategies of all aggregators are constructed, forming a non-cooperative game equilibrium constraint. The charging and discharging scheduling strategy is generated by collaboratively solving the master-slave game equilibrium constraint and the non-cooperative game equilibrium constraint, and by performing strategy initialization, iterative update, and convergence determination steps. The iterative update steps include: in each iteration, lower-level users update their charging plans based on the current electricity price, middle-level aggregators update their prices based on user feedback, and upper-level operators update their allocation weights based on the prices; the upper-level operator updating the allocation weights specifically includes: obtaining the voltage measurement value of the j-th aggregator at time t, calculating the equivalent voltage on the distribution network side at time t, and dynamically adjusting the power allocation ratio factor of the j-th aggregator based on the voltage-power density deviation; the convergence determination steps include: determining whether the change in the decision variable in two consecutive iterations is less than a preset threshold, and confirming iterative convergence based on the dual criteria of voltage-power density deviation and target power price deviation.
8. An electronic device, characterized in that, include: At least one processor; as well as The memory that is communicatively connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-6.
9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, Computer instructions are used to cause a computer to perform the method according to any one of claims 1-6.
Citation Information
Patent Citations
Large-scale electric vehicle charging and discharging optimization scheduling method based on multi-main-body double-layer game
CN114662759A
Multi-agent cooperative scheduling electric energy hybrid game method
CN121073093A