A graph theory-statistics joint modeling driven dynamic clustering control method for electric vehicles

By using graph theory-statistical joint modeling and multi-agent deep deterministic policy gradient algorithm, dynamic clustering and active and reactive power joint optimization of electric vehicles are achieved, which solves the problems of poor convergence speed and decision performance in the interaction between electric vehicles and the power grid, and improves the accuracy of carrying capacity assessment and system stability of new energy access to the distribution network.

CN121012034BActive Publication Date: 2026-01-27STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511534964.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-27
Publication Date
2026-01-27
Estimated Expiration
2045-10-27

AI Technical Summary

Technical Problem

Existing technologies suffer from poor convergence speed and decision-making performance in the interaction between electric vehicles and the power grid, as well as the inability to accurately reflect the actual carrying capacity of nodes. In particular, when large-scale new energy sources are connected to the distribution network, traditional methods are difficult to adapt to dynamic changes and high-dimensional scenarios.

Method used

A graph theory-statistical joint modeling-driven approach is adopted. By constructing graph theory topology and statistical distribution characteristics and combining them with a multi-agent deep deterministic policy gradient algorithm, a two-stage serial decision architecture is designed to realize dynamic grouping and joint optimization of active and reactive power of electric vehicles. The carrying capacity margin is quantified by combining a multi-dimensional space dynamic evaluation model.

Benefits of technology

It improves the adaptability and decision-making performance of electric vehicle grouping, enhances the accuracy of the carrying capacity assessment of new energy access to the distribution network, and improves the stability of the distribution network and the capacity for new energy absorption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121012034B_ABST
    Figure CN121012034B_ABST
Patent Text Reader

Abstract

The application relates to a kind of graph theory-statistics joint modeling driven electric vehicle dynamic grouping control method, constructs graph theory-statistics joint driven vehicle network two-way interaction model;Subsystems are divided to distribution network, and based on multi-agent deep deterministic policy gradient algorithm, using centralized training, distributed execution framework, constructs two-stage series decision architecture, to solve distribution network electric vehicle grouping and distribution network active and reactive power joint optimization problem;In the first stage of series decision, the electric vehicles in each subsystem are adaptively grouped in real time, and in the second stage, active-reactive power joint optimization control is carried out based on the grouping result;A multi-dimensional space dynamic evaluation model is constructed to integrate time series fluctuation, voltage stability and reverse power flow constraints, and the maximum access capacity margin boundary of the photovoltaic grid-connected point is obtained by solving the evaluation model.Compared with the prior art, the method can improve the new energy margin of the distribution network node level, optimize the voltage and reduce the network loss.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power system optimization scheduling technology, and in particular to a dynamic grouping control method for electric vehicles driven by graph theory-statistical joint modeling. Background Technology

[0002] The large-scale distributed renewable energy access to the distribution network has a huge impact on the distribution network due to the intermittency and fluctuation of its output. This makes the two-way interaction technology between electric vehicles and the grid (V2G) to achieve joint optimization of active and reactive power in the smart grid an important technical means to improve the renewable energy carrying capacity of the distribution network.

[0003] With the increasing demand for enhanced new energy vehicle capacity, EV grouping strategies have become a crucial aspect of vehicle-to-grid (V2G) interaction. Most research on EV grouping focuses on clustering algorithms; however, these algorithms often rely on static states and pre-defined rules, which may prove inadequate in dynamically changing scenarios. Furthermore, as the number of EVs connected to the grid grows exponentially, the scheduling problem of V2G interaction exhibits characteristics such as high dimensionality and strong uncertainty. While some research has focused on heuristic algorithms, these often heavily rely on static assumptions and low-dimensional, precise modeling, making it difficult to adapt to the complex state changes in dynamic scenarios. Deep reinforcement learning algorithms possess both adaptive policy learning mechanisms and strong expressive power, enabling intelligent decision-making under high-dimensional and state-changing conditions. Although deep reinforcement learning algorithms demonstrate strong modeling capabilities and solution efficiency in electric vehicle scheduling and distribution network optimization, they still face the "curse of dimensionality" in scenarios with large distribution network node scales or high state dimensions, severely limiting the convergence speed and decision-making performance of the algorithms.

[0004] Currently, research on renewable energy carrying capacity generally focuses on comprehensively assessing or improving the renewable energy acceptance capacity of distribution networks from multiple dimensions, such as system safety and stability and power quality assurance. Furthermore, traditional renewable energy carrying capacity assessment indicators and methods are all aimed at evaluating or improving the overall renewable energy carrying capacity of distribution networks, and cannot accurately reflect the actual carrying capacity of individual nodes. Summary of the Invention

[0005] The purpose of this invention is to solve the technical problems of poor convergence speed and decision performance and inability to accurately reflect the actual carrying capacity of nodes in the current technical solutions, and to provide a dynamic grouping control method for electric vehicles driven by graph theory-statistical joint modeling.

[0006] The objective of this invention can be achieved through the following technical solutions:

[0007] A graph theory-statistical joint modeling-driven dynamic grouping control method for electric vehicles includes the following steps:

[0008] A bidirectional vehicle-to-network interaction model driven by graph theory and statistics is constructed by integrating graph theory and statistical characteristics.

[0009] The distribution network is divided into subsystems, and a two-stage serial decision architecture is constructed based on the multi-agent deep deterministic policy gradient algorithm, using a centralized training and distributed execution framework, to solve the joint optimization problem of electric vehicle grouping and active and reactive power of the distribution network.

[0010] In the first stage of the two-stage cascaded decision-making architecture, electric vehicles in each subsystem are dynamically and adaptively grouped in real time; in the second stage, the vehicle-network active and reactive power joint optimization control is carried out based on the grouping results.

[0011] A multi-dimensional spatial dynamic evaluation model is constructed for photovoltaic grid-connected points. The carrying capacity margin is quantified by solving the high-dimensional geometric volume enclosed by the boundary of the evaluation model, thereby determining the maximum access capacity margin boundary of the photovoltaic grid-connected point.

[0012] As a preferred technical solution, the graph theory-statistics jointly driven vehicle-to-grid bidirectional interaction model uses graph theory to characterize the structured relationship of the interaction effect between electric vehicles and charging piles, establishing a topological model of vehicle-to-grid interaction; combined with statistical modeling, it calculates and analyzes the temporal-spatial distribution and interdependence among electric vehicle groups for each node. The specific model construction is as follows:

[0013] Electric vehicles and charging stations are defined as mobile nodes and static nodes, respectively. Mobile nodes represent the real-time behavior of electric vehicle users, and static nodes represent the real-time status of charging stations. Each electric vehicle forms a virtual connection with the charging station of its subsystem to describe the charging and discharging needs of the electric vehicle.

[0014] Based on the node attributes of electric vehicle nodes, the edge weights of the connections between each electric vehicle mobile node and the charging node, and the power distribution network requirements, electric vehicles with the same characteristics are divided into groups. Each electric vehicle node in the same group forms a group connection with the corresponding virtual group node, thus forming a group subgraph.

[0015] A set of virtual nodes is created for group division. The virtual nodes are used to connect the electric vehicle nodes in their corresponding groups, representing the group relationship of electric vehicles within the group. The group subgraph includes a reactive power optimization group subgraph, an off-peak charging group subgraph, a peak response group subgraph, a backup power supply group subgraph, and a new energy consumption group subgraph. Electric vehicles dynamically switch their groups based on real-time demand and grid status.

[0016] For each distribution network subsystem, a corresponding heterogeneous information subgraph is constructed. In the heterogeneous information subgraph, the node set within the subsystem includes: the charging node set, the electric vehicle mobile node set, and the group virtual node set; the edge set within the subsystem includes: the set of edges connecting electric vehicles and subsystem charging stations, and all group edges representing clusters within the subsystem.

[0017] The weights of mobile nodes include their geographical location, mileage, battery state of charge, and the start time of charging; the mileage weight uses the mileage probability function; the battery state of charge weight is the state of charge at the end of the last charge of the electric vehicle minus the quotient of the mileage function and the maximum mileage that the electric vehicle can travel in a fully charged state; the start time of charging weight uses the probability density function of the start time of charging.

[0018] When an electric vehicle connects to a charging station for charging, the charging duration of the connected edge is assigned as the edge weight.

[0019] As a preferred technical solution, for the power distribution network, based on the power grid load matching relationship between the charging station and other distribution network nodes, and comprehensively considering charging demand, power flow, and electrical distance factors, the subsystem is divided as follows:

[0020] Using the impedance modulus as the electrical distance, and taking each charging station node as the root node, the shortest electrical distance between each node and each root node is calculated. Considering the load demand and power generation capacity of each node, the power flow between each node is calculated using the optimal power flow method.

[0021] Using the shortest electrical distance, node geographic coordinates, and power flow as the basis for clustering, each charging station node is used as the initial centroid of the cluster, and the distribution network subsystem is divided through clustering algorithm.

[0022] As a preferred technical solution, the objective function of the combined active and reactive power optimization problem of the distribution network electric vehicle grouping includes an active power optimization objective, a reactive power optimization objective, and an economic performance objective. The active power optimization objective reduces system network losses by coordinating the charging and discharging behavior of electric vehicles. The reactive power optimization objective regulates the distribution of reactive power by utilizing the power transmission between electric vehicles and charging piles. The economic performance objective considers the charging and discharging costs of electric vehicles and the charging service costs of charging stations.

[0023] As a preferred technical solution, in the two-stage serial decision architecture based on the multi-agent deep deterministic policy gradient algorithm, the environment is a power distribution network that includes electric vehicles and new energy access, and the agent is the electric vehicle scheduler of each subsystem; the state of the power distribution network changes over time, and the agent adjusts the group allocation and charging behavior of electric vehicles by observing the state information of the power distribution network and electric vehicles based on the current observation information.

[0024] Real-time collection of operational status data from various nodes of the power distribution network, and acquisition of key attribute information of electric vehicles;

[0025] A centralized training method is adopted, and the electric vehicle grouping strategy is trained by a multi-agent deep deterministic policy gradient algorithm. Training data is shared among multiple agents to determine the optimal action strategy for each agent.

[0026] The agent's action is to match mobile nodes with similar attributes, form connections between them and corresponding group virtual nodes, and control the charging behavior of electric vehicle mobile nodes. The goal is to achieve power balance, voltage stability, and effective absorption of new energy in the power grid.

[0027] During the distributed execution phase, each agent makes independent decisions based on the trained policy network and adjusts the group charging strategy of the electric vehicle mobile nodes according to the real-time observed state information.

[0028] As a preferred technical solution, the multi-agent deep deterministic policy gradient algorithm constructs an evaluation network and action network for the agents based on an observable Markov decision process, and adopts a serial two-stage decision framework. In the first stage, the agents perform dynamic grouping decisions; in the second stage, power scheduling is optimized based on the grouping results. In the observable Markov decision process:

[0029] The state space includes the dynamic clustering decision-making state space and the state space of the vehicle-to-grid interaction control stage. The dynamic clustering decision-making state space includes the state information of the distribution network state information and the state information of the heterogeneous information graph of vehicle-to-grid interaction. The state space of the vehicle-to-grid interaction control stage includes the group state space after the electric vehicles are grouped.

[0030] The joint observation space includes the observation information upon which the agent's first and second-stage decisions depend;

[0031] In the joint action space, a multi-dimensional discrimination method for joint boundary decision-making is adopted in the first decision-making stage. For the grouping strategy, the action is defined as a dynamic adjustment operation on the group boundary, including the adjustment of the SOC boundary value, the electric distance boundary value, the EV load mode boundary value, and the charging demand boundary value. In the second stage, corresponding optimization strategies are formulated for different groups. When the electric vehicle flow node belongs to the reactive power optimization group, continuous control is executed. When the electric vehicle flow node belongs to other groups, a discrete charging and discharging strategy based on fast charging power and slow charging power is executed.

[0032] The state transition probability function, determined by the environment, is used to characterize the dynamic process of the distribution network's state evolution being affected by the current state and decision-making strategies. In the first stage, the state transition is driven by the cluster boundary adjustment decision and the environment, characterizing the impact of cluster boundary adjustment on the division of electric vehicle groups. In the second stage, the state transition is affected by the vehicle-grid interaction optimization decision, describing the effect of electric vehicle charging and discharging on the distribution network state.

[0033] The reward function of the intelligent agent comprehensively considers three control objectives: distribution network voltage stability, charge balance, and economic performance.

[0034] A discount factor is set to balance the conversion coefficients of immediate rewards and future rewards.

[0035] As a preferred technical solution, the evaluation network of the multi-agent deep deterministic policy gradient algorithm agent is used to evaluate the merits of the decision-making actions of each subsystem agent under different states, thereby guiding the selection of electric vehicle grouping strategies; in the second stage, the evaluation network combines the dispatchability of electric vehicles, the operating status of the distribution network and economic objectives to optimize the power scheduling strategy of each subsystem agent.

[0036] In the cascaded two-stage decision-making framework of the multi-agent deep deterministic policy gradient algorithm, the first-stage input of the evaluation network includes the observation state of the heterogeneous information graph of the local environment of the distribution network and the interaction between the vehicle and the network, as well as the action of the electric vehicle grouping strategy; the second-stage input is the observation state of the local environment after grouping and the scheduling action of the second stage; the output of the evaluation network is the state-action value estimate that measures the merits of the action strategy; the parameters to be optimized are the evaluation network parameters and the target evaluation network parameters.

[0037] The loss function of the evaluation network directly evaluates the entire decision chain, and iteratively updates the evaluation network parameters by minimizing the difference between the estimated value of the output and the expected cumulative reward.

[0038] As a preferred technical solution, the multi-agent deep deterministic policy gradient algorithm introduces a two-stage experience replay mechanism based on structural response labels in the evaluation network part, in which each agent combines the experience data accumulated during training in the state-action evaluation estimation process.

[0039] The first stage of the dual-stage experience replay mechanism is the structural decision stage, which is responsible for constructing the dynamic grouping structure of electric vehicles. The experience pool of the structural decision stage stores the observation values ​​of each agent in the electric vehicle grouping stage, the group division at the current time step, and the global performance response of the group structure under the control of the second stage.

[0040] The second stage is the strategy optimization stage, which optimizes the operation of the power distribution network by coordinating the active and reactive power of the grouped electric vehicles based on the group structure. The experience pool of the optimization stage stores the observations and actions of each agent in the vehicle-to-grid interaction control stage, the joint reward of each agent at the end of the two-stage chain, and the state at the end of the two-stage chain decision.

[0041] As a preferred technical solution, the action network of the multi-agent deep deterministic policy gradient algorithm agent is used to evaluate the effect of different scheduling strategies, thereby assisting in optimizing the scheduling of electric vehicle charging stations within the subsystem; during the training phase, the agent learns and optimizes the action strategy using the current environmental observation state; during the execution phase, the agent dynamically adjusts the grouping strategy of electric vehicles and optimizes the power distribution of the power grid based on the output strategy of the action network.

[0042] In the cascaded two-stage decision framework of the multi-agent deep deterministic policy gradient algorithm, the agent's observations in the first stage include the local observation state of the distribution network and the state information related to electric vehicles, and the output is the grouping strategy of electric vehicles; the local observations input in the second stage include the grouping results of the first stage and the further observed environmental state, and the output is the electric vehicle scheduling action; the parameters to be optimized in the action network are the action network parameters and the target action network parameters.

[0043] Each agent's action network incorporates a target network and a policy entropy term; the action network uses a policy gradient method to jointly optimize the action network parameters across the entire decision chain.

[0044] As a preferred technical solution, the multidimensional space dynamic evaluation model characterizes the maximum photovoltaic access margin boundary composed of multiple nodes by maximizing the Pareto front of the photovoltaic capacity adjustment amount connected to each node, and uses the expansion or contraction of the margin boundary in the multidimensional space to characterize the change in the system's carrying capacity.

[0045] The high-dimensional space model is reduced in dimension by approximation and projection using a uniform manifold, and the high-dimensional geometry enclosed by the margin boundary in the high-dimensional space is mapped to the low-dimensional space.

[0046] Discretize the time series and calculate the volume of the geometric body enclosed by the boundary map of the spatial dynamic evaluation model for each discrete time series in the low-dimensional space. Define the average value as the new energy carrying capacity margin.

[0047] Compared with the prior art, the present invention has the following beneficial effects:

[0048] 1) This invention addresses the uncertainty in group partitioning caused by the heterogeneity of electric vehicle behavior and spatial clustering. It constructs a joint modeling mechanism that integrates prior graph structure knowledge with statistical distribution characteristics to form a logically related graph model, used to characterize the structural coupling relationship between electric vehicles and charging infrastructure. Based on this, a dynamic grouping mechanism based on joint boundary discrimination is proposed. By dynamically adjusting the boundary weights in the graph structure, the temporal adaptive evolution of the grouping structure is achieved, effectively enhancing the group partitioning's ability to perceive behavioral dynamics and its structural flexibility.

[0049] 2) To achieve decoupled modeling and collaborative optimization of structural partitioning and group control, this invention designs a two-stage serial optimization architecture based on MADDPG. The first stage learns the optimal clustering strategy, and the second stage performs joint active and reactive power optimization control of the vehicle-network system based on the clustering results. To address the problem of information fragmentation between multiple stages, structural response labels are introduced as a bridge for the experience playback mechanism, ensuring the effective transmission of structural information in strategy optimization, thereby improving the feedback perceptibility of clustering decisions and the structural adaptability of the control strategy.

[0050] 3) This invention addresses the challenge of assessing the overall structure and dynamic boundaries of new energy capacity under specific multi-node access scenarios. It constructs a multi-dimensional feasible domain model considering factors such as time-series fluctuations, voltage stability, and reverse power flow risks. Furthermore, it combines uniform manifold approximation and projection (UMAP) techniques to reduce the dimensionality of the high-dimensional solution space, proposing a node-level new energy carrying capacity margin assessment index system. This system can be used to identify the spatial distribution impact of electric vehicle behavior changes on new energy access margin in the distribution network, providing a more refined and structurally aware quantitative basis for the coordinated planning of the main and distribution networks. Attached Figure Description

[0051] Figure 1 This is a flowchart of the "graph theory-statistics" joint-driven strategy method for improving the carrying capacity of new energy in V2G based on multi-agent deep reinforcement learning, as proposed in this invention.

[0052] Figure 2 This is a schematic diagram of the logical topology of the vehicle-to-network interaction process of the present invention.

[0053] Figure 3 This is a schematic diagram of the centralized training-distributed execution framework of the present invention.

[0054] Figure 4 This is a schematic diagram of the two-stage experience playback mechanism of the present invention.

[0055] Figure 5 This is a flowchart of the MADDPG algorithm of the present invention.

[0056] Figure 6 This is a schematic diagram of the topology and subsystem division of a 113-node distribution network in a certain location in an implementation example of the present invention.

[0057] Figure 7 This is a reward convergence graph obtained from an implementation example of the present invention.

[0058] Figure 8 This is a schematic diagram of the full-time load situation under disordered charging in a daily implementation example of the present invention.

[0059] Figure 9 This is a schematic diagram of the full-time load situation within a day under EV dynamic clustering in the implementation example of the present invention.

[0060] Figure 10 This is a comparison diagram of the full-time voltage in the implementation examples of this invention.

[0061] Figure 11 This is a full-time voltage heatmap for a day under dynamic clustering in an implementation example of the present invention.

[0062] Figure 12 This is a comparison diagram of network loss in the implementation examples of this invention.

[0063] Figure 13 This is a schematic diagram of the feasible domain boundary (unordered charging) after dimensionality reduction in the implementation examples of this invention.

[0064] Figure 14 This is a schematic diagram of the feasible domain boundary (EV dynamic clustering) after dimensionality reduction in the implementation examples of this invention. Detailed Implementation

[0065] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are based on the technical solution of the present invention and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.

[0066] Example 1

[0067] This invention addresses the large-scale, high-penetration characteristics of new energy sources in emerging power distribution systems. It constructs a bidirectional interaction model between electric vehicles (EVs) and the power grid driven by a graph theory-statistics approach, proposes a node-level EV carrying capacity margin assessment index, and designs an adaptive dynamic clustering strategy for EVs that considers the EV carrying capacity of the distribution network. First, by integrating graph theory topology and statistical distribution characteristics, a logical relationship graph model between EVs and charging facilities is constructed, achieving a balance between structural flexibility and quantitative accuracy. Second, the distribution network is partitioned using K-means clustering, and a two-stage serial decision-making architecture is adopted through a multi-agent deep deterministic strategy gradient algorithm. A two-stage experience replay mechanism based on structural response labels is designed to enable real-time adaptive dynamic clustering of EVs within each partition, jointly optimizing the active and reactive power of the distribution network to improve EV carrying capacity. Third, for the assessment of node-level EV carrying capacity, this invention constructs a high-dimensional dynamic assessment model integrating time-series fluctuations, voltage stability, and reverse power flow constraints. Based on the UMAP dimensionality reduction method, the capacity margin boundary under multi-node collaboration is characterized, effectively improving the accuracy and interpretability of EV access capacity assessment. Figure 1 As shown, the above steps are implemented as follows:

[0068] S1. Graph Theory-Statistics Jointly Driven Modeling of Bidirectional Interaction between Electric Vehicles and the Power Grid: First, Dijkstra's algorithm and optimal power flow are used to obtain the shortest electrical distance and power flow between each node and the charging station, and the distribution network is divided into subsystems using k-means clustering. Then, to meet the flexibility requirements of the vehicle-grid interaction control model, graph theory is introduced to characterize the structured relationship of the interaction effect between electric vehicles and charging piles, and a topological model of vehicle-grid interaction is established. Combined with statistical modeling, accurate computational support is provided for each node, and the temporal-spatial distribution and interdependence among electric vehicle groups are analyzed to establish a graph theory-statistics jointly driven model of bidirectional interaction between electric vehicles and the power grid.

[0069] S2. Modeling of Electric Vehicle Segmentation and Vehicle-to-Grid Interaction in a Novel Distribution System Based on Multi-Agent Collaborative Optimization: Based on a vehicle-to-grid interaction model jointly driven by graph theory and statistics, autonomous electric vehicle adaptive segmentation strategy learning agents are established in each subsystem. A Markov decision model and a multi-agent attention athlete-referee coordination governance architecture are established. A centralized training and decentralized learning method is adopted. Each agent uses a deep deterministic policy gradient algorithm and adopts a serial two-stage decision-making method. At the same time, a two-stage experience replay mechanism based on structural response labels is designed. With the objectives of minimizing active power loss, minimizing voltage deviation, and optimizing economic performance, online adaptive segmentation of electric vehicles in the distribution network is realized, simulating the charging and discharging behavior of electric vehicles under different charging demands, SOC states, and grid load levels.

[0070] S3. A Novel Distribution System Node-Level New Energy Carrying Capacity Margin Assessment Model: Based on the dynamic clustering of electric vehicles in the distribution network, and considering the volatility and randomness of new energy sources, this model analyzes the new energy capacity addition capability of the distribution network at the node level. First, considering multiple defined photovoltaic grid-connected points, and taking into account factors such as time-series changes, voltage stability, and counter-current flow, a multi-dimensional dynamic assessment model of the distribution network's new energy carrying capacity is established to characterize the maximum photovoltaic access margin boundary composed of multiple nodes. Then, the high-dimensional model is reduced in dimensionality through Uniform Manifold Approximation and Projection (UMAP) to simplify the representation of the multi-dimensional model. A novel node-level new energy carrying capacity margin assessment index is defined. By discretizing the time series, the dynamic boundary of the carrying capacity margin is calculated to more accurately reflect the power system's capacity to carry new energy at the node level, and its advantages in assessing new energy carrying capacity are analyzed.

[0071] Furthermore, the specific implementation of the graph theory-statistical joint-driven bidirectional interaction modeling of electric vehicles and power grids in step 1 is as follows:

[0072] 1.1 Distribution Network Subsystem Division

[0073] To effectively achieve electric vehicle dispatching, optimize distribution network operation, and improve the renewable energy carrying capacity of the distribution network, this invention divides the distribution network into subsystems based on the grid load matching relationship between charging stations and other distribution network nodes, comprehensively considering factors such as charging demand, power flow, and electrical distance. The specific steps are as follows:

[0074] (1.1.1) Formulate the basis for dividing the distribution network subsystem. The impedance modulus is taken as the electrical distance, and each charging station node is taken as the root node. The shortest electrical distance between each node and each root node is calculated by Dijkstra's algorithm. Considering the load demand and power generation capacity of each node, the power flow between each node is calculated by the optimal power flow method.

[0075] (1.1.2) Distribution network subsystem partitioning. The shortest electrical distance, node geographical coordinates, and power flow are used as the basis for clustering. At the same time, in order to avoid the number of nodes in a single partition being too small, a minimum number of nodes constraint is set for the subsystem. Each charging station node is used as the initial centroid of the cluster, and the distribution network subsystem is partitioned by the k-means clustering algorithm.

[0076] 1.2 Vehicle-Network Interaction Graph Theory Model

[0077] Electric vehicles interact with charging stations to achieve power exchange with the power distribution network. For example... Figure 2As shown, the vehicle-to-grid interaction process is logically topologically processed, and electric vehicles and charging stations are defined as mobile nodes and static nodes, respectively. Each electric vehicle forms a virtual connection with the charging station of its subsystem to describe the charging and discharging needs of the electric vehicle.

[0078] The interaction between electric vehicles and charging stations is defined as the original heterogeneous information graph. G=(V,E) .in V It is a set of nodes, including all electric vehicles and centralized charging stations; E It is an edge set, which includes all connections between electric vehicles and charging stations, as well as connections between electric vehicles.

[0079] For each distribution network subsystem i Each constructs a corresponding heterogeneous information subgraph. G i =(V i ,E i ) .in V i ⊆V It is a subsystem i The set of nodes within; E i ⊆E It is a subsystem i The edge set within.

[0080] Each zone has a centralized charging station, and the charging nodes are... v CS,i Create a group of virtual nodes Vgroup i Used for grouping. Each electric vehicle node v EV,i Initially, each electric vehicle is connected to a charging node, representing its charging demand. To depict the spatial distribution and interdependence among electric vehicle groups, virtual nodes are used as links. Groups are formed based on the node attributes of electric vehicle nodes and the edge weights of the electric vehicle-charging station connections. Electric vehicle nodes of the same type are immediately connected to their corresponding virtual nodes in the group, forming a group subgraph.

[0081] Node set V i As shown in Equation 1.

[0082] (1)

[0083] in, VCS i , VEV i Representing a subgraph of a certain group Ggroup i The charging node set and EV flow node set in the middle; Vgroup i Represents a subgraph of a certain group Ggroup i The set of virtual nodes in the group; vQ Optimize group virtual nodes for reactive power; voffpeak i It is a virtual node for off-peak charging groups; vpeak i For peak response group virtual nodes, vbackup i Virtual nodes for backup power groups; vrenewable i This is a virtual node for new energy consumption groups. A group virtual node is a non-existent virtual node used only to connect electric vehicle (EV) nodes within its corresponding group, representing the group relationship between EVs within that group.

[0084] Subsystem i Total number of electric vehicles (EVs) Ω v EV,i It changes as the flow node moves:

[0085] (2)

[0086] Among them, deg( v𝜑 i The ) represents the degree of a virtual node in a group, and its value indicates the group it represents. f The number of internal EVs.

[0087] The set of all edges connecting electric vehicles to charging stations in this zone. E EV-CS,i As shown in Equation 3:

[0088] (3)

[0089] Based on the node weights of electric vehicle (EV) nodes and the movement of each electric vehicle node v EV With charging nodes v CS Based on the edge weights of the connections and the distribution network requirements, electric vehicles with similar characteristics are grouped into groups. Each electric vehicle (EV) node in the same group forms a group edge with a virtual group node, creating a group subgraph that reflects the collaborative optimization of load scheduling and charging time among electric vehicles within the group. Based on real-time demand and grid status, tasks are assigned, allowing vehicles to dynamically switch groups.

[0090] (4)

[0091] (5)

[0092] in, E EVgroup,i Representation Subsystem i All group edges representing clusters within the same area; vgroup i Represents a virtual node in a certain group; GQ It is a reactive power optimization group subgraph; Goffpeak i This is a subgraph of the off-peak charging group; Gpeak i It is a peak response group subgraph; Gbackup i This is a sub-diagram of the backup power supply group; Greenable i This is a sub-diagram of new energy consumption groups; E i = E EV-CS,i ∪ E EVgroup,i This is the edge set of the partitioned subgraph.

[0093] When a certain electric vehicle flow node v EV,i From subsystem i Move to another subsystem j At that time, it and its subsystem i charging nodes v cs,i The connection between them is immediately broken, detaching them from the subsystem. i Scheduling and management, and also with subsystems j charging nodes v cs,j Instantly form connections and participate in subsystems j Cluster-based optimized scheduling.

[0094] 1.3 Joint Modeling of Vehicle-Network Interaction Using Graph Theory and Statistics

[0095] By logically topologicalizing the interaction process between electric vehicles (EVs) and charging stations, and representing the behavioral characteristics of EV users with dynamic flow nodes, the overall impact of EVs on the power distribution network can be more clearly revealed, rather than being limited to the charging behavior of individual EVs. Furthermore, the statistically driven modeling provides underlying computational support on top of the flexibility of graph theory models, ensuring model accuracy and adapting to the needs of complex systems.

[0096] Mobile nodes represent the real-time behavior of electric vehicle users, while static nodes represent the real-time status of charging stations. The weights of mobile nodes include their geographical location, mileage, state of charge (SOC), and the time they begin charging. Data research indicates that daily mileage... d day Approximately following a log-normal distribution, the mileage probability function is obtained through short-time decomposition and assigned as the mileage weight for each mobile node:

[0097] (6)

[0098] in, f ( t () is a normalized time distribution function that describes the proportion of vehicle travel during different time periods of a day; d dayDaily mileage; t Represents a specific time of day; m This represents the expected distance of the trip. m =3.54; s The standard deviation is taken after fitting. s =0.76.

[0099] The State of Charge (SOC) of an electric vehicle battery is defined as the ratio of remaining battery capacity to maximum battery capacity. The SOC weights for each flow node are shown in Equation 7. When the SOC drops to 40%, a charging demand is triggered.

[0100] (7)

[0101] in, D This refers to the maximum range that an electric vehicle can travel on a full charge. d ( t ) is a function for driving distance; t 0 This is the time when the last charge ended. SOC t0 The state of charge of the electric vehicle at the moment the last charge ended.

[0102] The probability density function at the start of charging at the EV mobile node is:

[0103] (8)

[0104] in, t c The moment charging of the electric vehicle begins; t trigger This is the SOC threshold trigger time. s =0.25.

[0105] In addition, use 0-1 variables α fast , β slow These represent the fast and slow charging requirements of the mobile nodes, respectively. α fast + β slow =1. When a mobile node connects to a static node, i.e., when an electric vehicle connects to a charging station for charging, an edge is assigned. E EV-CS The length of the charging time is shown in Equation 9:

[0106] (9)

[0107] in, t chargging Charging time; d This is the charge / discharge coefficient (1 for charging, -1 for discharging). SOC tgt Indicates the target SOC; C b For the battery capacity of electric vehicles; or For charging efficiency; P max and P slow These correspond to the standard charging power for fast charging and slow charging, respectively.

[0108] The charging and discharging power of the flow node can be obtained from Equation 4. P EV for:

[0109] (10)

[0110] Assuming the battery characteristics of the electric vehicles within the subsystem are identical, N m To connect to the charging node m The number of static nodes determines the total charging power of the vehicles. P total for:

[0111] (11)

[0112] 1.4 Active and Reactive Power Joint Optimization Model Based on Vehicle-to-Network Interaction

[0113] In addition to being a load, electric vehicles also possess the characteristics of mobile energy storage, capable of feeding back power to the distribution network via charging stations, thus acting as a power source. Furthermore, vehicle-to-grid interaction has strong grid regulation potential; its spatiotemporal elasticity allows loads to be shifted across time and space through orderly charging and discharging, thereby achieving peak shaving and valley filling, smoothing renewable energy fluctuations, and optimizing grid load and the economic benefits for EV users. Charging stations, through inverters, can compensate for reactive power in the grid, adjusting the reactive power distribution and reducing voltage deviations.

[0114] Therefore, this invention establishes a joint optimization model of active and reactive power based on vehicle-grid interaction to improve the operational stability of the distribution network and the capacity for renewable energy absorption.

[0115] objective function

[0116] 1) The active power optimization objective is to reduce system network losses and improve the load balance of the distribution network by coordinating the charging and discharging behavior of EVs.

[0117] (12)

[0118] in, G l(i,j) It is the firstl The conductivity of the branch circuit; U This refers to the node voltage amplitude. i ij For nodes i and j The voltage phase angle difference between them.

[0119] 2) Reactive power optimization objective: Utilize the power transmission between electric vehicles and charging piles to regulate the distribution of reactive power, thereby optimizing the reactive power regulation capability of the power distribution network and reducing voltage deviation.

[0120] (13)

[0121] in, U i For nodes i The voltage amplitude; Uspec i Set the voltage value for the node.

[0122] 3) Economic performance targets.

[0123] (14)

[0124] (15)

[0125] (16)

[0126] in, C EV It is the cost of charging and discharging electric vehicles; It is the cost of charging services at charging stations; Pch k , Pdis k They represent electric vehicles. k The charging and discharging power; T ch , T dis It refers to the charging and discharging time; l ch , l dis That is the corresponding electricity price; l service It is the charging service fee at the charging station.

[0127] In summary, the objective function is defined as:

[0128] (17)

[0129] in, m P , m Q and m C These are the weight coefficients for each objective function.

[0130] Constraints

[0131] 1) Power flow equation constraint

[0132] (18)

[0133] 2) Voltage constraint

[0134] (19)

[0135] 3) Reactive power output constraints of charging piles

[0136] (20)

[0137] 4) Branch flow constraints

[0138] (twenty one)

[0139] 5) Electric vehicle SOC constraints

[0140] (twenty two)

[0141] in, P i , Q i Representing nodes respectively i The active and reactive power; U i , U j Representing nodes respectively i , j The voltage amplitude; G ij and B ij They are nodes i and j The electrical conductance and susceptance between them; N bus It is a set of distribution network nodes.

[0142] Furthermore, the specific implementation of the series two-stage control model for electric vehicle grouping and power optimization based on the novel power distribution system multi-agent system in step 2 is as follows:

[0143] 2.1 MADDPG-based joint optimization architecture for electric vehicle grouping and distribution network active and reactive power

[0144] This invention utilizes the Multi-Agent Deep Deterministic Policy Gradient Algorithm (MADDPG) to solve the joint optimization problem of electric vehicle clustering and active and reactive power in a distribution network. It employs a centralized training-distributed execution learning framework to construct a serial two-stage decision-making architecture based on MADDPG. For example... Figure 3 As shown, the centralized training-distributed execution framework consists of agents and an environment. The environment is the power distribution network that includes electric vehicles and renewable energy access, while the agents are the electric vehicle schedulers for each subsystem. The state of the power distribution network changes over time. By observing the state information of the power distribution network and electric vehicles, the agents adjust the group allocation and charging behavior of electric vehicles based on the current observation information, thereby optimizing the load distribution and voltage stability of the power distribution network, improving the renewable energy carrying capacity, and reducing the power loss of the system.

[0145] The specific steps are as follows: First, by deploying sensing terminals and dispatch control equipment at various nodes of the distribution network, real-time data on system voltage, current, load power, and other operating status are collected to construct a dynamic monitoring environment. Simultaneously, for EV mobile nodes, key attribute information such as their location, SOC, and energy demand is acquired to establish a foundation for global perception and dynamic modeling of the vehicle-network status. Then, a centralized training method is used to train the electric vehicle grouping strategy using the MADDPG algorithm, sharing training data among multiple agents to determine the optimal action strategy for each agent. Each agent's action involves matching mobile nodes with similar attributes, forming edges between them and corresponding group virtual nodes, and controlling the charging behavior of EV mobile nodes. The goal is to achieve power balance, voltage stability, and effective absorption of new energy sources in the power grid.

[0146] Finally, each agent makes independent decisions during the distributed execution phase based on the trained policy network and adjusts the grouping charging strategy of EV mobile nodes according to real-time observed state information. Through this strategy, the agents can achieve online adaptive grouping to optimize the renewable energy access capacity of the distribution network and the load dispatching of electric vehicles, improve the renewable energy carrying capacity of the distribution network, and ensure the stability and economy of the power grid.

[0147] 2.2 Markov Decision Process Modeling

[0148] This invention models the active and reactive power joint optimization problem based on dynamic grouping of electric vehicles as a Decentralized Partially Observable Markov Decision Process (Dec-POMDP) ​​with multiple agents, and adopts a serial two-stage decision framework. First, dynamic grouping decision is performed; then, power scheduling is optimized based on the grouping results; and finally, the reward is determined by the vehicle-to-grid interaction strategy. Dec-POMDP is typically represented by an octet: ( M,S,O,A, P ,R, , ),in, M It is a collection of intelligent agents; S It is a set of state spaces; O It is a set of joint observation spaces; A It is the set of joint action spaces; P is the state transition probability function determined by the environment; R It is a reward function; It is the probability function of the initial state; (0,1) is a discount factor used to balance the discounting coefficients between immediate and future rewards.

[0149] 1) State and Observation Space: The state set contains global state information of the heterogeneous information graph of the distribution network and vehicle-grid interaction, denoted as...

[0150] (twenty three)

[0151] (twenty four)

[0152] in, S 1 represents the dynamic cluster decision-making state space; S grid For distribution network status information; L ={ p L , q L} represents the set of active and reactive loads at all nodes within the distribution network; P { p pv} represents the set of active power outputs from all photovoltaic systems within the distribution network; U { u} represents the set of voltage amplitudes at all nodes in the distribution network; S EV Status information for the heterogeneous information graph of vehicle-to-everything (V2X) interaction; oh EV Represents the attributes of all EV flow nodes; oh EV-CS This represents the set of edge weights connecting all EV mobile nodes and charging pile static nodes; S 2 represents the state space during the vehicle-to-grid (V2G) interaction and control phase; S cluster The group state space after grouping EVs; G EV ={ g EV} represents the EV flow node clustering status and group characteristics in the heterogeneous information graph of vehicle-to-everything (V2X) interaction; o 1,m and o 2,m Representing intelligent agents respectively m The observational information upon which the first and second stage decisions rely.

[0153] 2) Action Space: To avoid the computational complexity of traditional enumeration-based grouping methods while leveraging the advantages of DDPG in continuous control problems, this invention employs a multi-dimensional discrimination method with joint boundary decision-making in the first decision-making stage. For the grouping strategy, the action space is... a 1 is defined as a dynamic adjustment operation on the cluster boundary to achieve efficient and adaptive optimization.

[0154] (25)

[0155] in, A 1 represents the dynamic grouping decision-making action space; Δ i SOC This indicates an adjustment to the SOC threshold value; Δ i dist This indicates an adjustment to the electrical distance boundary value; Δ i load This indicates an adjustment to the EV load mode boundary value; Δ i c This is to adjust the threshold for charging demand.

[0156] Execute action a After step 1, the cluster boundary changes:

[0157] (26)

[0158] in, i t and i t+1 These are the cluster boundaries before and after the adjustment; i SOC , i dist , i load and i c They represent according to oh SOC The group boundaries are defined based on electrical distance, historical EV load data, and charging demand. After the boundaries are adjusted, EV flow nodes will fall into the corresponding group range according to these boundary parameters, forming connections with the corresponding group virtual nodes.

[0159] Based on the grouping results, the agent formulates corresponding optimization strategies for different groups in the second stage to achieve overall optimized scheduling of the power grid.

[0160] (27)

[0161] in, A 2 represents the action space for the second stage; g This is the group index to which the corresponding EV flow node belongs; g =1 represents the reactive power optimization group, which executes continuous control. P ∈[- P max , P max ], Q ∈[- Q max , Q max Other groups implement discrete charge / discharge strategies: g =2 represents the peak-shaving group; g =3 represents the valley filling group; g =4 represents the new energy consumption group; g =5 represents the backup power group; P max This is for fast charging power; P slow This is the slow charging power. Q max This represents the maximum reactive power transmission.

[0162] 3) State transition probability function: Since the state evolution of the distribution network is affected by the current state and decision-making strategies, the state transition function is used to characterize this dynamic process. In the first stage, the state transition is driven by the cluster boundary adjustment decision, which determines the grouping of electric vehicles; in the second stage, the state transition is affected by the vehicle-grid interaction optimization decision, describing the effect of electric vehicle charging and discharging on the state of the distribution network.

[0163] (28)

[0164] in, S′ This is the state space after the execution of the two-stage decision chain, that is, the state of the system after the second-stage action is executed.

[0165] The state transition probability function quantifies the uncertainty of electric vehicle (EV) group charging behavior and the volatility of distribution network load. The main uncertainties in state transition include fluctuations in EV charging demand, fluctuations in charging station power demand, and uncertainties in distribution network conditions. During model training, the state transition relationship considers the mutual influence between distribution network load and EV charging behavior.

[0166] 4) Reward Function: In designing the reward function for the intelligent agent, this invention comprehensively considers three control objectives: distribution network voltage stability, charge balance, and economic performance. Based on the established active and reactive power joint optimization model based on vehicle-grid interaction, the reward function is designed as follows:

[0167] (29)

[0168] 2.3 MADDPG Agent Network Structure Construction

[0169] Based on the Markov decision process modeling described above, the evaluation network and action network of the MADDPG agent are constructed. The specific steps for solving the joint optimization problem using the constructed MADDPG agent network are as follows: Figure 5 As shown.

[0170] (2.3.1) Evaluation of the network component

[0171] Under the proposed dynamic electric vehicle (EV) clustering architecture, an evaluation network based on the MADDPG algorithm is constructed to assess the merits of decision-making actions of each subsystem agent under different states, thereby guiding the selection of EV clustering strategies. Then, during the vehicle-to-grid (V2G) interaction and control phase, the evaluation network further optimizes the power scheduling strategies of each subsystem agent by considering the EV's dispatchability, the power grid's operating status, and economic objectives. Through joint training, this network ensures the coordination between the clustering strategy and the V2G interaction strategy, enabling the EV group to not only adapt to load fluctuations but also improve renewable energy absorption capacity, enhance voltage stability, and reduce system operating costs through reasonable power regulation.

[0172] The evaluation network input for the agent contains relevant information from two stages. First, the input for the clustering stage includes the observation status of the heterogeneous information graph of the local environment of the distribution network and the vehicle-grid interaction. o 1. and the actions of electric vehicle grouping strategies. a 1. Then, the input for the vehicle-to-grid (V2G) interaction control phase is the local environmental observation state after clustering. o 2. And the scheduling actions during the vehicle-to-grid interaction and control phase. a 2. The output is the state-action value function, which measures the merits of the action strategy. Q θ ( o 1, a 1 ,o 2, a 2) The estimated values, the parameters to be optimized are the evaluation network parameter φ and the target evaluation network parameter φ. ′ Each agent's evaluation network consists of a main evaluation network and a target evaluation network. The target evaluation network's role is to improve the stability and convergence of training.

[0173] The loss function of the MADDPG algorithm is shown in Equation 30. The two-stage loss functions are no longer separated but directly evaluate the entire decision chain, resulting in more stable training. The loss function is used to quantify the value of network evaluation. Q θ ( o 1, a 1 , o 2, a 2) and expected cumulative rewards y m The difference between them, including the expected cumulative reward y m The result is obtained using the Bellman equation, as shown in Equation 31. The network is evaluated by minimizing the loss function. L (A) Iterative updates are performed. Specifically, during the centralized training phase, based on the observed state of the distribution network and the action instructions of the electric vehicle grouping strategy, the merits of the current action of the strategy are measured, and the accuracy and stability of the action evaluation are continuously improved. The evaluation network parameter update process is as follows:

[0174] (30)

[0175] (31)

[0176] (32)

[0177] in, r m It is an intelligent agent m Perform the first phase actions sequentially. a 1,m With the second phase of action a 2,m The instant reward that the system receives afterwards; o′ and a′ These represent the state and action at the next moment after the two phases; c It is a discount factor; Q θ’ It is the value function of the target network; α It is the learning rate used to evaluate the network and control the step size for each parameter update; i To evaluate network parameters; ∇ 𝜃 L (𝜃) represents the gradient of the loss function with respect to the parameter 𝜃. By calculating the gradient and updating the parameters in the opposite direction of the gradient, the loss function can be gradually reduced, thereby optimizing the model.

[0178] The MADDPG algorithm introduces an experience replay mechanism in the evaluation network, enabling each agent to evaluate the state-action function. Q θ (o 1 ,a 1 ,o 2 ,a 2) The estimation process can incorporate empirical data accumulated during training, thereby enhancing the understanding of dynamic changes in the distribution network. This mechanism helps the agent to optimize decisions more stably and improve the effectiveness of dynamic grouping and power optimization scheduling of electric vehicles.

[0179] To achieve effective information coupling between structural decision-making (EV flow node grouping) and strategy optimization (vehicle-to-grid interaction control strategy), this invention designs a two-stage experience replay mechanism based on structural response labels to adapt to a serial two-stage deep reinforcement learning architecture. For example... Figure 4 As shown, in this architecture, the first stage is the structural decision stage, which is responsible for building the dynamic grouping structure of electric vehicles; the second stage is the strategy optimization stage, which performs active and reactive power coordination control on the grouped electric vehicles based on the group structure to optimize the operation of the power distribution network.

[0180] Traditional multi-stage learning methods often decouple structure partitioning from behavioral decision-making during training, making it difficult for behavioral strategies to perceive structural information, and structure partitioning cannot obtain behavioral outcome feedback. To address this, this invention introduces structured response labels into the experience replay pool to achieve the following key characteristics:

[0181] The storage format of the experience pool in the structural decision-making stage is shown in Equation 33:

[0182] (33)

[0183] in, o 1 i These are the observations of each agent in the EV clustering phase; G t Grouping for the current time step; F ( G t ) represents the global performance response of the group structure under the second-stage control.

[0184] The storage format of the experience pool during the strategy optimization phase is shown in Equation 34:

[0185] (34)

[0186] in, o 2 i and a 2 i These refer to the observations and actions of various intelligent agents in the vehicle-to-grid (V2G) interaction control phase. r i The joint reward at the end of the two phases is linked together for each agent; s′ i This represents the state at the end of the two-stage decision-making process.

[0187] By introducing structural response labels, experiences from the structural decision-making phase can carry cluster context information, thereby enabling cross-phase information transfer. Replaying experiences not only allows for group-based filtering but also avoids information loss caused by complete decoupling of structure and strategy. While ensuring independent optimization capabilities between modules, it achieves soft-coupled replay collaboration between structure and behavior.

[0188] (4.3.2) Action Network Part

[0189] In the proposed online optimization architecture for dynamic grouping of electric vehicles, an action network is built based on the MADDPG algorithm to evaluate the effectiveness of different scheduling strategies, thereby assisting in optimizing the scheduling of electric vehicle charging stations within the subsystem. During the training phase, the electric vehicle manager learns and optimizes action strategies using the current environmental observation state. During the execution phase, the manager issues scheduling instructions to charging stations within the subsystem based on the output of the action network, dynamically adjusting the electric vehicle grouping strategy and optimizing the power allocation of the distribution network.

[0190] In the cascaded two-stage decision-making framework of the MADDPG algorithm, the agent m The action network is divided into two phases. The first phase involves local observations of the input. o 1,m This includes local observations of the power distribution network and related state information of electric vehicles, outputting the grouping actions of the electric vehicles. a 1,m The second stage involves inputting local observations. o 2,m The input consists of the clustering results from the first stage and the further observed environmental conditions; the output is the electric vehicle scheduling action. a 2,m The parameters to be optimized are the action network parameters. ϕ With target action network parameters ϕ′ Each agent's action network incorporates a target network and a policy entropy term to promote policy stability and diversity.

[0191] The MADDPG algorithm uses the policy gradient method to train the action network, with each agent iteratively optimizing its own action network parameters based on gradient descent. ϕ The policy gradient function is shown in Equation 35. The two-stage policy gradients are no longer calculated separately, but are directly jointly optimized across the entire decision chain. During the centralized training phase, each electric vehicle manager continuously updates its electric vehicle scheduling strategy based on the current environmental observation status and the load demand information of electric vehicles within the subsystem, moving towards better electric vehicle grouping and distribution network load optimization. The action network parameter update process is as follows:

[0192] (35)

[0193] (36)

[0194] (37)

[0195] Among them, ∇ ϕ J ( ϕ ) indicates strategy π ϕ Regarding parameters ϕ The gradient is used to optimize the reward function. J ( ϕ ); t It is a state-action sequence trajectory; ∇ ϕ log π ϕ ( a 1 |o 1) Indicates the first-stage strategy π ϕ In state o 1. Select action a The gradient of the logarithmic probability of 1; ∇ ϕ log π ϕ ( a 2 |o 2) Indicates the second-stage strategy π ϕ In state o 2 Select Actions a The gradient of the logarithmic probability of 2; Q θ ( o 1 ,a 1 ,o 2 ,a 2) is the state-action value function, representing the state... o 1. Select action a 1 and reach the state o 2. Select action after a Expected return of 2; H ( π ϕ ( a 1 |o 1)) represents the policy entropy term in the first stage; l It is the entropy weighting coefficient, used to adjust the balance between exploration and utilization; β It is the learning rate of the action network.

[0196] (4.3.3) Target network parameter update section

[0197] For updating the target network parameters, a soft update mechanism is adopted to improve the stability of training and reduce the errors generated by the interaction between the agent and the environment.

[0198] (38)

[0199] (39)

[0200] in, e It is a very small constant used to control the rate at which the target network parameters are updated. In this way, the parameter changes of the target network are smoother, thus effectively avoiding training instability caused by fluctuations in the target Q-value during training, thereby improving the convergence and robustness of the model.

[0201] Furthermore, the specific implementation of the distribution network renewable energy carrying capacity modeling in step 3 is as follows:

[0202] Traditional concepts of renewable energy carrying capacity emphasize the maximum capacity of renewable energy connected to distributed generation (DG) at the system level, seeking the globally optimal renewable energy acceptance capacity at the system level. However, they are significantly weak in characterizing the carrying capacity margin of multiple specific DG grid-connected nodes. Due to regional differences and local constraints in the power grid, relying solely on system-level assessments is insufficient to accurately reflect the actual carrying capacity of each node.

[0203] While system-level methods can assess the carrying capacity boundary of the entire network, they are difficult to refine down to individual nodes and cannot directly guide the addition of local renewable energy capacity. To address this, this invention focuses on identified renewable energy grid-connected nodes in radial distribution networks with known topologies. It comprehensively considers factors such as time-series variations, voltage control, and local power flow distribution to establish a node-level renewable energy carrying capacity margin assessment model.

[0204] 3.1 Evaluation Indicators for the Carrying Capacity of New Energy Sources at the Node Level in New Power Systems

[0205] To verify the effect of the proposed strategy on the renewable energy carrying capacity of the distribution network, this invention establishes a multi-dimensional spatial dynamic evaluation model of the renewable energy carrying capacity of the distribution network for multiple determined photovoltaic grid connection points, and characterizes the maximum photovoltaic access margin boundary composed of multiple nodes.

[0206] (40)

[0207] in, x For the set of decision variables; Δ SPV* i ( t )for t Time Node i The per-unit value of the photovoltaic capacity regulation amount connected to the grid.

[0208] The constraints include node voltage constraints, branch power flow constraints, and reverse power flow constraints:

[0209] (41)

[0210] (42)

[0211] (43)

[0212] in, P reverse ( t )and P limit They represent t The reverse power at any given moment and the maximum allowable reverse power flow of the system.

[0213] The multidimensional spatial model characterizes the maximum capacity margin boundary for photovoltaic grid connection at a given location, with the expansion or contraction of the boundary representing changes in the system's carrying capacity. Taking m=3 as an example, the change in the three-dimensional volume enclosed by the boundary of the corresponding three-dimensional spatial model is calculated. If the volume increases, it indicates that the boundary has expanded and the carrying capacity has increased; conversely, the boundary has contracted and the carrying capacity has decreased.

[0214] However, as power distribution networks become increasingly complex and the number of renewable energy sources connected to the grid gradually increases, the dimensionality of the model also increases. Since the hypervolume enclosed by the boundaries of high-dimensional space is difficult to solve, it is difficult to intuitively represent the Pareto front under the mutual constraints of the carrying capacity among all nodes when the number of renewable energy grid-connected nodes exceeds three.

[0215] To address this, high-dimensional models can be reduced in dimensionality through uniform manifold approximation and projection. That is, the distribution of boundary points in high-dimensional space is mapped to low-dimensional space after passing through UMAP, thereby reducing the difficulty of visualization and representation.

[0216] UMAP is based on a key assumption that high-dimensional data is actually distributed on a low-dimensional manifold. First, for each data point, its neighborhood is calculated (the neighborhood is determined by a parameter). n neighbors (Control), and then use the exponential decay function to define the connection weights between points to form a fuzzy topology.

[0217] (44)

[0218] (45)

[0219] in, X and Y These are sets of points in high-dimensional space and low-dimensional space, respectively. l and c These are the dimensions of the high-dimensional space and the low-dimensional space, respectively. NThe number of points in space.

[0220] In higher-dimensional space, a point I and points J The similarity is defined as:

[0221] (46)

[0222] in, d ( x I ,x J () represents the distance between points; r i It is a point I The distance to its nearest neighbor; e i It is an adaptive bandwidth parameter that ensures that the neighborhood of each point contains approximately n neighbors One point.

[0223] In low-dimensional space, a point I and points J The similarity is defined as:

[0224] (47)

[0225] in, a and b These are hyperparameters, which are automatically optimized by minimizing the loss function.

[0226] By minimizing the cross-entropy between the high-dimensional and low-dimensional similarity distributions, the low-dimensional representation preserves the high-dimensional topology as much as possible:

[0227] (48)

[0228] By preserving the nonlinear results in the high-dimensional space through UMAP, the projected low-dimensional boundary points map the carrying capacity boundary characteristics of the original data. To facilitate data analysis and visualization, this invention reduces the dimension of the multi-dimensional dynamic evaluation model of the distribution network's renewable energy carrying capacity to three dimensions for network renewable energy carrying capacity assessment and calculation of the dynamic carrying capacity boundary.

[0229] Calculate the volume of the geometry enclosed by each discrete-time carrying capacity boundary map, and define the average value as the new energy carrying capacity margin. x :

[0230] (49)

[0231] in, T For time-series discrete quantities; f t ( s 1,s 2, s 3) Indicates the result after dimensionality reduction t A boundary diagram of the carrying capacity of new energy sources at any given time.

[0232] Example 2

[0233] As a specific implementation example of the present invention, this embodiment uses a 113-node system in a certain location for simulation to verify the effectiveness of the proposed method and strategy in improving the renewable energy margin at the node level of the distribution network, optimizing voltage, and reducing network losses. The topology, zoning, and photovoltaic grid-connection locations of the distribution network are as follows: Figure 6 As shown. To ensure the realism and diversity of the operating scenarios, this embodiment uses real load and historical photovoltaic output data, with a time granularity of half an hour and a total duration of three days. Considering that the training data is historical data, this invention incorporates random noise perturbation during the training process to enhance the proposed model's ability to explore unknown environments. In the example, the upper and lower limits of the voltage amplitude are set to 0.95 pu and 1.05 pu, respectively. When power flow non-convergence occurs, a fault is identified in the distribution network, the loop is terminated, and a large negative number is assigned to the reward function.

[0234] Table 1 Algorithm Parameters

[0235]

[0236] The power distribution network includes 4,000 EVs. Considering the heterogeneity of different vehicle models, the battery capacity of the EVs is set to 20–60 kWh, covering the capacity range of mainstream passenger electric vehicles. Based on typical operating conditions, the maximum charging power per vehicle is set to 7 kW in slow charging mode and 40 kW in fast charging mode. The charging mode of the EVs is switched on demand, and the actual charging process is jointly regulated by factors such as the current grid status and the user's travel flexibility. The training process and power flow calculation of the proposed control method are performed in Python 3.9 with the PyTorch deep learning framework. Specific algorithm parameter configurations are shown in Table 1.

[0237] 1. Training effect analysis

[0238] To comprehensively evaluate the effectiveness of the proposed method, simultaneous comparative experiments were conducted. Baseline algorithms included the MAAC algorithm with an attention mechanism and the MADDPG algorithm with a dual-policy decoupling structure. The figure below shows the performance of each algorithm during training. The light-colored portion represents the standard deviation band, used to approximately represent the fluctuation range of reward values ​​in five independent training iterations; the dark line is the moving average curve of the reward values, used to smooth the volatility of rewards in a single round. During training, the reward values ​​fluctuated within a certain range and gradually converged to a certain value, indicating that the agent could effectively converge after a set number of iterations. Due to the uncertainty of the environmental state, the reward values ​​also fluctuated within a certain range.

[0239] Depend on Figure 7 As can be seen, the method proposed in this invention outperforms the MAAC algorithm in both convergence speed and final performance. The MADDPG algorithm, which uses a non-contiguous two-stage decision framework (i.e., a dual-policy decoupled structure), converges faster than the method in this invention due to the introduction of parallel optimization and modular multi-task learning. However, it suffers from greater fluctuations after convergence and significantly lower overall stability due to the weak coupling and insufficient policy coordination between the EV grouping and vehicle-to-network interaction stages. The method proposed in this invention simplifies model complexity by using a contiguous two-stage decision framework, enhances the policy correlation between stages, reduces performance fluctuations during training, and improves the final average reward value. In summary, the method in this invention demonstrates superior performance in both convergence and stability, validating its effectiveness and superiority in complex tasks.

[0240] 2. Analysis of the regulatory effect

[0241] To verify the effectiveness and superiority of the proposed strategy in active-reactive power optimization of distribution networks, this invention uses 100 consecutive days of distribution network and charging station data as a training set, and another 3 consecutive days of data as a test set to evaluate the comprehensive performance of the proposed strategy. This invention aims to develop a scheduling strategy based on dynamic clustering of electric vehicle groups participating in distribution network interaction optimization. It temporarily disregards individual differences in user behavior responses, assuming that vehicles receive unified control within the allowable scheduling range, in order to evaluate the potential of the clustering mechanism to improve system optimization.

[0242] 2.1) Load regulation effect of the test set

[0243] To analyze the distribution characteristics of the load in the time and space dimensions, a heat map of the active power load of the nodes throughout the time series was plotted, as shown below. Figure 8 and Figure 9As shown in the figure. The results indicate that after dynamic clustering optimization of electric vehicles, the overall load level of the nodes that were originally high-loaded decreased, and the color of the load heatmap became significantly lighter. From a temporal perspective, the load intensity during the morning and evening peak hours was significantly alleviated after optimization, and the peak color weakened, while the load level increased during periods of lower load, such as midday, and the image height rose. This phenomenon shows that the proposed strategy achieves a partial redistribution of load at some nodes in space, and effectively smooths peaks and fills valleys in time, demonstrating a strong ability to regulate load in the spatiotemporal migration.

[0244] 2.2) Voltage regulation effect of the test set

[0245] Figure 10 This paper presents a comparison of distribution network voltage distribution under the proposed EV dynamic clustering strategy, EV static clustering, and unordered EV charging scenarios. In the figure, the upper, middle, and lower surfaces correspond to the EV dynamic clustering strategy, the EV static clustering strategy, and the unordered charging scenario before optimization, respectively. To more intuitively compare the voltage trends over time and nodes under different strategies, this invention introduces a vertical visual offset to the voltage surfaces of each scheme in the 3D visualization. This processing does not change the original data; it only reduces the impact of surface overlap on information display, thereby improving image clarity and contrast.

[0246] from Figure 10 It can be seen that the EV dynamic grouping strategy proposed in this invention significantly improves the level of voltage dynamic change, and the improvement in the curvature of the voltage surface is particularly obvious in the distribution network end node range, which is better than the performance before optimization and under the EV static grouping strategy. During the evening peak period (17:00–20:00), for the regional nodes (70–79 and 98–113) at the end of the distribution network, the local average voltage under the disordered charging scheme is only 0.9597 pu, which shows a significant problem of low voltage. After adopting the static grouping strategy, the local average voltage is increased to 0.9610 pu, and the dynamic grouping strategy proposed in this invention further increases it to 0.9618 pu. Overall, the global average voltage of the system is improved from 0.9884 pu before optimization to 0.9910 pu after static grouping optimization, and further to 0.9917 pu under the method of this invention, which fully demonstrates the effectiveness of the proposed strategy in terms of voltage improvement and voltage stability assurance.

[0247] To more clearly highlight the optimization effect, the optimization results for one day were extracted and plotted separately, allowing for a more intuitive observation of the voltage level improvement. For example... Figure 11As shown, the optimized system did not exhibit voltage exceedance issues across the entire timeframe and all nodes. Compared to the minimum voltage of 0.9460 pu before optimization and the minimum voltage of 0.9502 pu under the static clustering strategy, the proposed dynamic clustering strategy effectively increased the minimum voltage to 0.9519 pu, further validating its advantages in ensuring the voltage stability of the distribution network.

[0248] Before optimization, the disorderly charging of electric vehicles further amplified peak loads, leading to voltage exceeding limits at the end nodes of the distribution network, with voltage levels falling below the normal range. Although static clustering strategies can smooth out peaks and valleys to some extent, their fixed group divisions limit flexibility and scheduling reserves, making it difficult to continuously adapt to dynamic load changes, resulting in insufficient optimization effects. In contrast, the dynamic clustering method proposed in this invention can more fully leverage system flexibility and significantly improve voltage levels and operational safety.

[0249] 2.3) Active power loss control effect of the test set

[0250] In the optimization of power distribution network operation, active power loss is one of the important indicators for measuring system operating efficiency. Addressing new challenges such as dynamic load fluctuations and electric vehicle (EV) integration, this invention verifies the practical value of the proposed method in suppressing network losses through comparative analysis. The average daily network losses under three strategies—disordered EV charging, static clustering, and dynamic clustering—are 0.1258 pu, 0.1218 pu, and 0.1206 pu, respectively. To comprehensively evaluate the effects of different optimization strategies at different time periods, this invention selects typical time segments every 3 hours from three days of dynamic simulation results, extracting the corresponding active power losses for comparative analysis.

[0251] like Figure 12 As shown, the curves corresponding to "disordered charging" are distributed on the outermost edge, "static clustering" is in the middle, and the curve corresponding to "dynamic clustering" is basically enveloped by the former two, located on the innermost edge, demonstrating superior network loss control capability. 6:00 and 18:00 correspond to the morning and evening peak load periods, respectively, when active power loss levels are high, making the optimization strategies particularly effective. At the 6:00 time segment, the average active power loss for the three strategies is 0.1674 pu, 0.1344 pu, and 0.1189 pu, respectively. The dynamic clustering strategy reduced network loss by 28.97% compared to disordered charging, while the static clustering strategy reduced loss by 19.71%. Similarly, at the 18:00 time segment, the average active power loss for the three strategies is 0.1932, 0.1683, and 0.1397, respectively. The dynamic clustering strategy reduced network loss by 27.69% compared to disordered charging, while the static clustering strategy reduced loss by 12.89%.

[0252] During off-peak hours at night (0:00~5:00), the differences in network losses among the three strategies tend to be consistent due to a significant reduction in charging activity by electric vehicle users. In addition, the dynamic clustering strategy, by flexibly adjusting charging timing and shifting the load to off-peak periods in some times, although causing slightly higher instantaneous network losses than disordered charging in certain periods, performs better overall.

[0253] The results show that during typical peak load periods, the strategy proposed in this invention can effectively regulate the load of charging stations, thereby achieving a good match with the load demand of the distribution network, and thus significantly improving the load balance and system stability of the distribution network.

[0254] 3. Analysis of Node-Level New Energy Margin Improvement

[0255] As shown in the analysis of the regulation effect, the EV dynamic grouping strategy proposed in this invention has significant advantages over disordered charging and static grouping in terms of load spatiotemporal migration regulation, network loss reduction, and voltage support under the background of high proportion of new energy access to the distribution network. Therefore, it can be seen that the flexibility and adaptability of the dynamic grouping strategy give it the potential to globally smooth out the fluctuations of new energy, making up for deficiencies and reducing excesses.

[0256] To further evaluate the performance of the strategy proposed in this invention in terms of node-level renewable energy capacity addition capability, the maximum grid access margin boundary of photovoltaic power composed of multiple nodes is solved. A Pareto front was randomly selected from each of the photovoltaic maximum access margin boundaries solved under the unordered charging and EV dynamic clustering strategies for comparison. Under the unordered charging scenario, the performance vector for the addition of renewable energy capacity at each node is: [0.4629, 0.6538, 0.9947, 0.8164, 0.6739, 0.6720, 0.9968, 0.2608, 0.6513, 0.5248, 0.7927, 0.9217, 0.3128, 0.9705, 0.2567]; under the EV dynamic clustering strategy, the performance vector for renewable energy capacity margin is: [0.8461, 1.5802, 0.4820, 1.6298, 0.4608, 0.5107, 1.5541, 1.2077]. [0.5549, 0.4401, 1.4187, 0.6107, 0.4061, 0.7132, 0.7389]. The performance vectors show that, under the condition that the photovoltaic capacity of each node is mutually constrained, the maximum photovoltaic capacity margin in the Pareto front of disordered charging is 0.9947 pu, and the total margin of all nodes is 9.9618 pu; while in the Pareto front of the dynamic clustering strategy scenario, the maximum photovoltaic capacity margin reaches 1.6298 pu, and the total margin of all nodes is 13.1540 pu. The maximum photovoltaic capacity that can be added to a single node is increased by 63.85% compared to before optimization, and the total margin is increased by 32.04% compared to before optimization.

[0257] Figure 13 After UMAP dimensionality reduction, the photovoltaic capacity margin solution set for the disordered charging scenario is mapped to a feasible domain graph in a low-dimensional space. The Pareto front constitutes the boundary of the entire feasible solution space. The calculated volume of the geometric volume enclosed by the bearing capacity boundary graph is 3.2272.

[0258] Figure 14 The figure shows the dimensionality-reduced feasible solution space for photovoltaic capacity margin in the EV dynamic clustering strategy scenario. The orange-yellow points represent the dimensionality-reduced feasible solution set points in the disordered charging scenario. As shown in the figure, the photovoltaic capacity margin boundary of the strategy proposed in this invention completely encloses the orange-yellow points. The calculated volume of the geometric volume enclosed by the capacity boundary map is 16.7829, which is 520.05% of the volume enclosed by the capacity margin boundary in the disordered charging scenario. The node-level new energy carrying capacity margin of the distribution network is significantly improved.

[0259] This invention addresses the problems of joint active and reactive power optimization and node-level renewable energy carrying capacity assessment in distribution systems with large-scale renewable energy integration. It proposes an adaptive dynamic clustering strategy for electric vehicles based on a graph theory-statistics joint-driven modeling approach using multi-agent deep reinforcement learning, as well as a node-level renewable energy carrying capacity assessment index. Based on numerical examples, the following conclusions can be drawn:

[0260] 1) This invention establishes a vehicle-to-grid interaction model driven by both graph theory and statistics. By logically topologicalizing the interaction process between EVs and charging piles, and using dynamic flowing nodes to represent the behavioral characteristics of EV users, it can more clearly reveal the overall impact of EVs on the power distribution network, rather than being limited to the charging behavior of individual electric vehicles. Simultaneously, the statistical model, based on the flexibility of the graph theory model, provides accurate support for underlying calculations, adapting to the needs of complex systems.

[0261] 2) The MADDPG algorithm adopts a learning framework of "centralized training-distributed execution" and a two-stage decision architecture. At the same time, it designs a two-stage experience playback mechanism based on structural response labels to realize the coupling of dual strategies in the EV dynamic grouping stage and the vehicle-network interaction control stage after grouping. By maximally simulating the global state of the real distribution network environment and the heterogeneous information graph of vehicle-network interaction, it improves the EV dynamic grouping effect, smooths the voltage fluctuation of nodes within the distribution network, improves the absorption of new energy, and reduces network losses.

[0262] 3) This invention establishes a multi-dimensional spatial dynamic evaluation model for the renewable energy carrying capacity of a distribution network system under multi-source constraints. It reduces the dimensionality of a large number of high-dimensional feasible points using UMAP and employs the three-dimensional convex hull volume as a representation of carrying capacity to characterize the maximum grid connection margin boundary of photovoltaic systems at the node level. The results show that the dimensionality-reduced three-dimensional convex hull volume can reflect the potential carrying capacity upper limit of the system under different strategies, and has certain quantitative analysis value.

[0263] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0264] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.

Claims

1. A graph theory-statistical joint modeling-driven dynamic grouping control method for electric vehicles, characterized in that, include: A bidirectional vehicle-to-network interaction model driven by graph theory and statistics is constructed by integrating graph theory and statistical characteristics. The distribution network is divided into subsystems, and a two-stage serial decision architecture is constructed based on the multi-agent deep deterministic policy gradient algorithm, using a centralized training and distributed execution framework, to solve the joint optimization problem of electric vehicle grouping and active and reactive power of the distribution network. In the first stage of the two-stage cascaded decision-making architecture, electric vehicles in each subsystem are dynamically and adaptively grouped in real time; in the second stage, the vehicle-network active and reactive power joint optimization control is carried out based on the grouping results. A multi-dimensional spatial dynamic evaluation model is constructed for photovoltaic grid-connected points. The carrying capacity margin is quantified by solving the high-dimensional geometric volume enclosed by the boundary of the evaluation model, thereby determining the maximum access capacity margin boundary of the photovoltaic grid-connected point.

2. The method for dynamic grouping control of electric vehicles driven by graph theory-statistical joint modeling according to claim 1, characterized in that, The graph theory-statistics jointly driven vehicle-to-grid bidirectional interaction model uses graph theory to characterize the structured relationship of the interaction effect between electric vehicles and charging piles, establishing a topological model of vehicle-to-grid interaction. Combined with statistical modeling, it calculates and analyzes the temporal-spatial distribution and interdependencies among electric vehicle groups for each node. The specific model construction is as follows: Electric vehicles and charging stations are defined as mobile nodes and static nodes, respectively. Mobile nodes represent the real-time behavior of electric vehicle users, and static nodes represent the real-time status of charging stations. Each electric vehicle forms a virtual connection with the charging station of its subsystem to describe the charging and discharging needs of the electric vehicle. Based on the node attributes of electric vehicle nodes, the edge weights of the connections between each electric vehicle mobile node and the charging node, and the power distribution network requirements, electric vehicles with the same characteristics are divided into groups. Each electric vehicle node in the same group forms a group connection with the corresponding virtual group node, thus forming a group subgraph. A set of virtual nodes is created for group division. The virtual nodes are used to connect the electric vehicle nodes in their corresponding groups, representing the group relationship of electric vehicles within the group. The group subgraph includes a reactive power optimization group subgraph, an off-peak charging group subgraph, a peak response group subgraph, a backup power supply group subgraph, and a new energy consumption group subgraph. Electric vehicles dynamically switch their groups based on real-time demand and grid status. For each distribution network subsystem, a corresponding heterogeneous information subgraph is constructed. In the heterogeneous information subgraph, the node set within the subsystem includes: the charging node set, the electric vehicle mobile node set, and the group virtual node set; the edge set within the subsystem includes: the set of edges connecting electric vehicles and subsystem charging stations, and all group edges representing clusters within the subsystem. The weights of mobile nodes include their geographical location, mileage, battery state of charge, and the start time of charging; the mileage weight uses the mileage probability function; the battery state of charge weight is the state of charge at the end of the last charge of the electric vehicle minus the quotient of the mileage function and the maximum mileage that the electric vehicle can travel in a fully charged state; the start time of charging weight uses the probability density function of the start time of charging. When an electric vehicle connects to a charging station for charging, the charging duration of the connected edge is assigned as the edge weight.

3. The method for dynamic grouping control of electric vehicles driven by graph theory-statistical joint modeling according to claim 1, characterized in that, For the aforementioned distribution network, based on the grid load matching relationship between the charging station and other distribution network nodes, and taking into account charging demand, power flow, and electrical distance factors, the subsystem is divided as follows: Using the impedance modulus as the electrical distance, and taking each charging station node as the root node, the shortest electrical distance between each node and each root node is calculated. Considering the load demand and power generation capacity of each node, the power flow between each node is calculated using the optimal power flow method. Using the shortest electrical distance, node geographic coordinates, and power flow as the basis for clustering, each charging station node is used as the initial centroid of the cluster, and the distribution network subsystem is divided through clustering algorithm.

4. The method for dynamic grouping control of electric vehicles driven by graph theory-statistical joint modeling according to claim 1, characterized in that, The objective function of the combined active and reactive power optimization problem of the distribution network electric vehicle grouping includes an active power optimization objective, a reactive power optimization objective, and an economic performance objective. The active power optimization objective reduces system network losses by coordinating the charging and discharging behavior of electric vehicles. The reactive power optimization objective regulates the distribution of reactive power by utilizing the power transmission between electric vehicles and charging piles. The economic performance objective considers the charging and discharging costs of electric vehicles and the charging service costs of charging stations.

5. The method for dynamic grouping control of electric vehicles driven by graph theory-statistical joint modeling according to claim 1, characterized in that, In the two-stage serial decision architecture based on the multi-agent deep deterministic policy gradient algorithm, the environment is a power distribution network that includes electric vehicles and new energy access, and the agent is the electric vehicle scheduler of each subsystem. The state of the power distribution network changes over time. The agent observes the state information of the power distribution network and electric vehicles, and adjusts the group allocation and charging behavior of electric vehicles according to the current observation information. Real-time collection of operational status data from various nodes of the power distribution network, and acquisition of key attribute information of electric vehicles; A centralized training method is adopted, and the electric vehicle grouping strategy is trained by a multi-agent deep deterministic policy gradient algorithm. Training data is shared among multiple agents to determine the optimal action strategy for each agent. The agent's action is to match mobile nodes with similar attributes, form connections between them and corresponding group virtual nodes, and control the charging behavior of electric vehicle mobile nodes. The goal is to achieve power balance, voltage stability, and effective absorption of new energy in the power grid. During the distributed execution phase, each agent makes independent decisions based on the trained policy network and adjusts the group charging strategy of the electric vehicle mobile nodes according to the real-time observed state information.

6. The method for dynamic grouping control of electric vehicles driven by graph theory-statistical joint modeling according to claim 1, characterized in that, The multi-agent deep deterministic policy gradient algorithm constructs an evaluation network and action network for the agents based on an observable Markov decision process, and adopts a serial two-stage decision framework; in the first stage, the agents perform dynamic grouping decisions, and in the second stage, power scheduling is optimized based on the grouping results; in the observable Markov decision process: The state space includes the dynamic clustering decision-making state space and the state space of the vehicle-to-grid interaction control stage. The dynamic clustering decision-making state space includes the state information of the distribution network state information and the state information of the heterogeneous information graph of vehicle-to-grid interaction. The state space of the vehicle-to-grid interaction control stage includes the group state space after the electric vehicles are grouped. The joint observation space includes the observation information upon which the agent's first and second-stage decisions depend; In the joint action space, a multi-dimensional discrimination method for joint boundary decision-making is adopted in the first decision-making stage. For the grouping strategy, the action is defined as a dynamic adjustment operation on the group boundary, including the adjustment of the SOC boundary value, the electric distance boundary value, the EV load mode boundary value, and the charging demand boundary value. In the second stage, corresponding optimization strategies are formulated for different groups. When the electric vehicle flow node belongs to the reactive power optimization group, continuous control is executed. When the electric vehicle flow node belongs to other groups, a discrete charging and discharging strategy based on fast charging power and slow charging power is executed. The state transition probability function, determined by the environment, is used to characterize the dynamic process of the state evolution of the distribution network being affected by the current state and decision-making strategy; In the first stage, state transitions are driven by cluster boundary adjustment decisions and the environment, characterizing the impact of cluster boundary adjustments on electric vehicle grouping; in the second stage, state transitions are affected by vehicle-grid interaction optimization decisions, describing the effect of electric vehicle charging and discharging on the distribution network state. The reward function of the intelligent agent comprehensively considers three control objectives: distribution network voltage stability, charge balance, and economic performance. A discount factor is set to balance the conversion coefficients of immediate rewards and future rewards.

7. The method for dynamic grouping control of electric vehicles driven by graph theory-statistical joint modeling according to claim 6, characterized in that, The evaluation network of the multi-agent deep deterministic policy gradient algorithm agents is used to evaluate the merits of the decision-making actions of each subsystem agent under different states, thereby guiding the selection of electric vehicle grouping strategies. In the second stage, the evaluation network combines the dispatchability of electric vehicles, the operating status of the distribution network, and economic objectives to optimize the power scheduling strategies of each subsystem agent. In the cascaded two-stage decision-making framework of the multi-agent deep deterministic policy gradient algorithm, the first-stage input of the evaluation network includes the observation state of the heterogeneous information graph of the local environment of the distribution network and the interaction between the vehicle and the network, as well as the action of the electric vehicle grouping strategy; the second-stage input is the observation state of the local environment after grouping and the scheduling action of the second stage; the output of the evaluation network is the state-action value estimate that measures the merits of the action strategy; the parameters to be optimized are the evaluation network parameters and the target evaluation network parameters. The loss function of the evaluation network directly evaluates the entire decision chain, and iteratively updates the evaluation network parameters by minimizing the difference between the estimated value of the output and the expected cumulative reward.

8. The method for dynamic grouping control of electric vehicles driven by graph theory-statistical joint modeling according to claim 7, characterized in that, The multi-agent deep deterministic policy gradient algorithm introduces a two-stage experience replay mechanism based on structural response labels in the evaluation network part. Each agent combines the experience data accumulated during training in the state-action evaluation estimation process. The first stage of the dual-stage experience replay mechanism is the structural decision stage, which is responsible for constructing the dynamic grouping structure of electric vehicles. The experience pool of the structural decision stage stores the observation values ​​of each agent in the electric vehicle grouping stage, the group division at the current time step, and the global performance response of the group structure under the control of the second stage. The second stage is the strategy optimization stage, which optimizes the operation of the power distribution network by coordinating the active and reactive power of the grouped electric vehicles based on the group structure. The experience pool of the optimization stage stores the observations and actions of each agent in the vehicle-to-grid interaction control stage, the joint reward of each agent at the end of the two-stage chain, and the state at the end of the two-stage chain decision.

9. The method for dynamic grouping control of electric vehicles driven by graph theory-statistical joint modeling according to claim 6, characterized in that, The multi-agent deep deterministic policy gradient algorithm uses an action network to evaluate the effectiveness of different scheduling strategies, thereby assisting in optimizing the scheduling of electric vehicle charging stations within the subsystem. During the training phase, the agents learn and optimize action strategies using the current environmental observation state. During the execution phase, the agents dynamically adjust the grouping strategy of electric vehicles and optimize the power allocation of the power distribution network based on the output strategy of the action network. In the cascaded two-stage decision framework of the multi-agent deep deterministic policy gradient algorithm, the agent's observations in the first stage include the local observation state of the distribution network and the state information related to electric vehicles, and the output is the grouping strategy of electric vehicles; the local observations input in the second stage include the grouping results of the first stage and the further observed environmental state, and the output is the electric vehicle scheduling action; the parameters to be optimized in the action network are the action network parameters and the target action network parameters. Each agent's action network incorporates a target network and a policy entropy term; the action network uses a policy gradient method to jointly optimize the action network parameters across the entire decision chain.

10. The graph theory-statistical joint modeling-driven dynamic grouping control method for electric vehicles according to claim 1, characterized in that, The multidimensional space dynamic evaluation model characterizes the maximum photovoltaic access margin boundary composed of multiple nodes by maximizing the Pareto front of the photovoltaic capacity adjustment on each node, and uses the expansion or contraction of the margin boundary in the multidimensional space to characterize the change in the system's carrying capacity. The high-dimensional space model is reduced in dimension by approximation and projection using a uniform manifold, and the high-dimensional geometry enclosed by the margin boundary in the high-dimensional space is mapped to the low-dimensional space. Discretize the time series and calculate the volume of the geometric body enclosed by the boundary map of the spatial dynamic evaluation model for each discrete time series in the low-dimensional space. Define the average value as the new energy carrying capacity margin.

Citation Information

Patent Citations

  • Method and system for predicting power of microgrid group

    CN105826944A

  • Large-scale electric vehicle grouped participating in power grid frequency modulation-based control method

    CN110048406A