Game reinforcement learning-based reactive power dispatching method for electric vehicle participating in power distribution network
By combining multi-agent game-theoretic reinforcement learning with an improved particle swarm optimization algorithm, real-time collaborative reactive power dispatching of electric vehicle groups in the power distribution network was achieved, solving the problems of voltage stability and response lag in traditional methods and improving the dispatching accuracy and stability of the power grid.
Patent Information
- Application Number
- CN202510996777.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2025-10-31
AI Technical Summary
Traditional methods struggle to achieve real-time collaborative decision-making in dynamic load fluctuations and random charging/discharging scenarios of electric vehicles, leading to voltage overshooting, delayed or excessive compensation. Furthermore, existing multi-agent reinforcement learning schemes have slow convergence speeds and are prone to getting trapped in local optima.
A game-theoretic reinforcement learning-based method for electric vehicles to participate in reactive power dispatching in power distribution networks is proposed. By combining a multi-agent game-theoretic reinforcement learning framework with an improved particle swarm optimization algorithm, the policy gradient is updated in real time and the adaptive inertia coefficient is adjusted to achieve collaborative reactive power dispatching of electric vehicle groups.
It improves the convergence speed and scheduling accuracy of the algorithm, enhances the voltage stability and power quality of the distribution network, solves the problem of insufficient dynamic response of electric vehicle groups in traditional methods, and realizes efficient and reliable distributed reactive power scheduling.
Smart Images

Figure CN120879636A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of reactive power optimization technology in power system distribution networks, and in particular to a method for reactive power dispatching of electric vehicles in distribution networks based on game-theoretic reinforcement learning. Background Technology
[0002] In recent years, with the integration of distributed renewable energy and large-scale electric vehicles, the reactive power dispatching problem in power distribution networks has become increasingly complex. Traditional approaches often employ centralized optimization or rule-based hierarchical control: centralized optimization utilizes a static power flow sensitivity matrix and heuristic search to uniformly configure reactive power sources; rule-based hierarchical control first determines the reactive power compensation capacity of nodes through offline optimization, and then the local controller starts and stops the compensation devices according to the voltage deviation threshold. Although the technical path is clear and the engineering implementation is simple, it has three shortcomings in scenarios involving dynamic load fluctuations and random charging and discharging of electric vehicles:
[0003] (1) Centralized optimization relies on global measurement and high-speed communication, making it difficult to complete large-scale variable solutions within a second-level scheduling cycle. Once communication is interrupted, voltage over-limit is likely to occur.
[0004] (2) The rule base or linear sensitivity method lacks online self-learning ability and cannot correct the control strategy in real time according to the asynchronous access characteristics of electric vehicle groups, resulting in compensation lag or over-compensation.
[0005] (3) Although existing multi-agent reinforcement learning schemes are adaptive, they usually treat the electric vehicle set as a single agent or adopt a one-way coupling method, which only provides a reference for the optimization algorithm in the initial stage. They do not have bidirectional collaboration and closed-loop performance evaluation, and the convergence speed is slow and they are prone to getting trapped in local optima.
[0006] Meanwhile, swarm intelligence algorithms such as particle swarm optimization (PSO) have been used for reactive power optimization in power distribution networks due to their low parameter count and ease of coding. However, classical PSO algorithms are prone to premature convergence in high-dimensional search spaces and lack the ability to handle real-time constraints. Some researchers have attempted to introduce adaptive inertia weights, local neighborhood topology, and random jumps to improve diversity, but these still limit the core logic of the algorithm to empirical iterations of speed and position, lacking a description of the game-theoretic behavior of the coupling between electric vehicles and the power distribution network.
[0007] Therefore, how to provide a method for electric vehicles to participate in reactive power dispatching of power distribution networks based on game-theoretic reinforcement learning is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0008] One objective of this invention is to propose a method for reactive power dispatching of electric vehicles in power distribution networks based on game-theoretic reinforcement learning. This invention has the advantages of achieving real-time collaborative decision-making, accelerating algorithm convergence, and improving the voltage stability and reactive power compensation accuracy of power distribution networks under the condition of large-scale random access of electric vehicles.
[0009] The method for electric vehicles participating in reactive power dispatching in a power distribution network based on game-theoretic reinforcement learning according to an embodiment of the present invention includes the following steps:
[0010] Data on distribution network node voltage, line current, electric vehicle charging and discharging status, and network topology are collected. Based on the graph attention mechanism, local embedding vectors and global embedding vectors are constructed to obtain the first embedding dataset.
[0011] In the multi-agent game reinforcement learning framework, for each electric vehicle, the master policy network in the multi-agent game reinforcement learning framework is updated based on the first embedded dataset and the current collaborative payoff matrix, and an initial set of reactive power adjustment actions and corresponding policy gradients are generated.
[0012] The policy gradient is recalculated by introducing a dual reward shaping based on node voltage deviation penalty and inter-vehicle cooperation reward, and the updated policy gradient is obtained.
[0013] An improved particle swarm optimization algorithm is initialized to update the policy gradient to determine the initial position and velocity of the particles, and an adaptive inertia coefficient is set according to the variance of the cooperative benefit to obtain the first particle swarm.
[0014] For the first particle swarm, a composite fitness function is constructed based on node voltage sensitivity, line reactive power margin, and expected convergence steps. The particle position and velocity are iteratively updated to obtain the second particle swarm.
[0015] An online collaborative training mechanism is implemented to map the particle velocities of the second particle swarm to the network parameters of a multi-agent game reinforcement learning framework, and the adaptive inertia coefficient is adjusted synchronously to obtain a collaborative optimization strategy.
[0016] The reactive power scheduling instructions are generated based on the collaborative optimization strategy and sent to each electric vehicle for execution. Voltage stability indicators, reactive power compensation rate and power quality indicators are collected, and the parameters of the multi-agent game reinforcement learning framework and the improved particle swarm optimization algorithm are updated.
[0017] Optionally, the step of collecting distribution network node voltage, line current, electric vehicle charging and discharging status, and network topology information, and constructing local and global embedding vectors based on a graph attention mechanism to obtain the first embedding dataset specifically includes:
[0018] Time synchronization calibration of the measurement devices at the distribution network nodes is performed, and a unified sampling period is set;
[0019] According to the sampling period, the voltage of each distribution network node, the current of each line, the charging power and discharging power of each electric vehicle, and the network topology connection relationship are collected to form the original data set.
[0020] The original dataset is normalized and then arranged according to node index and time order to obtain formatted input data;
[0021] Based on the network topology connections, a directed weighted graph is constructed, and the voltage of each node and the line current are used as node features to obtain graph structure data;
[0022] A graph attention mechanism is used to calculate the attention weights between nodes on the graph structure data, and node features are aggregated based on the attention weights to obtain a local embedding vector;
[0023] Perform full-graph pooling on the local embedding vector to obtain the global embedding vector;
[0024] The local embedding vector and the global embedding vector are concatenated in node order to form the first embedding dataset.
[0025] Optionally, in the multi-agent game reinforcement learning framework, updating the master policy network in the multi-agent game reinforcement learning framework for each electric vehicle based on the first embedded dataset and the current collaborative payoff matrix, and generating an initial set of reactive power adjustment actions and corresponding policy gradients specifically includes:
[0026] Based on the first embedded dataset, a composite state vector is constructed for each electric vehicle, which includes node local embedding vectors, global embedding vectors, and row vectors of the collaborative benefit matrix.
[0027] The composite state vector is input into the main policy network, and a soft maximization transformation with a learnable temperature parameter is introduced into the input layer to regulate the exploration degree of the policy output.
[0028] A two-layer value evaluation network is connected in parallel in the main policy network. The first layer outputs the state value, and the second layer outputs the action advantage. The outputs of the two layers are linearly combined to generate a normalized advantage value.
[0029] Based on the normalized advantage value, the Monte Carlo backtracking method is used to estimate the expected cumulative revenue of each action, and symmetric normalization is performed on the collaborative revenue matrix to obtain the corrected revenue matrix;
[0030] Based on the correction benefit matrix, the trainable parameters of the main policy network are updated using the policy gradient descent algorithm with the driving term, and the updated network output is used as the initial reactive power adjustment action set.
[0031] The Euclidean norm is calculated for the difference between the two most recent trainable parameter vectors of the main policy network, and the norm is compared with a preset threshold. If it is less than the threshold, the parameters of the last two layers of the policy network are frozen; otherwise, all parameters of the main policy network are iteratively updated to obtain the corresponding policy gradient.
[0032] Optionally, the construction and operation of the multi-agent game reinforcement learning framework specifically includes:
[0033] During the scheduling initialization phase, a collaborative benefit matrix row index consistent with its node index is established for each electric vehicle intelligent agent, and the row index is bound to the corresponding composite state vector;
[0034] The row vectors of the collaborative benefit matrix are symmetrically normalized and then concatenated with the corresponding composite state vectors, serving as the inputs to the main strategy network and the parallel two-layer value evaluation network.
[0035] At the input end, a soft maximization transformation controlled by a learnable temperature parameter is performed on the spliced input data to uniformly regulate the action exploration degree of each agent;
[0036] After the main strategy network outputs the initial set of reactive power adjustment actions and the two-layer value evaluation network generates normalized advantage values, the initial set of reactive power adjustment actions is sorted according to the normalized advantage values, and the sorting results are written into the row vector of the corresponding collaborative benefit matrix.
[0037] The collaborative benefit matrix after writing the sorting results is symmetric normalized again to obtain the updated corrected benefit matrix. In the next loop, the row vector of the corrected benefit matrix is reconstructed with the corresponding composite state vector as input.
[0038] The policy gradient descent algorithm with driving terms is used to update the trainable parameters of each main policy network and the two-layer value evaluation network simultaneously. The algorithm determines whether to freeze the parameters of the last two layers of the main policy network based on the comparison between the difference of the two most recent trainable parameter vectors and a preset threshold.
[0039] Optionally, the step of introducing dual reward shaping based on node voltage deviation penalty and inter-vehicle cooperation reward to the policy gradient, and recalculating the policy gradient to obtain the updated policy gradient specifically includes:
[0040] The difference between the real-time voltage value and the rated voltage value of each electric vehicle node is squared and multiplied by an adaptive penalty coefficient updated based on exponential moving average to generate a node voltage deviation vector.
[0041] Calculate the reciprocal of the difference in reactive power regulation between any two electric vehicles, multiply it by their network topology connection weights, and then sum them after hyperbolic tangent transformation to generate the inter-vehicle cooperative reward vector.
[0042] The node voltage deviation vector and the inter-vehicle cooperation reward vector are concatenated by node index, and a first-order graph smoothing operation is performed based on the distribution network Laplace matrix to obtain the smoothed reward vector.
[0043] The smoothed reward vector is standardized with zero mean and unit variance, and the standardized vector is orthogonally projected onto the original policy gradient vector in Euclidean space to obtain the orthogonal reward component.
[0044] The adaptive mixing coefficients are determined based on the reciprocal of the policy gradient variance of the previous scheduling cycle. The orthogonal reward components and the original policy gradient are weighted and superimposed according to the adaptive mixing coefficients to obtain the updated policy gradient candidate vector.
[0045] Apply the Nesterov momentum update rule with look-ahead term to the candidate gradient vector of the update policy to output the final gradient of the update policy.
[0046] Optionally, the initialization of the improved particle swarm optimization algorithm, which updates the policy gradient to determine the initial position and velocity of the particles, and sets an adaptive inertia coefficient according to the cooperative benefit variance to obtain the first particle swarm, specifically includes:
[0047] The update policy gradient is split into several sub-vectors according to the node index, and each sub-vector is normalized by L2 norm and then mapped to the unit hypersphere. The particle initial position matrix is obtained by scaling it according to the preset radius.
[0048] For each particle in the initial position matrix, extract the orthogonal complementary basis vector of the normalized subvector corresponding to the particle, randomly select a set of orthogonal complementary basis vectors according to Gaussian distribution and weight them to generate a random exploration vector orthogonal to the gradient direction, and then perform a linear combination to obtain the initial velocity matrix of the particle.
[0049] Calculate the variance of the collaborative benefit matrix and the mean square deviation of the node voltage, and calculate the adaptive inertia coefficient;
[0050] The individual learning factor is calculated based on the voltage sensitivity coefficient of each node, and the social learning factor is calculated based on the network betweenness centrality of the corresponding node. The individual learning factor and the social learning factor are then normalized.
[0051] A dynamic K-nearest neighbor topology is constructed for each particle in the initial position matrix using Mahalanobis distance as the metric, and a neighborhood index table is recorded.
[0052] The initial particle position matrix, initial particle velocity matrix, adaptive inertia coefficient, individual learning factor, social learning factor, and neighborhood index table are bound together according to the particle index to form the first particle swarm.
[0053] Optionally, the iterative process of the improved particle swarm optimization algorithm specifically includes:
[0054] At the beginning of each iteration, a composite guiding vector field is constructed based on the update strategy gradient vector and the preset distribution network node impedance sensitivity matrix, and the current particle velocity is orthogonally decomposed on the vector field to obtain the guiding velocity component.
[0055] The fractional-order Caputo derivative is used to perform memory weighting on the particle velocity sequence, outputting the memory velocity component, which is then linearly superimposed with the guiding velocity component to form the predicted velocity.
[0056] The predicted velocity is subjected to Hamiltonian dynamic drift-diffusion decomposition to obtain reversible and irreversible velocity components, and a quantum tunneling perturbation vector is applied only to the reversible velocity components.
[0057] The cross-correlation entropy scaling factor is calculated based on the reciprocal of the mean square deviation of the voltage at each node, and the amplitude of the quantum tunneling perturbation vector is adjusted using the cross-correlation entropy scaling factor to generate the perturbation correction velocity.
[0058] The perturbation correction velocity is superimposed with the current particle velocity according to the weights of the adaptive inertia coefficient, individual learning factor and social learning factor. After updating the particle velocity, the particle position is corrected. For particle dimensions that exceed the feasible range of reactive power adjustment, the spectral Waffle method is used to map them back to the feasible region.
[0059] After all particle positions are updated, the particle swarm spectrum radius is calculated and written into the multi-agent game reinforcement learning framework as a policy entropy regularization term. At the same time, the variance of collaborative payoffs and the mean square deviation of node voltages are recalculated, and the adaptive inertia coefficient, individual learning factor and social learning factor are updated.
[0060] Optionally, the step of constructing a composite fitness function based on node voltage sensitivity, line reactive power margin, and expected convergence steps, and iteratively updating particle position and velocity, specifically includes:
[0061] For each particle in the first particle swarm, calculate the node voltage sensitivity vector and the line reactive power margin vector, and perform interval linear normalization on the voltage sensitivity vector and the line reactive power margin vector to obtain the normalized voltage sensitivity vector and the normalized line reactive power margin vector.
[0062] Within the same iteration round, the node voltage sensitivity information entropy and the line reactive power margin information entropy are obtained based on the element distribution of the normalized voltage sensitivity vector and the normalized line reactive power margin vector, respectively, and the node voltage sensitivity weight and the line reactive power margin weight are determined.
[0063] The time modulation factor is calculated based on the expected number of convergence steps and the current iteration count, and a composite fitness function is constructed.
[0064] The gradient of the composite fitness function with respect to the current position of the particle is calculated using the finite difference method, and the velocity correction term is formed by the gradient influence coefficient.
[0065] After updating the particle velocity and position, node voltage sensitivity reflection boundary mapping is performed on dimensions that exceed the feasible range of reactive power adjustment.
[0066] After all the particles have been corrected, the set of particles with the best composite fitness function value is used to form the second particle swarm.
[0067] Optionally, the online collaborative training mechanism, which maps the particle velocities of the second particle swarm to the network parameters of the multi-agent game reinforcement learning framework and synchronously adjusts the adaptive inertia coefficient to obtain the collaborative optimization strategy, specifically includes:
[0068] After the second particle swarm is formed, all particles are sorted from high to low according to the fluctuation amplitude of the particle velocity vector, and the sorting results are divided into several micro-batches, which are then entered into the mapping update queue in sequence.
[0069] For the micro-batch at the head of the queue, according to the preset deterministic mapping rules, the particle velocity vector is mapped one by one to the parameter block of equal length in the main strategy network, and an exclusive write time slice is allocated for the parameter block to complete a lock-free parameter update.
[0070] After each micro-batch parameter update is completed, a fast consistency check is immediately performed on the policy output before and after the update. If the check result exceeds the safety boundary, the update is revoked and the corresponding particle velocity vector is marked as a state to be rescaled.
[0071] Record all consistency test indicators of completed micro-batches, construct a rolling window monitoring curve, and lower the adaptive inertia coefficient step value when the indicator is continuously in a stable range within the window, and raise the adaptive inertia coefficient step value when the fluctuation intensifies.
[0072] After all micro-batches have completed parameter mapping and consistency checks, the latest adaptive inertia coefficients and the updated master strategy network parameters are simultaneously broadcast to the particle swarm side, and the collaborative reward matrix is reconstructed to reorder the particle members.
[0073] Before the start of the next scheduling cycle, the particle velocity vector rescaling label table is cleared, and the final adaptive inertia coefficient of this cycle is frozen as the adaptive inertia coefficient of the next cycle, thus obtaining the cooperative optimization strategy.
[0074] The beneficial effects of this invention are:
[0075] (1) This invention achieves collaborative reactive power scheduling of electric vehicle groups by deeply coupling multi-agent game reinforcement learning with improved particle swarm optimization algorithm, mapping particle velocity to policy network in real time and adjusting inertia coefficient in closed loop, effectively improving the algorithm convergence speed and scheduling accuracy, and enhancing the voltage stability and power quality maintenance capabilities of power distribution network.
[0076] (2) Through graph attention embedding, dual reward shaping and reflection boundary based on node voltage sensitivity, this invention can achieve rapid state perception and adaptive strategy update in the case of random access of electric vehicles, significantly improve the robustness of the scheduling system to load fluctuations, and show better real-time adaptability in large-scale distribution network operation scenarios.
[0077] (3) In terms of reactive power optimization and electric vehicle collaborative control in distribution networks, this invention effectively solves the problems of premature convergence and delayed response of traditional centralized or one-way coupling methods by using online collaborative training mechanism and topological Laplace smoothing social learning method. It breaks through the bottleneck of lack of bidirectional collaboration and dynamic closed-loop verification in the existing technology, realizes efficient and reliable distributed reactive power scheduling, and thus effectively improves the intelligent operation level of the new distribution network. Attached Figure Description
[0078] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0079] Figure 1 This is a flowchart of the method for electric vehicles to participate in reactive power dispatching of power distribution networks based on game-theoretic reinforcement learning proposed in this invention. Detailed Implementation
[0080] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0081] refer to Figure 1 A method for electric vehicles to participate in reactive power dispatching in power distribution networks based on game-theoretic reinforcement learning includes the following steps:
[0082] Data on distribution network node voltage, line current, electric vehicle charging and discharging status, and network topology are collected. Based on the graph attention mechanism, local embedding vectors and global embedding vectors are constructed to obtain the first embedding dataset.
[0083] In the multi-agent game reinforcement learning framework, for each electric vehicle, the master policy network in the multi-agent game reinforcement learning framework is updated based on the first embedded dataset and the current collaborative payoff matrix, and an initial set of reactive power adjustment actions and corresponding policy gradients are generated.
[0084] The policy gradient is recalculated by introducing a dual reward shaping based on node voltage deviation penalty and inter-vehicle cooperation reward, and the updated policy gradient is obtained.
[0085] An improved particle swarm optimization algorithm is initialized to update the policy gradient to determine the initial position and velocity of the particles, and an adaptive inertia coefficient is set according to the variance of the cooperative benefit to obtain the first particle swarm.
[0086] For the first particle swarm, a composite fitness function is constructed based on node voltage sensitivity, line reactive power margin, and expected convergence steps. The particle position and velocity are iteratively updated to obtain the second particle swarm.
[0087] An online collaborative training mechanism is implemented to map the particle velocities of the second particle swarm to the network parameters of a multi-agent game reinforcement learning framework, and the adaptive inertia coefficient is adjusted synchronously to obtain a collaborative optimization strategy.
[0088] The reactive power scheduling instructions are generated based on the collaborative optimization strategy and sent to each electric vehicle for execution. Voltage stability indicators, reactive power compensation rate and power quality indicators are collected, and the parameters of the multi-agent game reinforcement learning framework and the improved particle swarm optimization algorithm are updated.
[0089] By combining a multi-agent game-theoretic reinforcement learning framework with an improved particle swarm optimization (PSO) algorithm, this study addresses the shortcomings of traditional reactive power dispatching methods for electric vehicle (EV) groups in dynamic response. Through a graph attention mechanism, it effectively processes voltage, line current, EV charging / discharging status, and network topology information within the distribution network, enabling real-time optimization of EV reactive power regulation strategies. Within the multi-agent game-theoretic reinforcement learning framework, policy gradient updates and dual-reward shaping further enhance the collaborative efficiency between EVs and the distribution network. The introduction of the improved PSO algorithm, through the setting of adaptive inertia coefficients and collaborative reward variance, achieves more efficient global search, enabling rapid convergence and optimization of reactive power dispatching in complex power grid environments. Furthermore, the design of an online collaborative training mechanism allows for seamless cooperation between the PSO and game-theoretic reinforcement learning frameworks, ultimately yielding an optimized reactive power dispatching strategy. This improves voltage stability, reactive power compensation rate, and power quality, overcoming the limitations of existing dispatching methods.
[0090] In this embodiment, the process of collecting distribution network node voltages, line currents, electric vehicle charging and discharging status, and network topology information, and constructing local and global embedding vectors based on a graph attention mechanism to obtain the first embedding dataset specifically includes:
[0091] Time synchronization calibration of the measurement devices at the distribution network nodes is performed, and a unified sampling period is set;
[0092] According to the sampling period, the voltage of each distribution network node, the current of each line, the charging power and discharging power of each electric vehicle, and the network topology connection relationship are collected to form the original data set.
[0093] The original dataset is normalized and then arranged according to node index and time order to obtain formatted input data;
[0094] Based on the network topology connections, a directed weighted graph is constructed, and the voltage of each node and the line current are used as node features to obtain graph structure data;
[0095] A graph attention mechanism is used to calculate the attention weights between nodes on the graph structure data, and node features are aggregated based on the attention weights to obtain a local embedding vector;
[0096] Perform full-graph pooling on the local embedding vector to obtain the global embedding vector;
[0097] The local embedding vector and the global embedding vector are concatenated in node order to form the first embedding dataset.
[0098] By collecting data on distribution network node voltages, line currents, electric vehicle charging and discharging status, and network topology, and utilizing a graph attention mechanism, this invention effectively constructs local and global embedding vectors, which are then used as input datasets for optimizing reactive power regulation strategies for electric vehicles within a multi-agent game-theoretic reinforcement learning framework. The graph attention mechanism accurately captures the spatiotemporal characteristics of the distribution network and the dynamic state of electric vehicles, improving their responsiveness to load fluctuations. Normalizing the original dataset and arranging it according to node index and time order ensures data consistency and real-time performance, providing robust data support for subsequent reactive power scheduling strategy updates. The directed weighted graph constructed using network topology information further enhances the collaborative regulation capability of electric vehicle groups under different voltage environments, thereby achieving more precise reactive power scheduling. In this invention, the application of the graph attention mechanism overcomes the limitations of traditional methods in perceiving the distribution network state, effectively enhancing the interaction between electric vehicles and the power grid and improving scheduling accuracy.
[0099] In this embodiment, the step of updating the master policy network in the multi-agent game reinforcement learning framework for each electric vehicle based on the first embedded dataset and the current collaborative payoff matrix, and generating an initial set of reactive power adjustment actions and corresponding policy gradients, specifically includes:
[0100] Based on the first embedded dataset, a composite state vector is constructed for each electric vehicle, which includes node local embedding vectors, global embedding vectors, and row vectors of the collaborative benefit matrix.
[0101] The composite state vector is input into the main policy network, and a soft maximization transformation with a learnable temperature parameter is introduced into the input layer to regulate the exploration degree of the policy output.
[0102] A two-layer value evaluation network is connected in parallel in the main policy network. The first layer outputs the state value, and the second layer outputs the action advantage. The outputs of the two layers are linearly combined to generate a normalized advantage value.
[0103] Based on the normalized advantage value, the Monte Carlo backtracking method is used to estimate the expected cumulative revenue of each action, and symmetric normalization is performed on the collaborative revenue matrix to obtain the corrected revenue matrix;
[0104] Based on the correction benefit matrix, the trainable parameters of the main policy network are updated using the policy gradient descent algorithm with the driving term, and the updated network output is used as the initial reactive power adjustment action set.
[0105] The Euclidean norm is calculated for the difference between the two most recent trainable parameter vectors of the main policy network, and the norm is compared with a preset threshold. If it is less than the threshold, the parameters of the last two layers of the policy network are frozen; otherwise, all parameters of the main policy network are iteratively updated to obtain the corresponding policy gradient.
[0106] By introducing a policy gradient update and dual-reward shaping mechanism within a multi-agent game-theoretic reinforcement learning framework, this invention effectively improves the reactive power regulation strategy of electric vehicles (EVs). During the decision-making process of each EV, the parameters of the master policy network are updated based on the embedded dataset and collaborative reward matrix generated by the graph attention mechanism, combined with real-time grid load information and the charging / discharging status of the EVs. In the game-theoretic reinforcement learning process, the introduction of dual-reward shaping based on node voltage deviation and inter-vehicle cooperative rewards not only improves the response speed of EVs in reactive power regulation but also further optimizes the cooperative effect. Through real-time policy gradient updates and reward shaping, the EV group can autonomously adjust its strategy to adapt to grid load fluctuations and voltage changes, avoiding the slow convergence and lag response problems inherent in traditional methods. The introduction of this mechanism enhances the cooperative capability of the EV group, improves the reactive power regulation accuracy and overall stability of the distribution network, and effectively solves the problems of slow response and inaccurate scheduling in existing multi-agent systems.
[0107] In this embodiment, the construction and operation of the multi-agent game reinforcement learning framework specifically includes:
[0108] During the scheduling initialization phase, a collaborative benefit matrix row index consistent with its node index is established for each electric vehicle intelligent agent, and the row index is bound to the corresponding composite state vector;
[0109] The row vectors of the collaborative benefit matrix are symmetrically normalized and then concatenated with the corresponding composite state vectors, serving as the inputs to the main strategy network and the parallel two-layer value evaluation network.
[0110] At the input end, a soft maximization transformation controlled by a learnable temperature parameter is performed on the spliced input data to uniformly regulate the action exploration degree of each agent;
[0111] After the main strategy network outputs the initial set of reactive power adjustment actions and the two-layer value evaluation network generates normalized advantage values, the initial set of reactive power adjustment actions is sorted according to the normalized advantage values, and the sorting results are written into the row vector of the corresponding collaborative benefit matrix.
[0112] The collaborative benefit matrix after writing the sorting results is symmetric normalized again to obtain the updated corrected benefit matrix. In the next loop, the row vector of the corrected benefit matrix is reconstructed with the corresponding composite state vector as input.
[0113] The policy gradient descent algorithm with driving terms is used to update the trainable parameters of each main policy network and the two-layer value evaluation network simultaneously. The algorithm determines whether to freeze the parameters of the last two layers of the main policy network based on the comparison between the difference of the two most recent trainable parameter vectors and a preset threshold.
[0114] By constructing a multi-agent game-theoretic reinforcement learning framework and combining it with an improved particle swarm optimization algorithm, this invention achieves refined collaborative optimization in the reactive power dispatching of electric vehicles in power distribution networks. During the dispatch initialization phase, each electric vehicle agent is assigned a row index of a collaborative payoff matrix consistent with its node index, and this index is bound to the corresponding composite state vector, ensuring that each agent can update its strategy based on the actual power grid state. Based on this, the row vector of the collaborative payoff matrix, after symmetric normalization, is concatenated with the composite state vector and used as input to the main strategy network and the parallel two-layer value evaluation network, further optimizing the strategy output. By introducing a learnable temperature parameter soft maximization transformation at the input, the action exploration degree of each agent is effectively adjusted, improving the flexibility of strategy updates. The two-layer value evaluation network generates a normalized advantage value for each reactive power adjustment action and sorts the initial reactive power adjustment actions according to this advantage value, optimizing dispatch accuracy. Furthermore, this invention ensures the consistency of strategy updates by performing a dynamic symmetric normalization operation on the collaborative payoff matrix. Ultimately, by using a policy gradient descent algorithm that drives the quantity term to simultaneously update the trainable parameters of the main policy network and the two-layer value evaluation network, efficient coordinated scheduling between electric vehicles and the power grid was achieved, significantly improving the stability of the distribution network and the reactive power compensation effect.
[0115] In this embodiment, the step of introducing dual reward shaping based on node voltage deviation penalty and inter-vehicle cooperation reward to the policy gradient, and recalculating the policy gradient to obtain the updated policy gradient specifically includes:
[0116] The difference between the real-time voltage value and the rated voltage value of each electric vehicle node is squared and multiplied by an adaptive penalty coefficient updated based on exponential moving average to generate a node voltage deviation vector.
[0117] Calculate the reciprocal of the difference in reactive power regulation between any two electric vehicles, multiply it by their network topology connection weights, and then sum them after hyperbolic tangent transformation to generate the inter-vehicle cooperative reward vector.
[0118] The node voltage deviation vector and the inter-vehicle cooperation reward vector are concatenated by node index, and a first-order graph smoothing operation is performed based on the distribution network Laplace matrix to obtain the smoothed reward vector.
[0119] The smoothed reward vector is standardized with zero mean and unit variance, and the standardized vector is orthogonally projected onto the original policy gradient vector in Euclidean space to obtain the orthogonal reward component.
[0120] The adaptive mixing coefficients are determined based on the reciprocal of the policy gradient variance of the previous scheduling cycle. The orthogonal reward components and the original policy gradient are weighted and superimposed according to the adaptive mixing coefficients to obtain the updated policy gradient candidate vector.
[0121] Apply the Nesterov momentum update rule with look-ahead term to the candidate gradient vector of the update policy to output the final gradient of the update policy.
[0122] By introducing a dual reward shaping mechanism based on node voltage deviation penalties and inter-vehicle cooperative rewards into a game-theoretic reinforcement learning framework, this invention effectively enhances the collaborative capability of electric vehicle groups in reactive power dispatching of power distribution networks. Specifically, during each policy gradient update, the voltage deviation is first calculated based on the difference between the node voltage value and the rated voltage value, and then multiplied by an adaptive penalty coefficient to generate a node voltage deviation vector. Then, the reciprocal of the reactive power adjustment difference between any two electric vehicles is calculated and weighted according to their network topology connection weights to generate an inter-vehicle cooperative reward vector. By merging these two methods and performing graph smoothing, the adjustment differences between individuals can be reduced, enhancing the overall coordination of the group. Furthermore, this invention employs zero-mean, unit-variance standardization to ensure the consistency and stability of reward information. Next, the standardized reward information is orthogonally projected onto the original policy gradient vector to obtain orthogonal reward components, which are then combined with the reciprocal of the policy gradient variance from the previous scheduling cycle to adjust the adaptive mixing coefficients, achieving dynamic policy gradient updates. Finally, the Nesterov momentum update rule with a look-ahead term ensures the smoothness and stability of the update process, further improving the speed and accuracy of policy convergence. This method effectively solves the problems of reactive power scheduling lag and low accuracy in traditional methods, significantly improving the coordinated scheduling effect between electric vehicle groups and the power grid.
[0123] In this embodiment, the initialization of the improved particle swarm optimization algorithm, which updates the policy gradient to determine the initial position and velocity of the particles, and sets an adaptive inertia coefficient according to the cooperative benefit variance to obtain the first particle swarm, specifically includes:
[0124] The update policy gradient is split into several sub-vectors according to the node index, and each sub-vector is normalized by L2 norm and then mapped to the unit hypersphere. The particle initial position matrix is obtained by scaling it according to the preset radius.
[0125] For each particle in the initial position matrix, extract the orthogonal complementary basis vector of the normalized subvector corresponding to the particle, randomly select a set of orthogonal complementary basis vectors according to Gaussian distribution and weight them to generate a random exploration vector orthogonal to the gradient direction, and then perform a linear combination to obtain the initial velocity matrix of the particle.
[0126] Calculate the variance σ of the collaborative benefit matrix s and the mean square deviation of node voltage δ v And calculate the adaptive inertia coefficient ω:
[0127]
[0128] Where, ω min ω is the preset minimum inertia constant. max To preset the maximum constant of inertia, σ s For the variance of collaborative revenue, Let δ be the exponential moving average of the variance of collaborative revenue over the most recent m scheduling periods. v The mean square deviation of the node voltage. Let λ1 be the exponential moving average of the mean square deviation of node voltage over the most recent m scheduling cycles, λ2 be the constant for adjusting the slope of the cooperative revenue variance, θ be the constant for adjusting the slope of the node voltage deviation, τ be the current particle swarm iteration number, and T be the constant for adjusting the slope of the node voltage deviation. max This represents the total number of iterations in the particle swarm optimization.
[0129] The individual learning factor c1 is calculated based on the voltage sensitivity coefficient of each node, and the social learning factor c2 is calculated based on the network betweenness centrality of the corresponding node. Then, c1 and c2 are normalized and c1+c2=2 are maintained.
[0130] A dynamic K-nearest neighbor topology is constructed for each particle in the initial position matrix using Mahalanobis distance as the metric, and a neighborhood index table is recorded.
[0131] The initial particle position matrix, initial particle velocity matrix, adaptive inertia coefficient ω, individual learning factor c1, social learning factor c2, and neighborhood index table are bound according to particle index to form the first particle swarm.
[0132] By improving the initialization of the particle swarm optimization (PSO) algorithm, this invention effectively solves the problems of low search efficiency and slow convergence speed of traditional PSO algorithms in high-dimensional complex power grid environments. In the initialization phase, the policy gradient is used as the initial velocity vector of the PSO, and the initial position and velocity of each particle are dynamically adjusted based on the node position and velocity components of the policy gradient. Furthermore, by normalizing the particle velocity and position and setting an adaptive inertia coefficient, the search capability of the PSO under complex constraints in the power grid is further improved. To accelerate the global search process of the PSO, this invention designs a fitness evaluation mechanism based on cooperative revenue variance and node voltage sensitivity. By adjusting the adaptive inertia coefficient in real time, the responsiveness of the PSO to changes in the power grid state is improved. By introducing a Lévy transition random vector and combining it with voltage deviation as a scaling factor, this invention enhances the ability of the PSO to escape local optima, thereby achieving more efficient search in complex power grid environments. Finally, through a policy gradient descent algorithm with a driving term, the PSO can converge to the global optimum in a short time and update the search strategy of the PSO in real time, greatly improving the efficiency and accuracy of reactive power dispatching in the distribution network. The improved particle swarm optimization algorithm of this invention significantly improves the optimization speed while ensuring computational accuracy, breaking through the bottleneck of traditional algorithms in the application of large-scale electric vehicles to power distribution networks.
[0133] In this embodiment, the iterative process of the improved particle swarm optimization algorithm specifically includes:
[0134] At the beginning of each iteration, a composite guiding vector field is constructed based on the update strategy gradient vector and the preset distribution network node impedance sensitivity matrix, and the current particle velocity is orthogonally decomposed on the vector field to obtain the guiding velocity component.
[0135] The fractional-order Caputo derivative is used to perform memory weighting on the particle velocity sequence, outputting the memory velocity component, which is then linearly superimposed with the guiding velocity component to form the predicted velocity.
[0136] The predicted velocity is subjected to Hamiltonian dynamic drift-diffusion decomposition to obtain reversible and irreversible velocity components, and a quantum tunneling perturbation vector is applied only to the reversible velocity components.
[0137] The cross-correlation entropy scaling factor is calculated based on the reciprocal of the mean square deviation of the voltage at each node, and the amplitude of the quantum tunneling perturbation vector is adjusted using the cross-correlation entropy scaling factor to generate the perturbation correction velocity.
[0138] The perturbation correction velocity is superimposed with the current particle velocity according to the weights of the adaptive inertia coefficient, individual learning factor and social learning factor. After updating the particle velocity, the particle position is corrected. For particle dimensions that exceed the feasible range of reactive power adjustment, the spectral Waffle method is used to map them back to the feasible region.
[0139] After all particle positions are updated, the particle swarm spectrum radius is calculated and written into the multi-agent game reinforcement learning framework as a policy entropy regularization term. At the same time, the variance of collaborative payoffs and the mean square deviation of node voltages are recalculated, and the adaptive inertia coefficient, individual learning factor and social learning factor are updated.
[0140] By implementing an online collaborative training mechanism, the particle velocities of the second particle swarm are mapped to the network parameters of a multi-agent game-theoretic reinforcement learning framework. This invention achieves seamless collaboration between the particle swarm and the reinforcement learning framework, overcoming the limitations of unidirectional feedback in traditional methods. In this process, the velocity vector of each particle in the second particle swarm is first standardized to ensure a reasonable velocity distribution. Then, according to the mapping rules, the particle velocities are bound to the corresponding network parameters in the main policy network, thereby accurately updating the policy. By synchronously adjusting the adaptive inertia coefficient and other learning parameters, this invention enhances the adaptability of the particle swarm to changes in the power grid state. To ensure the stability of policy updates, this invention also introduces a real-time consistency detection mechanism, monitoring the behavior of the particle swarm during the update process to ensure that each updated policy improves power grid dispatch. Through this mechanism, the electric vehicle swarm can maintain high coordination with the power grid during dispatch, thereby optimizing the reactive power dispatch strategy and further improving voltage stability, reactive power compensation accuracy, and power quality. This invention overcomes the bottlenecks of traditional dispatch methods and improves adaptability and effectiveness in distribution network environments with large-scale electric vehicle integration.
[0141] In this embodiment, the step of constructing a composite fitness function based on node voltage sensitivity, line reactive power margin, and expected convergence steps, and iteratively updating particle position and velocity, specifically includes:
[0142] Calculate the node voltage sensitivity vector s for each particle in the first particle swarm. i With the reactive power margin vector m of the line i and for s i m i Perform interval linear normalization to obtain the normalized vector.
[0143] Within the same iteration round, based on The information entropy H is obtained from the element distribution. s H m Determine the weight w s w m :
[0144]
[0145] Among them, H s H represents the entropy of node voltage sensitivity information. mFor the reactive power margin information entropy of the line, w s For node voltage sensitivity weights, w m For line reactive power margin weighting;
[0146] Based on the expected convergence step number T exp Calculate the time modulation factor w with the current iteration count τ. t =1-τ / T exp And construct a composite fitness function:
[0147]
[0148] Among them, f i w represents the particle fitness value. s For node voltage sensitivity weights, w m For the reactive power margin weight of the line, w t For time modulation factor, The average voltage sensitivity of the particle nodes. This represents the average reactive power margin of the particle circuit.
[0149] Using the finite difference method to study f i Regarding particle position x i Calculate gradient And the rate correction term is formed using the gradient influence coefficient κ.
[0150] After updating the particle velocity and position, nodal voltage sensitivity reflection boundary mapping is performed on dimensions that exceed the feasible range of reactive power regulation:
[0151]
[0152] x i ←x i +v i ;
[0153] Among them, v i Let ω be the particle velocity vector, c1 be the adaptive inertia coefficient, r1 be a uniformly random number in the interval [0,1], and p be the particle velocity vector. best,i x represents the optimal position in the history of an individual particle. i Let c1 be the particle's current position, c2 be the social learning factor, r2 be a uniformly random number in the interval [0,1], and g be the particle's current position. best The position is the global optimal position of the particle swarm, and κ is the gradient influence coefficient. This represents the gradient of the composite fitness function with respect to the particle position.
[0154] After all the particles have been corrected, the set of particles with the best composite fitness function value is used to form the second particle swarm.
[0155] By constructing a composite fitness function based on node voltage sensitivity, line reactive power margin, and expected convergence steps, and iteratively updating particle position and velocity, this invention introduces a more accurate fitness evaluation mechanism into the improved particle swarm optimization algorithm. First, the node voltage sensitivity vector and line reactive power margin vector corresponding to each particle are calculated and normalized to ensure information standardization and consistency. Next, the weights obtained through information entropy calculation further optimize the coordination between particles, avoiding excessive concentration or dispersion during the particle swarm search process. By introducing a time modulation factor, this invention can dynamically adjust the behavior of the particle swarm according to different iteration cycles, effectively balancing the relationship between exploration and utilization, and improving the efficiency of the global search. Finite difference gradient calculation of the composite fitness function further enhances the adaptability and accuracy of the particle swarm in complex power grid environments. Furthermore, by combining the policy gradient update mechanism and the setting of the adaptive inertia coefficient, this invention significantly improves the search capability and optimization efficiency of the particle swarm in the reactive power dispatching problem for electric vehicles and distribution networks. This method not only optimizes the reactive power dispatching effect but also solves the performance bottleneck of traditional particle swarm optimization algorithms in high-dimensional constrained environments, improving the flexibility and response speed of distribution network dispatching.
[0156] In this embodiment, the online collaborative training mechanism, which maps the particle velocities of the second particle swarm to the network parameters of the multi-agent game reinforcement learning framework and synchronously adjusts the adaptive inertia coefficient to obtain the collaborative optimization strategy, specifically includes:
[0157] After the second particle swarm is formed, all particles are sorted from high to low according to the fluctuation amplitude of the particle velocity vector, and the sorting results are divided into several micro-batches, which are then entered into the mapping update queue in sequence.
[0158] For the micro-batch at the head of the queue, according to the preset deterministic mapping rules, the particle velocity vector is mapped one by one to the parameter block of equal length in the main strategy network, and an exclusive write time slice is allocated for the parameter block to complete a lock-free parameter update.
[0159] After each micro-batch parameter update is completed, a fast consistency check is immediately performed on the policy output before and after the update. If the check result exceeds the safety boundary, the update is revoked and the corresponding particle velocity vector is marked as a state to be rescaled.
[0160] Record all consistency test indicators of completed micro-batches, construct a rolling window monitoring curve, and lower the adaptive inertia coefficient step value when the indicator is continuously in a stable range within the window, and raise the adaptive inertia coefficient step value when the fluctuation intensifies.
[0161] After all micro-batches have completed parameter mapping and consistency checks, the latest adaptive inertia coefficients and the updated master strategy network parameters are simultaneously broadcast to the particle swarm side, and the collaborative reward matrix is reconstructed to reorder the particle members.
[0162] Before the start of the next scheduling cycle, the particle velocity vector rescaling label table is cleared, and the final adaptive inertia coefficient of this cycle is frozen as the adaptive inertia coefficient of the next cycle, resulting in a collaborative optimization strategy. By executing an online collaborative training mechanism, this invention achieves seamless collaboration between the particle swarm and the game-theoretic reinforcement learning framework, further enhancing the collaborative optimization capability between the electric vehicle group and the power distribution network. In the specific implementation process, the particle velocities in the second particle swarm are first mapped to the network parameters of the game-theoretic reinforcement learning framework, ensuring that the updates of the particle swarm and the network strategy are synchronized. To improve the efficiency of parameter updates, the particle velocities are standardized and then precisely bound to the corresponding parameter blocks in the main strategy network according to preset mapping rules. In addition, a real-time consistency detection mechanism ensures that each update of particle velocities and network parameters meets the preset stability standard, preventing instability caused by excessive parameter changes. During the update process, this invention also introduces dynamic adjustment of the adaptive inertia coefficient, enabling the particle swarm to dynamically adjust the magnitude of the strategy update according to changes in the power grid state, thereby improving the adaptability of the particle swarm in complex power grid environments. With the collaborative operation of particle swarm optimization and game-theoretic reinforcement learning frameworks, the interaction between electric vehicles and the power grid becomes more efficient, thereby optimizing the reactive power dispatch strategy. At the end of each dispatch cycle, the updated network parameters and inertia coefficients are broadcast to the entire particle swarm, ensuring timely updates and optimizations of the dispatch strategy. Ultimately, this online collaborative training mechanism not only improves reactive power dispatch accuracy but also significantly enhances voltage stability, reactive power compensation capability, and power quality, breaking through the limitations of traditional dispatch methods and adapting to the needs of large-scale electric vehicle grid integration.
[0163] Example 1:
[0164] To verify the feasibility of this invention in the reactive power dispatching of electric vehicles in distribution networks, this embodiment selects a distribution network in a certain city as the application scenario, and conducts experiments based on the city's electric vehicle charging and discharging data, distribution network voltage, current, and reactive power compensation data. This city has rapidly increased the use of electric vehicles in recent years, and traditional reactive power dispatching methods suffer from problems such as delayed dispatch response, high computational resource consumption, and difficulty in adapting to rapidly changing grid conditions when electric vehicles are widely integrated. Against this backdrop, the method of this invention can provide an efficient and accurate reactive power dispatching solution in environments where the charging and discharging status of electric vehicles and the distribution network status change rapidly.
[0165] Traditional methods for scheduling electric vehicles (EVs) typically rely on static rules or centralized optimization strategies. These methods cannot respond in real-time to dynamic factors such as fluctuations in the number of EVs and changes in grid load, leading to delayed or insufficient reactive power compensation during the scheduling process. This invention, however, combines game-theoretic reinforcement learning with an improved particle swarm optimization algorithm, enabling each EV to make autonomous decisions based on its local voltage status and scheduling information from surrounding vehicles, thus achieving a more precise and flexible scheduling strategy. Furthermore, the online collaborative training mechanism of this invention fosters closer interaction among different EVs during the scheduling process, thereby improving overall scheduling efficiency.
[0166] Specifically, during the experiment, real-time charging and discharging data, node voltage, current status, and network topology information of each electric vehicle were first collected from the power distribution network's dispatch system. This information was processed using a graph attention mechanism to generate embedded datasets of node voltage and line current, which served as input to the game-theoretic reinforcement learning framework. To ensure dispatch efficiency, this invention introduces a dual-reward shaping mechanism based on node voltage deviation and inter-vehicle cooperative rewards to update the policy gradient, enabling electric vehicles to rapidly adjust their dispatch strategies according to grid load fluctuations and electric vehicle charging and discharging demands.
[0167] Next, the system uses an improved particle swarm optimization algorithm to update the policy gradient. After each update, the particle swarm algorithm adjusts the position and velocity of the particles based on factors such as node voltage sensitivity and line reactive power margin, iteratively optimizing the particle swarm to continuously approach the global optimum. Through the method of this invention, the particle swarm can quickly adjust the reactive power dispatch scheme according to voltage fluctuations and load changes in the distribution network during each iteration.
[0168] During the scheduling process, an online collaborative training mechanism was employed to map the particle velocities of the second particle swarm to the network parameters of a multi-agent game-theoretic reinforcement learning framework, and to simultaneously adjust the adaptive inertia coefficient. This process ensures that at each stage of scheduling, the electric vehicle group can adjust its scheduling strategy in real time based on the global optimal solution, thereby achieving more precise reactive power compensation and voltage stability.
[0169] The effectiveness of this invention can be analyzed by comparing the scheduling efficiency of traditional methods and the method of this invention under different electric vehicle access scenarios. Through experiments, we obtained the following comparative results:
[0170] Table 1: Comparison of scheduling efficiency between traditional methods and the method of this invention under different electric vehicle access scenarios.
[0171]
[0172] Table 1 shows that, compared with the traditional method, the method of this invention significantly reduces scheduling time, voltage deviation, reactive power compensation efficiency, and power quality optimization rate. Specifically, under the traditional method, the scheduling time is 300 seconds and the voltage deviation is 7.2%, while after using the method of this invention, the scheduling time is reduced to 180 seconds and the voltage deviation is reduced to 4.0%. The reactive power compensation efficiency also increases from 60% to 85%, and the power quality optimization rate increases from 55% to 80%. This result fully demonstrates that the present invention can effectively improve reactive power dispatch efficiency and maintain voltage stability and power quality in complex power grid environments with large-scale electric vehicle integration.
[0173] Further experiments demonstrate that the method of this invention can guarantee the real-time performance and stability of the dispatch strategy under conditions of drastic fluctuations in grid load. During testing, when the number of electric vehicles connected to the distribution network suddenly increased, the method of this invention could respond to the dispatch change within seconds, while traditional methods, due to computational lag, could not effectively adapt to this change, resulting in increased voltage fluctuations and compensation delays.
[0174] Table 2: Comparison of Dispatch Response Time under Power Grid Load Fluctuation Conditions
[0175]
[0176] As shown in Table 2, when the grid load change is 50kW, the traditional method has a dispatch response time of 60 seconds, a voltage fluctuation of 5.6%, and a compensation lag of 10 seconds; while the method of the present invention has a dispatch response time of only 15 seconds, a voltage fluctuation of only 2.3%, and a compensation lag time reduced to 2 seconds. This indicates that the method of the present invention demonstrates significant advantages in handling sudden load fluctuations.
[0177] In summary, this invention solves the problems of scheduling lag and insufficient reactive power compensation in the process of electric vehicles connecting to the power distribution network by introducing a combination of game-theoretic reinforcement learning and improved particle swarm optimization algorithm. It significantly improves the operating efficiency and power quality of the power grid, and demonstrates stronger adaptability and optimization capabilities in complex scenarios with dynamic load and fluctuations in electric vehicle charging demand.
[0178] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A method for electric vehicles participating in reactive power dispatching in power distribution networks based on game-theoretic reinforcement learning, characterized in that, Includes the following steps: Data on distribution network node voltage, line current, electric vehicle charging and discharging status, and network topology are collected. Based on the graph attention mechanism, local embedding vectors and global embedding vectors are constructed to obtain the first embedding dataset. In the multi-agent game reinforcement learning framework, for each electric vehicle, the master policy network in the multi-agent game reinforcement learning framework is updated based on the first embedded dataset and the current collaborative payoff matrix, and an initial set of reactive power adjustment actions and corresponding policy gradients are generated. The policy gradient is recalculated by introducing a dual reward shaping based on node voltage deviation penalty and inter-vehicle cooperation reward, and the updated policy gradient is obtained. An improved particle swarm optimization algorithm is initialized to update the policy gradient to determine the initial position and velocity of the particles, and an adaptive inertia coefficient is set according to the variance of the cooperative benefit to obtain the first particle swarm. For the first particle swarm, a composite fitness function is constructed based on node voltage sensitivity, line reactive power margin, and expected convergence steps. The particle position and velocity are iteratively updated to obtain the second particle swarm. An online collaborative training mechanism is implemented to map the particle velocities of the second particle swarm to the network parameters of a multi-agent game reinforcement learning framework, and the adaptive inertia coefficient is adjusted synchronously to obtain a collaborative optimization strategy. The reactive power scheduling instructions are generated based on the collaborative optimization strategy and sent to each electric vehicle for execution. Voltage stability indicators, reactive power compensation rate and power quality indicators are collected, and the parameters of the multi-agent game reinforcement learning framework and the improved particle swarm optimization algorithm are updated.
2. The method for electric vehicles participating in reactive power dispatching in power distribution networks based on game-theoretic reinforcement learning according to claim 1, characterized in that, The process of collecting data on distribution network node voltages, line currents, electric vehicle charging and discharging status, and network topology, and constructing local and global embedding vectors based on a graph attention mechanism to obtain the first embedding dataset specifically includes: Time synchronization calibration of the measurement devices at the distribution network nodes is performed, and a unified sampling period is set; According to the sampling period, the voltage of each distribution network node, the current of each line, the charging power and discharging power of each electric vehicle, and the network topology connection relationship are collected to form the original data set. The original dataset is normalized and then arranged according to node index and time order to obtain formatted input data; Based on the network topology connections, a directed weighted graph is constructed, and the voltage of each node and the line current are used as node features to obtain graph structure data; A graph attention mechanism is used to calculate the attention weights between nodes on the graph structure data, and node features are aggregated based on the attention weights to obtain a local embedding vector; Perform full-graph pooling on the local embedding vector to obtain the global embedding vector; The local embedding vector and the global embedding vector are concatenated in node order to form the first embedding dataset.
3. The method for electric vehicles participating in reactive power dispatching in power distribution networks based on game-theoretic reinforcement learning according to claim 1, characterized in that, In the multi-agent game reinforcement learning framework, the process of updating the master policy network in the multi-agent game reinforcement learning framework for each electric vehicle based on the first embedded dataset and the current collaborative payoff matrix, and generating an initial set of reactive power adjustment actions and corresponding policy gradients, specifically includes: Based on the first embedded dataset, a composite state vector is constructed for each electric vehicle, which includes node local embedding vectors, global embedding vectors, and row vectors of the collaborative benefit matrix. The composite state vector is input into the main policy network, and a soft maximization transformation with a learnable temperature parameter is introduced into the input layer to regulate the exploration degree of the policy output. A two-layer value evaluation network is connected in parallel in the main policy network. The first layer outputs the state value, and the second layer outputs the action advantage. The outputs of the two layers are linearly combined to generate a normalized advantage value. Based on the normalized advantage value, the Monte Carlo backtracking method is used to estimate the expected cumulative revenue of each action, and symmetric normalization is performed on the collaborative revenue matrix to obtain the corrected revenue matrix; Based on the correction benefit matrix, the trainable parameters of the main policy network are updated using the policy gradient descent algorithm with the driving term, and the updated network output is used as the initial reactive power adjustment action set. The Euclidean norm is calculated for the difference between the two most recent trainable parameter vectors of the main policy network, and the norm is compared with a preset threshold. If it is less than the threshold, the parameters of the last two layers of the policy network are frozen; otherwise, all parameters of the main policy network are iteratively updated to obtain the corresponding policy gradient.
4. The method for electric vehicles participating in reactive power dispatching in power distribution networks based on game-theoretic reinforcement learning according to claim 3, characterized in that, The construction and operation of the multi-agent game-theoretic reinforcement learning framework specifically includes: During the scheduling initialization phase, a collaborative benefit matrix row index consistent with its node index is established for each electric vehicle intelligent agent, and the row index is bound to the corresponding composite state vector; The row vectors of the collaborative benefit matrix are symmetrically normalized and then concatenated with the corresponding composite state vectors, serving as the inputs to the main strategy network and the parallel two-layer value evaluation network. At the input end, a soft maximization transformation controlled by a learnable temperature parameter is performed on the spliced input data to uniformly regulate the action exploration degree of each agent; After the main strategy network outputs the initial set of reactive power adjustment actions and the two-layer value evaluation network generates normalized advantage values, the initial set of reactive power adjustment actions is sorted according to the normalized advantage values, and the sorting results are written into the row vector of the corresponding collaborative benefit matrix. The collaborative benefit matrix after writing the sorting results is symmetric normalized again to obtain the updated corrected benefit matrix. In the next loop, the row vector of the corrected benefit matrix is reconstructed with the corresponding composite state vector as input. The policy gradient descent algorithm with driving terms is used to update the trainable parameters of each main policy network and the two-layer value evaluation network simultaneously. The algorithm determines whether to freeze the parameters of the last two layers of the main policy network based on the comparison between the difference of the two most recent trainable parameter vectors and a preset threshold.
5. The method for electric vehicles participating in reactive power dispatching in power distribution networks based on game-theoretic reinforcement learning according to claim 1, characterized in that, The process of introducing a dual-reward shaping approach to the policy gradient, based on a node voltage deviation penalty and an inter-vehicle cooperation reward, and recalculating the policy gradient to obtain the updated policy gradient, specifically includes: The difference between the real-time voltage value and the rated voltage value of each electric vehicle node is squared and multiplied by an adaptive penalty coefficient updated based on exponential moving average to generate a node voltage deviation vector. Calculate the reciprocal of the difference in reactive power regulation between any two electric vehicles, multiply it by their network topology connection weights, and then sum them after hyperbolic tangent transformation to generate the inter-vehicle cooperative reward vector. The node voltage deviation vector and the inter-vehicle cooperation reward vector are concatenated by node index, and a first-order graph smoothing operation is performed based on the distribution network Laplace matrix to obtain the smoothed reward vector. The smoothed reward vector is standardized with zero mean and unit variance, and the standardized vector is orthogonally projected onto the original policy gradient vector in Euclidean space to obtain the orthogonal reward component. The adaptive mixing coefficients are determined based on the reciprocal of the policy gradient variance of the previous scheduling cycle. The orthogonal reward components and the original policy gradient are weighted and superimposed according to the adaptive mixing coefficients to obtain the updated policy gradient candidate vector. Apply the Nesterov momentum update rule with look-ahead term to the candidate gradient vector of the update policy to output the final gradient of the update policy.
6. The method for electric vehicles participating in reactive power dispatching of power distribution networks based on game-theoretic reinforcement learning according to claim 1, characterized in that, The initialization of the improved particle swarm optimization algorithm, which updates the policy gradient to determine the initial position and velocity of particles, and sets an adaptive inertia coefficient according to the cooperative benefit variance, to obtain the first particle swarm specifically includes: The update policy gradient is split into several sub-vectors according to the node index, and each sub-vector is normalized by L2 norm and then mapped to the unit hypersphere. The particle initial position matrix is obtained by scaling it according to the preset radius. For each particle in the initial position matrix, extract the orthogonal complementary basis vector of the normalized subvector corresponding to the particle, randomly select a set of orthogonal complementary basis vectors according to Gaussian distribution and weight them to generate a random exploration vector orthogonal to the gradient direction, and then perform a linear combination to obtain the initial velocity matrix of the particle. Calculate the variance of the collaborative benefit matrix and the mean square deviation of the node voltage, and calculate the adaptive inertia coefficient; The individual learning factor is calculated based on the voltage sensitivity coefficient of each node, and the social learning factor is calculated based on the network betweenness centrality of the corresponding node. The individual learning factor and the social learning factor are then normalized. A dynamic K-nearest neighbor topology is constructed for each particle in the initial position matrix using Mahalanobis distance as the metric, and a neighborhood index table is recorded. The initial particle position matrix, initial particle velocity matrix, adaptive inertia coefficient, individual learning factor, social learning factor, and neighborhood index table are bound together according to the particle index to form the first particle swarm.
7. The method for electric vehicles participating in reactive power dispatching in power distribution networks based on game-theoretic reinforcement learning according to claim 6, characterized in that, The iterative process of the improved particle swarm optimization algorithm specifically includes: At the beginning of each iteration, a composite guiding vector field is constructed based on the update strategy gradient vector and the preset distribution network node impedance sensitivity matrix, and the current particle velocity is orthogonally decomposed on the vector field to obtain the guiding velocity component. The fractional-order Caputo derivative is used to perform memory weighting on the particle velocity sequence, outputting the memory velocity component, which is then linearly superimposed with the guiding velocity component to form the predicted velocity. The predicted velocity is subjected to Hamiltonian dynamic drift-diffusion decomposition to obtain reversible and irreversible velocity components, and a quantum tunneling perturbation vector is applied only to the reversible velocity components. The cross-correlation entropy scaling factor is calculated based on the reciprocal of the mean square deviation of the voltage at each node, and the amplitude of the quantum tunneling perturbation vector is adjusted using the cross-correlation entropy scaling factor to generate the perturbation correction velocity. The perturbation correction velocity is superimposed with the current particle velocity according to the weights of the adaptive inertia coefficient, individual learning factor and social learning factor. After updating the particle velocity, the particle position is corrected. For particle dimensions that exceed the feasible range of reactive power adjustment, the spectral Waffle method is used to map them back to the feasible region. After all particle positions are updated, the particle swarm spectrum radius is calculated and written into the multi-agent game reinforcement learning framework as a policy entropy regularization term. At the same time, the variance of collaborative payoffs and the mean square deviation of node voltages are recalculated, and the adaptive inertia coefficient, individual learning factor and social learning factor are updated.
8. The method for electric vehicles participating in reactive power dispatching of power distribution networks based on game-theoretic reinforcement learning according to claim 1, characterized in that, The specific steps of constructing a composite fitness function based on node voltage sensitivity, line reactive power margin, and expected convergence steps, and iteratively updating particle position and velocity, include: For each particle in the first particle swarm, calculate the node voltage sensitivity vector and the line reactive power margin vector, and perform interval linear normalization on the voltage sensitivity vector and the line reactive power margin vector to obtain the normalized voltage sensitivity vector and the normalized line reactive power margin vector. Within the same iteration round, the node voltage sensitivity information entropy and the line reactive power margin information entropy are obtained based on the element distribution of the normalized voltage sensitivity vector and the normalized line reactive power margin vector, respectively, and the node voltage sensitivity weight and the line reactive power margin weight are determined. The time modulation factor is calculated based on the expected number of convergence steps and the current iteration count, and a composite fitness function is constructed. The gradient of the composite fitness function with respect to the current position of the particle is calculated using the finite difference method, and the velocity correction term is formed by the gradient influence coefficient. After updating the particle velocity and position, node voltage sensitivity reflection boundary mapping is performed on dimensions that exceed the feasible range of reactive power adjustment. After all the particles have been corrected, the set of particles with the best composite fitness function value is used to form the second particle swarm.
9. The method for electric vehicles participating in reactive power dispatching in power distribution networks based on game-theoretic reinforcement learning according to claim 1, characterized in that, The online collaborative training mechanism maps the particle velocities of the second particle swarm to the network parameters of the multi-agent game reinforcement learning framework and synchronously adjusts the adaptive inertia coefficient to obtain the collaborative optimization strategy, specifically including: After the second particle swarm is formed, all particles are sorted from high to low according to the fluctuation amplitude of the particle velocity vector, and the sorting results are divided into several micro-batches, which are then entered into the mapping update queue in sequence. For the micro-batch at the head of the queue, according to the preset deterministic mapping rules, the particle velocity vector is mapped one by one to the parameter block of equal length in the main strategy network, and an exclusive write time slice is allocated for the parameter block to complete a lock-free parameter update. After each micro-batch parameter update is completed, a fast consistency check is immediately performed on the policy output before and after the update. If the check result exceeds the safety boundary, the update is revoked and the corresponding particle velocity vector is marked as a state to be rescaled. Record all consistency test indicators of completed micro-batches, construct a rolling window monitoring curve, and lower the adaptive inertia coefficient step value when the indicator is continuously in a stable range within the window, and raise the adaptive inertia coefficient step value when the fluctuation intensifies. After all micro-batches have completed parameter mapping and consistency checks, the latest adaptive inertia coefficients and the updated master strategy network parameters are simultaneously broadcast to the particle swarm side, and the collaborative reward matrix is reconstructed to reorder the particle members. Before the start of the next scheduling cycle, the particle velocity vector rescaling label table is cleared, and the final adaptive inertia coefficient of this cycle is frozen as the adaptive inertia coefficient of the next cycle, thus obtaining the cooperative optimization strategy.
Citation Information
Patent Citations
Reactive power grid capacity configuration method for random inertia factor particle swarm optimization algorithm
CN104037776A
Power grid optimal carbon energy composite flow obtaining method based on swarm intelligence reinforcement learning
CN105023056A
Reactive power optimization compensation method and device for electric vehicle charging station participating in power distribution network
CN117578494A
Multi-agent federated reinforcement learning-based vehicle-road collaborative control system and method under complex intersection
WO2024016386A1
Cited By
Inland river port unmanned shore tackle automatic scheduling method and system based on artificial intelligence
CN121169038A
Heavy-load unmanned helicopter cooperative hoisting method based on multi-agent learning
CN121411489A
Municipal road intelligent construction and collaborative management method based on digital twinning
CN121436922A
Electric vehicle V2G intelligent power distribution method based on sequence control
CN121485051A
A V2G Smart Power Distribution Method for Electric Vehicles Based on Sequence Control
CN121485051B