A mutual supply and mutual assistance hybrid control method based on a honeycomb-shaped horizontal topological architecture

Through the mutual supply and mutual assistance hybrid control method based on a honeycomb-shaped horizontal topology architecture, the problem of load imbalance in traditional distribution networks is solved, load balancing and grid stability are improved, and the needs of refined development of modern distribution networks are met.

CN120184981BActive Publication Date: 2025-07-25NANJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510661041.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-07-25
Estimated Expiration
2045-05-22

AI Technical Summary

Technical Problem

When the distribution network under the traditional vertical management model faces the grid connection of renewable resources and changes in load characteristics, it is difficult to achieve load balancing, resulting in instability of the power system and safety risks. The existing horizontal collaborative mutual supply and mutual assistance strategies are insufficient.

Method used

A hybrid control method for mutual supply and mutual assistance based on a honeycomb lateral topology architecture is adopted, and effective node sets are screened through graph theory abstract representation and complex network theory, and load balancing optimization is used for load balancing, combined with a dual-delay deep deterministic strategy gradient algorithm to achieve optimal collaborative mutual supply and mutual assistance of a honeycomb distribution network.

Benefits of technology

It improves the load balancing capability of the distribution network, ensures the reliability and stability of the power grid, reduces the safety risks of the power system, and improves the level of refined management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120184981B_ABST
    Figure CN120184981B_ABST
Patent Text Reader

Abstract

The present invention proposes a mutual supply and mutual assistance hybrid control method based on a honeycomb-shaped lateral topological architecture, including the following steps: abstractly representing the lateral topological architecture of the honeycomb-shaped distribution network through graph theory, where each vertex represents the main substation of each regional grid in the honeycomb-shaped distribution network, and the regional grids are connected through flexible interconnection devices; constructing an effective node set through a screening method based on the honeycomb-shaped effective loop topology; characterizing the hybrid discrete-continuous action space model through a conditional variational autoencoder and an embedding table, and embedding an approximate power flow graph neural network into the reward function of the parameterized action Markov decision process to achieve the optimal collaborative mutual supply and mutual assistance of the honeycomb-shaped lateral topology. The present invention can greatly improve the overall load balance rate of the honeycomb-shaped distribution network on the basis of meeting the power flow constraints, ensure the safe and optimized operation of the power grid, and thereby expand the application of the deep reinforcement learning method in the power grid hybrid problem.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of power system operation and control, and particularly to a mutual supply and mutual assistance hybrid control method based on a honeycomb-shaped horizontal topological structure. Background Art

[0002] Due to the relatively simple power supply structure, rigid load characteristics, and mostly unidirectional power receiving operation characteristics of traditional extensive distribution networks, a "top-down" vertical management and planning mode is often adopted. With the large-scale grid connection of renewable resources, the resource density on each side of the power source, grid, load, and energy storage has increased significantly, which easily leads to an increase in power flow uncertainty. The power users' requirements for power supply reliability are also continuously increasing, and the construction of distribution networks has gradually entered the stage of refined development. Under this background, due to the relatively rough granularity of the traditional vertical management mode, it is easy to cause a decline in resource interactivity, a reduction in the operation effect of the distribution network, and even possible safety problems such as power fluctuations and voltage over-limit, making it difficult to achieve refined management of the distribution network. Therefore, on the basis of retaining the vertical management mode of the distribution network, adding a horizontal management mode is an inevitable trend in the construction and development of distribution networks.

[0003] In this mode, the power supply area of the distribution network is divided into several non-overlapping unit power supply grids of the same voltage level, and different power supply grids can achieve coordinated mutual supply and mutual assistance regulation through formulated strategies, so as to achieve full-range and refined horizontal management of the distribution network. The emergence of the new horizontal management mode can effectively enhance power supply reliability, improve dispatching flexibility, and enhance the lean management level. However, most existing studies are based on vertical management concepts, ignoring the possible problems that may occur in the distribution network during horizontal control, and there are few studies on the horizontal collaborative mutual supply and mutual assistance strategy of the distribution network. Therefore, for the research on the horizontal collaborative mutual supply and mutual assistance regulation strategy, the research ideas under the vertical architecture cannot be directly migrated.

[0004] At the same time, due to the extremely uneven spatio-temporal distribution of distributed renewable energy, it is easy to cause uneven load rates of regional power supply grids, which will not only affect the utilization rate of grid substations, but may even lead to uneconomical operation of lightly loaded transformers and unsafe operation of heavily loaded transformers. If the grid load rate is not regulated in time, it may further cause serious impacts such as large-scale power outage in the regional power grid, thus making it fall into a fault state and bringing huge potential safety hazards to the distribution system. Therefore, the load balancing problem among multi-region grids in the horizontal control mode is the content that needs to be focused on.

[0005] Different from petal-shaped distribution networks and diamond-shaped distribution networks, honeycomb-shaped distribution networks do not require traditional switching operations, but use the hexagonal characteristics of topology to achieve power supply guarantee of the power system with high economic benefits. In view of this, the present invention proposes a mutual supply and mutual assistance hybrid control method based on a honeycomb-shaped horizontal topological structure. Summary of the Invention

[0006] In view of the increasingly refined development needs of modern distribution networks, the present invention proposes a mutual supply and mutual assistance hybrid control method based on a honeycomb-shaped horizontal topological architecture. Based on the flexible interconnection devices at the distribution layer, the highly interconnected honeycomb-shaped distribution network structure can take into account excellent scalability and flexibility while ensuring reliability and stability, meet the increasingly refined development needs of modern distribution networks, achieve load balancing of the honeycomb-shaped distribution network, and solve the problems raised in the above background technology. The technical solutions provided by the present invention are as follows:

[0007] A mutual supply and mutual assistance hybrid control method based on a honeycomb-shaped horizontal topological architecture includes the following steps:

[0008] Step 1, abstractly represent the horizontal topological architecture of the honeycomb-shaped distribution network through graph theory. Each vertex represents the main substation of each regional grid in the honeycomb-shaped distribution network, and the regional grids are connected through flexible interconnection devices, expressed as where V is the set of all nodes in the honeycomb-shaped distribution network, E is the set of connecting lines between nodes, and A is the adjacency matrix representing the connection relationship of nodes;

[0009] Step 2, construct an effective node set through a screening method for honeycomb-shaped effective loop topologies based on complex network theory and entropy-TOPSIS;

[0010] Step 3, characterize the hybrid discrete-continuous action space model through a conditional variational autoencoder and an embedding table, and embed an approximate power flow graph neural network into the reward function of the parameterized action Markov decision process to achieve optimal cooperative mutual supply and mutual assistance of the honeycomb-shaped horizontal topology;

[0011] The hybrid discrete-continuous action space model includes two processes: encoding and decoding. Before encoding, construct a latent space for G discrete actions through a learnable embedding table where each row of the embedding table represents a discrete action g t corresponding to a one-dimensional continuous vector with a total dimension of d1; use a conditional variational autoencoder for the encoding operation. With the state s t and the embedding table as conditions, map the continuous action parameters t corresponding to the discrete action g to the latent variable to construct the latent representation space. The encoding process is expressed as:

[0012]

[0013] The decoding process is expressed as:

[0014]

[0015] Among them, ξ is the embedding table parameter, is the encoding process, Φ is the encoding parameter, is the decoding process, ψ is the decoding parameter, is the predicted continuous action parameter, and respectively represent the conversion network and the reconstruction fully connected layer; the embedding table and the conditional variational autoencoder network and their parameters are trained by minimizing the loss function:

[0016]

[0017] Among them, the first term on the right side of the CVAE equation is the L2 reconstruction loss, and the second term is the KL divergence between the variational posterior of the latent representation variable and the Gaussian prior, and the two together constitute the loss function of the conditional variational autoencoder; represents the mathematical expectation; D is the replay buffer, represents the continuous action parameter from the replay buffer.

[0018] Preferably, step 2 specifically includes the following steps:

[0019] Step 2.1, select the following indicators to construct a screening system, including the degree centrality index DC, the closeness centrality index CC, the betweenness centrality index LBC, and the electrical betweenness centrality index EBC;

[0020] Step 2.2, based on the entropy value - TOPSIS honeycomb effective topology screening process, the specific steps are as follows:

[0021] Step 2.2.1, construct a multi - attribute decision - making matrix: X=(x ij ) N×4 =(DC i , CC i , LBC i , EBC i );

[0022] Step 2.2.2, standardize and normalize X:

[0023]

[0024] Obtain the matrix R′=(r ij ′) N×4 ;

[0025] Step 2.2.3, calculate the information entropy of the j - th indicator:

[0026]

[0027] Step 2.2.4, calculate the difference coefficient g of each index j = 1 - e j , and calculate the index weight

[0028] Step 2.2.5, construct the weighted normalized decision matrix Z = (z ij ) N×4 , z ij = r ij × ω j ;

[0029] Step 2.2.6, calculate the positive ideal solution and the negative ideal solution:

[0030] Z + = {max z ij | j = 1, 2, 3, 4}

[0031] Z - = {min z ij | j = 1, 2, 3, 4}

[0032] Step 2.2.7, calculate the distance between the index of each node and the optimal solution and the worst solution:

[0033]

[0034] Step 2.2.8, calculate the relative closeness of each node:

[0035]

[0036] Rank the importance according to the relative closeness, and select the top 70% of the nodes to construct an effective node set.

[0037] Preferably, in step 3, the hybrid discrete - continuous action space model is characterized by a conditional variational auto - encoder and an embedding table, specifically:

[0038] The hybrid discrete - continuous action space model includes two processes: encoding and decoding. Before encoding, construct a latent space of G discrete actions through a learnable embedding table , where each row of the embedding table represents a one - dimensional continuous vector corresponding to the discrete action g t , and the total dimension is d1; use the conditional variational auto - encoder for encoding operations, with the state s t and the embedding table as conditions, map the continuous action parameters t corresponding to the discrete action g to the latent variable , thereby constructing a latent representation space. The encoding process is expressed as:

[0039]

[0040] The decoding process is expressed as:

[0041]

[0042] where ξ is the embedding table parameter, is the encoding process, Φ is the encoding parameter, is the decoding process, ψ is the decoding parameter, is the predicted continuous action parameter, and respectively represent the transformation network and the reconstruction fully connected layer.

[0043] Preferably, the embedding table and the conditional variational autoencoder network and their parameters are trained by minimizing the loss function:

[0044]

[0045] where the first term on the right is the L2 reconstruction loss, and the second term is the KL divergence between the variational posterior and the Gaussian prior of the latent representation variable which together constitute the loss function of the conditional variational autoencoder; represents the mathematical expectation; D is the replay buffer, represents the continuous action parameter from the replay buffer.

[0046] Preferably, the hybrid discrete - continuous action space model further uses a sub - network after the decoder of the conditional variational autoencoder to predict the state residual. For any sample its state residual and its predicted state residual are defined as:

[0047]

[0048] The L2 - norm squared prediction loss is used as the regularization term:

[0049]

[0050] The total training loss function of the final hybrid action representation model is:

[0051] Loss HDCAR (Φ, ψ, ζ) = Loss CVAE (Φ, ψ, ζ) + θLoss Dyn (Φ, ψ, ζ)

[0052] where θ is the weight corresponding to the regularization term.

[0053] Preferably, the approximate power flow graph neural network consists of a masked encoder and an approximate power flow convolutional layer. The binary masked encoder represents known features and unknown features, sets the known features to 0, and sets the unknown features to 1. The mapping operation of the masked encoder is as follows:

[0054]

[0055] where w0 and w1 are weight matrices, b0 and b1 are biases, and they are trained through iteration; σ(·) represents the activation function;

[0056] Combining the input node feature matrix X with the masked encoder M can obtain the encoded graph feature matrix of the k-th layer:

[0057]

[0058] The approximate power flow convolutional layer has K layers, and iteratively processes the effective graph features to predict the node feature matrix X after the final action K , in the message passing stage, the iterative update process of each node is described by the following formula:

[0059]

[0060] where t is the number of iterations in the message passing stage; is the neighbor node of node i; is the node feature of the hidden layer, is the feature of the connection line between node i and node j, and is updated through the transfer function and the node update function U t (·); is the node feature after each iteration of message passing update; is the feature vector after message passing; R(·) is the final node feature update function, is all sum;

[0061] feature vector is fed into the multi-layer perceptron to obtain the graph feature matrix after message passing combined with the encoded graph feature matrix X k and sent to the topological adaptive graph convolutional operator layer for processing to obtain the new encoded graph feature X k+1 :

[0062]

[0063] where the matrix S k is the normalized adjacency matrix; is a trainable weight matrix; the hyperparameter D is a convolutional layer that takes into account neighboring nodes at a distance D from the target node.

[0064] Preferably, step 3 further uses the Twin Delayed Deep Deterministic Policy Gradient algorithm TD3 to train the reinforcement learning agent. TD3 uses a total of six network structures, where the Actor network and the Target-Actor network are used to train the policy network, and the Critic1 network, the Crtic2 network, the Target-Critic1 network, and the Target-Crtic2 network are used to estimate the action value function Q of the current state and action. A neural network is used to approximate the six networks to obtain the corresponding parameters φ, ω1, ω2, and

[0065] Preferably, the training process of TD3 is as follows:

[0066] During each policy iteration, the honeycomb distribution network generates different states s t , which are input into the Actor network to obtain the action a t ′. Gaussian noise ∈~N(0,σ) is added to the action to obtain a t = a t ′t + ∈ ; then, a t is input into the environment to obtain the next state s t+1 and the reward r t+1 ; the latent policy π φ outputs the latent action vectors e and z x , and the conditional variational autoencoder is used to decode the latent variables to obtain the corresponding mixed actions g and x g , and these data are stored in the replay buffer in the form of Done is the final state of the environment;

[0067] The Critic1 network and the Crtic2 network take the latent action as input to approximate the mixed discrete-continuous action value function The corresponding Target-Critic1 network and Target-Crtic2 network are used to estimate the target values y1 and y2; during the training process, the following formula is used to calculate the target values:

[0068]

[0069] Strategy Indicates the Target-Critic network used to estimate the target value; selects the minimum value between the two estimated values as the target value y t :

[0070]

[0071] Regularize the target value by introducing the method of target policy smoothing:

[0072]

[0073] where r t is the reward function, γ is the discount factor, with a value range of 0 to 1, ∈’ is the clipped noise, and c is the boundary of the Gaussian function N(0,σ);

[0074] The training objectives of the Critic1 network and the Crtic2 network are to regress to the learned target value y t , obtain the loss functions of the two networks, and use the gradient descent method to train the two neural networks:

[0075]

[0076] where, represents the mathematical expectation, is the learning rate of the network parameters; τ i is the exponential smoothing parameter, used to update the learning rate of the target network parameters

[0077] The ultimate goal of the Actor network is to find the action that maximizes the Q value, approximate the policy function using a parameterized neural network, and use the gradient descent method to train the Actor network. The TargetActor network is also updated using the exponential smoothing parameter. The detailed formula is as follows:

[0078]

[0079] where, M is the total number of training iterations; λ φ is the learning rate of the Actor network; η is the exponential smoothing parameter, used to update the learning rate of the target Actor network and the target soft Actor network parameters

[0080] Before training the Actor network, multiple rounds of training of the TD3 Critic network are required.

[0081] Compared with the prior art, the beneficial effects achieved by the present invention are:

[0082] By introducing a hybrid discrete - continuous action representation model, the reward volatility is reduced, and the final convergent reward value is higher than that of the algorithm before the model was introduced, indicating the effectiveness of the proposed hybrid discrete - continuous action representation model in solving the hybrid discrete - continuous action problem.

[0083] Based on the flexible interconnection devices at the distribution level, the highly interconnected honeycomb - shaped distribution network structure can balance excellent scalability and flexibility while ensuring reliability and stability. Different from the petal - shaped distribution network and the diamond - shaped distribution network, the honeycomb - shaped distribution network does not require traditional switching operations, but uses the topological hexagonal characteristics to achieve power supply guarantee of the power system with high economic efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0084] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, but do not constitute a limitation to the present invention. In the drawings:

[0085] Figure 1 is the flowchart of the method provided by the present invention;

[0086] Figure 2 is the framework diagram of the method provided by the present invention;

[0087] Figure 3 is the structural diagram of the hybrid discrete - continuous action representation model provided by the present invention;

[0088] Figure 4 is the structural diagram of the approximate power flow graph neural network provided by the present invention;

[0089] Figure 5 is the performance comparison result graph of the HDCAR - TD3 algorithm provided by the present invention and other mainstream deep reinforcement learning algorithms;

[0090] Figure 6 is the performance comparison result graph of the HDCAR - TD3 algorithm provided by the present invention and other mainstream deep reinforcement learning algorithms for the hybrid discrete - continuous action space. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0091] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0092] To make the above - mentioned objects, features, and effects of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the drawings and specific embodiments.

[0093] Embodiment 1: A mutual supply and mutual assistance hybrid control method based on a honeycomb-shaped lateral topological architecture, as Figure 1 shown, includes:

[0094] Step 1, abstractly represent the lateral topological architecture of the honeycomb-shaped distribution network through graph theory. Each vertex represents the main substation of each regional grid in the honeycomb-shaped distribution network, and the regional grids are connected through flexible interconnection devices;

[0095] Step 2, construct an effective node set through a screening method for honeycomb-shaped effective circular topologies based on complex network theory and entropy-TOPSIS;

[0096] Step 3, model the hybrid discrete-continuous action space through a conditional variational autoencoder and an embedding table, and embed an approximate power flow graph neural network into the reward function of the parameterized action Markov decision process to achieve the optimal cooperative mutual supply and mutual assistance of the honeycomb-shaped lateral topology.

[0097] The specific framework diagram is as Figure 2 shown. Based on graph theory, the honeycomb-shaped distribution network is represented by . The node set V represents the main substation of each regional grid in the honeycomb-shaped distribution network, the edge set E represents the branch with a flexible interconnection device, and the feature matrix F reflects the load rate feature α in each node i i .

[0098] The screening method for honeycomb-shaped effective circular topologies based on complex network theory and entropy-TOPSIS in Step 2 specifically includes the following steps:

[0099] Step 2.1, in order to eliminate invalid honeycomb topological structures, using the idea of multi-attribute decision-making, in terms of index selection, consider constructing a multi-index screening system by combining the topological structure and electrical characteristics of the honeycomb-shaped distribution network, including degree centrality index, closeness centrality index, betweenness centrality index, and electrical betweenness centrality index.

[0100] Degree centrality index: The larger the degree of a node, the more crucial the position of the node in the network topological structure. Its calculation formula is:

[0101]

[0102] where k i is the degree of node i, which is the total number of edges connecting node i and other nodes j in the network, that is where when there is an edge between node i and node j, a ij = 1, otherwise a ij = 0; N is the number of nodes in the network topological structure.

[0103] Closeness centrality index: It is the reciprocal of the product of the number of nodes N in the network and the average shortest electrical distance L, reflecting the degree of a node being at the center in the topological structure. Its calculation formula is:

[0104]

[0105] where the shortest electrical distance is the line with the minimum sum of impedance values between two nodes, L is the average shortest electrical distance of the network, and d ij is the shortest distance between node i and node j.

[0106] Betweenness centrality index: It is the number of times the shortest electrical distance between two non - adjacent nodes in the power grid after normalization passes through a certain node. The more times a certain node is passed through in the power network transmission, the more important the position of this point in the network. Its calculation formula is:

[0107]

[0108] where, B i represents the number of times the shortest electrical distance between two non - adjacent nodes passes through node i.

[0109] Electrical betweenness centrality index: In addition to considering the honeycomb - shaped network topology and the importance of associated lines, the particularity of the engines and loads inside the nodes also needs to be considered. Its calculation formula is:

[0110]

[0111] where, and respectively represent the weight coefficients of the generator and the load inside node i; n is the total number of nodes inside node i; n G and n L respectively represent the number of generators and loads inside node i; is the active power of the k - th generator in node i, is the active power of the k - th load in node i; is the betweenness centrality of node i; W i * is the weight coefficient of the node transmission load rate, and its calculation formula is:

[0112]

[0113] where, α i is the load rate of node i, α max and α min respectively represent the upper and lower limits of the load rate of node i.

[0114] Step 2.2, the honeycomb-shaped effective topology screening process based on entropy value - TOPSIS, and the specific steps are as follows:

[0115] Step 2.2.1, construct a multi-attribute decision matrix: X = (x ij ) N×4 = (DC i , CC i , LBC i , EBC i );

[0116] Step 2.2.2, standardize and normalize X according to formulas (6) and (7):

[0117]

[0118] Obtain the matrix R' = (r ij ) N×4 ;

[0119] Step 2.2.3, calculate the information entropy of the j-th index:

[0120]

[0121] Step 2.2.4, calculate the information redundancy value of each index, that is, the difference coefficient g j = 1 - e j , and calculate the index weight

[0122] Step 2.2.5, construct a weighted normalized decision matrix Z = (z ij ) N×4 , z ij = r ij × ω j ;

[0123] Step 2.2.6, calculate the positive ideal solution and the negative ideal solution:

[0124] Z + = {max z ij | j = 1, 2, 3, 4} (9)

[0125] Z - = {min z ij | j = 1, 2, 3, 4} (10)

[0126] Step 2.2.7, calculate the distance between the index of each node and the optimal solution and the worst solution:

[0127]

[0128] Step 2.2.8, calculate the relative closeness of each node:

[0129]

[0130] Sort the importance according to the relative proximity, and select the top 70% of the nodes to construct an effective node set.

[0131] Step 3 specifically includes the optimal honeycomb network reconstruction method based on the hybrid discrete-continuous action representation model, the optimal collaborative mutual supply and mutual assistance strategy based on the approximate power flow graph neural network, and the optimal hybrid control strategy based on the HDCAR-TD3 algorithm, which are described in detail below.

[0132] Optimal honeycomb network reconstruction method based on the hybrid discrete-continuous action representation model:

[0133] Since the transition between states in the honeycomb network topology reconstruction completely depends on the control decision and the current state, this process can be described as a Markov decision process. The basic Markov decision process can be described by a 5-tuple:

[0134] M = (S, A, P(s t+1 |s t , a t ), R t , γ) (13)

[0135] where S is the state space, and s t ∈S is the state of the honeycomb distribution network at time t, mainly including topological structure information and node feature information.

[0136] A is the action space. The reinforcement learning agent executes the action a t according to s t . Assume a learning stochastic policy π(a t |s t ), which is a state-based action probability distribution. The state value function and action value function under π can be described as follows:

[0137]

[0138] P is the state transition probability matrix. The agent observes the state s t of the current environment at each moment t, and simulates the state s t+1 of the next moment environment after being affected by this decision action to take an action a t .

[0139] R is the reward function, which is the goal that the reinforcement learning agent wants to achieve. In the present invention, the learning goal of the reinforcement learning agent is to achieve the best load balancing of the honeycomb distribution network after the action.

[0140] γ is the discount factor, with a value range from 0 to 1. The higher the γ value, the more the Markov decision process focuses on long-term cumulative rewards; conversely, it emphasizes more on the rewards of current actions.

[0141] Since the reconstructed topology is represented by discrete actions and the load rate supply is represented by continuous actions, and there is a certain strong dependence between them, the action space of the Markov decision process is expanded. The expanded Markov decision process is called the parametric action Markov decision process.

[0142] In the parametric action Markov decision process, it is assumed that there exists a finite set of discrete actions G, and each discrete action g t corresponds to a set of continuous actions As Figure 3 shown, the Hybrid Discrete-Continuous Action Representation (HDCAR) model mainly consists of two processes: encoding and decoding. Before encoding, a latent space of G discrete actions is constructed through a learnable embedding table where d1 is the total dimension of the corresponding continuous vector. Each row of the embedding table represents a one-dimensional continuous vector corresponding to the Gth discrete action; then, an encoding operation is performed on the processed hybrid action latent space. Here, a conditional variational autoencoder is used, with the state and the embedding table as conditions, to map the continuous action parameters t corresponding to the discrete action g to the latent variable thereby constructing a latent representation space. The encoding process is expressed as:

[0143]

[0144] The decoding process is expressed as:

[0145]

[0146] where ξ is the embedding table parameter, is the encoding process, Φ is the encoding parameter, is the decoding process, ψ is the decoding parameter, is the predicted continuous action parameter, and represent the transformation network and the reconstruction fully connected layer respectively.

[0147] The embedding table, the conditional variational autoencoder network, and their parameters are trained by minimizing the loss function:

[0148]

[0149] where the first term on the right side of Equation (17) is the L2 reconstruction loss, and the second term is the latent representation variable The KL divergence between the variational posterior and the Gaussian prior, which together constitute the loss function of the conditional variational autoencoder; denotes the mathematical expectation; D is the replay buffer, and samples are stored in the replay buffer in the form of and represents the continuous action parameter from the replay buffer.

[0150] The hybrid discrete-continuous action representation model uses a subnetwork after the decoder of the conditional variational autoencoder to predict the state residual. For any sample its state residual and its predicted state residual can be defined as:

[0151]

[0152] Furthermore, the following squared L2 norm prediction loss is used as a regularization term:

[0153]

[0154] Combining Equation (18) and Equation (21), the final total training loss function of the hybrid action representation model is:

[0155] Loss HDCAR (Φ, ψ, ζ) = Loss CVAE (Φ, ψ, ζ) + θLoss Dyn (Φ, ψ, ζ) (22)

[0156] where θ is the weight corresponding to the regularization term.

[0157] Optimal collaborative mutual supply and mutual assistance strategy based on approximate power flow graph neural network:

[0158] As mentioned above, the learning objective of the agent is to achieve the best load balance rate after taking actions. At each moment, the immediate reward provided by the environment to the agent is:

[0159]

[0160] where is the actual load rate of each regional grid substation in the honeycomb-shaped distribution network; is the average load rate of the honeycomb-shaped distribution network.

[0161] The topological information of the honeycomb-shaped distribution network can be represented by an undirected connected graph where V is the set of all nodes in the honeycomb-shaped distribution network, and E is the set of connection lines between nodes. The adjacency matrix A represents the connection relationship between nodes. Since the adjacency matrix can reflect the physical topological information of the honeycomb-shaped distribution network, it is part of the input of the approximate power flow graph neural network (APFC-GNN).

[0162] In addition, the node feature matrix X and the edge feature matrix X E are another part of the input. The node feature matrix contains the voltage amplitude feature, voltage amplitude feature, active power feature, reactive power feature, and load rate feature of each node, and the edge feature matrix contains the resistance and reactance features of each power line. The specific APFC-GNN structure is as Figure 4 shown, consisting of a masked encoder and an approximate power flow convolutional layer:

[0163] Masked encoder: In traditional power flow calculations, the known features and unknown features of different types of nodes are different. APFC-GNN needs to predict the values of unknown features while fixing the known features. Therefore, a binary masked encoder is introduced to represent the known features and unknown features, enabling APFC-GNN to identify the feature values that need to be predicted. In the present invention, the known features are set to 0 and the unknown features are set to 1. The encoding mask layer is composed of two fully connected layer binary masks, and its mapping operation is:

[0164]

[0165] where w0 and w1 are weight matrices, and b0 and b1 are biases, all of which can be trained iteratively; σ(·) represents the activation function.

[0166] Combining the input node feature matrix X with the masked encoder M can obtain the encoded graph feature matrix of the k-th layer:

[0167]

[0168] Approximate power flow convolutional layer: This layer consists of K interconnected approximate power flow convolutional layers, whose function is to iteratively process the effective graph features obtained through the masked encoder M to predict the node feature matrix X K after the final action. The output node feature matrix is similar to the input node feature matrix, both containing the voltage amplitude feature, power feature, and load rate feature of each node.

[0169] Each approximate power flow convolutional layer utilizes the characteristic of message passing in graph neural networks that it can aggregate and update graph feature information by using the features of neighboring nodes and the graph structure. Each approximate power flow convolutional layer utilizes the characteristics of the message passing network in graph neural networks to aggregate and update the graph's feature information by using the features of adjacent nodes and the graph structure. In the message passing stage, the iterative update process of each node can be described by the following formula:

[0170]

[0171] where t is the number of iterations in the message passing phase; are the neighbor nodes of node i; is the node feature of the hidden layer, is the feature of the connection line between node i and node j, which can be updated through the transfer function and the node update function U t (·); is the node feature after the message passing update in each iteration; is the feature vector after message passing, R(·) is the final node feature update function, is all sum.

[0172] These vectors are fed into a multi-layer perceptron (MLP) to obtain the message transmitted to the node, where the ReLU activation function is used.

[0173]

[0174] where and are weight matrices, and are biases.

[0175] The graph feature matrix after message passing is combined with the encoded graph feature matrix X k and then sent to the topological adaptive graph convolution operator layer for processing to obtain a new encoded graph feature X k+1 ;

[0176]

[0177] where the matrix S k is the normalized adjacency matrix; is a trainable weight matrix; the hyperparameter D represents the convolutional layer considering adjacent nodes at a distance of D from the target node. The larger the value of D, the more adjacent nodes the convolutional layer considers.

[0178] To better evaluate the difference between the predicted unknown feature value of the model and the actual power flow attribute feature corresponding to the agent's action, the error of the unknown feature is used as the evaluation index.

[0179]

[0180] To verify the feasibility of the proposed APFC-GNN approximate power flow, a transformation topology is selected in this embodiment for example analysis. During the training process of the model, the AdamW optimizer is used with a learning rate of 0.001 and a dropout rate of 0.2. In addition, five third-order approximate power flow convolutional layers are used with a hidden layer dimension of 512. According to the data in Table 1, it can be clearly seen that the proposed APFC-GNN method has similar performance and higher accuracy compared with the traditional power flow calculation method. At the same time, the execution times of the two methods are 0.7073 s and 1.4946 s respectively, which indicates that the proposed approximate power flow graph neural network method is reduced by about half compared with the traditional method, meeting the real-time requirements of the hybrid control strategy.

[0181] Table 1 Training Results of Approximate Power Flow Graph Neural Network

[0182]

[0183] Optimal Hybrid Control Strategy Based on HDCAR-TD3 Algorithm: Considering the complexity of the honeycomb-shaped distribution network, the twin-delayed deep deterministic policy gradient algorithm (TD3) is used to train the reinforcement learning agent. As Figure 2 shown in the lower left part, TD3 adopts a total of six network structures, where the Actor network and the Target-Actor network are used to train the policy network, while the Critic1 network, the Crtic2 network, the Target-Critic1 network, and the Target-Crtic2 network are used to estimate the action value function of the current state and action, that is, the Q value. The above six networks are approximated using neural networks, and their corresponding parameters are φ, ω1, ω2, and The TD3 training process considering the HDCAR model is as follows:

[0184] Training Sample Generation:

[0185] During each policy iteration, the honeycomb-shaped distribution network generates different states s t , which are input into the Actor network to obtain the action a t '. To ensure a certain degree of exploration, Gaussian noise ∈~N(0,σ) is added to the action to obtain a t = a t 't + ∈; then, a t is input into the environment to obtain the next state s t+1 and the reward r t+1 . The data [s t , a t , r t , s t+1, [Done] is called experience and stored in the replay buffer D; Done is the final state of the environment, with only two values: True and False. When the experience data generated by the interaction between the agent and the environment is a trajectory, Done = True means that the number of episodes is up to the set number of episodes for the trajectory, that is, this interaction process ends, and it is necessary to return to the initial state. The agent reselects actions to collect new experiences; Done = False means that this piece of experience is in the middle of the trajectory and the interaction process has not ended yet. Due to the HDCAR model, the potential policy π is obtained in this process φ The output potential action vectors e and z x , use equations (15) and (16) to decode the latent variables to obtain the corresponding mixed actions g and x g . Therefore, the form of the experience stored in the replay buffer is

[0186] Training of the Critic network:

[0187] The Critic network aims to approximate the action value function, also known as the Q function. Therefore, we use a neural network parameterized by ω to approximate the mixed discrete - continuous action value function. Using only one Q network for training easily leads to an overestimation bias of the Q value. Therefore, we use a target value network and the clipped double Q - learning algorithm (CDQL) to solve this problem. In short, the double Critic network and take the potential action as the input to approximate the mixed discrete - continuous action value function The corresponding Target - Critic networks are and used to estimate the target values y1 and y2. During training, use and adopt the following formula to calculate the target values of the double Critic network:

[0188]

[0189] Since the Actor - critic policy changes slowly, the current network and the target network are too similar to make independent estimates and there is little improvement. To separate action selection and Q - value update, thus reducing the possibility of overestimation, they are alternately used for action selection and value update. Specifically, when selecting actions from Q1, Q2 is used to update the Q value, and similarly, when selecting actions from Q2, Q1 is used to update the Q value. The policy represents the Target - Critic network used to estimate the target values.

[0190] To avoid introducing additional overestimation in the target update, the minimum value between the two estimates is selected as the target update y of CDQL t :

[0191]

[0192] By introducing the method of target policy smoothing to regularize the target value, it can be expressed as follows:

[0193]

[0194] where ∈’ represents the clipped noise. Adding noise to the target value is equivalent to adding a regularization term in the Critic. c is the boundary of the Gaussian function N(0,σ).

[0195] For the two Critic networks, their training objectives are both to regress to this learned target value y t , so the loss functions of the two networks can be obtained, and the two neural networks are trained using the gradient descent method:

[0196]

[0197] where is the learning rate of the network parameters; τ i is the exponential smoothing parameter, which is used to update the learning rate of the target network parameters

[0198] Training of the Actor network:

[0199] As mentioned before, the Actor network for offline training needs to provide deterministic actions according to the environmental conditions. Since there are two Q networks in TD3, and the ultimate goal of the Actor network is to find the action that maximizes the Q value, the Actor can freely choose one of the Q networks to maximize. Generally, the Actor makes the output of maximize. Therefore, a parameterized neural network is used to approximate the policy function, and the Actor network is trained using the gradient descent method. The Target Actor network is also updated using the exponential smoothing parameter. The detailed formulas are as follows:

[0200]

[0201] where M is the total number of training iterations; λ φ is the learning rate of the Actor network; η is the exponential smoothing parameter, which is used to update the learning rate of the target Actor network and the target soft Actor network parameters

[0202] The Actor network continuously corrects the policy according to the evaluation results of the Critic network. If there is a large error in the trained Critic network, the learning process of the Actor network will also be affected by the previous error. Therefore, a delayed policy update method is adopted in the TD3 algorithm. Specifically, before training the Actor network, multiple rounds of training of the TD3 Critic network are required.

[0203] In this embodiment, the construction of the typical honeycomb-shaped distribution network in the upper left first part and the collection of data samples are completed through Pandapower and Simbetch. Figure 2 The writing and testing of the proposed policy method are completed through Python. To verify the superiority of the proposed hybrid discrete-continuous action representation model, the proposed algorithm is compared with mainstream deep reinforcement learning algorithms as baselines, including algorithms based on DDPG and algorithms based on TD3. For fair comparison, all baseline algorithms have the same network structure as the proposed HDCAR-TD3 algorithm. For each baseline, the policy interacts with the environment 500,000 times, and 5 experiments are conducted simultaneously and the average results are shown. To test the optimization performance of the proposed hybrid discrete-continuous action representation model, the present invention introduces it into the DDPG algorithm, which is called HDCAR-DDPG. Figure 5 Shows the training curves of the agent under different algorithms, where the solid line represents the average cumulative reward every 500 time steps, and the light shaded area represents the standard deviation of five trials.

[0204] As Figure 5 shown, the proposed HDCAR-TD3 algorithm shows a smoother training process with the least fluctuations, and the reward value gradually converges to around -1100, while the performance of other baselines deteriorates to varying degrees. During the training process, the learning speed of TD3 is faster than that of HDCAR-TD3, but the reward value drops from about 250,000 times to about -1400. The learning speeds of DDPG and HDCAR-DDPG are slower than those of TD3 and HDCAR-TD3, and at the same time show obvious reward fluctuations, making it difficult for the agent to converge to the optimal stable value. The reason is that both the TD3 and DDPG algorithms are deep reinforcement learning algorithms only applicable to solving continuous action space problems. Therefore, their learning performance for the optimal policy in the strongly dependent hybrid discrete-continuous action space is poor. However, after introducing the proposed hybrid discrete-continuous action representation model, both HDCAR-DDPG and HDCAR-TD3 show a certain degree of performance improvement, with reduced volatility, and the final convergence reward value is higher than that of the algorithms before the model is introduced, indicating the effectiveness of the proposed hybrid discrete-continuous action representation (HDCAR) model in solving hybrid discrete-continuous action problems.

[0205] Furthermore, the proposed HDCAR-TD3 algorithm is compared with existing mainstream deep reinforcement learning baselines for hybrid discrete-continuous action spaces (such as PDQN-TD3, HHQN-TD3, and HPPO) to evaluate its effectiveness in terms of training performance. As Figure 6 shown, the hybrid control strategy trained by HDCAR-TD3 is superior to the strategies trained using the other three mainstream deep reinforcement learning baselines for hybrid discrete-continuous action spaces. The proposed method has the highest reward value and the lowest fluctuation value. The learning speed of HHQN-TD3 is faster than that of HDCAR-TD3, but the reward value drops from about 200,000 times. Except that PDQN-TD3 failed to train successfully, the other algorithms converged, but HPPO and HHQN-TD3 converged to a lower value of -1500, which means that they could not learn the optimal strategy in the strongly correlated hybrid discrete-continuous action space, proving the poor generality of these two algorithms. The above comparison results strongly prove the effectiveness and reliability of the method proposed in the present invention in the strongly correlated hybrid discrete-continuous action space.

[0206] In addition, the present invention compares the performance of the proposed HDCAR-TD3 algorithm with mainstream deep reinforcement learning algorithms (DRL) and advanced intelligent algorithms (ADIA) through two specific cases, such as genetic algorithms (GA) and differential evolution algorithms (DE). The present invention considers two cases of load imbalance, including the following situations.

[0207] Case 1: The load rate of the honeycomb-shaped distribution network grid S1 is lower than the normal range [60%, 75%];

[0208] Case 2: The load rate of the honeycomb-shaped distribution network grid S1 is higher than the normal range [60%, 75%].

[0209] Table 2 Optimal strategy solutions under different algorithms

[0210]

[0211] Table 2 shows the results of the optimal strategy solutions of various algorithms in the two cases. The 0 and 1 in the optimal network reconstruction correspond to Figure 2In the on / off state of the SOP in the first part, the load rate supply is the load rate supply between the honeycomb-shaped distribution network area grids under different discrete actions, where the positive and negative signs respectively represent supplying / receiving the load rate. As shown in bold in Table 2, the action strategy adopted by HDCAR-TD3 has the best improvement effect on the load imbalance degree of the honeycomb-shaped distribution network in two different cases. Advanced intelligent algorithms and other deep reinforcement learning baselines for the hybrid discrete-continuous action space can quickly alleviate the load imbalance of the honeycomb-shaped distribution network to a certain extent, but they do not adopt the optimal strategy, and their calculation time is too long to meet the real-time requirements of load balancing and regulation services. The above conclusions prove that the proposed HDCAR-TD3 algorithm has strong effectiveness and superiority in the hybrid control problem of the cellular horizontal architecture with a strongly correlated hybrid discrete-continuous action space.

[0212] Example 2: The computer-readable storage medium of this example stores a computer program, and when the program is executed by a processor, it implements the steps in a mutual supply and mutual assistance hybrid control method based on a honeycomb-shaped horizontal topology architecture in Example 1.

[0213] The computer-readable storage medium of this example can be an internal storage unit of the terminal, such as the hard disk or memory of the terminal; the computer-readable storage medium of this example can also be an external storage device of the terminal, such as a plug-in hard disk, smart memory card, secure digital card, flash card, etc. equipped on the terminal; further, the computer-readable storage medium can also include both the internal storage unit and the external storage device of the terminal.

[0214] The computer-readable storage medium of this example is used to store the computer program and other programs and data required by the terminal, and the computer-readable storage medium can also be used to temporarily store the data that has been output or will be output.

[0215] Example 3: The computer device of this example includes a processor, a memory, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in a mutual supply and mutual assistance hybrid control method based on a honeycomb-shaped horizontal topology architecture in Example 1.

[0216] In this example, the processor can be a central processing unit, or other general-purpose processors, digital signal processors, application-specific integrated circuits, off-the-shelf programmable gate arrays, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.; the memory can include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory can also include a non-volatile random access memory. For example, the memory can also store information about the device type.

[0217] Those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0218] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A mutual supply and mutual assistance hybrid control method based on a honeycomb-shaped horizontal topological architecture, characterized in that, It includes the following steps: Step 1: Abstractly represent the horizontal topological structure of the honeycomb-shaped distribution network through graph theory. Each vertex represents the main substation of each regional grid in the honeycomb-shaped distribution network, and the regional grids are connected through flexible interconnection devices, expressed as where V is the set of all nodes in the honeycomb-shaped distribution network, E is the set of connection lines between nodes, and A is the adjacency matrix representing the connection relationship of nodes; Step 2: Construct an effective node set through a screening method of a honeycomb-shaped effective loop topology based on complex network theory and entropy-TOPSIS; Step 3: Characterize the hybrid discrete-continuous action space model through a conditional variational autoencoder and an embedding table, and embed an approximate power flow graph neural network into the reward function of the parameterized action Markov decision process to achieve the optimal collaborative mutual supply and mutual assistance of the honeycomb-shaped lateral topology; The hybrid discrete - continuous action space model consists of two processes: encoding and decoding. Before encoding, through a learnable embedding table a latent space for G discrete actions is constructed, where each row of the embedding table represents a discrete action g t corresponding to a one - dimensional continuous vector, with a total dimension of d1; Perform an encoding operation using a conditional variational autoencoder, with the state s t and the embedding table as conditions, map the discrete action g t to the corresponding continuous action parameters to the latent variables in order to construct a latent representation space. The encoding process is expressed as: The decoding process is expressed as: Among them, ξ is the embedding table parameter, is the encoding process, Φ is the encoding parameter, is the decoding process, ψ is the decoding parameter, is the predicted continuous action parameter, and respectively represent the transformation network and the reconstruction fully-connected layer; The embedding table, the conditional variational autoencoder network and its parameters are trained through a minimized loss function: Among them, the first term on the right side of the CVAE equation is the L2 reconstruction loss, and the second term is the KL divergence between the variational posterior of the latent representation variable and the Gaussian prior. The two together constitute the loss function of the conditional variational autoencoder; denotes the mathematical expectation; D is a replay buffer, representing consecutive action parameters from the replay buffer.

2. The mutual supply and mutual assistance hybrid control method based on a honeycomb-shaped horizontal topological structure according to claim 1, wherein, Step 2 specifically includes the following steps: Step 2.1: Select the following indicators to construct a screening system, including the degree centrality index DC, the closeness centrality index CC, the betweenness centrality index LBC, and the electrical betweenness centrality index EBC; Step 2.2: The screening process of the honeycomb-shaped effective topology based on entropy-TOPSIS, the specific steps are: Step 2.2.1, construct a multi-attribute decision matrix: X = (x ij ) N×4 = (DC i , CC i , LBC i , EBC i ); Step 2.2.2: Standardize and normalize X; Obtain matrix R′ = (r ij ′) N×4 ; Step 2.2.3: Calculate the information entropy of the jth indicator; Step 2.2.4, calculate the difference coefficient g of each index j = 1 - e j , and calculate the index weight Step 2.2.5, construct the weighted normalized decision matrix Z = (z ij ) N×4 , z ij = r ij × ω j ; Step 2.2.6: Calculate the positive ideal solution and the negative ideal solution; Z + = {maxz ij | j = 1, 2, 3, 4} Z - = {minz ij | j = 1, 2, 3, 4} Step 2.2.7: Calculate the distance between the indicators of each node and the optimal solution and the worst solution; Step 2.2.8: Calculate the relative closeness of each node; Rank the importance according to the relative closeness, and select the top 70% of the nodes to construct an effective node set.

3. A mutual supply and mutual assistance hybrid control method based on a honeycomb-shaped horizontal topology architecture according to claim 1, characterized in that, The hybrid discrete - continuous action space model further uses a sub - network after the decoder of the conditional variational auto - encoder to predict the state residual, for any sample whose state residual and its predicted state residual are defined as: Use the L2 norm square prediction loss as a regularization term: The total training loss function of the final hybrid action representation model is: Among them, is the weight corresponding to the regularization term.

4. A mutual supply and mutual assistance hybrid control method based on a honeycomb-shaped horizontal topology architecture according to claim 1, characterized in that The approximate power flow graph neural network consists of a masked encoder and an approximate power flow convolutional layer. The binary masked encoder characterizes the known features and the unknown features, sets the known features to 0, and sets the unknown features to 1. The mapping operation of the masked encoder is: Where w0 and w1 are weight matrices, b0 and b1 are biases, and are trained through iteration; σ(·) represents an activation function; combining the input node feature matrix X with the masked encoder M can obtain the encoded graph feature matrix of the kth layer: The approximate current convolution layer has K layers, and iteratively processes the effective graph features in sequence to predict the node feature matrix X after the final action. K , in the message passing phase, the iterative update process of each node is described by the following formula: where t is the number of iterations in the message passing phase; is the neighbor node of node i; is the node feature of the hidden layer, is the feature of the connection line between node i and node j, updated through the transfer function and the node update function U t (·); is the node feature after message passing update in each iteration; is the feature vector after message passing; R(·) is the final node feature update function, is all sum; Feature vector is fed into a multi-layer perceptron to obtain the graph feature matrix after message passing is combined with the encoded graph feature matrix X k and sent to the topological adaptive graph convolution operator layer for processing to obtain the new encoded graph feature X k+1 : Among them, matrix S k is the normalized adjacency matrix; is the trainable weight matrix; the hyperparameter D is the convolutional layer considering adjacent nodes at a distance of D from the target node.

5. A mutual supply and mutual assistance hybrid control method based on a honeycomb-shaped horizontal topology architecture according to claim 1, characterized in that, Step 3 further uses the Twin Delayed Deep Deterministic Policy Gradient algorithm TD3 to train the reinforcement learning agent. TD3 adopts a total of six network structures. Among them, the Actor network and the Target-Actor network are used to train the policy network, and the Critic1 network, the Crtic2 network, the Target-Critic1 network, and the Target-Crtic2 network are used to estimate the action value function Q of the current state and action. Neural networks are used to approximate the six networks to obtain the corresponding parameters φ, ω1, ω2, and 6. A mutual supply and mutual assistance hybrid control method based on a honeycomb-shaped horizontal topology architecture according to claim 5, characterized in that, The training process of TD3 is as follows: During each policy iteration, the honeycomb-shaped distribution network generates different states s t , and these states are input into the Actor network to obtain the action a t ′. Gaussian noise ∈~N(0,σ) is added to the action to obtain a t = a t ′t + ∈; then, a t is input into the environment to obtain the next state s t+1 and the reward r t+1 ; the latent policy π φ output by the latent action vectors e and z x is decoded using the conditional variational autoencoder for the latent variables to obtain the corresponding mixed actions g and x g , and these data are stored in the replay buffer in the form of Done is the final state of the environment; Critic1 network and Crtic2 network take potential actions as inputs to approximate the mixed discrete - continuous action value function The corresponding Target - Critic1 network and Target - Crtic2 network are used to estimate the target values y1 and y2; during training, use the following formula to calculate the target values: Strategy Indicates the Target-Critic network for estimating the target value; selects the minimum value between the two estimated values as the target value y t : Regularize the target value by introducing a method of target policy smoothing: where r t is the reward function, γ is the discount factor with a value range from 0 to 1, ∈’ is the clipped noise, and c is the boundary of the Gaussian function N(0,σ); The training objectives of the Critic1 network and the Critic2 network are to regress to the learning target value y t , obtain the loss functions of the two networks, and use the gradient descent method to train the two neural networks: Among them, represents the mathematical expectation, is the learning rate of the network parameters; τ i is the exponential smoothing parameter for updating the learning rate of the target network parameters The ultimate goal of the Actor network is to find the action that maximizes the Q value, approximate the policy function using a parameterized neural network, and train the Actor network using the gradient descent method. The TargetActor network is also updated using an exponential smoothing parameter. The detailed formula is as follows: where M is the total number of training iterations; λ φ is the learning rate of the Actor network; η is the exponential smoothing parameter and is the learning rate for updating the parameters of the target Actor network and the target soft Actor network Before training the Actor network, it is necessary to go through multiple rounds of training of the Critic network of TD3.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps in a mutual supply and mutual assistance hybrid control method based on a honeycomb-shaped lateral topology architecture as described in any one of claims 1-6.

8. A computer device, comprising a processor, a memory, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in a mutual supply and mutual assistance hybrid control method based on a honeycomb-shaped lateral topology architecture as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Digital-analog combined drive graph depth reinforcement learning power system optimization scheduling method

    CN118523284A

  • Microgrid network topology flexible configuration method based on matrix topology

    CN119651751A