Power distribution network multi-domain protection method and system based on multi-agent reinforcement learning

Through the method of multi-agent reinforcement learning, a multi-domain protection system is built, and the node characteristics are extracted and the agent's action decisions are generated by using elastic optical networks and network coding technology. The reliability and resource utilization problems of the traditional distribution communication network protection mechanism in a multi-domain network environment are solved, and efficient and secure multi-domain protection is achieved.

CN120433136APending Publication Date: 2025-08-05STATE GRID ECONOMIC TECH RES INST CO LTD +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411758099.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-03
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

The traditional power distribution communication network protection mechanism is difficult to meet the reliable bearing requirements in a multi-domain network environment, and cannot effectively realize multi-domain protection, resulting in low security and resource utilization of the power distribution communication network.

Method used

Using a method based on multi-agent reinforcement learning, using elastic optical network as the bearer technology, combining network function virtualization and network coding technology, an enable distribution communication network model is built, and node features and edge features are extracted through edge node switching convolutional graph neural network, and an agent action set is constructed. Multi-agent depth deterministic strategy gradient algorithm is used for iterative learning to generate action protection decisions.

Benefits of technology

It realizes efficient multi-domain protection in the event of single node or link failure, improves the security, reliability and resource utilization of the power distribution communication network, reduces backup redundancy, and improves the efficient and reliable bearing capacity of network services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120433136A_ABST
    Figure CN120433136A_ABST
Patent Text Reader

Abstract

The invention discloses a power distribution network multi-domain protection method and system based on multi-agent reinforcement learning, and the method comprises the steps: employing an elastic optical network as a bearing technology of a mobile access network, and constructing an enabling power distribution communication network model based on network function virtualization and network coding; processing node features and edge features extracted by using the edge node switching convolutional graph neural network into node edge feature vectors; the control model slices the service request into a plurality of working slices and a backup protection slice; modeling the action set of each slice into a plurality of agents; and carrying out iterative learning on the node edge feature vector by adopting a multi-agent depth deterministic strategy gradient algorithm according to the network state of each agent so as to carry out parameter optimization on each agent, and deploying a generated action protection decision in an enabling power distribution communication network model. According to the method provided by the invention, the multi-domain protection requirement of the power distribution communication network and the safety, reliability and high efficiency of the power distribution communication network service are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of power grid protection technology, and in particular to a distribution network multi-domain protection method and system based on multi-agent reinforcement learning. Background Art

[0002] With the influx of distributed power sources and bidirectional charging stations connected to distribution networks, fault current levels and directions have changed significantly, increasing the complexity and uncertainty of distribution networks. This change requires distribution communication networks to have higher carrying capacity to ensure real-time and accurate transmission of fault information, thereby ensuring the reliability and accuracy of information transmission.

[0003] With the continuous expansion of distribution network scale and the improvement of interconnection level, multi-domain protection has become the focus of the safe operation of distribution communication network. Traditional distribution communication network protection mechanisms are mostly oriented towards the protection of a single network domain, which is difficult to meet the reliable carrying requirements of distribution communication network in an environment where multiple domain networks such as optical and wireless networks are simultaneously accessed.

[0004] It can be seen that how to realize the multi-domain protection requirements of the distribution communication network protection mechanism and improve the security, reliability and network resource utilization of the distribution communication network has become a technical problem that technical personnel in this field need to solve urgently. Summary of the Invention

[0005] The present invention provides a distribution network multi-domain protection method and system based on multi-agent reinforcement learning to solve the technical problems of how to realize the multi-domain protection requirements of the distribution communication network protection mechanism and improve the security, reliability and network resource utilization of the distribution communication network, realize the multi-domain protection requirements of the distribution communication network protection mechanism, and improve the security, reliability and efficiency of the distribution communication network service.

[0006] In a first aspect, the present invention provides a multi-domain protection method for a distribution network based on multi-agent reinforcement learning, which is applied to a distribution communication network of a 5G / B5G mobile access network. The method comprises:

[0007] Adopting elastic optical network as the bearer technology of the 5G / B5G mobile access network, and constructing an enabling distribution communication network model based on network function virtualization technology and network coding technology, the enabling distribution communication network model having a certain number of network coding nodes;

[0008] Extracting node features and edge features of the power distribution communication network model using an edge node switching convolutional graph neural network, and flattening the node features and edge features into node edge feature vectors;

[0009] Controlling the enabled power distribution communication network model to slice the received service request of the external power distribution communication network into a plurality of working slices, and using a depth-first search algorithm to construct a backup protection slice corresponding to each of the working slices;

[0010] Modeling the action sets of the working slice and the backup protection slice as several intelligent agents;

[0011] Iteratively learning the node edge feature vectors using a multi-agent deep deterministic policy gradient algorithm based on network status data of each of the intelligent agents collected through the enabled power distribution communication network to optimize parameters of each of the intelligent agents and generate action protection decisions for each of the intelligent agents;

[0012] All of the action protection decisions are deployed in the distribution communication enabled network model.

[0013] Preferably, the extracting node features and edge features of the power distribution communication network enabled by using an edge node switching convolutional graph neural network includes:

[0014] extracting a node adjacency matrix of the power distribution enabled communication network by a node embedding technique with node and edge features;

[0015] extracting a link graph of the power distribution enabled communication network by edge embedding technology with edge and node features;

[0016] Extracting the node features according to the node adjacency matrix;

[0017] The edge features are extracted according to the link graph.

[0018] Preferably, the controlling the enabling power distribution communication network model to slice the received service request of the external power distribution communication network into a plurality of working slices, and constructing a backup protection slice corresponding to each working slice using a depth-first search algorithm, including:

[0019] Controlling the enabling power distribution communication network model, and using a K shortest path algorithm to slice a received service request of the external power distribution communication network into a plurality of working slices;

[0020] The network coding node is controlled to construct a backup protection slice corresponding to each working slice using a depth-first search algorithm based on the plurality of working slices.

[0021] Preferably, the intelligent agent includes at least a first intelligent agent, a second intelligent agent and a third intelligent agent;

[0022] The action set includes at least: a DU-CU wireless baseband function deployment action set, a network coding routing action set, and a spectrum allocation action set;

[0023] The first agent corresponds to the DU-CU wireless baseband function deployment action set;

[0024] The second agent corresponds to the network coding routing action set;

[0025] The third agent corresponds to the spectrum allocation action set.

[0026] Preferably, the multi-agent deep deterministic policy gradient algorithm is used to iteratively learn the node edge feature vector to optimize the parameters of each agent and generate the action protection decision of each agent, including:

[0027] Initializing the policy network, action value network, goal network, and experience replay buffer of each agent upon entering learning mode;

[0028] Each of the intelligent agents generates a corresponding protection decision for an action to be executed using its own strategy network based on its own network status data and the network status data of other intelligent agents;

[0029] Each of the intelligent agents executes the corresponding protection decision of the action to be executed, generating a new network state and an observation reward value;

[0030] Each of the intelligent agents inputs the new network state and the protection decision of the action to be executed to the intelligent agent, as well as the new network state and the protection decision of the action to be executed to the intelligent agent, into its own action value network to calculate the expected reward value;

[0031] Updating the policy network parameters of the policy network according to the observed reward value and the expected reward value;

[0032] According to the expected reward value, the action value network parameters of the action value network are updated using a policy gradient algorithm;

[0033] Interacting the policy network parameters between different intelligent agents to enable the different intelligent agents to perform collaborative policy learning;

[0034] The learning mode is repeated until a preset termination condition is reached to generate action protection decisions for each of the intelligent agents.

[0035] In a second aspect, the present invention further provides a distribution network multi-domain protection system based on multi-agent reinforcement learning, which implements the above-mentioned distribution network multi-domain protection method based on multi-agent reinforcement learning, and the system is applied to the distribution communication network of the 5G / B5G mobile access network. The system includes: the system includes: an enabling distribution communication network architecture unit, a feature extraction unit, a reliable slice generation unit, an agent modeling unit, an agent action protection decision generation unit, and an action protection decision deployment unit;

[0036] The power distribution communication network architecture unit is configured to adopt an elastic optical network as a bearer technology for the 5G / B5G mobile access network, and to construct an power distribution communication network model based on network function virtualization technology and network coding technology, wherein the power distribution communication network model has a plurality of network coding nodes;

[0037] The feature extraction unit is configured to extract node features and edge features of the enabled power distribution communication network model using an edge node switching convolutional graph neural network, and flatten the node features and edge features into node edge feature vectors;

[0038] The reliable slice generation unit is used to control the enabled power distribution communication network model to slice the received service request of the external power distribution communication network into a plurality of working slices, and to construct a backup protection slice corresponding to each working slice using a depth-first search algorithm;

[0039] The agent modeling unit is used to model the action sets of the working slice and the backup protection slice into multiple agents;

[0040] The intelligent agent action protection decision generating unit is configured to iteratively learn the node edge feature vector using a multi-agent deep deterministic policy gradient algorithm based on the network status data of each intelligent agent collected through the enabled power distribution communication network, so as to optimize the parameters of each intelligent agent and generate an action protection decision for each intelligent agent;

[0041] The action protection decision deployment unit is used to deploy all the action protection decisions in the enabled power distribution communication network model.

[0042] Preferably, the extracting node features and edge features of the power distribution communication network enabled by using an edge node switching convolutional graph neural network includes:

[0043] extracting a node adjacency matrix of the power distribution enabled communication network by a node embedding technique with node and edge features;

[0044] extracting a link graph of the power distribution enabled communication network by edge embedding technology with edge and node features;

[0045] Extracting the node features according to the node adjacency matrix;

[0046] The edge features are extracted according to the link graph.

[0047] Preferably, the controlling the enabling power distribution communication network model to slice the received service request of the external power distribution communication network into a plurality of working slices, and constructing a backup protection slice corresponding to each working slice using a depth-first search algorithm, including:

[0048] Controlling the enabling power distribution communication network model, and using a K shortest path algorithm to slice a received service request of the external power distribution communication network into a plurality of working slices;

[0049] The network coding node is controlled to construct a backup protection slice corresponding to each working slice using a depth-first search algorithm based on the plurality of working slices.

[0050] Preferably, the intelligent agent includes at least a first intelligent agent, a second intelligent agent and a third intelligent agent;

[0051] The action set includes at least: a DU-CU wireless baseband function deployment action set, a network coding routing action set, and a spectrum allocation action set;

[0052] The first agent corresponds to the DU-CU wireless baseband function deployment action set;

[0053] The second agent corresponds to the network coding routing action set;

[0054] The third agent corresponds to the spectrum allocation action set.

[0055] Preferably, the multi-agent deep deterministic policy gradient algorithm is used to iteratively learn the node edge feature vector to optimize the parameters of each of the agents and generate the action protection decision of each of the agents, including:

[0056] Initializing the policy network, action value network, goal network, and experience replay buffer of each agent upon entering learning mode;

[0057] Each of the intelligent agents generates a corresponding protection decision for an action to be executed using its own strategy network based on its own network status data and the network status data of other intelligent agents;

[0058] Each of the intelligent agents executes the corresponding protection decision of the action to be executed, generating a new network state and an observation reward value;

[0059] Each of the intelligent agents inputs the new network state and the protection decision of the action to be executed to the intelligent agent, as well as the new network state and the protection decision of the action to be executed to the intelligent agent, into its own action value network to calculate the expected reward value;

[0060] Updating the policy network parameters of the policy network according to the observed reward value and the expected reward value;

[0061] According to the expected reward value, the action value network parameters of the action value network are updated using a policy gradient algorithm;

[0062] Interacting the policy network parameters between different intelligent agents to enable the different intelligent agents to perform collaborative policy learning;

[0063] The learning mode is repeated until a preset termination condition is reached to generate action protection decisions for each of the intelligent agents.

[0064] The present invention provides a method and system for multi-domain protection of a distribution network based on multi-agent reinforcement learning. Compared with the prior art, the embodiments of the present invention have the following advantages:

[0065] (1) The "N+1" protection method based on network coding models the network slice deployment problem of multi-domain protection as an interactive process of multiple intelligent agents in the distribution network mobile communication environment. Each intelligent agent continuously optimizes its decision-making according to the reward value fed back by the distribution communication network environment through distributed execution and centralized training. This method can achieve more intelligent and efficient multi-domain protection when a single node or link fails, without the need for backup redundancy, and significantly improves the efficient and reliable carrying capacity of the distribution communication network service.

[0066] (2) Ensure the security, reliability and efficiency of distribution communication network services. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1 This is a schematic diagram of the steps of a distribution network multi-domain protection method based on multi-agent reinforcement learning provided by a preferred embodiment of the present invention;

[0068] Figure 2 It is a structural diagram of a distribution network multi-domain protection system based on multi-agent reinforcement learning provided by a preferred embodiment of the present invention. DETAILED DESCRIPTION

[0069] The following is a detailed explanation of the embodiments of the present invention in conjunction with the accompanying drawings. The embodiments are provided for illustrative purposes only and cannot be understood as limitations on the present invention. The accompanying drawings are for reference and illustration purposes only and do not constitute a limitation on the scope of protection of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. In the description of the present invention, the terms "first", "second", "third", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first", "second", "third", etc. may explicitly or implicitly include one or more of the features. In the description of the present invention, unless otherwise specified, the meaning of "multiple" is two or more.

[0070] In the description of the present invention, it should be noted that, unless otherwise expressly specified and limited, the terms "installed", "connected" and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or an indirect connection through an intermediate medium, or it can be a communication between the two components. The terms "vertical", "horizontal", "left", "right", "up", "down" and similar expressions used herein are for illustrative purposes only, and do not indicate or imply that the device or component referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention. The term "and / or" used herein includes any and all combinations of one or more related listed items. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0071] In describing the present invention, it should be noted that, unless otherwise defined, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art. The terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. Those skilled in the art will understand the specific meanings of the above terms in the present invention in specific circumstances.

[0072] It's important to note that fifth-generation mobile communications (5G) and the post-5G (B5G) era are becoming increasingly effective communications technologies for distribution networks, profoundly impacting the development of the distribution network industry. 5G / B5G mobile access networks will also become a key enabling technology for the rapid construction of new power systems. 5G / B5G wireless access networks utilize a three-layer architecture, consisting of a centralized unit (CU), distributed unit (DU), and active antenna unit (AAU).

[0073] In view of this, the present invention provides a multi-domain protection method for distribution network based on multi-agent reinforcement learning, which is applied to the distribution communication network of 5G / B5G mobile access network. For details, please refer to Figure 1 , the method comprising:

[0074] S1. Adopt elastic optical network as the bearer technology of the 5G / B5G mobile access network, and build an enabling distribution communication network model based on network function virtualization technology and network coding technology. The enabling distribution communication network model has a certain number of network coding nodes.

[0075] S2. Utilize an edge node switching convolutional graph neural network to extract node features and edge features of the power distribution communication network model, and flatten the node features and edge features into node edge feature vectors.

[0076] S3. Control the enabled power distribution communication network model to slice the received service request of the external power distribution communication network into several working slices, and use a depth-first search algorithm to construct a backup protection slice corresponding to each of the working slices.

[0077] S4. Model the action sets of the working slice and the backup protection slice into several intelligent agents.

[0078] S5. Based on the network status data of each intelligent agent collected through the enabled power distribution communication network, a multi-agent deep deterministic policy gradient algorithm is used to iteratively learn the node edge feature vector to optimize the parameters of each intelligent agent and generate an action protection decision for each intelligent agent.

[0079] S6. Deploy all the action protection decisions in the enabled power distribution communication network model.

[0080] Optical networks, due to their abundant bandwidth resources and low cost, are considered one of the ideal bearer technologies for 5G / B5G mobile access networks in power distribution networks. In a preferred embodiment of the present invention, an elastic optical network is used as the bearer technology for the 5G / B5G mobile access network. To achieve reliable and efficient bearering of the 5G / B5G mobile access network carried by the elastic optical network for power distribution communication network protection services, network coding technology is introduced into some nodes of the power distribution communication network to implement network coding operations. The nodes performing network coding operations are set as network coding nodes, and the network coding nodes perform encoding operations on the incoming data streams.

[0081] Furthermore, based on network function virtualization technology, virtual network function units (NFUs) for active antenna units (AAUs), distributed units (DUs), and centralized units (CUs) are deployed in general-purpose processors (GPPs). Each fiber link provides multiple frequency slots, each with a bandwidth of 6.25 GHz. Bandwidth-variable transponders (BVTs), bandwidth-variable optical cross-connects (BV-OXCs), optoelectronic optical converters (OEOs), and general-purpose processors are deployed at the same site to build an enabling distribution communication network. The data streams carrying service requests in the enabling distribution communication network flow from the base station (AAU), through network nodes and network coding nodes, and ultimately converge into the multi-access edge computing (MEC) for data processing.

[0082] Furthermore, an edge node switching convolutional graph neural network is used to extract node and edge features of the power distribution communication network. Specifically, a node embedding technique with node and edge features is used to extract the node adjacency matrix of the power distribution communication network, which is represented as G. An edge embedding technique with edge and node features is used to extract the link graph of the power distribution communication network, which is represented as a line graph L(G).

[0083] Node embedding technology with node and edge features is an important concept in graph neural networks, which aims to map nodes and edges in a graph into a low-dimensional vector space while preserving the structural and feature information in the graph.

[0084] Node embedding technology represents each node in a graph as a low-dimensional vector that captures the similarities and relationships between nodes. Node embedding technology makes similar nodes closer in the vector space, while dissimilar nodes are farther apart.

[0085] In a graph structure, edges not only represent connections between nodes but can also carry additional information, namely edge features. Edge features can describe the type, strength, or other properties of relationships between nodes. For example, in a social network, edge features can represent the type of friendship between users (such as friends, colleagues, family members, etc.); in a transportation network, edge features can represent the type of road (such as highways, city streets, etc.) or traffic volume.

[0086] Node embedding technology with node and edge features can capture the information in the graph more comprehensively. Node features provide the inherent properties of nodes, while edge features describe the type and strength of the relationship between nodes. Combining the two can generate richer and more accurate node features.

[0087] Edge embedding, which uses both edge and node features, is also a key concept in graph neural networks. Edge embedding represents each edge in a graph as a low-dimensional vector that captures the similarities and relationships between edges. Edge embedding places similar edges closer together in vector space, while dissimilar edges are further apart. This helps understand the connection patterns and relationship types within a graph and is crucial for graph data analysis and mining.

[0088] In a graph, both edges and nodes carry important information. Edge features describe the type, strength, or other properties of the relationship between nodes, while node features provide intrinsic properties of the nodes. Edge embedding techniques that combine edge and node features can more comprehensively capture information in the graph, improving the accuracy and practicality of edge embeddings.

[0089] Two forward-pass feature propagation rules are adopted for the node adjacency matrix and link graph to alternately update the network feature embeddings of nodes and edges.

[0090] Furthermore, node features are extracted based on the node adjacency matrix, and edge features are extracted based on the link graph. The edge node switching convolutional graph neural network flattens node and edge features into a single vector, which is composed of both node and edge features and is known as the node edge feature vector. Using this method to extract the node edge feature vector effectively allows for the perception of network conditions and the acquisition of network status information.

[0091] In the preferred embodiment of the present invention, it is assumed that only one fault occurs at a time in the enabled distribution communication network, namely, a node fault or a link fault, i.e., a PP node (i.e., BV-OXC, BVT, GPP) fault or a fiber link fault. In the enabled distribution communication network, each service sends a request to the protection terminal, and the protection terminal shares the protection action signal through the 5G / B5G mobile access network.

[0092] In a preferred embodiment of the present invention, there is a situation where optical networks and wireless multi-domain networks are jointly connected to the power grid system. The traditional distribution network reliable protection mechanism cannot be dynamically adaptively adapted to the communication technology, making the distribution network protection equipment prone to frequent locking due to fluctuations in communication quality, increasing the difficulty of operation and maintenance of distribution network protection services. Therefore, the preferred embodiment of this application adopts a multi-domain protection method. For each service request, a multi-domain network slice is constructed according to the connection order of the enabled distribution communication network composed of AAU, DU, CU, and MEC, and then embedded into the enabled distribution communication network.

[0093] How to build a multi-domain protection reliable slice based on network coding to ensure the reliability of the distribution communication network, realize multi-domain protection of data transmission and network functions, and minimize the total backup redundancy are the main technical problems to be solved by the deployment of multi-domain protection reliable slices.

[0094] To solve the above technical problems, in a preferred embodiment of the present invention, an "N+1" protection method based on network coding is adopted. In terms of data transmission and network functions, working slices and reliable slices are executed through DU-CU wireless baseband function deployment actions, network coding routing actions and spectrum allocation actions of N working slices and 1 backup protection slice in multi-domain protection.

[0095] In the "N+1" protection method based on network coding, only one path needs to be added as a backup protection path for N working paths. If any node or link fails, the destination node can decode the encoded information received from the backup protection path to restore the original N signals. Compared with the "1+1" protection method, the "N+1" protection method can effectively save spectrum resources and eliminate a large number of redundant signals.

[0096] Using the "N+1" protection method based on network coding, a service request is split into three parts, x, y, and z, at source node s0. These are transmitted to destination node D via three working paths. The protection path is used to transmit the data x⊕y⊕z, which is the result of an XOR operation on x, y, and z. The number of spectrum slots consumed by x⊕y⊕z on a protection path is the same as the number of spectrum slots consumed by x, y, or z on each working path. However, the protection path carries all the information about x, y, and z. If a link from s0 to D fails, destination node D can decode x⊕y⊕z and recover the information about x, y, and / or z.

[0097] After slicing the service request into several working slices and one backup protection slice, the action sets of the working slices and the backup protection slices are modeled as several intelligent agents. The action sets include at least: DU-CU wireless baseband function deployment action set, network coding routing action set and spectrum allocation action set.

[0098] An intelligent agent is an agent that can perceive its environment, make decisions, and take actions to achieve specific goals. Intelligent agents possess the following characteristics: autonomy: agents can make decisions independently, without relying on external instructions; adaptability: agents can adjust their behavior to environmental changes and maintain effectiveness; interactivity: agents can communicate and cooperate with their environment and other agents, enabling them to work together in multi-agent systems to complete tasks; and learning ability: agents can optimize their decision-making processes and improve the efficiency of task completion through continuous experience accumulation, allowing them to continuously adapt to new environments and tasks and achieve continuous evolution.

[0099] Based on the working slice and backup protection slice, several intelligent agents are constructed, including at least a first intelligent agent, a second intelligent agent, and a third intelligent agent. Each intelligent agent is responsible for different action protection decisions. Specifically, the first intelligent agent corresponds to the DU-CU wireless baseband function deployment action set, the second intelligent agent corresponds to the network coding routing action set, and the third intelligent agent corresponds to the spectrum allocation action set. Each intelligent agent includes a policy network, an action value network, a target network, and an experience replay buffer. The policy network is used to generate the actions of the intelligent agent, while the action value network is used to evaluate the reward value and future expected reward value obtained by the action protection decision taken by the intelligent agent.

[0100] Furthermore, by enabling the distribution communication network to collect the network status data of each intelligent agent, a multi-agent deep deterministic policy gradient algorithm is used to iteratively learn the node edge feature vector based on the network status data of each intelligent agent to optimize the parameters of each intelligent agent and generate the action protection decision of each intelligent agent.

[0101] The network status data includes at least: PP node capacity, MEC node capacity, AAU node spectrum resource block, frequency slot capacity and fiber link distance.

[0102] The Multi-Agent Deep Deterministic Policy Gradient (MDPG) algorithm is a reinforcement learning algorithm used in multi-agent environments, designed to address collaborative decision-making in multi-agent systems. The core concept of the MDPG algorithm is that in a multi-agent environment, the behavior of each agent depends not only on the state of the environment but also on the policies of other agents. Therefore, the algorithm uses an independent policy network and action-value network architecture for each agent and considers the policy information of other agents during training to improve learning performance and stability.

[0103] The multi-agent deep deterministic policy gradient algorithm consists of the following steps:

[0104] When entering learning mode, each agent's policy network, action-value network, target network, and experience replay buffer are initialized.

[0105] Furthermore, each agent uses its own strategy network to generate corresponding protection decisions for actions to be executed based on its own network status data and the network status data of other agents.

[0106] Furthermore, each agent executes its corresponding pending action protection decision to generate a new network state and observation reward value.

[0107] Furthermore, each agent inputs its own new network state and pending action protection decision, as well as the new network state and pending action protection decision of other agents into its own action value network to calculate the expected reward value.

[0108] In a preferred embodiment of the present invention, the reward function equation is:

[0109]

[0110] Among them, a t Indicates the protection strategy for the action to be executed, F t Indicates a t The reward value when all constraints are met, r max Indicates a t The reward value when the constraint is not met.

[0111] F t is a large reward value, indicating that multi-domain reliable deployment consumes bandwidth resources. t When all constraints are met, it is a valid action and generates a larger reward value; r max Is a small reward value or negative reward as punishment, when action a t When the requirements are not met, it is an invalid action, and a smaller reward value or a negative reward is generated as a punishment.

[0112] Furthermore, based on the observed and expected rewards, the policy network parameters are updated to improve the prediction accuracy of the value function. During the parameter update process, other loss terms, such as regularization terms, also need to be considered.

[0113] Furthermore, according to the expected reward value, the policy gradient algorithm is used to update the action value network parameters of the action value network to optimize the action protection decision of the intelligent agent.

[0114] In a policy gradient algorithm, the agent's action-protection decisions are represented by a parameterized function of the policy network, which directly outputs the probability distribution of the action-protection decision to be taken in a given state. The policy gradient algorithm updates the action-value network parameters by calculating the gradient of the target network function with respect to the policy network parameters, thereby improving the decision.

[0115] Furthermore, the policy network parameters are interacted between different agents to enable them to make collaborative decisions.

[0116] Agents can exchange their own policy network parameters or gradient information, or indirectly achieve collaborative decision-making learning between different agents by sharing the experience in the experience replay buffer. This interaction or sharing helps agents better adapt to the multi-agent environment during training.

[0117] Finally, the above learning mode is repeated until the preset termination condition is reached to generate the action protection decision of each agent.

[0118] The "N+1" protection method based on network coding models the network slicing deployment problem of multi-domain protection as an interactive process of multiple intelligent agents in the distribution network mobile communication network environment. Each intelligent agent continuously optimizes its decision-making through distributed execution and centralized training based on the reward value feedback from the distribution communication network environment. This method can achieve more intelligent and efficient multi-domain protection when a single node or link fails, without the need for backup redundancy, significantly improving the efficient and reliable carrying capacity of the distribution communication network service.

[0119] The present invention proposes a multi-domain protection of distribution network based on multi-agent reinforcement learning. Specifically, elastic optical network is adopted as the bearer technology of 5G / B5G mobile access network, and an enabling distribution communication network model is constructed based on network function virtualization technology and network coding technology. The enabling distribution communication network model has a number of network coding nodes; the edge node switching convolutional graph neural network is used to extract the node features and edge features of the enabling distribution communication network model, and the node features and edge features are flattened into node edge feature vectors; the enabling distribution communication network model is controlled to slice the service request received from the external distribution communication network into several working slices, and a depth-first search algorithm is used to construct a backup protection slice corresponding to each working slice; the action sets of the working slices and the backup protection slices are modeled as several agents; according to the network status data of each agent collected through the enabling distribution communication network, a multi-agent deep deterministic policy gradient algorithm is used to iteratively learn the node edge feature vectors to optimize the parameters of each agent and generate action protection decisions for each agent; all action protection decisions are deployed in the enabling distribution communication network model. The multi-domain protection method for distribution networks based on multi-agent reinforcement learning provided in this application realizes the multi-domain protection requirements of the distribution communication network protection mechanism, as well as the security, reliability and efficiency of the distribution communication network services.

[0120] Accordingly, if Figure 2 As shown, based on a distribution network multi-domain protection method based on multi-agent reinforcement learning, an embodiment of the present invention further provides a distribution network multi-domain protection system based on multi-agent reinforcement learning, which implements the distribution network multi-domain protection method based on multi-agent reinforcement learning disclosed in an embodiment of the present invention. The system is applied to the distribution communication network of the 5G / B5G mobile access network, and the system includes: an enabling distribution communication network architecture unit 1, a feature extraction unit 2, a reliable slice generation unit 3, an agent modeling unit 4, an agent action protection decision generation unit 5 and an action protection decision deployment unit 6;

[0121] The power distribution communication network architecture unit 1 is configured to adopt an elastic optical network as a bearer technology for the 5G / B5G mobile access network, and to construct an power distribution communication network model based on network function virtualization technology and network coding technology. The power distribution communication network model has a plurality of network coding nodes.

[0122] The feature extraction unit 2 is used to extract node features and edge features of the enabled power distribution communication network model using an edge node switching convolutional graph neural network, and flatten the node features and edge features into node edge feature vectors;

[0123] The reliable slice generation unit 3 is used to control the enabled power distribution communication network model to slice the received service request of the external power distribution communication network into a plurality of working slices, and to construct a backup protection slice corresponding to each working slice using a depth-first search algorithm;

[0124] The agent modeling unit 4 is used to model the action sets of the working slice and the backup protection slice into several agents;

[0125] The agent action protection decision generating unit 5 is configured to iteratively learn the node edge feature vectors using a multi-agent deep deterministic policy gradient algorithm based on the network status data of each agent collected through the enabled power distribution communication network, so as to optimize the parameters of each agent and generate an action protection decision for each agent;

[0126] The action protection decision deployment unit 6 is configured to deploy all the action protection decisions in the enabled power distribution communication network model.

[0127] Furthermore, in the feature extraction unit 2, the extracting of node features and edge features of the power distribution communication network by using the edge node switching convolutional graph neural network includes:

[0128] extracting a node adjacency matrix of the power distribution enabled communication network by a node embedding technique with node and edge features;

[0129] extracting a link graph of the power distribution enabled communication network by edge embedding technology with edge and node features;

[0130] Extracting the node features according to the node adjacency matrix;

[0131] The edge features are extracted according to the link graph.

[0132] Furthermore, in the reliable slice generation unit 3, the control unit 3 slices the service request received from the external power distribution communication network into a plurality of working slices, and constructs a backup protection slice corresponding to each working slice using a depth-first search algorithm, including:

[0133] Controlling the enabling power distribution communication network model, and using a K shortest path algorithm to slice a received service request of the external power distribution communication network into a plurality of working slices;

[0134] The network coding node is controlled to construct a backup protection slice corresponding to each working slice using a depth-first search algorithm based on the plurality of working slices.

[0135] Furthermore, the intelligent agent includes at least a first intelligent agent, a second intelligent agent and a third intelligent agent;

[0136] The action set includes at least: a DU-CU wireless baseband function deployment action set, a network coding routing action set, and a spectrum allocation action set;

[0137] The first agent corresponds to the DU-CU wireless baseband function deployment action set;

[0138] The second agent corresponds to the network coding routing action set;

[0139] The third agent corresponds to the spectrum allocation action set.

[0140] Furthermore, in the action protection decision deployment unit 6, the multi-agent deep deterministic policy gradient algorithm is used to iteratively learn the node edge feature vector to optimize the parameters of each of the agents, and generate the action protection decision of each of the agents, including:

[0141] Initializing the policy network, action value network, goal network, and experience replay buffer of each agent upon entering learning mode;

[0142] Each of the intelligent agents generates a corresponding protection decision for an action to be executed using its own strategy network based on its own network status data and the network status data of other intelligent agents;

[0143] Each of the intelligent agents executes the corresponding protection decision of the action to be executed, generating a new network state and an observation reward value;

[0144] Each of the intelligent agents inputs the new network state and the protection decision of the action to be executed to the intelligent agent, as well as the new network state and the protection decision of the action to be executed to the intelligent agent, into its own action value network to calculate the expected reward value;

[0145] Updating the policy network parameters of the policy network according to the observed reward value and the expected reward value;

[0146] According to the expected reward value, the action value network parameters of the action value network are updated using a policy gradient algorithm;

[0147] Performing interaction of the policy network parameters between different intelligent agents to enable the different intelligent agents to perform collaborative policy learning;

[0148] The learning mode is repeated until a preset termination condition is reached to generate action protection decisions for each of the intelligent agents.

[0149] For the specific definition of a distribution network multi-domain protection system based on multi-agent reinforcement learning, please refer to the above-mentioned definition of a distribution network multi-domain protection method based on multi-agent reinforcement learning, which will not be repeated here. A person of ordinary skill in the art will appreciate that the modules and steps described in conjunction with the embodiments disclosed in the present invention can be implemented in hardware, software, or a combination of both. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0150] A distribution network multi-domain protection method and system based on multi-agent reinforcement learning is provided in this embodiment to solve the technical problems of how to realize the multi-domain protection requirements of the distribution communication network protection mechanism, and to improve the security, reliability and network resource utilization of the distribution communication network. Elastic optical network is adopted as the bearer technology of 5G / B5G mobile access network. Based on network function virtualization technology and network coding technology, an enabling distribution communication network model is constructed, and the enabling distribution communication network model has a certain number of network coding nodes; the edge node switching convolutional graph neural network is used to extract the node features and edge features of the enabling distribution communication network model, and the node features and edge features are flattened into node edge feature vectors; the enabling distribution communication network model is controlled to slice the service requests received from the external distribution communication network into several working slices, and a depth-first search algorithm is used to construct a backup protection slice corresponding to each working slice; the action sets of the working slices and the backup protection slices are modeled as several intelligent agents; according to the network status data of each intelligent agent collected through the enabling distribution communication network, a multi-agent deep deterministic policy gradient algorithm is used to iteratively learn the node edge feature vectors to optimize the parameters of each intelligent agent and generate the action protection decision of each intelligent agent; all action protection decisions are deployed in the enabling distribution communication network model. The multi-domain protection method for distribution networks based on multi-agent reinforcement learning provided in this application realizes the multi-domain protection requirements of the distribution communication network protection mechanism, as well as the security, reliability and efficiency of the distribution communication network services.

[0151] Each embodiment in this specification is described in a progressive manner, and the same or similar parts of each embodiment can be directly referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. It should be noted that the various technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the various technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0152] The above-described embodiments merely represent several preferred embodiments of the present invention. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that a person skilled in the art could make several improvements and substitutions without departing from the technical principles of the present invention, and such improvements and substitutions should also be considered within the scope of the present invention. Therefore, the scope of the present invention should be determined by the scope of the claims.

Claims

1. A multi-domain protection method for distribution network based on multi-agent reinforcement learning, characterized in that: The method is applied to a power distribution communication network of a 5G / B5G mobile access network, and the method comprises: Adopting elastic optical network as the bearer technology of the 5G / B5G mobile access network, and constructing an enabling distribution communication network model based on network function virtualization technology and network coding technology, the enabling distribution communication network model having a certain number of network coding nodes; Extracting node features and edge features of the power distribution communication network model using an edge node switching convolutional graph neural network, and flattening the node features and edge features into node edge feature vectors; Controlling the enabled power distribution communication network model to slice the received service request of the external power distribution communication network into a plurality of working slices, and using a depth-first search algorithm to construct a backup protection slice corresponding to each of the working slices; Modeling the action sets of the working slice and the backup protection slice as several intelligent agents; Iteratively learning the node edge feature vectors using a multi-agent deep deterministic policy gradient algorithm based on network status data of each of the intelligent agents collected through the enabled power distribution communication network to optimize parameters of each of the intelligent agents and generate action protection decisions for each of the intelligent agents; All of the action protection decisions are deployed in the distribution communication enabled network model.

2. The distribution network multi-domain protection method based on multi-agent reinforcement learning according to claim 1 is characterized in that: The method of extracting node features and edge features of the power distribution communication network by using an edge node switching convolutional graph neural network includes: extracting a node adjacency matrix of the power distribution enabled communication network by a node embedding technique with node and edge features; extracting a link graph of the power distribution enabled communication network by edge embedding technology with edge and node features; Extracting the node features according to the node adjacency matrix; The edge features are extracted according to the link graph.

3. The distribution network multi-domain protection method based on multi-agent reinforcement learning according to claim 1 is characterized in that: The controlling the enabling power distribution communication network model to slice the received service request of the external power distribution communication network into a plurality of working slices, and constructing a backup protection slice corresponding to each working slice using a depth-first search algorithm, including: Controlling the enabling power distribution communication network model, and using a K shortest path algorithm to slice a received service request of the external power distribution communication network into a plurality of working slices; The network coding node is controlled to construct a backup protection slice corresponding to each working slice using a depth-first search algorithm based on the plurality of working slices.

4. The distribution network multi-domain protection method based on multi-agent reinforcement learning according to claim 1 is characterized in that: The intelligent agents include at least a first intelligent agent, a second intelligent agent and a third intelligent agent; The action set includes at least: a DU-CU wireless baseband function deployment action set, a network coding routing action set, and a spectrum allocation action set; The first agent corresponds to the DU-CU wireless baseband function deployment action set; The second agent corresponds to the network coding routing action set; The third agent corresponds to the spectrum allocation action set.

5. The distribution network multi-domain protection method based on multi-agent reinforcement learning according to claim 1 is characterized in that: The multi-agent deep deterministic policy gradient algorithm is used to iteratively learn the node edge feature vector to optimize the parameters of each of the agents and generate action protection decisions for each of the agents, including: Initializing the policy network, action value network, goal network, and experience replay buffer of each agent upon entering learning mode; Each of the intelligent agents generates a corresponding protection decision for an action to be executed using its own strategy network based on its own network status data and the network status data of other intelligent agents; Each of the intelligent agents executes the corresponding protection decision of the action to be executed, generating a new network state and an observation reward value; Each of the intelligent agents inputs the new network state and the protection decision of the action to be executed to the intelligent agent, as well as the new network state and the protection decision of the action to be executed to the intelligent agent, into its own action value network to calculate the expected reward value; Updating the policy network parameters of the policy network according to the observed reward value and the expected reward value; According to the expected reward value, the action value network parameters of the action value network are updated using a policy gradient algorithm; Interacting the policy network parameters between different intelligent agents to enable the different intelligent agents to perform collaborative policy learning; The learning mode is repeated until a preset termination condition is reached to generate action protection decisions for each of the intelligent agents.

6. A distribution network multi-domain protection system based on multi-agent reinforcement learning, which implements the distribution network multi-domain protection method based on multi-agent reinforcement learning according to any one of claims 1 to 5, characterized in that: The system is applied to the power distribution communication network of the 5G / B5G mobile access network, and the system includes: an enabling power distribution communication network architecture unit, a feature extraction unit, a reliable slice generation unit, an intelligent agent modeling unit, an intelligent agent action protection decision generation unit, and an action protection decision deployment unit; The power distribution communication network architecture unit is configured to adopt an elastic optical network as a bearer technology for the 5G / B5G mobile access network, and to construct an power distribution communication network model based on network function virtualization technology and network coding technology, wherein the power distribution communication network model has a plurality of network coding nodes; The feature extraction unit is configured to extract node features and edge features of the enabled power distribution communication network model using an edge node switching convolutional graph neural network, and flatten the node features and edge features into node edge feature vectors; The reliable slice generation unit is used to control the enabled power distribution communication network model to slice the received service request of the external power distribution communication network into a plurality of working slices, and to construct a backup protection slice corresponding to each working slice using a depth-first search algorithm; The agent modeling unit is used to model the action sets of the working slice and the backup protection slice into multiple agents; The intelligent agent action protection decision generating unit is configured to iteratively learn the node edge feature vector using a multi-agent deep deterministic policy gradient algorithm based on the network status data of each intelligent agent collected through the enabled power distribution communication network, so as to optimize the parameters of each intelligent agent and generate an action protection decision for each intelligent agent; The action protection decision deployment unit is used to deploy all the action protection decisions in the enabled power distribution communication network model.

7. The distribution network multi-domain protection system based on multi-agent reinforcement learning according to claim 6, characterized in that: The method of extracting node features and edge features of the power distribution communication network by using an edge node switching convolutional graph neural network includes: extracting a node adjacency matrix of the power distribution enabled communication network by a node embedding technique with node and edge features; extracting a link graph of the power distribution enabled communication network by edge embedding technology with edge and node features; Extracting the node features according to the node adjacency matrix; The edge features are extracted according to the link graph.

8. The distribution network multi-domain protection system based on multi-agent reinforcement learning according to claim 6, characterized in that: The controlling the enabling power distribution communication network model to slice the received service request of the external power distribution communication network into a plurality of working slices, and constructing a backup protection slice corresponding to each working slice using a depth-first search algorithm, including: Controlling the enabling power distribution communication network model, and using a K shortest path algorithm to slice a received service request of the external power distribution communication network into a plurality of working slices; The network coding node is controlled to construct a backup protection slice corresponding to each working slice using a depth-first search algorithm based on the plurality of working slices.

9. The distribution network multi-domain protection system based on multi-agent reinforcement learning according to claim 6, characterized in that: The intelligent agents include at least a first intelligent agent, a second intelligent agent and a third intelligent agent; The action set includes at least: a DU-CU wireless baseband function deployment action set, a network coding routing action set, and a spectrum allocation action set; The first agent corresponds to the DU-CU wireless baseband function deployment action set; The second agent corresponds to the network coding routing action set; The third agent corresponds to the spectrum allocation action set.

10. The distribution network multi-domain protection system based on multi-agent reinforcement learning according to claim 6, characterized in that: The multi-agent deep deterministic policy gradient algorithm is used to iteratively learn the node edge feature vector to optimize the parameters of each of the agents and generate action protection decisions for each of the agents, including: Initializing the policy network, action value network, goal network, and experience replay buffer of each agent upon entering learning mode; Each of the intelligent agents generates a corresponding protection decision for an action to be executed using its own strategy network based on its own network status data and the network status data of other intelligent agents; Each of the intelligent agents executes the corresponding protection decision of the action to be executed, generating a new network state and an observation reward value; Each of the intelligent agents inputs the new network state and the protection decision of the action to be executed to the intelligent agent, as well as the new network state and the protection decision of the action to be executed to the intelligent agent, into its own action value network to calculate the expected reward value; Updating the policy network parameters of the policy network according to the observed reward value and the expected reward value; According to the expected reward value, the action value network parameters of the action value network are updated using a policy gradient algorithm; Interacting the policy network parameters between different intelligent agents to enable the different intelligent agents to perform collaborative policy learning; The learning mode is repeated until a preset termination condition is reached to generate action protection decisions for each of the intelligent agents.

Citation Information

Patent Citations

  • Mobile edge computing unloading method based on multi-agent reinforcement learning

    CN112367353A

  • Power distribution network multi-region cooperative reactive power optimization method based on multi-agent reinforcement learning

    CN115483703A