Substation communication routing path automatic planning method based on shortest path

By combining the Dijkstra algorithm and deep reinforcement learning model, dynamically optimizing the communication routing path of the substation is solved, and the problem that the routing planning method in the existing technology is difficult to adapt to the complex dynamic environment, achieving efficient and reliable communication network management.

CN120342938APending Publication Date: 2025-07-18CHUZHOU POWER SUPPLY CO OF STATE GRID ANHUI ELECTRIC POWER CORP
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510501522.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The routing planning method of existing substation communication networks relies on manual experience and is difficult to adapt to complex and dynamic communication environments, resulting in low routing efficiency, uneven resource allocation, and slow fault response, which cannot meet the needs of smart grids for high reliability, low latency and adaptive adjustment of communication.

Method used

The deep reinforcement learning model based on the shortest path is adopted, combined with the Dijkstra algorithm and the deep reinforcement learning model, and the communication path is dynamically adjusted to adapt to changes in network state by modeling the substation communication network, calculating and optimizing the routing path in real time.

Benefits of technology

It realizes automatic planning of substation communication networks, which can adapt to network needs of different scales and complexities, reduce communication delays, improve network efficiency, ensure communication reliability and stability, and reduce manual intervention and operation and maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120342938A_ABST
    Figure CN120342938A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of electric power communication, and particularly discloses a substation communication routing path automatic planning method based on a shortest path, which is used for solving the problems of low routing efficiency, non-uniform resource allocation and slow fault response due to the fact that an existing routing planning method always depends on artificial experience and is difficult to adapt to a complex and dynamic communication environment. According to the method, modeling is carried out on the substation communication network, and the substation communication routing path automatic planning algorithm based on the combination of the shortest path algorithm and the deep reinforcement learning is combined with the high efficiency of the shortest path algorithm and the complex data processing capability of the deep reinforcement learning; forward propagation calculation is carried out on the network state in real time by using the deep reinforcement learning model, the optimal action is dynamically selected according to the current state, automatic planning of the substation communication routing path is realized, and substation communication network routing planning requirements of different scales and complexities can be met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power communication, and more specifically, to an automatic planning method for substation communication routing paths based on the shortest path. Background Art

[0002] With the rapid development of smart grids, the complexity and data volume of substation communication networks have been increasing continuously, posing higher requirements for the planning of communication routing paths. The communication planning between substations can be effectively expressed using graph theory, which is a branch of mathematics that studies graphs. A graph consists of several given points and the edges (lines) connecting two points, and is used to describe a specific relationship between certain things. The points represent things, and the lines connecting two points indicate that there is a certain relationship between the corresponding two things. Each substation and line junction box can be regarded as a point, and the communication between two junction boxes can be regarded as the line connecting two points. At the same time, the line weight can be set according to the line length, signal attenuation, or the number of fiber cores used in the optical cable and the service level to analyze the carrying capacity and importance of the optical cable, thus forming a graph. The realization of routing path planning is to solve the shortest path under relevant constraint conditions, which is a typical optimization problem and is widely used to solve many problems in production practice, such as routing restoration after earthquake disasters, pipeline laying, line arrangement, plant layout, equipment renewal, etc., and is often used as a basic tool to solve other optimization problems. The existing literature (Liu Baoju. Research on Service-Driven High-Reliability Routing Algorithm in Smart Grid Communication Network [D]. Beijing University of Posts and Telecommunications, 2021. DOI: 10.26969 / d.cnki.gbydu.2021.000045.) gives Figure 2 The schematic diagram of routing restoration after earthquake disasters is shown as follows.

[0003] Modern substation communication networks need to process a large amount of real-time data from various devices (such as intelligent terminals, relay protection devices, measurement and control devices, etc.) to ensure the safe, stable and efficient operation of the power system. However, with the increase in communication devices and the sharp increase in data traffic, the complexity of substation communication networks is constantly rising, posing higher requirements for the planning of communication routing paths. Traditional communication routing planning methods mainly rely on manual experience or static configuration-based strategies, and often have difficulty adapting to complex and dynamic communication environments, with problems such as low routing efficiency, uneven resource allocation, and slow fault response. This method lacks flexibility and real-time performance in the face of emergencies, traffic fluctuations or equipment failures, and is difficult to meet the requirements of high reliability, low latency and adaptive adjustment of communication in smart grids. Therefore, developing an algorithm that can automatically plan substation communication routing paths has important practical significance and application value. To solve the above problems, a technical solution is provided herein. Summary of the Invention

[0004] To overcome the above defects of the prior art, the present invention provides an automatic planning method for substation communication routing paths based on the shortest path, which uses a deep reinforcement learning model to perform forward propagation calculations on the network state in real time, dynamically selects the optimal action according to the current state, and solves the problems that existing routing planning methods often rely on manual experience, are often difficult to adapt to complex and dynamic communication environments, and have low routing efficiency, uneven resource allocation, and slow fault response, so as to solve the problems raised in the above background technology.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] An automatic planning method for substation communication routing paths based on the shortest path, comprising the following steps:

[0007] Step 1, network modeling: Model the substation communication network, regarding substations, control centers, and other communication nodes as nodes in the graph, and the communication links between nodes as edges in the graph;

[0008] Step 2, initial path planning: Use the Dijkstra algorithm to calculate the initial shortest path from the source node to the target node in the substation communication network, where the Dijkstra algorithm is used to calculate the shortest path of the entire network;

[0009] Step 3, path dynamic optimization: Adopt a deep reinforcement learning model, and dynamically adjust the routing path through policy optimization based on the real-time state of the network;

[0010] Step 4, path adjustment and execution: When the network state changes, adjust the communication path in real time according to the prediction result of the deep reinforcement learning model;

[0011] Step 5, result evaluation optimization and output: Output the result of adjusting the communication path in text form for users to reference and use, and the output result includes path information (nodes and links on the path), path performance indicators (length, stability, bandwidth).

[0012] As a further solution of the present invention, in Step 1, the specific steps of network modeling are:

[0013] Step 11, Represent the substation communication network as a graph G=(V, E), where V is the set of nodes, E is the set of edges, and for each edge e∈E, and each edge e has a weight, which is used to find the shortest path from the source node s to the target node t;

[0014] Step 12, calculate the weight w(e) of edge e according to the latency, bandwidth, and reliability of the link, where the latency d(e) of the link represents the latency of the link corresponding to edge e, the bandwidth b(e) of the link represents the bandwidth of the link corresponding to edge e, the reliability r(e) of the link represents the reliability of the link corresponding to edge e, and the weight w(e) of edge e is calculated by weighted sum. The calculation formula is:

[0015]

[0016] In the formula: w(e) is the weight of edge e, α is the weight coefficient of the latency of the link corresponding to edge e, d(e) is the latency of the link corresponding to edge e, β is the weight coefficient of the bandwidth of the link corresponding to edge e, b(e) is the bandwidth of the link corresponding to edge e, γ is the weight coefficient of the reliability of the link corresponding to edge e, and r(e) is the reliability of the link corresponding to edge e.

[0017] As a further solution of the present invention, in step 2, the specific steps for calculating the shortest path of the entire network by the Dijkstra algorithm are as follows:

[0018] Step 21, initialize the distance from the source node to all nodes to infinity, and the distance from the source node to itself to 0. The formula is expressed as:

[0019]

[0020] In the formula: d(s,v) is the distance function between the source node s and the remaining node v, v is the remaining node, and s is the source node;

[0021] Step 22, use the priority queue to select the node with the smallest current distance and update the distances of its neighbor nodes;

[0022] Step 23, repeat step 22 until the shortest paths of all nodes are determined.

[0023] As a further solution of the present invention, in step 3, the specific steps for adopting the deep reinforcement learning model in path dynamic optimization are as follows:

[0024] Step 31, define the state space: The state space S is composed of the real-time state information of each node and link in the network. The real-time state information includes the latency, bandwidth, and reliability of the link. There are n nodes and m links in the network. The state space is expressed as:

[0025] s = {s1, s2, …, s i , …, s n};

[0026] In the formula: s is the state space, s1 is the real-time state of the first link, s2 is the real-time state of the second link, s i is the real-time state of the i-th link, sn is the real-time state of the nth link;

[0027] s i The real-time state of the ith link is expressed as:

[0028] s i = {d(e), b(e), r(e)};

[0029] In the formula: s i is the real-time state of the ith link, d(r) is the link delay corresponding to edge e, b(e) is the link bandwidth corresponding to edge e, and r(e) is the link reliability corresponding to edge e;

[0030] Step 32, define the action space: Each node in the network has k actions, and the action space is expressed as:

[0031] A = {a1, a2, …, a j , …, a k};

[0032] In the formula: A is the action space, a1 is the routing selection from the first node to the adjacent node, a2 is the routing selection from the second node to the adjacent node, a j is the routing selection from the jth node to the adjacent node, and a k is the routing selection from the kth node to the adjacent node;

[0033] Step 33, design the reward function: Design the reward function by comprehensively considering link delay, path load, and fault recovery ability. The formula of the reward function is as follows:

[0034] R(s i , a) = -α1·d(s i , a) - β1·l(s i , a) + γ1·r(s i , a);

[0035] In the formula: R(s i , a) is the reward function, α1 is the weight coefficient of the path delay after executing action a, d(s i , a) is the path delay after executing action a, β1 is the weight coefficient of the path load of the link after executing action a, l(s i , a) is the path load of the link after executing action a, γ1 is the weight coefficient of the reliability of the path after executing action a, and r(s i , a) is the reliability of the path after executing action a;

[0036] Step 34, training the deep reinforcement learning model using a deep Q-network: In the deep Q-network algorithm, the Q-value represents the cumulative reward that can be obtained by executing action a in a given state s. The Q-value function is expressed as:

[0037]

[0038] In the formula: Q(s,a) is the cumulative reward that can be obtained by executing action a in a given state s. E represents the expectation, which means calculating the expected cumulative reward under all possible state transitions. T is the time step, representing the time range for reward accumulation, and μ t is the discount factor at time t, representing the weight of future rewards relative to the current reward, with a value between (0,1). R(s t ,a t ) is the immediate reward obtained by taking action a in state s t at time step t. s t is the state at time step t, and a t is the action at time step t. t

[0039] As a further solution of the present invention, in step 34, during the training of the deep reinforcement learning model, the optimal routing selection strategy is learned by continuously interacting with the environment. Among them, the update formula of the Q-value is:

[0040] Q'(s t ,a t )←Q(s t ,a t )+ω[r t +μ t ·maxQ(s t+1 ,a t )-Q(s t ,a t )];

[0041] In the formula: Q'(s t ,a t ) is the updated Q-value, Q(s t ,a t ) is the Q-value of executing action a t in the current state s t , representing the expected cumulative reward under the current policy. ω is the learning rate, r t is the immediate reward at the current time t, μ t is the discount factor at time t, and maxQ(s t+1 ,a t ) is the maximum Q-value among all possible actions a t+1 in the next state s t , representing the optimal choice for the next step. Q(s​t+1 , a t ) is the state s at time t + 1 t+1 Execute action a t The Q-value of s t+1 is the state at time t + 1, representing the state transition of the system after taking the current action, Q(s t , a t ) is the cumulative reward that can be obtained by executing action a given the state s at time t.

[0042] As a further solution of the present invention, in step 4, according to the prediction result of the deep reinforcement learning model, the specific steps for real-time adjustment of the communication path are:

[0043] Step 41, Policy Selection: Perform forward propagation calculation on the state s, and output the Q-value set {Q(s, a1), Q(s, a2), …, Q(s, a j , …, a k )} corresponding to the action space A = {a1, a2, …, a j ), …, Q(s, a k )}, Q(s, a j ) is the cumulative reward that can be obtained by executing action a j given the state s, and select the action a * with the maximum Q-value:

[0044]

[0045] In the formula: a * is the action with the maximum Q-value, Q(s, a j ) is the cumulative reward that can be obtained by executing action a j given the state s, is the maximum Q(s, a j ) value corresponding to all candidate actions a j ∈ A, and argmax is the action that makes Q(s, a j ) obtain the maximum value;

[0046] Step 42, Result Prediction and Path Reconstruction: After obtaining the action a * with the maximum Q-value through the deep reinforcement learning model, based on the original path P old perform local update to re-plan a new path P new , and the formula for re-planning the new path is:

[0047]

[0048] In the formula: P new is the re-planned new path, w(e') is the weight of edge e', ∑ e'∈Pw(e') is the cumulative weight of the path, which represents the sum of the weights w(e') of all edges e' in the candidate path P. is the path with the minimum cumulative weight;

[0049] Step 43, path confirmation and verification: For each link e' in the new path, confirm whether its delay d(e'), bandwidth b(e')_, and reliability r(e') meet the requirements through the path check model. The formula of the path check model is:

[0050]

[0051] In the formula: d max is the delay threshold of link e', b max is the bandwidth threshold of link e', r max is the reliability threshold of link e'.

[0052] The technical effects and advantages of an automatic substation communication routing path planning method based on the shortest path in the present invention: By modeling the substation communication network, the present invention uses an automatic substation communication routing path planning algorithm that combines the shortest path algorithm and deep reinforcement learning. Combining the efficiency of the shortest path algorithm and the ability of deep reinforcement learning to process complex data, the deep reinforcement learning model is used to perform forward propagation calculations on the network state in real time, and the optimal action is dynamically selected according to the current state, realizing the automatic planning of the substation communication routing path, and being able to adapt to the routing planning requirements of substation communication networks of different scales and complexities; The present invention can adjust the routing path in real time according to the change of the network state, adapt to the needs of network dynamic changes. By combining the shortest path algorithm and the deep reinforcement learning algorithm, it can efficiently find the optimal path, reduce communication delay, and improve network efficiency; In case of link failure or network congestion, it can quickly adjust the path to ensure the reliability and stability of communication. Description of the Drawings

[0053] Figure 1 is a schematic diagram of the Dijkstra algorithm in step 2 of an automatic substation communication routing path planning method based on the shortest path provided by the present invention;

[0054] Figure 2 is a schematic diagram of routing recovery after an earthquake disaster in the prior art;

[0055] Figure 3 is a schematic flow diagram of an automatic substation communication routing path planning method based on the shortest path provided by the present invention. Detailed Embodiments

[0056] The technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the present invention. Obviously, the described technical solutions are only a part of the present invention, rather than all of it. Based on the technical solutions of the present invention, all other technical solutions obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.

[0057] An automatic planning method for substation communication routing paths based on the shortest path includes the following steps:

[0058] Step 1, network modeling: Model the substation communication network, regarding substations, control centers, and other communication nodes as nodes in the graph, and the communication links between nodes as edges in the graph;

[0059] Step 2, initial path planning: Use the Dijkstra algorithm or the A algorithm to calculate the initial shortest path from the source node to the target node in the substation communication network. Among them, the Dijkstra algorithm is used to calculate the shortest path of the entire network, and the A algorithm is used for target-oriented path calculation;

[0060] Step 3, path dynamic optimization: Adopt a deep reinforcement learning (DQN) model, and based on the real-time network state (including link traffic, fault conditions, bandwidth occupancy rate), dynamically adjust the routing path through policy optimization;

[0061] Step 4, path adjustment and execution: When the network state changes, adjust the communication path in real time according to the prediction results of the deep reinforcement learning model;

[0062] Step 5, result evaluation, optimization and output: Output the result of adjusting the communication path in text form for users to refer to and use. The output results include path information (including nodes and links on the path), and path performance indicators (including length, stability, bandwidth).

[0063] Quickly calculate the initial optimal path through the Dijkstra algorithm or A algorithm to ensure that communication data can reach the target in the shortest time and improve the overall response speed of the communication network; adopt a deep reinforcement learning (DQN) model, which can make intelligent adjustments according to the real-time network state (such as link traffic, fault conditions, bandwidth occupancy, etc.), enhance the adaptability of the network, and avoid interruptions caused by communication congestion or faults; when the network state changes, perform rapid path switching according to the prediction results to ensure the stability of substation communication and reduce service interruptions caused by link congestion or faults; through automated path planning and dynamic adjustment, reduce the need for manual intervention, lower the operation and maintenance costs, and improve the automation and intelligence levels of network management; through real-time optimization and dynamic adjustment of the path, optimize the allocation of network resources (such as bandwidth, links), avoid resource waste, and enhance the utilization efficiency of the overall communication network; combine deep reinforcement learning technology to autonomously learn and optimize in a complex network environment, have strong adaptive capabilities, be able to handle sudden network changes, and improve the fault tolerance of the network; by outputting detailed path information (including nodes, links) and performance indicators (such as path length, stability, bandwidth, etc.), provide intuitive decision-making basis for operation and maintenance personnel to facilitate monitoring and optimizing network operation; during the path optimization process, fully consider the link state, select high-quality links for communication, reduce packet delay and loss, and ensure the real-time requirements of the power system.

[0064] Figure 3 The flow diagram of an automatic substation communication routing path planning method based on the shortest path provided by the present invention. As Figure 3 shown, the flow of an automatic substation communication routing path planning method based on the shortest path is as follows:

[0065] Data preprocessing: Extract data such as network topology structure, node attributes, and link attributes from the substation communication network, clean and transform the data to meet the input requirements of deep reinforcement learning;

[0066] Train the deep reinforcement model: Use the processed data set to train the deep reinforcement learning model so that it can learn and represent the potential laws and patterns in the network. During the training process, adjust the parameters and structure of the model according to the actual situation to improve the prediction accuracy and generalization ability of the model;

[0067] Calculate the shortest path: Use the Dijkstra algorithm to calculate the shortest path between the source node and the target node, and save the result in the distance table. During the calculation process, set different weight functions (such as hop count, bandwidth, delay, etc.) according to actual needs to reflect the advantages and disadvantages of the path in different network environments;

[0068] Network status perception: Input the preprocessed data into the trained deep reinforcement learning model to obtain the perception results of the network status. The perception results include information such as the overall network condition, the status of each node, and the availability of links.

[0069] Routing planning: Combine the shortest path calculation results and the network status perception results, and perform routing planning according to business rules and constraints. During the planning process, adopt a greedy strategy to gradually determine the nodes and links on the path, and at the same time consider factors such as the length, stability, and bandwidth of the path, evaluate and optimize the planning results to ensure the feasibility and efficiency of the routing.

[0070] Result output: Output the planning results in a visual or text form for users to refer to and use. The output results include the detailed information of the path (such as the nodes and links on the path), the performance metrics of the path (such as length, stability, bandwidth, etc.), and possible risks or suggestions.

[0071] Specifically, in step 1, the specific steps of network modeling are as follows:

[0072] Step 11: Represent the substation communication network as a graph G=(V, E), where V is the set of nodes and E is the set of edges. Each edge e∈E and each edge e has a weight, which is used to find the shortest path from the source node s to the target node t.

[0073] Step 12: Calculate the weight w(e) of edge e according to the delay, bandwidth, and reliability of the link. The delay d(e) of the link represents the link delay corresponding to edge e, the bandwidth b(e) of the link represents the link bandwidth corresponding to edge e, and the reliability r(e) of the link represents the link reliability corresponding to edge e. Calculate the weight w(e) of edge e through a weighted sum, and the calculation formula is:

[0074]

[0075] In the formula: w(r) is the weight of edge e, α is the weight coefficient of the link delay corresponding to edge e, d(e) is the link delay corresponding to edge e, β is the weight coefficient of the link bandwidth corresponding to edge e, b(e) is the link bandwidth corresponding to edge e, γ is the weight coefficient of the link reliability corresponding to edge e, and r(e) is the link reliability corresponding to edge e.

[0076] By modeling the substation communication network as a graph G=(V, E), the topological structure of the network can be clearly defined, making the relationship between nodes and links intuitive and explicit, providing a solid foundation for subsequent path planning. When calculating the edge weights, link delay, bandwidth, and reliability are comprehensively considered, which can balance different network performance requirements. The delay d(e) is used to ensure that the communication delay meets the real-time requirements of power dispatching, the bandwidth b(e) is used to ensure the data transmission speed to meet the large-volume data demand, and the reliability r(e) is used to ensure the stability of the link and reduce the impact of communication failures on the system. By adjusting the weight coefficients, the path can be optimized for different application scenarios. For real-time control, the weight coefficient of the link delay corresponding to edge e can be increased to reduce the delay. For data transmission, the weight coefficient of the link bandwidth corresponding to edge e can be increased to ensure sufficient bandwidth. For critical tasks, the weight coefficient of the link reliability corresponding to edge e can be increased to improve the reliability. When calculating the path, network resources can be reasonably allocated to avoid network congestion caused by a single metric (such as the shortest path), and the overall utilization efficiency of the network can be improved. It is applicable to various topological structures such as star, ring, and mesh, ensuring that the substation communication network can still operate efficiently in a complex environment. When the network state changes (such as an increase in the delay or a decrease in the reliability of some links), the weight model can be dynamically updated to provide data support for subsequent path optimization.

[0077] Specifically, in step 2, the specific steps for calculating the shortest path of the entire network by the Dijkstra algorithm are as follows:

[0078] Step 21, initialize the distance from the source node to all nodes to infinity, and the distance from the source node to itself to 0. The formula is expressed as:

[0079]

[0080] In the formula: d(s, v) is the distance function between the source node s and the remaining node v, v is the remaining node, and s is the source node;

[0081] Step 22, use a priority queue to select the node with the smallest current distance and update the distances of its neighbor nodes;

[0082] Step 23, repeat step 22 until the shortest paths of all nodes are determined.

[0083] Such as Figure 1Schematic diagram of the Dijkstra algorithm in step 2 as shown. Set node V1 as the current node. At this time, the neighbor nodes of node V1 include V2, V3, and V4, and the weights are 9, 14, and 15 respectively. Update their shortest distances: d(V1, V2) = 9 (V1→V2), d(V1, V3) = 14 (V1→V3), d(V1, V4) = 15 (V1→V2). At this time, V2 is the unvisited node with the shortest distance and is selected as the next current node;

[0084] Set node V2 as the current node. At this time, the neighbor node of node V2 is V7, and the weight is 24. Calculate the new path: d(V1, V7) = d(V1, V2) + d(V2, V7) = 9 + 24 = 33. Update the shortest path of V7 to 33. The unvisited node of the current shortest path is V3 and is selected as the next current node;

[0085] Set node V3 as the current node. The neighbor node of V3 is V5, and the weight is 30. Calculate the new path: d(V1, V5) = d(V1, V3) + d(V3, V5) = 14 + 30 = 44. Update the shortest path of V5 to 44. The unvisited node of the current shortest path is V4 and is selected as the next current node;

[0086] Set node V4 as the current node. The neighbor node of V4 is V5, and the weight is 17. Calculate the new path: d(V1, V5) = d(V1, V4) + d(V4, V5) = 15 + 17 = 32. Update the shortest path of V5 to 32. The unvisited node of the current shortest path is V5 and is selected as the next current node;

[0087] Set node V5 as the current node. The neighbor nodes of V5 are V6 and V8, and the weights are 11 and 16 respectively. Calculate the new paths: d(V1, V6) = d(V1, V5) + d(V5, V6) = 32 + 11 = 43, d(V1, V8) = d(V1, V5) + d(V5, V8) = 32 + 16 = 48. Update the shortest paths of V6 and V8 to 43 and 48 respectively. The unvisited node of the current shortest path is V6 and is selected as the next current node;

[0088] Set node V6 as the current node. The neighbor node of V6 is V8, and the weight is 6. Calculate the new path: d(V1, V8) = d(V1, V6) + d(V6, V8) = 43 + 6 = 48. The unvisited node of the current shortest path is V7 and is selected as the next current node;

[0089] Set node V7 as the current node. The neighbor nodes of V7 are V6 and V8, with weights of 2 and 19 respectively. Calculate the new paths as follows: d(V1, V6) = d(V1, V7) + d(V7, V6) = 33 + 2 = 35 (better than 43, so it is updated to 35), d(V1, V8) = d(V1, V7) + d(V7, V8) = 33 + 19 = 52 (greater than 48, so it is not updated). The unvisited node of the current shortest path is V8, which is selected as the next current node;

[0090] Set node V8 as the current node. Node V8 has no unvisited neighbors, and the algorithm ends;

[0091] Based on the above process, the shortest path distances from node V1 to all nodes are as follows: d(V1, V1) = 0, d(V1, V2) = 9, d(V1, V3) = 14, d(V1, V4) = 15, d(V1, V5) = 32, d(V1, V6) = 35, d(V1, V7) = 33, d(V1, V8) = 48;

[0092] The code to implement the Dijkstra algorithm in step 2 is as follows:

[0093]

[0094]

[0095] The Dijkstra algorithm is based on the greedy strategy. Each time, it selects the node on the current shortest path and updates it to ensure that the shortest path from the source node to all nodes is finally obtained, guaranteeing the optimality of the path; combined with a priority queue (implemented by a binary heap, for example), the time complexity of the algorithm can reach, and it performs well in sparse graphs, being able to efficiently solve practical problems; the Dijkstra algorithm is not only applicable to the single-source shortest path problem but can also be extended to various network optimization scenarios, such as routing protocols (OSPF) and transportation and logistics optimization; since the Dijkstra algorithm does not generate randomness during execution and the results are consistent each time, this determinacy helps maintain reliability in critical tasks (such as air navigation and network communication); by maintaining a priority queue, the algorithm always selects the optimal path for update, avoiding unnecessary calculations, thereby reducing the computational overhead and improving the execution efficiency; in a dynamic network environment, the Dijkstra algorithm can quickly adapt to changes in the network topology through incremental update methods to assist in traffic adjustment and congestion management.

[0096] Specifically, in step 3, the specific steps for using a deep reinforcement learning model in path dynamic optimization are as follows:

[0097] Step 31, define the state space: The state space S is composed of the real-time state information of each node and link in the network. The real-time state information includes the delay, bandwidth, and reliability of the link. There are n nodes and m links in the network, and the state space is expressed as:

[0098] s = {s1, s2, …, s i , …, s n};

[0099] Where: s is the state space, s1 is the real-time state of the first link, s2 is the real-time state of the second link, s i is the real-time state of the i-th link, s n is the real-time state of the n-th link;

[0100] s i The real-time state of the i-th link is expressed as:

[0101] s i = {d(e), b(e), r(e)};

[0102] Where: s i is the real-time state of the i-th link, d(e) is the link delay corresponding to edge e, b(e) is the link bandwidth corresponding to edge e, and r(e) is the link reliability corresponding to edge e;

[0103] Step 32, define the action space: Each node in the network has k actions, and the action space is expressed as:

[0104] A = {a1, a2, …, a j , …, a k};

[0105] Where: A is the action space, a1 is the routing selection from the first node to the adjacent node, a2 is the routing selection from the second node to the adjacent node, a j is the routing selection from the j-th node to the adjacent node, a k is the routing selection from the k-th node to the adjacent node;

[0106] Step 33, design the reward function: Design the reward function by comprehensively considering the link delay, path load, and fault recovery ability. The formula of the reward function is as follows:

[0107] R(s i , a) = -α1·d(s i , a) - β1·(s i , a) + γ1·r(s i , a);

[0108] Where: R(s i, a) is the reward function, α1 is the weight coefficient of the path delay after executing action a, d(s i , a) is the path delay after executing action a, β1 is the weight coefficient of the path load of the link after executing action a, l(s i , a) is the path load of the link after executing action a, γ1 is the weight coefficient of the reliability of the path after executing action a, r(s i , a) is the reliability of the path after executing action a;

[0109] Step 34, use the deep Q-network to train the deep reinforcement learning model: The Q value in the deep Q-network algorithm represents the cumulative reward that can be obtained by executing action a under a given state s. The Q value function is expressed as:

[0110]

[0111] In the formula: Q(s, a) is the cumulative reward that can be obtained by executing action a under a given state s. E is the expectation, which means calculating the expected cumulative reward under all possible state transitions. T is the time step, which represents the time range for reward accumulation. μ t is the discount factor at time t, which represents the weight of future rewards relative to the current reward, and its value is between (0, 1). R(s t , a t ) is the immediate reward obtained by taking action a in state s at time step t. s t takes action a t at time step t. s t is the state at time step t, and a t is the action at time step t.

[0112] By taking the delay, bandwidth, and reliability of the link as the state space, the changes in the network state can be perceived in real time, ensuring that the communication path is dynamically adjusted according to the current network conditions, thus effectively coping with emergencies (such as traffic peaks and link failures); based on the Deep Q-Network (DQN), the path planning is automatically learned and optimized, reducing manual intervention and experience dependence, avoiding the inefficiency and errors caused by manual adjustment in traditional methods, and improving the intelligence level of path planning; the reward function comprehensively considers link delay, path load, and fault recovery ability, ensuring that the path selection achieves an optimal balance among low latency, high reliability, and load balancing, and enhancing the overall performance of the substation communication network; by introducing a reliability parameter into the reward function, the DQN model can tend to select high-reliability links during the training process, reducing data loss caused by unstable links and enhancing the security of power dispatching; through the experience replay mechanism of deep Q-learning, the network can continuously learn and optimize from historical data without having to recalculate the shortest path every time the path changes, improving the decision-making efficiency and saving computing resources; by controlling the balance between short-term and long-term rewards through the discount factor, DQN can make a trade-off between short-term performance (low latency) and long-term performance (network stability and resource balance), ensuring the sustainable and stable operation of the network; when a link failure or congestion occurs in the network, the DQN model can adjust the path based on the real-time state, quickly switch to the backup link, reduce the communication interruption time, and improve the disaster tolerance ability of the network.

[0113] Specifically, in step 34, during the training of the deep reinforcement learning model, by continuously interacting with the environment to learn the optimal routing selection strategy, where the update formula for the Q value is:

[0114] Q'(s t ,a t )←Q(s t ,a t )+ω[r t +μ t ·maxQ(s t+1 ,a t )-q(s t ,a t )];

[0115] In the formula: q'(s t ,a t ) is the updated Q value, Q(s t ,a t ) is the Q value of executing action a t under the current state s t , representing the expected cumulative reward under the current policy, ω is the learning rate, r t is the immediate reward at the current moment t, μ t is the discount factor at moment t, maxQ(s t+1, a t ) is for all possible actions a t+1 in the next state s t The maximum Q value among them represents the optimal choice for the next step. Q(s t+1 , a t ) is the Q value of the state s t+1 when performing action a t at time t + 1, s t+1 is the state at time t + 1, representing the state transition of the system after taking the current action. Q(s t , a t ) is the cumulative reward that can be obtained by performing action a given the state s at time t.

[0116] By continuously interacting with the environment, DRL can adaptively adjust the routing selection strategy according to the real-time state of the network (such as traffic changes, node failures), adapting to different network conditions; the Q-value update process continuously approaches the optimal value, enabling the model to converge to the globally optimal routing strategy in long-term operation, reducing the overall network delay and congestion; due to the introduction of the discount factor, DRL not only considers the current immediate reward in the decision-making process but also can weigh the possible future benefits to achieve long-term optimal routing decisions; by controlling the learning rate of the model, it can be gradually optimized in the continuous iteration process and quickly adjust the decision when the network environment changes; by maximizing the Q value of the next state, the model can explore the possibilities of different paths, find the optimal or sub-optimal paths, thereby improving the robustness and reliability of the network; after training, DRL can achieve autonomous decision-making, reducing the dependence on traditional manual configuration of the routing table, reducing the maintenance cost, and improving the intelligent level of network management; through parallel computing and an efficient Q-value update strategy, DRL can quickly respond to changes in the network state, ensure real-time routing decisions, and improve data transmission efficiency.

[0117] Specifically, in step 4, the specific steps for real-time adjustment of the communication path according to the prediction result of the deep reinforcement learning model are as follows:

[0118] Step 41, policy selection: Perform forward propagation calculation on the state s, and output the Q-value set {Q(s, a1), Q(s, a2), …, Q(s, a j , …, a k} corresponding to the action space A = (a1, a2, …, a j ), …, Q(s, a k ). Q(s, a j ) is the cumulative reward that can be obtained by performing action a j given the state s, and select the action a * with the maximum Q value:

[0119]

[0120] Where: a * is the action with the maximum Q value, Q(s, a j ) is the cumulative reward that can be obtained by executing action a j in the given state s, is the maximum Q(s, a j ) value corresponding to all candidate actions a j ∈A, and argmax is the action that makes Q(s, a j ) obtain the maximum value;

[0121] Step 42, result prediction and path reconstruction: After obtaining the action a with the maximum Q value through the deep reinforcement learning model, * based on the original path P old perform local update to re-plan the new path P new , and the formula for re-planning the new path is:

[0122]

[0123] Where: P new is the re-planned new path, w(e') is the weight of edge e', ∑ e'∈P w(e') is the path cumulative weight, indicating the sum of the weights w(e') of all edges e' in the candidate path P, is the path with the minimum cumulative weight;

[0124] Step 43, path confirmation and verification: For each link e' in the new path, confirm whether its delay d(e'), bandwidth b(e')_, and reliability r(e') meet the requirements through the path check model. The formula of the path check model is:

[0125]

[0126] Where: d max is the delay threshold of link e', b max is the bandwidth threshold of link e', r max is the reliability threshold of link e'.

[0127] Using a deep reinforcement learning model, the Q-values of all candidate actions are calculated through forward propagation, enabling real-time determination of the optimal action in the current network state and achieving intelligent and adaptive routing decisions; the argmax operation is adopted to ensure that the selected action is optimal in terms of cumulative reward, which helps to obtain better network performance in long-term operation; the ε-greedy strategy can be combined to, while ensuring the use of the current optimal action, retain a certain exploration probability to adapt to environmental changes and avoid getting stuck in local optima; through local updates, only the affected parts are re-planned for the path, which can not only quickly respond to network state changes but also take into account the global communication performance and reduce unnecessary global replanning; the edge weights are calculated using comprehensive metrics such as latency, bandwidth, and reliability to ensure that the selected path achieves the best balance among multiple performance metrics, improving network transmission efficiency and stability. When the deep reinforcement learning model selects the optimal action, the original path is locally updated based on this action, which can quickly adapt to local link state changes and ensure that the overall path always maintains the best state; after the path is adjusted, by checking the latency, bandwidth, and reliability of each link in the new path, it is ensured that the new path fully meets the preset performance criteria, thus preventing the overall communication quality from being affected by insufficient performance of a certain link; a path inspection model is introduced to verify each link, realizing real-time monitoring and verification of the newly planned path and ensuring the reliability and security of network switching; if a certain link does not meet the conditions, the fault tolerance mechanism can be triggered immediately or rolled back to the backup path, enhancing the robustness and stability of the entire system.

[0128] In the embodiment of the present invention, an automatic planning algorithm for substation communication routing paths based on the combination of the shortest path algorithm and deep reinforcement learning combines the efficiency of the shortest path algorithm and the ability of deep reinforcement learning to process complex data to achieve the automatic planning of substation communication routing paths. It has the advantages of high efficiency, flexibility, and scalability, and can adapt to the routing planning requirements of substation communication networks with different scales and complexities; the present invention can adjust the routing path in real time according to changes in the network state to meet the needs of network dynamic changes; by combining the shortest path algorithm and the deep reinforcement learning algorithm, it can efficiently find the optimal path, reduce communication latency, and improve network efficiency; in case of link failures or network congestion, it can quickly adjust the path to ensure the reliability and stability of communication.

[0129] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

[0130] Finally, the above are only the preferred solutions of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An automatic planning method for substation communication routing paths based on the shortest path, characterized in that It includes the following steps: Step 1, network modeling: Model the substation communication network. Consider substations, control centers, and other communication nodes as nodes in a graph, and consider the communication links between nodes as edges in the graph. Calculate the weight w(e) of edge e according to the delay, bandwidth, and reliability of the link. The weight calculation formula for edge e is: Where: w(e) is the weight of edge e, α is the weight coefficient of the link delay corresponding to edge e, d(e) is the link delay corresponding to edge e, β is the weight coefficient of the link bandwidth corresponding to edge e, b(e) is the link bandwidth corresponding to edge e, γ is the weight coefficient of the link reliability corresponding to edge e, and r(e) is the link reliability corresponding to edge e; Step 2, initial path planning: Use the Dijkstra algorithm to calculate the initial shortest path from the source node to the target node in the substation communication network. Initialize the distance from the source node to all nodes to infinity, and the distance from the source node to itself to 0. The formula is expressed as: Where: d(s,v) is the distance function between the source node s and the remaining node v, v is the remaining node, and s is the source node; Step 3, path dynamic optimization: Adopt a deep reinforcement learning model, and based on the real-time state of the network, dynamically adjust the routing path through policy optimization; Step 4, path adjustment and execution: When the network state changes, adjust the communication path in real time according to the prediction results of the deep reinforcement learning model; Step 5, result evaluation, optimization, and output: Output the result of adjusting the communication path in text form for users to refer to and use. The output results include path information and path performance indicators.

2. The automatic planning method for substation communication routing path based on the shortest path according to claim 1, wherein In Step 1, the specific steps of network modeling are: Step 11, Represent the substation communication network as a graph G=(V,E), where V is the set of nodes and E is the set of edges. Each edge e∈E, and each edge e has a weight, which is used to find the shortest path from the source node s to the target node t; Step 12, Calculate the weight w(e) of edge e according to the delay, bandwidth, and reliability of the link. Among them, the link delay d(e) represents the link delay corresponding to edge e, the link bandwidth b(e) represents the link bandwidth corresponding to edge e, the link reliability r(e) represents the link reliability corresponding to edge e, and calculate the weight w(e) of edge e through weighted sum.

3. The automatic planning method for substation communication routing path based on the shortest path according to claim 1, wherein In Step 3, the specific steps of adopting a deep reinforcement learning model in path dynamic optimization are: Step 31, Define the state space: The state space S is composed of the real-time state information of each node and link in the network. The real-time state information includes the delay, bandwidth, and reliability of the link. There are n nodes and m links in the network. The state space is expressed as: s = {s1, s2, …, s i , …, s n}; Where: s is the state space, s1 is the real-time state of the first link, s2 is the real-time state of the second link, s i is the real-time state of the i-th link, s n is the real-time state of the n-th link; s i The real-time status of the i-th link is represented as: s i = {d(e), b(e), r(e)}; where: s i is the real-time status of the i-th link, d(e) is the link delay corresponding to edge e, b(e) is the link bandwidth corresponding to edge e, and r(e) is the link reliability corresponding to edge e; Step 32, Define the action space: Each node in the network has k actions. The action space is expressed as: A = {a1, a2, …, a j , …, a k}; Where: A is the action space, A1 is the routing selection from the first node to the adjacent node, a2 is the routing selection from the second node to the adjacent node, a j is the routing selection from the j-th node to the adjacent node, a k is the routing selection from the k-th node to the adjacent node; Step 33, Design the reward function: Design the reward function by comprehensively considering link delay, path load, and fault recovery ability. The reward function formula is as follows: R(s i , a) = -α1·d(s i , a) - β1·l(s i , a) + γ1·r(s i , a); where: R(s i , a) is the reward function, α1 is the weight coefficient of the path delay after executing action a, d(s i , a) is the path delay after executing action a, β1 is the weight coefficient of the path load of the link after executing action a, l(s i , a) is the path load of the link after executing action a, γ1 is the weight coefficient of the reliability of the path after executing action a, r(s i , a) is the reliability of the path after executing action a; Step 34, Use the deep Q-network to train the deep reinforcement learning model: The Q value in the deep Q-network algorithm represents the cumulative reward that can be obtained by executing action a under the given state s. The Q value function is expressed as: Where: Q(s,a) is the cumulative reward that can be obtained by executing action a in the given state s. E represents the expectation, indicating that the expected cumulative reward is calculated under all possible state transitions. T is the time step, representing the time range for reward accumulation, μ t is the discount factor at time t, representing the weight of future rewards relative to the current reward, with a value between (0,1). R(s t ,a t ) is the immediate reward obtained by taking action a in state s at time step t. s t takes action a t at time step t. s t is the state at time step t, and a t is the action at time step t.

4. The method for automatically planning a substation communication routing path based on the shortest path according to claim 3, characterized in that, In step 34, during the training process of the deep reinforcement learning model, the optimal routing selection strategy is learned by continuously interacting with the environment. Among them, the update formula of the Q value is as follows: Q'(s t ,a t ) ← Q(s t ,a t ) + ω[r t + μ t · maxQ(s t+1 ,a t ) - Q(s t ,a t )]; Where: Q'(s t ,a t ) is the updated Q value, Q(s t ,a t ) is the Q value when the current state s t performs the action a t , representing the expected cumulative reward under the current policy. ω is the learning rate, r t is the immediate reward at the current time t, μ t is the discount factor at time t, maxQ(s t+1 ,a t ) is the maximum Q value among all possible actions a t+1 in the next state s t , representing the optimal choice for the next step. Q(s t+1 ,a t ) is the Q value when the state s t+1 performs the action a t at time t + 1, s t+1 is the state at time t + 1, representing the state transition of the system after taking the current action. Q(s t ,a t ) is the cumulative reward that can be obtained by performing the action a given the state s at time t.

5. The automatic planning method for substation communication routing path based on the shortest path according to claim 1, wherein In step 4, the specific steps for real-time adjustment of the communication path according to the prediction results of the deep reinforcement learning model are as follows: Step 41, Policy Selection: Perform forward propagation calculation on state s, and output the Q-value set {Q(s, a1), Q(s, a2), …, Q(s, a j , …, a k )} corresponding to the action space A = {a1, a2, …, a j ), …, Q(s, a k ). Q(s, a j ) is the cumulative reward that can be obtained by executing action a j under the given state s. Select the action a * with the maximum Q-value: where: a * is the action with the maximum Q value, Q(s, a j ) is the cumulative reward that can be obtained by executing the action a j in the given state s, is the maximum Q(s, a j ) value corresponding to all candidate actions a j ∈ A, and argmax is the action that makes Q(s, a j ) reach the maximum value; Step 42, result prediction and path reconstruction: Obtain the action a with the maximum Q value through the deep reinforcement learning model * After that, based on the original path P old Perform local update to re-plan the new path P new The formula for re-planning the new path is: Where: P new is the newly planned path, w(e') is the weight of edge e', ∑ e'∈ Pw(e') is the cumulative path weight, indicating the sum of the weights w(e') of all edges e' in the candidate path P, is the path with the minimum cumulative weight; Step 43, path confirmation and verification: For each link e' in the new path, confirm whether its delay d(e'), bandwidth b(e'), and reliability r(e') meet the requirements through the path inspection model. The formula of the path inspection model is: where: d max is the delay threshold of link e', b max is the bandwidth threshold of link e', r max is the reliability threshold of link e'.

Citation Information

Cited By

  • Multi-hop path construction method of passive sensor network

    CN120614664A