SRv6 progressive deployment and traffic engineering method based on deep reinforcement learning

By training deep reinforcement learning agents in the simulated network, optimizing the SRv6 traffic path, and gradually upgrading the key nodes to SRv6 nodes, the problems of high deployment costs and inaccurate routing optimization in the existing technology are solved, and low-cost and efficient SRv6 node deployment and precise traffic engineering routing optimization are achieved.

CN119652806BActive Publication Date: 2025-05-16UNIV OF SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510165843.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-05-16
Estimated Expiration
2045-02-14

AI Technical Summary

Technical Problem

In a wide area network environment, it is difficult for the existing technology to deploy SRv6 nodes at low cost and efficiently to combine routing optimization of traffic engineering, and ignore the impact of SRv6 technology on bandwidth consumption, resulting in inaccurate routing optimization decisions.

Method used

SRv6 progressive deployment and traffic engineering methods based on deep reinforcement learning are adopted. By training deep reinforcement learning agents in the simulated network, the traffic path scheduling strategy is learned, the maximum link utilization is optimized, and the key nodes are gradually upgraded to SRv6 nodes according to the optimization results.

Benefits of technology

It realizes low-cost and efficient SRv6 node deployment, accurately identify key nodes that contribute the most to traffic engineering optimization, reduces deployment costs, improves routing optimization effects, and takes into account the additional bandwidth consumption of SRv6 technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119652806B_ABST
    Figure CN119652806B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of communication technology, and discloses a SRv6 progressive deployment and traffic engineering method based on deep reinforcement learning, including: training a deep reinforcement learning agent in a simulated network where all SRv6 nodes are deployed; obtaining a pre-set number of nodes that transmit the most SRv6 traffic and upgrading them to SRv6 nodes based on the optimization results of the current network historical traffic matrix by the trained agent, and obtaining a hybrid network; optimizing the traffic matrix of the hybrid network by the agent, and using a local search method to correct the optimization results of the agent. Based on the optimization results of the representative traffic matrix by the deep reinforcement learning agent, the present invention accurately identifies the key nodes that contribute the most to the optimization of traffic engineering, and upgrades them to SRv6 nodes first. In the case of a small number of deployments, the maximum link utilization optimization effect close to the entire SRv6 network can be achieved, reducing the cost of equipment upgrades.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of communication technology, and in particular to an SRv6 progressive deployment and traffic engineering method based on deep reinforcement learning. Background Art

[0002] In the current WAN environment, operators use traffic engineering (TE) technology to effectively manage network resources. Segment Routing IPv6 (SRv6) technology based on the IPv6 forwarding plane is considered a key technology in traffic engineering because it can customize traffic transmission paths. Although SRv6 has demonstrated excellent traffic scheduling capabilities in traffic engineering, upgrading traditional IPv6 networks to SRv6 networks will face huge economic and technical challenges, such as high equipment upgrade costs, heterogeneous network equipment, and large-scale upgrades that easily cause equipment downtime.

[0003] Since SRv6 can coexist with IPv6, gradually deploying SRv6 nodes and gradually transitioning traditional IPv6 networks to SRv6 networks has become a feasible solution. In a network with mixed IP nodes and SRv6 nodes, the deployment location and number of SRv6 nodes directly affect the available paths and routing optimization effects of traffic in traffic engineering. The deployment of SRv6 is coupled with the routing optimization of traffic engineering. However, existing research usually deploys SRv6 nodes based on the overall distribution of network traffic (Most Loaded Link, MLL), without combining routing optimization of traffic engineering well. How to combine routing optimization of traffic engineering to deploy SRv6 nodes at low cost and high efficiency to help routing optimization is a technical problem that needs to be solved.

[0004] In further technical issues, since the use of SRv6 technology will encapsulate additional headers for data packets, it will bring additional bandwidth consumption. Existing work usually ignores this part of consumption, resulting in inaccurate routing optimization decisions. This problem will cause the deployed SRv6 nodes to be less helpful for traffic engineering, and it will cost more to deploy more SRv6 nodes to improve routing optimization effects. Summary of the invention

[0005] In order to solve the above technical problems, the present invention provides an SRv6 progressive deployment and traffic engineering method based on deep reinforcement learning.

[0006] In order to solve the above technical problems, the present invention adopts the following technical solutions:

[0007] A SRv6 progressive deployment and traffic engineering method based on deep reinforcement learning, including:

[0008] In a simulated network where all SRv6 nodes are deployed, a proximal policy optimization algorithm is used to train a deep reinforcement learning agent to learn a traffic path scheduling strategy to reduce the maximum link utilization of the simulated network.

[0009] According to the optimization results of the current network historical traffic matrix by the trained intelligent agent, the nodes with the largest SRv6 traffic transmission are obtained and upgraded to SRv6 nodes to obtain a hybrid network.

[0010] The traffic matrix in the hybrid network is optimized by intelligent agents to generate a preliminary traffic path scheduling strategy. The default route is corrected for the paths involving undeployed SRv6 nodes in the preliminary traffic path scheduling strategy. The path of the critical traffic is iteratively adjusted using a local search algorithm until the maximum link utilization of the critical traffic can no longer be reduced.

[0011] In one embodiment, the use of a proximal policy optimization algorithm to train a deep reinforcement learning agent to learn a traffic path scheduling strategy to reduce the maximum link utilization of a simulated network specifically includes:

[0012] Initialize the traffic matrix of the simulated network;

[0013] Find the first α% of the most congested links in the simulated network, and select the first β% of the traffic with the largest bandwidth consumption from the traffic passing through these most congested links as the key traffic;

[0014] The intelligent agent optimizes the path of each key traffic in turn according to the bandwidth size and selects the optimized path.

[0015] In one embodiment, the agent optimizes each key traffic path in turn according to the bandwidth size and selects the optimized path, specifically including:

[0016] For each key flow to be optimized, the agent first removes the impact of the key flow on the network bandwidth; each key flow has multiple available paths. When the key flow selects one of the available paths, the capacity of the link in the simulated network, the utilization of the link when the current key flow is not transmitted, the link utilization that the current key flow needs to occupy after using this path, and the additional bandwidth used by the current key flow using SRv6 technology are used as the state of reinforcement learning. , the agent is based on the state Output Action After the action is executed in the environment, the current key traffic path changes, the maximum link utilization in the simulated network changes, and a reward is returned The intelligent agent stores the current experience in the experience pool and then optimizes the next key flow.

[0017] In one of the embodiments, the action is: selecting one of the available paths of the current key flow as the path of the optimized key flow.

[0018] In one embodiment, the reward The calculation method is:

[0019] ;

[0020] in, It represents the maximum link utilization before the action is executed, that is, at the t-1th time; It represents the maximum link utilization after the action is executed, that is, at the tth moment.

[0021] In one embodiment, the additional bandwidth used by the current critical traffic using the SRv6 technology is calculated as follows:

[0022] ;

[0023] ;

[0024] in, Indicates the current critical flow Using Paths Additional bandwidth consumed when Indicates the current critical flow The rate at which packets are transmitted; Indicates that when using SRv6 technology, the current key traffic The additional length of the data packet encapsulation; It is the basic overhead of the segment routing header in the SRv6 protocol. The length of the segment identifier in the segment routing header. Indicates the length of the external additional encapsulated IPv6 header.

[0025] In one embodiment, the optimization result of the current network historical traffic matrix by the trained intelligent agent is used to obtain a preset number of nodes that transmit the most SRv6 traffic and upgrade them to SRv6 nodes, which specifically includes:

[0026] Generate a representative traffic matrix based on the historical traffic matrix of the current network, use the trained intelligent agent to optimize the path of the representative traffic matrix, and calculate the deployment index of each node in the current network based on the optimization result;

[0027] The deployment indicators are sorted and a set number of nodes with the highest deployment indicator values ​​are upgraded to SRv6 nodes.

[0028] In one embodiment, the calculation of the deployment index of each node in the current network based on the optimization result specifically includes:

[0029] Traverse the SRv6 traffic optimized by the intelligent agent, extract the nodes contained in the segment routing header of the SRv6 traffic, accumulate the SRv6 traffic passing through the node, and obtain the deployment indicator value of each node.

[0030] Compared with the prior art, the beneficial technical effects of the present invention are:

[0031] 1. Low-cost and efficient progressive deployment strategy: Based on the optimization results of the representative traffic matrix by the deep reinforcement learning agent, the present invention accurately identifies the key nodes that contribute the most to the traffic engineering (TE) optimization, and upgrades them to SRv6 nodes first. SRv6 nodes can be deployed in more critical locations, and the maximum link utilization optimization effect close to the full SRv6 network can be achieved with less deployment cost. In one embodiment, the present invention can calculate the deployment index (DI) in a fine-grained manner according to each SRv6 node specified in the segment routing header, quantify the contribution of the node to the traffic optimization through the deployment index, and find out the nodes that really play a key role in routing optimization.

[0032] 2. Fast routing optimization: The algorithm proposed in this paper can quickly calculate the optimized traffic routing and make adjustments quickly when the traffic changes, which is suitable for dynamic network environments. Deep reinforcement learning agents can make decisions quickly, local search only optimizes a small number of traffic paths, and can be stopped at any time.

[0033] 3. Considering the additional bandwidth used by SRv6 technology: When optimizing the traffic path, the intelligent agent of the present invention incorporates the additional bandwidth used by SRv6 technology for the current key traffic into the state of reinforcement learning. On the premise of considering the additional bandwidth used by SRv6 technology, it can perform more accurate traffic routing optimization and effectively reduce the cost of deploying SRv6 nodes. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 is a flow chart of a method in one embodiment of the present invention;

[0035] Figure 2 A schematic diagram of traffic routing optimization for progressive deployment of SRv6 in one of the embodiments of the present invention;

[0036] Figure 3 The figure is a schematic diagram of an algorithm flow used when executing a method in one embodiment of the present invention. DETAILED DESCRIPTION

[0037] A preferred embodiment of the present invention is described in detail below with reference to the accompanying drawings.

[0038] like Figure 1 As shown, the present invention provides an SRv6 progressive deployment and traffic engineering method based on deep reinforcement learning, comprising the following steps:

[0039] S1: Training a deep reinforcement learning agent in a simulated network with all SRv6 nodes deployed;

[0040] S2: Based on the optimization results of the current network historical traffic matrix by the trained intelligent agent, the nodes with the largest SRv6 traffic transmission are obtained and upgraded to SRv6 nodes to obtain a hybrid network.

[0041] S3: The traffic matrix of the hybrid network is optimized through the intelligent agent, and the optimization result of the intelligent agent is corrected using the local search method.

[0042] The present invention achieves traffic optimization performance close to that of a full SRv6 network by upgrading a small number of SRv6 devices, effectively reducing deployment costs. The algorithm proposed in the present invention can quickly calculate the optimized traffic routing and make rapid adjustments when traffic changes, and is suitable for dynamic network environments.

[0043] In a preferred embodiment, the training of the deep reinforcement learning agent in step S1 specifically includes:

[0044] A deep reinforcement learning agent is trained using a proximal policy optimization algorithm to learn traffic path scheduling policies to reduce the maximum link utilization of the simulated network.

[0045] In a preferred embodiment, a proximal policy optimization algorithm is used to train a deep reinforcement learning agent to learn a traffic path scheduling strategy to reduce the maximum link utilization of a simulated network, specifically including:

[0046] Initialize the traffic matrix of the simulated network;

[0047] Find the first α% of the most congested links in the simulated network, and select the first β% of the traffic with the largest bandwidth consumption from the traffic passing through these links as the key traffic;

[0048] The intelligent agent optimizes the path of each key traffic in turn according to the bandwidth size and selects the optimized path.

[0049] The agent is trained in a simulated network environment where all SRv6 nodes are deployed. First, the traffic matrix to be optimized is initialized according to the default routing protocol, and the congestion of the links in the network is observed. The first α% of the most congested links in the network are found, and then the first β% of the traffic that consumes the most bandwidth is selected from the traffic passing through these links, which is called the critical traffic. The agent will optimize the paths of the critical traffic in descending order according to the bandwidth.

[0050] SRv6 nodes are network devices that support SRv6 functions and can identify and process SRv6 packets. In an SRv6 network, packets specify their forwarding paths by carrying an SRv6 Segment Routing Header (SRH), and SRv6 nodes forward packets based on the information in the Segment Routing Header.

[0051] In a preferred embodiment, the agent optimizes each key traffic path in turn according to the bandwidth size and selects the optimized path, specifically including:

[0052] For each key flow to be optimized, the agent first removes the impact of the key flow on the network bandwidth; each key flow has multiple available paths. When the key flow selects one of the available paths, the capacity of the link in the simulated network, the utilization of the link when the current key flow is not transmitted, the link utilization that the current key flow needs to occupy after using this path, and the additional bandwidth used by the current key flow using SRv6 technology are used as the state of reinforcement learning. , the agent is based on the state , output action After the action is executed in the environment, the current key traffic path changes, the maximum link utilization in the simulated network changes, and a reward is returned The intelligent agent stores the current experience in the experience pool and then optimizes the next key flow.

[0053] In a preferred embodiment, the reward The calculation method is

[0054] ;

[0055] in, Indicates the maximum link utilization before the action is executed; It represents the maximum link utilization after the action is executed. The multiplication by 10 in the above formula is for scale adjustment.

[0056] At the tth moment, for a flow to be optimized, the agent first removes the impact of this flow on the network bandwidth. There are multiple available paths for the flow. If the flow chooses one of the available paths, observe the capacity of the link in the current network, the utilization of the link when this flow is not transmitted, the link utilization that this flow needs to occupy after using this path, and the additional bandwidth used by the current flow using SRv6 technology. This information is combined into a matrix. Then the matrices corresponding to all available paths are aggregated into a larger matrix as input, which is the state of reinforcement learning. The intelligent body will output an action based on the state The action here is: select one of the available paths for this traffic flow as the optimized traffic path. After the action is executed in the environment, the traffic path changes, the maximum link utilization in the network changes, and a reward is returned .

[0057] Current key traffic The calculation method for the additional bandwidth used by using SRv6 technology is as follows:

[0058] ;

[0059] ;

[0060] in, Indicates the current critical flow Using Paths The additional bandwidth consumed when SRv6 is used. If SRv6 technology is not used, the value is 0. Indicates the current critical flow The rate at which data packets are transmitted, assuming that the data packet transmission rate remains constant in the end-to-end transmission path. Indicates that when using SRv6 technology, the current key traffic The additional length of the packet encapsulation. This is the basic overhead of the segment routing header in the SRv6 protocol. The length of the Type-Length-Value (TLV) field is ignored here. According to the SRv6 technical specification, byte. The length of the segment identifier (Segement ID, SID) in the segment routing header. Each SID occupies 16 bytes. , The number of SIDs when using this path. Bytes, indicating the length of the additionally encapsulated IPv6 header. An additional IPv6 header needs to be encapsulated here to help data packets be transmitted smoothly in a mixed IP and SRv6 network.

[0061] The present invention incorporates the additional bandwidth used by the current critical traffic using the SRv6 technology into the state of reinforcement learning, which can further reduce the cost of deploying SRv6 nodes and improve deployment efficiency.

[0062] In a preferred embodiment, the optimization result of the current network historical traffic matrix by the trained agent in step S2 is used to obtain a preset number of nodes that transmit the most SRv6 traffic and upgrade them to SRv6 nodes, which specifically includes:

[0063] A representative traffic matrix is ​​generated based on the historical traffic matrix of the current network, and the trained intelligent agent is used to optimize the path of the key traffic in the representative traffic matrix. The deployment index of each node in the current network is calculated based on the optimization result. The nodes can be sorted according to the deployment index, and a set number of nodes with the highest deployment index value are upgraded to SRv6 nodes.

[0064] The historical traffic matrix in the current network is clustered using a clustering method to obtain a representative traffic matrix. The representative traffic matrix is ​​optimized by the agent. After the optimization is completed, the deployment index (DI) of all nodes is initialized to 0, and then the traffic that is not transmitted using SRv6 technology is removed, and the segment routing header (SRH) carried in the data packet of the SRv6 traffic is read. The nodes contained in the segment routing header of a specific SRv6 traffic are found, and the deployment index of these nodes is added to the size of this SRv6 traffic. Repeat this operation for all SRv6 traffic. Then sort the nodes in the current network according to the size of the deployment index. According to the number of SRv6 nodes that the current operator wants to upgrade, the nodes with larger deployment indexes, that is, the nodes that transmit more SRv6 traffic in traffic engineering (TE), are upgraded to SRv6 nodes.

[0065] In a preferred embodiment, the deployment index of each node in the current network is calculated based on the optimization result, specifically including:

[0066] Traverse the SRv6 traffic optimized by the intelligent agent, extract the nodes contained in the segment routing header of the SRv6 traffic, accumulate the SRv6 traffic passing through the node, and obtain the deployment indicator value of each node.

[0067] In a preferred embodiment, the flow matrix of the hybrid network is optimized by an agent in step S3, and the optimization result of the agent is corrected by a local search method, which specifically includes:

[0068] The traffic matrix in the hybrid network is optimized by the intelligent agent to generate a preliminary traffic scheduling plan;

[0069] Modify the default routes for the paths without SRv6 nodes deployed in the preliminary traffic scheduling plan;

[0070] A local search algorithm is used to iteratively adjust the path of the critical traffic until the maximum link utilization of the critical traffic cannot be further reduced.

[0071] After the SRv6 nodes are deployed, the agent is applied to the hybrid network to optimize the current traffic matrix and find the key traffic in the network at this time. Since the agent is trained in a full SRv6 network, the output of the agent may schedule ordinary IPv6 nodes as SRv6 nodes for traffic. Therefore, these traffic schedules are restored to use the default path transmission. Then find the key traffic in the network at this time and use the local search method to optimize the path. Traverse the available paths of the traffic from large to small, and select the path that can minimize the maximum link utilization as the transmission path of the traffic until the local search (LS) is stopped when the maximum link utilization of the network can no longer be reduced. Complete the path optimization of network traffic.

[0072] In one of the embodiments, the present invention can use the proximal policy optimization (PPO) algorithm to train a deep reinforcement learning (DRL) agent in a simulated network environment where all SRv6 nodes are deployed, so that the agent can learn how to use the paths provided by SRv6 technology to schedule traffic to reduce the maximum link utilization. After the training is completed, the nodes that transmit more SRv6 traffic are upgraded to SRv6 nodes based on the agent's optimization of traffic. After the deployment is completed, the agent optimizes the network traffic matrix (traffic demand in the network) under the actual SRv6 deployment ratio. Then the local search (LS) method is used to quickly optimize the agent's traffic scheduling decisions to effectively reduce the maximum link utilization. In one of the embodiments, the internal neural network of the agent uses a fully connected network (FCN) to reduce computational complexity and improve decision-making speed. In one of the embodiments, the algorithm flow used when the method of the present invention is running is as follows. Figure 3 shown.

[0073] The traffic matrix is ​​a matrix that describes the traffic demand between all source nodes and destination nodes in the network. It fully reflects the composition and flow of all traffic in the network.

[0074] In one embodiment, if Figure 2 As shown, in a network with 5 nodes, two nodes are SRv6 nodes and the other nodes are IPv6 nodes. Link congestion will occur when using the default route to transmit traffic; after using the method in the present invention to change the traffic transmission path, link congestion is eliminated.

[0075] It is obvious to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential features of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting, and the scope of the present invention is defined by the appended claims rather than the above description, and it is intended that all changes falling within the meaning and scope of the equivalent elements of the claims are included in the present invention, and any reference numerals in the claims should not be regarded as limiting the claims involved.

[0076] In addition, it should be understood that although this specification is described in terms of implementation methods, not every implementation method contains only one independent technical solution. This narrative method of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment may also be appropriately combined to form other implementation methods that can be understood by those skilled in the art.

Claims

1. A SRv6 progressive deployment and traffic engineering method based on deep reinforcement learning, characterized in that: include: In a simulated network where all SRv6 nodes are deployed, a proximal policy optimization algorithm is used to train a deep reinforcement learning agent to learn a traffic path scheduling strategy to reduce the maximum link utilization of the simulated network. The agent optimizes the paths of the selected key flows in order from large to small according to bandwidth: for each key flow to be optimized, the agent first removes the impact of the key flow on the network bandwidth; each key flow has multiple available paths. When the key flow selects one of the available paths, the capacity of the link in the simulated network, the utilization of the link when the current key flow is not transmitted, the link utilization that the current key flow needs to occupy after using this path, and the additional bandwidth used by the current key flow using SRv6 technology are used as the state of reinforcement learning. , the agent is based on the state Output Action After the action is executed in the environment, the current key traffic path changes, the maximum link utilization in the simulated network changes, and a reward is returned ; The agent stores the current experience in the experience pool, and then optimizes the next key flow; According to the optimization results of the current network historical traffic matrix by the trained intelligent agent, the nodes with the largest SRv6 traffic transmission are obtained and upgraded to SRv6 nodes to obtain a hybrid network. The traffic matrix in the hybrid network is optimized by intelligent agents to generate a preliminary traffic path scheduling strategy. The default route is corrected for the paths involving undeployed SRv6 nodes in the preliminary traffic path scheduling strategy. The path of the critical traffic is iteratively adjusted using a local search algorithm until the maximum link utilization of the critical traffic can no longer be reduced.

2. The SRv6 progressive deployment and traffic engineering method based on deep reinforcement learning according to claim 1, characterized in that: The method of using a proximal policy optimization algorithm to train a deep reinforcement learning agent and learn a traffic path scheduling strategy to reduce the maximum link utilization of a simulated network specifically includes: Initialize the traffic matrix of the simulated network; Find the first α% of the most congested links in the simulated network, and select the first β% of the traffic with the largest bandwidth consumption from the traffic passing through these most congested links as the key traffic; The intelligent agent optimizes the path of each key traffic in turn according to the bandwidth size and selects the optimized path.

3. The SRv6 progressive deployment and traffic engineering method based on deep reinforcement learning according to claim 1, characterized in that: The action is: selecting one of the available paths of the current key flow as the path of the optimized key flow.

4. The SRv6 progressive deployment and traffic engineering method based on deep reinforcement learning according to claim 1, characterized in that: The Reward The calculation method is: ; in, It represents the maximum link utilization before the action is executed, that is, at the t-1th time; It represents the maximum link utilization after the action is executed, that is, at the tth moment.

5. The SRv6 progressive deployment and traffic engineering method based on deep reinforcement learning according to claim 1, characterized in that: The additional bandwidth used by the SRv6 technology for the current critical traffic is calculated as follows: ; ; in, Indicates the current critical flow Using Paths Additional bandwidth consumed when Indicates the current critical flow The rate at which packets are transmitted; Indicates that when using SRv6 technology, the current key traffic The additional length of the data packet encapsulation; It is the basic overhead of the segment routing header in the SRv6 protocol. The length of the segment identifier in the segment routing header. Indicates the length of the external additional encapsulated IPv6 header.

6. The SRv6 progressive deployment and traffic engineering method based on deep reinforcement learning according to claim 1, characterized in that: The optimization result of the current network historical traffic matrix by the trained intelligent agent is used to obtain a preset number of nodes that transmit the most SRv6 traffic and upgrade them to SRv6 nodes, specifically including: Generate a representative traffic matrix based on the historical traffic matrix of the current network, use the trained intelligent agent to optimize the path of the representative traffic matrix, and calculate the deployment index of each node in the current network based on the optimization result; The deployment indicators are sorted and a set number of nodes with the highest deployment indicator values ​​are upgraded to SRv6 nodes.

7. The SRv6 progressive deployment and traffic engineering method based on deep reinforcement learning according to claim 6, characterized in that: The calculation of the deployment index of each node in the current network based on the optimization result specifically includes: Traverse the SRv6 traffic optimized by the intelligent agent, extract the nodes contained in the segment routing header of the SRv6 traffic, accumulate the SRv6 traffic passing through the node, and obtain the deployment indicator value of each node.

Citation Information

Patent Citations

  • Flow scheduling optimization method and system in SRv6 network, and medium

    CN117041133A

  • Key flow routing optimization method based on deep reinforcement learning

    CN119363647A