Fast range routing method based on key flow

By identifying and simplifying calculations based on a reinforcement learning algorithm based on key flows, the complexity of the range routing model is reduced, the computational bottleneck when network nodes increase is resolved, efficient routing strategy generation is achieved, and network performance is improved.

CN120602394AActive Publication Date: 2025-09-05BEIJING INST OF TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511046393.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-09-05
Estimated Expiration
2045-07-29

AI Technical Summary

Technical Problem

Existing range routing models suffer from high computational complexity when the number of network nodes increases, leading to hardware resource limitations and reduced scalability.

Method used

A fast range routing method based on key flows is adopted. Key flows and links are identified through reinforcement learning algorithm, non-key flow calculations are simplified, only key flows are solved, model variables and constraints are reduced, and routing strategies are generated using SMORE algorithm and Gurobi optimizer.

Benefits of technology

The solution speed of the range routing model has been significantly improved, scalability has been optimized, and the problem solving speed has been increased by 21 times.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120602394A_ABST
    Figure CN120602394A_ABST
Patent Text Reader

Abstract

The invention discloses a fast range routing method based on a key flow, which completes training of a flow identification agent by constructing a reinforcement learning algorithm, realizes judgment of flow and link attributes according to flow data, selects a key flow and a key link in a regional manner, and on the basis of the judgment, performs routing on the key flow and the key link. And inputting a key flow selection result and a key link selection result into a key flow range routing model, and generating routing strategies for the key flow and the non-key flow by adopting a key flow routing generation sub-model and a non-key flow routing generation sub-model respectively. According to the model, the range routing model is adopted to only calculate the key flow, and meanwhile, only the constraint is constructed for the key link, so that the number of model variables and constraints is reduced, the solving speed of the range routing model is remarkably improved, and the expandability of the range routing model is effectively optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of computer networks, and in particular relates to a fast range routing method based on key flows. Background Art

[0002] With the rapid development of the internet and communications technologies, and the emergence of new applications such as the Internet of Things and autonomous driving, traffic fluctuations in wide area networks (WANs) are becoming increasingly severe. Dynamic traffic within the network can cause congestion, severely degrading network performance. Traffic engineering is a network management and optimization method that precisely controls network traffic routing and distribution, optimizing resource utilization and improving network performance. Effectively addressing network fluctuations is a major challenge in the field of traffic engineering.

[0003] Traditional traffic engineering algorithms typically generate routing strategies in two steps: first, modeling the routing problem as a linear programming problem, such as a multi-commodity flow model; second, using a linear programming solver to solve the linear programming problem and generate a routing strategy for each traffic matrix. However, the routing strategies generated by the algorithms proposed in existing research are typically based on a fixed traffic matrix. When network fluctuations occur, the traffic matrix undergoes significant changes, and the routing strategy generated based on the historical traffic matrix is ​​no longer suitable for the current traffic matrix, resulting in degraded network performance.

[0004] Range routing is a routing modeling approach that mitigates the impact of network fluctuations. Unlike multi-commodity flow models, the optimization goal of range routing is to search for a suboptimal routing solution for all traffic matrices within a range. This ensures acceptable network performance across a wide range of traffic matrices, thereby mitigating the negative impact of network fluctuations. However, as the number of nodes in the network topology increases, the time required to solve the range routing problem increases dramatically.

[0005] The scalability of range routing models is a major obstacle hindering the practical application of this technology. Existing methods attempt to solve routing policies by fitting range routing models with neural networks. However, obtaining the labels of the routing policies used as training samples in these methods still requires solving the range routing model. Therefore, as the number of network nodes increases, the model solution may be limited by hardware resources, reducing the scalability of the range routing model. Summary of the Invention

[0006] In view of this, the present invention provides a fast range routing method based on key flows, establishes a key flow range routing model, reduces the computational complexity of the model solution by simplifying the calculation of non-key flows, and achieves a balance between performance and computational overhead in range routing.

[0007] The present invention provides a fast range routing method based on key flows, which specifically includes the following steps:

[0008] Step 1: In a wide area network environment, a traffic matrix is ​​obtained by sampling at a certain sampling interval to establish a training sample data set;

[0009] Step 2: A policy-based reinforcement learning architecture is used to establish a processing model including a traffic identification agent, a key flow range routing model, and a wide area network environment. The traffic identification agent outputs a key flow selection result based on the state. The input of the key flow range routing model is the action and state, and the output is the routing strategy. The routing strategy is deployed in the wide area network environment to calculate the reward, thereby updating the traffic identification agent. The state is the traffic matrix of the wide area network environment, the action is the selection result of the key flow and the key link, and the reward is the inverse of the maximum link utilization of the wide area network environment. The selection result of the key flow is either a key flow or a non-key flow, and the selection result of the key link is a key link.

[0010] Step 3: In actual use, the traffic matrix is ​​used as the input of the traffic identification agent to generate the selection results of key flows and key links. The selection results of key flows and key links and the historical traffic matrix are then input into the key flow range routing model to generate a routing strategy. Finally, the routing strategy is deployed in the wide area network environment to complete routing.

[0011] Furthermore, the critical flow range routing model includes a non-critical flow routing generation sub-model and a critical flow routing generation sub-model. The input of the non-critical flow routing generation sub-model is the non-critical flow and flow matrix, and the output is the non-critical flow routing strategy. The input of the critical flow routing generation sub-model is the critical flow and flow matrix, and the output is the critical flow routing strategy.

[0012] Furthermore, the non-critical flow routing generation sub-model is constructed using the SMORE algorithm, which is expressed as:

[0013] Bk e =SMORE(TM t )

[0014] Among them, Bk e is the traffic of non-critical flow on link e, TM t is the traffic matrix at the current moment.

[0015] Furthermore, the key flow routing generation sub-model is constructed using a range routing model, which is expressed as:

[0016] min PR

[0017]

[0018] Among them, PR is the performance gap between the current routing strategy and the optimal routing strategy. P is the proportion of the traffic demand from source s to destination d allocated to path p.s,d is the path set between source s and target d, V is the network node set, Bk e is the traffic of non-critical flow on link e, is the path splitting ratio, D s,d is the traffic demand from source s to destination d, c e is the capacity of link e, U opt (TM) is the maximum link utilization that can be achieved under the traffic matrix condition, E c is the selected set of key links, is the training sample dataset s t The flow from source s to destination d in .

[0019] Furthermore, the performance gap PR between the current routing strategy and the optimal routing strategy is calculated as follows:

[0020]

[0021] Among them, U R (TM) is the maximum link utilization of routing policy R, U opt (TM) is the maximum link utilization that can be achieved under the traffic matrix conditions.

[0022] Furthermore, the key flow routing generation sub-model is solved by constructing an auxiliary linear programming problem under the condition of a given routing strategy. The auxiliary linear programming problem is expressed as:

[0023]

[0024] η>0

[0025] in, is the total flow allocated to path p from source s to destination d, and η is the flow multiplier to ensure that the linear programming problem has a solution.

[0026] Furthermore, the key flow routing generation sub-model uses the mathematical programming solver Gurobi to complete the solution and obtain the optimal routing strategy.

[0027] Furthermore, the flow matrix is ​​used as the input of the flow identification agent to generate the selection results of key flows and key links, specifically in the following manner:

[0028] For each state s tAfter being processed by the policy network of the traffic identification agent, a vector I with a length of N×(N-1) and a vector J with a length of M are output. Each element in vector I corresponds to the importance score of a flow, and each element in vector J corresponds to the importance score of a link. On this basis, K samples are taken as the selection probability of each flow and link based on the importance score to generate K actions, where N is the number of nodes in the network topology and M is the number of links in the network topology.

[0029] Beneficial effects:

[0030] This invention trains a traffic identification agent through a reinforcement learning algorithm. It uses traffic data to determine flow and link attributes to identify key flows and key links. Based on this, the key flow and link selection results are input into a key flow range routing model. The key flow routing generation sub-model and the non-key flow routing generation sub-model generate routing strategies for key flows and non-key flows, respectively. This model uses a range routing model to calculate only key flows and establish constraints only for key links. This reduces the number of model variables and constraints, significantly improves the solution speed of the range routing model, and effectively optimizes the scalability of the range routing model. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 A schematic diagram of the overall structure of a model of a fast range routing method based on key flows provided by the present invention.

[0032] Figure 2 A schematic diagram of the model structure of a key flow range routing model constructed by a fast range routing method based on key flows provided by the present invention.

[0033] Figure 3 The present invention provides a practical application process of a fast range routing method based on key flows. DETAILED DESCRIPTION

[0034] The present invention is described in detail below with reference to the accompanying drawings and with reference to the embodiments.

[0035] This paper provides a fast range routing method based on key flows. Its core concept is to use a policy-based deep reinforcement learning framework to learn how to select key flows and key links from a historical traffic matrix. Based on this, a key flow range routing model is proposed based on the existing range routing model. This method simplifies the calculation of non-key flows to reduce model complexity, and only solves for key flows to obtain routing strategies, accelerating the solution of the range routing model. Furthermore, by constraining only key links, the number of constraints in the linear programming problem is reduced.

[0036] The present invention provides a fast range routing method based on key flows, which specifically includes the following steps:

[0037] Step 1: Establish a training sample dataset based on the wide area network environment.

[0038] In a WAN environment, the traffic matrix is ​​sampled at a fixed sampling interval, TM t Represents the traffic matrix obtained by the t-th sampling. Assuming that the current moment is the moment of sampling the traffic matrix, the training sample data set is constructed with l historical traffic matrices, that is, s t ={TM t-l+1 ,TM t-l+2 ,…,TM t}.

[0039] Step 2: The fast range routing method based on key flows proposed in this invention adopts a policy-based reinforcement learning architecture. The overall process of the method is as follows: Figure 1 As shown in the figure, it mainly includes traffic identification agent, key flow range routing model and wide area network environment, among which the training sample data set s t The traffic identification agent outputs the key flow and key link selection results, i.e., actions, based on the state, the key flow range routing model takes the action and state as input, and the output is the routing strategy. Finally, the routing strategy is deployed in the wide area network environment to test the performance of the routing strategy and calculate the reward, which is used to further update the relevant parameters of the traffic identification agent.

[0040] Specifically, the traffic identification agent outputs key flows and key link selections (actions) based on the input historical traffic matrix (state). For key flows, a key flow routing strategy is generated based on the key flow routing sub-model. For other non-critical flows, a non-critical flow routing strategy is generated based on the non-critical flow routing sub-model. Together, these two strategies form a complete routing strategy. Further deploying this routing strategy in a wide area network environment maximizes link utilization, which in turn allows for the calculation of rewards for parameter updates of the traffic identification agent.

[0041] Furthermore, the strategy network adopted by the traffic identification agent in the present invention is composed of a convolutional neural network backbone network and a fully connected layer, wherein the function of the fully connected layer is dimensionality reduction.

[0042] Actions represent the selection results of key flows and key links. Each source-destination pair represented by two nodes represents a flow. For a network topology with N nodes and M links, there are N×(N-1) flows. The action acquisition process is: for each state s tAfter being processed by the strategy network of the traffic identification agent, the output is a vector I of length N×(N-1) and a vector J of length M. Each element in vector I corresponds to the importance score of a flow, and each element in vector J corresponds to the importance score of a link. On this basis, the importance score is used as the selection probability of each flow and link to perform K sampling to generate K actions The definition of each action is as follows: first, the action space is defined as {0,1,...,(N*(N-1)-1)} and {0,1,...,M}, each number represents a flow and link, and each sampling selects the number of key flows k in the above two spaces. n =N×(N-1)×cr n and the number of critical links k l =N×(N-1)×cr l Non-repeating numbers are used as the numbers of the selected key flows and key links. The actions are the above two numbers with length k. n and k l vector of cr n and cr l is the compression ratio, which represents the ratio of the number of key flows and key links to the total number of flows.

[0043] The process of obtaining the reward r corresponding to the action is: select the result of the key flow With state s t The range of the key flow range routing model is input, and the routing strategy is obtained by solving the range routing model Gurobi, which is equivalent to the intermediate amount of the reward. The strategy is deployed in the wide area network environment. t+1 Calculate the maximum link utilization (MLU) U t and with U t The reciprocal of is the reward r t .

[0044] U t =CFRR(a t ,TM t+1 ) (1)

[0045]

[0046] On this basis, the present invention adopts a policy gradient-based reinforcement learning method (REINFORCE algorithm) to directly optimize the policy parameters to maximize the expected cumulative reward, with the goal of improving the expected return as the goal of deep reinforcement learning. The specific training process is as follows.

[0047] First, the action selection probability is expressed as the product of the importance scores of each key flow contained in the action:

[0048]

[0049] Where θ represents the traffic recognition agent parameters and π represents the action selection probability. Based on this, the average reward is used as the baseline value to update the parameters of the policy network:

[0050]

[0051] Among them, α represents the learning rate, b(s t ) represents the average reward baseline.

[0052] Step 3: The key flow range routing model takes state and action as input, divides the entire flow set into key flows and non-key flows based on the action, and divides the entire link into key links and non-key links. On this basis, the non-key flow routing generation sub-model and the key flow routing generation sub-model are used to generate routing strategies for the current state for non-key flows and key flows respectively. The key flow range routing model structure is as follows: Figure 2 shown.

[0053] The routing strategy for non-critical flows is based on the current traffic matrix TM t In the present invention, the non-critical flow routing generation sub-model is constructed using the SMORE algorithm, which is specifically expressed as follows:

[0054] Bk e =SMORE(TM t ) (5)

[0055] Among them, Bk e ,e∈E is the non-critical flow on each link after the SMORE algorithm obtains the routing strategy, TM t Represents the traffic matrix at the current moment.

[0056] Furthermore, the critical flow routing generation sub-model in the present invention is constructed using the scope routing model. Specifically, the routing strategy optimization process for critical flows in the present invention is similar to that of the scope routing model, except that the traditional scope routing model does not need to consider the background traffic generated by non-critical flows. In addition, the present invention also reduces the number of constraints in the scope routing model through non-critical flows. Critical flow scope routing considers scope routing of critical flows when non-critical flows generate background traffic.

[0057] First, given a matrix and routing strategy R:

[0058]

[0059] Among them U R (TM) is the maximum link utilization of routing policy R, U opt (TM) is the maximum link utilization that can be achieved under the traffic matrix conditions, and PR represents the performance gap between the current routing strategy and the optimal routing strategy.

[0060] The specific modeling of the key flow range routing model is as follows: Under the given topology structure G(V,E), each flow has a certain number of paths p, and the routing strategy needs to take the traffic demand D of each flow into account. s,d according to The system needs to optimize the proportion of the multiple paths of the flow. The allocation strategy reduces PR, and the specific modeling is as follows:

[0061] Table 1 Meaning of modeling symbols for key flow range routing model

[0062]

[0063]

[0064] min PR (7a)

[0065]

[0066] Among them, formulas (7b) and (7c) are the path splitting rates , (7d) is the related constraint of PR, the model is optimized for all TMs within the range of (7d), E c Indicates the selected set of critical links.

[0067] The present invention is a modification of the existing model. Because the range routing model has high computational complexity when the number of flows is very large, the present invention distinguishes between critical and non-critical flows and applies the range routing model only to critical flows, effectively reducing the number of flows participating in the range routing model and thus effectively reducing computational complexity.

[0068] The above constraint (7d) can be further written as follows by scaling:

[0069]

[0070] Under the given routing strategy, this constraint can be solved by constructing an auxiliary linear programming problem:

[0071]

[0072]

[0073] η>0 (9g)

[0074] Model (9) undergoes dual operation and combines the constraints in the original model (7) to obtain the final key flow range routing model:

[0075] min PR (10a)

[0076]

[0077] e∈E c ,s,d∈V,s≠d (10f)

[0078]

[0079] e∈E c p∈P s,d ,s,d∈V,s≠d

[0080] π(e,e * )≥0 e,e * ∈E c (10h)

[0081]

[0082] The model can be solved by the linear programming optimizer Gurobi to obtain the optimal path splitting rate The final routing strategy obtained by the algorithm is to route non-critical flows through SMORE and use Routing policy.

[0083] Step 4: In actual use, use Figure 3 The process shown takes the historical traffic matrix as input, generates key flow selection, then generates routing strategies based on the key flow range routing model, and finally deploys the routing strategies in the WAN environment.

[0084] The present invention constructs a fast range routing method based on key flows, which first selects key flows and key links through a policy network. Then, the key flows, key link selections and traffic matrix ranges are input into a key flow range routing model to obtain rewards to evaluate the quality of key flow and key link selections. Finally, the average reward is used as the baseline value, and the policy network parameters are updated by maximizing the expected reward. For the construction of the key flow range routing model, the present invention reduces the number of flows that require range routing through key flow selection, reduces the number of model constraints through key link selection, and adopts the SMORE routing algorithm with lower complexity for non-key flows, thereby greatly reducing the number of variables and constraints in the original range routing model and greatly improving the scalability of the algorithm. The performance of the method proposed by the present invention was evaluated through experimental simulation. Compared with the original range routing model, the present invention can increase the problem-solving speed by 21 times under the premise of little difference in performance.

[0085] Finally, it should be noted that it is apparent that those skilled in the art may make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, to the extent such modifications and variations fall within the scope of the claims and their equivalents, the present invention is intended to encompass such modifications and variations.

[0086] The above is only an embodiment of the present invention, but it is not intended to limit the scope of the present invention. Any structural changes made according to the present invention, as long as they do not lose the essence of the present invention, should be considered to fall within the scope of protection of the present invention and be subject to restrictions. Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working process and related instructions of the method described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0087] The term "comprise," "comprising," or any other similar term is intended to cover a non-exclusive inclusion such that a process, method, article, or apparatus / method that comprises a list of elements includes not only those elements but also other elements not expressly listed or inherent to such process, method, article, or apparatus / method.

[0088] Thus far, further embodiments have been listed to describe the technical solutions of the present invention. However, it is readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present invention.

[0089] In summary, the above are only preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A fast range routing method based on key flows, characterized in that: The specific steps include: Step 1: In a wide area network environment, a traffic matrix is ​​obtained by sampling at a certain sampling interval to establish a training sample data set; Step 2: A policy-based reinforcement learning architecture is used to establish a processing model including a traffic identification agent, a key flow range routing model, and a wide area network environment. The traffic identification agent outputs a key flow selection result based on the state. The input of the key flow range routing model is the action and state, and the output is the routing strategy. The routing strategy is deployed in the wide area network environment to calculate the reward, thereby updating the traffic identification agent. The state is the traffic matrix of the wide area network environment, the action is the selection result of the key flow and the key link, and the reward is the inverse of the maximum link utilization of the wide area network environment. The selection result of the key flow is either a key flow or a non-key flow, and the selection result of the key link is a key link. Step 3: In actual use, the traffic matrix is ​​used as the input of the traffic identification agent to generate the selection results of key flows and key links. The selection results of key flows and key links and the historical traffic matrix are then input into the key flow range routing model to generate a routing strategy. Finally, the routing strategy is deployed in the wide area network environment to complete routing.

2. The fast range routing method according to claim 1, characterized in that The critical flow range routing model includes a non-critical flow routing generation sub-model and a critical flow routing generation sub-model. The input of the non-critical flow routing generation sub-model is the non-critical flow and flow matrix, and the output is the non-critical flow routing strategy. The input of the critical flow routing generation sub-model is the critical flow and flow matrix, and the output is the critical flow routing strategy.

3. The fast range routing method according to claim 2, characterized in that The non-critical flow routing generation sub-model is constructed using the SMORE algorithm and is expressed as: Bk e =SMORE(TM t ) Among them, Bk e is the traffic of non-critical flow on link e, TM t is the traffic matrix at the current moment.

4. The fast range routing method according to claim 2, wherein: The key flow routing generation sub-model is constructed using the range routing model, which is expressed as: min PR Among them, PR is the performance gap between the current routing strategy and the optimal routing strategy. P is the proportion of the traffic demand from source s to destination d allocated to path p. s,d is the path set between source s and target d, V is the network node set, Bk e is the traffic of non-critical flow on link e, is the path splitting ratio, D s,d is the traffic demand from source s to destination d, c e is the capacity of link e, U opt (TM) is the maximum link utilization that can be achieved under the traffic matrix condition, E c is the selected set of key links, is the training sample data set s t The flow from source s to destination d in .

5. The fast range routing method according to claim 4, characterized in that: The performance gap PR between the current routing strategy and the optimal routing strategy is calculated as follows: Among them, U R (TM) is the maximum link utilization of routing policy R, U opt (TM) is the maximum link utilization that can be achieved under the traffic matrix conditions.

6. The fast range routing method according to claim 4, characterized in that: The key flow routing sub-model is solved by constructing an auxiliary linear programming problem under the condition of a given routing strategy. The auxiliary linear programming problem is expressed as: η>0 in, is the total flow allocated to path p from source s to destination d, and η is the flow multiplier to ensure that the linear programming problem has a solution.

7. The fast range routing method according to claim 6, characterized in that: The key flow routing generation sub-model adopts the mathematical programming solver Gurobi to complete the solution and obtain the optimal routing strategy.

8. The fast range routing method according to claim 1, wherein: The flow matrix is ​​used as the input of the flow identification agent to generate the selection results of key flows and key links. The specific method is as follows: For each state s t After being processed by the policy network of the traffic identification agent, a vector I with a length of N×(N-1) and a vector J with a length of M are output. Each element in vector I corresponds to the importance score of a flow, and each element in vector J corresponds to the importance score of a link. On this basis, K samples are taken as the selection probability of each flow and link based on the importance score to generate K actions, where N is the number of nodes in the network topology and M is the number of links in the network topology.

Citation Information

Patent Citations

  • Extensible routing method for realizing load balancing in software-defined wide area network

    CN112187657A

  • Method and device for evaluating availability of link in network, electronic equipment and storage medium

    CN119966859A

  • Network intercommunication method and system, electronic equipment, storage medium and program product

    CN120281602A

  • Efficient machine learning for network optimization

    US20200136957A1