Flow distribution method and related equipment

By training the agent of deep learning algorithms in the training environment, the problem of the existing SR technology being too long in traffic allocation in complex networks is solved, fast and autonomous traffic allocation is achieved, and network load is reduced.

CN120301833APending Publication Date: 2025-07-11BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510204050.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing traffic engineering scheme based on SR technology is too long to calculate under complex network topology and large-scale traffic matrix. The linear planning scheme only optimizes one goal and is difficult to solve, while the heuristic algorithm is difficult to verify, resulting in the inability to quickly and efficiently allocate traffic.

Method used

In the construction training environment, deep learning algorithms are used to train the agent, deploy the actual network topology according to the pre-set reward function, respond to business allocation needs, determine the traffic allocation ratio of each segment of the routing path, and use deep reinforcement learning for path selection and allocation.

Benefits of technology

It significantly reduces decision time, improves the efficiency of path planning, reduces network link load, and achieves fast and autonomous traffic allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120301833A_ABST
    Figure CN120301833A_ABST
Patent Text Reader

Abstract

The invention provides a flow distribution method and related equipment. The method comprises the following steps: in a constructed training environment, according to a preset reward function, training to obtain an intelligent agent based on a deep learning algorithm; deploying an actual network topology according to the training environment; in response to a service distribution demand, inputting a traffic matrix corresponding to the service into the trained agent to obtain a target action; the target action comprises a flow distribution proportion for each segment routing path; and determining a traffic distribution proportion of each segment routing path based on the target action, and performing traffic distribution processing on the service according to the traffic distribution proportion. According to the embodiment of the invention, the intelligent agent capable of quickly and autonomously making a decision is obtained through deep reinforcement learning training, and reasonable path selection and traffic distribution are performed on each traffic request, so that the global network link load is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of traffic allocation, and in particular, to a traffic allocation method and related devices. Background Art

[0002] With the development of Software-Defined Networking (SDN), the flexibility and programmability of the network have been improved. Network administrators can dynamically adjust network topologies, traffic flows, etc. according to actual needs to meet the requirements of different application scenarios. Traditional network IGP (Interior Gateway Protocol) protocols such as IS-IS and OSPF protocols use unified IGP metrics to determine fixed routing policies, lacking flexibility and conflicting with the concept of flexible control in SDN. Therefore, the combination of traffic engineering and SDN has gradually become the mainstream solution to reasonably allocate and utilize network bandwidth and resources according to actual needs.

[0003] Compared with traditional SDN technologies that require deploying SDN nodes throughout the network to achieve the function of flexible routing, SR (Segment Routing) technology can determine the entire forwarding path at the routing source end and does not require the deployment of SDN nodes, facilitating network upgrade and maintenance. SRv6 (Segment Routing IPv6) is based on SR technology, facing the IPv6 addressing space and extensible packet headers, and encapsulating fields on the extensible packet headers to identify SID (Segment Identifier) to guide the specific path of packet forwarding, facilitating traffic engineering control across the network to achieve balanced load of the entire network traffic.

[0004] Existing traffic engineering solutions based on SR technology, in addition to specifying the forwarding path through SR technology, also need to allocate traffic among multiple forwarding paths to fully utilize the characteristics of the SR-guided path and break away from the limitations of the traditional IGP forwarding path. The current mainstream solution is to establish a linear programming / heuristic algorithm solution, and determine the traffic allocation ratio of each path of the entire network traffic through a linear programming problem or a heuristic algorithm, thereby reducing network load.

[0005] However, the common problem of linear programming problems and heuristic algorithms is that the calculation time is very slow. Especially for complex network topologies and large-scale traffic matrices, it takes seconds to determine the traffic allocation ratio, which is unacceptable for low-latency network environments. In addition, due to the complexity of the allocation problem, the linear programming solution is usually limited to optimizing one objective, otherwise it will fall into a difficult-to-solve situation. And heuristic algorithms are difficult to be theoretically verified and can only obtain local optimal solutions. Summary of the Invention

[0006] In view of this, the purpose of the present application is to propose a traffic allocation method and related devices.

[0007] Based on the above purpose, the present application provides a traffic allocation method, including:

[0008] In the constructed training environment, according to a preset reward function, an intelligent agent based on a deep learning algorithm is trained;

[0009] Deploy the actual network topology according to the training environment;

[0010] In response to the service allocation requirement, input the traffic matrix corresponding to the service into the trained intelligent agent to obtain a target action; the target action includes the traffic allocation ratio for each segment routing path;

[0011] Based on the target action, determine the traffic allocation ratio for each segment routing path, and perform traffic allocation processing on the service according to the traffic allocation ratio.

[0012] In a possible implementation manner, the training environment includes a first network topology and a first traffic matrix set;

[0013] The training environment is constructed through the following steps:

[0014] Use an adjacency matrix to represent the topological structure of the first network topology;

[0015] Use a triple to represent the first traffic matrix set; the triple includes the source node, destination node, and requested bandwidth size of the traffic request.

[0016] In a possible implementation manner, the reward function is calculated through the following steps:

[0017] Obtain the first maximum link utilization rate and the first average link utilization rate in the initial state of the preset state space;

[0018] Obtain the second maximum link utilization rate and the second average link utilization rate in the current state of the preset state space;

[0019] Based on the corresponding action in the preset action space, the first maximum link utilization rate, the first average link utilization rate, the second maximum link utilization rate, and the second average link utilization rate, determine the reward coefficient;

[0020] Calculate the reward function according to the state of the reward coefficient.

[0021] In a possible implementation manner, the reward coefficient is calculated by the following formula:

[0022]

[0023] Among them, α represents the reward coefficient, and U max (b1) represents the first maximum link utilization rate, and U avg (b1) represents the first average link utilization rate, and U max (b t ) represents the second maximum link utilization rate, and U avg (b t ) represents the second average link utilization rate.

[0024] In a possible implementation, the reward function is calculated by the following formula:

[0025]

[0026] Among them, r t represents the reward function, and α represents the reward coefficient.

[0027] In a possible implementation, determining the traffic allocation ratio of each segment routing path based on the target action and performing traffic allocation processing on the service according to the traffic allocation ratio includes:

[0028] Determining the traffic allocation ratio corresponding to the segment routing node and the transmission path based on the target action;

[0029] Configuring a subnet mask for each of the transmission paths;

[0030] Performing traffic allocation processing on the service based on the traffic allocation ratio and the transmission path.

[0031] Based on the same inventive concept, an embodiment of the present application further provides a traffic allocation device, including:

[0032] A training module, configured to train an intelligent agent based on a deep learning algorithm according to a preset reward function in a constructed training environment;

[0033] A deployment module, configured to deploy an actual network topology according to the training environment;

[0034] An input module, configured to input a traffic matrix corresponding to the service into the trained intelligent agent in response to a service allocation requirement, and obtain a target action; the target action includes the traffic allocation ratio for each segment routing path;

[0035] An allocation module, configured to determine the traffic allocation ratio of each segment routing path based on the target action and perform traffic allocation processing on the service according to the traffic allocation ratio.

[0036] Based on the same inventive concept, an embodiment of the present application further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the traffic allocation method described in any one of the above.

[0037] Based on the same inventive concept, an embodiment of the present application further provides a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the computer to execute the traffic allocation method described in any one of the above.

[0038] Based on the same inventive concept, an embodiment of the present application further provides a computer program product, which includes computer program instructions. The computer instructions are used to cause the computer program product to execute the traffic allocation method described in any one of the above.

[0039] As can be seen from the above, for the traffic allocation method and related devices provided by the present application, in a constructed training environment, an agent based on a deep learning algorithm is trained according to a preset reward function; an actual network topology is deployed according to the training environment; in response to a service allocation requirement, a traffic matrix corresponding to the service is input into the trained agent to obtain a target action; the target action includes the traffic allocation ratio for each segment routing path; the traffic allocation ratio for each segment routing path is determined based on the target action, and traffic allocation processing is performed on the service according to the traffic allocation ratio. In the embodiment of the present application, an agent capable of making quick autonomous decisions is trained through Deep Reinforcement Learning (DRL), and reasonable path selection and traffic allocation are performed for each traffic request, thereby reducing the global network link load. It should be noted that the difference between the present application and the prior art is that the segment routing (SR) decision is directly made through reinforcement learning without determining the SR through linear programming. Therefore, compared with the prior art, the present application greatly reduces the decision-making time and effectively reduces the cost of path planning. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the technical solutions in the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0041] Figure 1 It is a schematic flowchart of the traffic allocation method according to the embodiment of the present application;

[0042] Figure 2 Structural schematic diagram of the flow distribution device according to an embodiment of the present application;

[0043] Figure 3 Structural schematic diagram of the electronic device according to an embodiment of the present application. Detailed implementation manners

[0044] To make the objectives, technical solutions and advantages of the present application clearer and more understandable, the following further details the present application with reference to specific embodiments and the accompanying drawings.

[0045] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present application should have the ordinary meanings understood by those of ordinary skill in the art to which the present application belongs. The "first", "second" and similar terms used in the embodiments of the present application do not denote any order, quantity or importance, but are only used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left" and "right" are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0046] It can be understood that, before using the technical solutions of the various embodiments in the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved will be informed to the user in an appropriate manner and the user's authorization will be obtained.

[0047] For example, when responding to receiving an active request from the user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server or a storage medium that executes the operations of the technical solutions of the present disclosure according to the prompt message.

[0048] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving an active request from the user may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0049] It can be understood that the above-mentioned notification and the process of obtaining user authorization are only illustrative and do not limit the implementation manner of the present disclosure. Other manners that comply with relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0050] As described in the background art section, with the development of Software-Defined Networking (SDN), the flexibility and programmability of the network have been improved. Network administrators can dynamically adjust the network topology, traffic flow direction, etc. according to actual needs to meet the requirements of different application scenarios. Traditional network IGP (Interior Gateway Protocol) protocols such as IS-IS and OSPF protocols use a unified IGP metric to determine a fixed routing strategy, lacking flexibility and conflicting with the concept of flexible control of SDN. Therefore, the combination of traffic engineering and SDN has gradually become the mainstream solution to reasonably allocate and utilize network bandwidth and resources according to actual needs.

[0051] Compared with traditional SDN technologies that require deploying SDN nodes throughout the network to achieve the function of flexible routing, SR (Segment Routing) technology can determine the entire forwarding path at the routing source end and does not require the deployment of SDN nodes, facilitating network upgrade and maintenance. SRv6 (Segment Routing IPv6) is based on SR technology and faces the IPv6 addressing space and extensible packet header. It encapsulates fields to identify SID (Segment Identifier) on the extensible packet header to guide the specific path of packet forwarding, facilitating the implementation of traffic engineering control throughout the network to achieve balanced load of the entire network traffic.

[0052] Existing traffic engineering solutions based on SR technology, in addition to specifying the forwarding path through SR technology, also need to allocate traffic among multiple forwarding paths to fully utilize the characteristics of the SR-guided path and break away from the limitations of the traditional IGP forwarding path. The current mainstream solution is to establish a linear programming / heuristic algorithm solution to determine the traffic allocation ratio of each path of the entire network traffic through a linear programming problem or a heuristic algorithm, thereby reducing network load.

[0053] However, the common problem of linear programming problems and heuristic algorithms is that the calculation time is very slow. Especially for complex network topologies and large-scale traffic matrices, it takes seconds to determine the traffic allocation ratio, which is unacceptable for low-latency network environments. In addition, due to the complexity of the allocation problem, the linear programming solution is usually limited to optimizing one goal, otherwise it will fall into a difficult situation of solution. And it is difficult to conduct theoretical verification for heuristic algorithms, and only local optimal solutions can be obtained.

[0054] In view of the above considerations, an embodiment of the present application proposes a traffic allocation method. In a constructed training environment, an intelligent agent based on a deep learning algorithm is trained according to a preset reward function; an actual network topology is deployed according to the training environment; in response to a service allocation requirement, a traffic matrix corresponding to the service is input into the trained intelligent agent to obtain a target action; the target action includes a traffic allocation ratio for each segment routing path; based on the target action, the traffic allocation ratio for each segment routing path is determined, and traffic allocation processing is performed on the service according to the traffic allocation ratio. In an embodiment of the present application, an intelligent agent capable of making quick autonomous decisions is trained through deep reinforcement learning (DRL), and reasonable path selection and traffic allocation are performed for each traffic request, thereby reducing the global network link load. It should be noted that the difference between the present application and the prior art is that the segment routing (SR) decision is directly made through reinforcement learning without determining the SR through linear programming. Therefore, compared with the prior art, the present application greatly reduces the decision-making time and effectively reduces the cost of path planning.

[0055] Hereinafter, the technical solutions of the embodiments of the present application will be described in detail through specific embodiments.

[0056] Refer to Figure 1 , the traffic allocation method of the embodiment of the present application includes the following steps:

[0057] Step S101, in a constructed training environment, an intelligent agent based on a deep learning algorithm is trained according to a preset reward function;

[0058] Step S102, an actual network topology is deployed according to the training environment;

[0059] Step S103, in response to a service allocation requirement, a traffic matrix corresponding to the service is input into the trained intelligent agent to obtain a target action; the target action includes a traffic allocation ratio for each segment routing path;

[0060] Step S104, based on the target action, the traffic allocation ratio for each segment routing path is determined, and traffic allocation processing is performed on the service according to the traffic allocation ratio.

[0061] Regarding step S101, first, a training environment needs to be constructed.

[0062] In some embodiments, the training environment includes a first network topology and a first traffic matrix set;

[0063] The training environment is constructed through the following steps: representing the topological structure of the first network topology using an adjacency matrix; representing the first traffic matrix set using triples; the triples include the source node, destination node, and requested bandwidth size of the traffic request.

[0064] In this embodiment, it is necessary to establish a training data set, that is, the first network topology and the first traffic matrix set. The topological structure is represented in the form of an adjacency matrix, and each element of the traffic matrix is a traffic request.

[0065] Specifically, the network topological structure is represented by an adjacency matrix M, and the adjacency matrix element m ij represents the link bandwidth size between the i-th node and the j-th node. At the same time, a node list N is maintained, and the i-th element n i of the list represents the node name of the i-th node, which is used to correspond to the node names in the training data set.

[0066] The training traffic matrix set L, each traffic matrix contains traffic requests between network nodes, and the traffic matrix is represented by L i where l ij =(s ij ,t ij ,d ij ) is the j-th traffic request of the i-th traffic matrix, which is a triple containing the source node, destination node, and requested bandwidth size of the traffic request.

[0067] The source node represents the originating node of the traffic request, that is, the starting point of the traffic.

[0068] The destination node represents the terminating node of the traffic request, that is, the end point of the traffic.

[0069] The required bandwidth size represents the bandwidth size required for this traffic request (for example, in Mbps or Gbps).

[0070] The traffic matrix set usually contains multiple traffic matrices, which are used to represent multiple different traffic patterns or scenarios. The traffic matrix can be extracted from historical network data or generated based on a statistical model to cover as many traffic distribution situations as possible and ensure the generality of training.

[0071] The adjacency matrix is a two-dimensional array, where the dimension is N×N (N represents the number of network nodes). The elements in the matrix are defined as follows:

[0072] If there is a direct link between node i and node j, the element is the link capacity (such as 10 Gbps).

[0073] If there is no direct link between node i and node j, the element is 0.

[0074] Definition of the traffic matrix:

[0075] The traffic matrix is ​​also a two-dimensional array with a dimension of N × N. The elements d in the matrix ij Indicates the traffic request size from node i to node j (such as 10Mbps).

[0076] In a feasible embodiment, assume that we have a simple network topology, including 5 routing nodes, the nodes are numbered 0, 1, 2, 3, 4. The actual topology of the network is as follows:

[0077] There is a link with a capacity of 10 Gbps between node 0 and node 1.

[0078] There is a link with a capacity of 5 Gbps between node 0 and node 4.

[0079] There is a link with a capacity of 15 Gbps between Node 1 and Node 2.

[0080] There is a link with a capacity of 20 Gbps between node 2 and node 3.

[0081] There is a link with a capacity of 10 Gbps between node 3 and node 4.

[0082] The adjacency matrix of the network topology is as follows:

[0083]

[0084] There is a link with a capacity of 10 Gbps between node 0 and node 1, a link with a capacity of 5 Gbps between node 0 and node 4, and so on for other elements.

[0085] Prepare two traffic matrices, representing different traffic patterns, denoted as L1 and L2. Each traffic matrix consists of several traffic requests, each of which is a triplet l ij =(s ij ,t ij ,d ij ).

[0086] Traffic matrix L1:

[0087] The traffic request is from node 0 to node 1, and the requested bandwidth is 5 Mbps.

[0088] The traffic request is from node 1 to node 2, and the requested bandwidth is 8 Mbps.

[0089] The traffic request is from node 2 to node 3, and the requested bandwidth is 10 Mbps.

[0090] The traffic request is from node 3 to node 4, and the requested bandwidth is 5 Mbps.

[0091] The traffic request is from node 0 to node 4, and the requested bandwidth is 2 Mbps.

[0092] It is expressed as:

[0093] L1 = {(0, 1, 5), (1, 2, 8), (2, 3, 0), (3, 4, 5), (0, 4, 2)}

[0094] Traffic matrix L2:

[0095] The traffic request is from node 0 to node 4, and the requested bandwidth is 5 Mbps.

[0096] The traffic request is from node 1 to node 2, and the requested bandwidth is 10 Mbps.

[0097] The traffic request is from node 2 to node 3, and the requested bandwidth is 8 Mbps.

[0098] The traffic request is from node 3 to node 4, and the requested bandwidth is 3 Mbps.

[0099] The traffic request is from node 0 to node 1, and the requested bandwidth is 2 Mbps.

[0100] It is expressed as:

[0101] L2 = {(0, 4, 5), (1, 2, 10), (2, 3, 8), (3, 4, 3), (0, 1, 2)}

[0102] In the above embodiments, the adjacency matrix represents the network topology as a two-dimensional matrix, where the rows and columns respectively represent the nodes in the network, and the elements in the matrix represent the link information between the nodes. This structure is intuitive and easy to store, and is particularly suitable for being represented by a two-dimensional array in computer processing. The elements of the adjacency matrix directly represent the connection situation between nodes and the link bandwidth size. Querying the link information between any two nodes only requires one access through matrix indexing, with high efficiency. The adjacency matrix can directly support matrix operations. For example, the number of paths or link weights can be calculated through matrix multiplication, which is convenient for data analysis or network optimization tasks. The adjacency matrix form is compatible with many algorithms (such as graph traversal, shortest path calculation), and can be conveniently docked with machine learning models. For example, it can be input into a deep learning model as a structural feature.

[0103] In the above embodiments, the traffic matrix is represented by triples. The triple form can dynamically express the specific information of each traffic request without the need to fix the matrix size. For example, in some networks, there may be multiple traffic requests between the same node pair. The triples can flexibly express these multiple traffic flows, while traditional matrices may require additional processing. Since the triples clearly mark the starting point, ending point, and demand size of the traffic, they are more convenient to parse. Especially in a dynamic network, using triples can more easily handle changes between nodes or the addition and deletion of traffic requests. In a large-scale network, the traffic matrix may be very sparse (there is no traffic between most node pairs). If it is represented in the form of a complete matrix, it will cause an increase in computational complexity and storage overhead. The triple form only records the actual traffic requests, which can significantly reduce the consumption of computing resources. In an actual network, traffic requests usually exist in a specific demand form (such as requesting a bandwidth of 10 Gbps from node A to node B). The triples directly correspond to these requests, facilitating docking with the actual scenario and improving the applicability of the model.

[0104] Therefore, combining the use of an adjacency matrix to represent the network topology and triples to represent the traffic matrix can well utilize the advantages of the two representation methods: The adjacency matrix provides the overall topology of the network and supports fast query and operation. The triple traffic matrix focuses on dynamic traffic requests, saving storage and computing resources. This combination method is particularly suitable for machine learning or optimization algorithms. It can not only reflect the structural characteristics of the network but also accurately express dynamic traffic information, having high practical engineering value.

[0105] Furthermore, taking the network topology adjacency matrix and the traffic matrix as the training environment, setting a proprietary reward function, action space, and state space, an intelligent agent capable of autonomous decision-making is trained.

[0106] In some embodiments, the reward function is calculated through the following steps: Obtain the first maximum link utilization rate and the first average link utilization rate in the initial state of the preset state space; Obtain the second maximum link utilization rate and the second average link utilization rate in the current state of the preset state space; Determine the reward coefficient based on the corresponding action in the preset action space, the first maximum link utilization rate, the first average link utilization rate, the second maximum link utilization rate, and the second average link utilization rate; Calculate the reward function according to the state of the reward coefficient.

[0107] In some embodiments, the reward coefficient is calculated by the following formula:

[0108]

[0109] where α represents the reward coefficient, U max (b1) represents the first maximum link utilization rate, Uavg (b1) represents the first average link utilization rate, U max (b t ) represents the second maximum link utilization rate, U avg (b t ) represents the second average link utilization rate.

[0110] In some embodiments, the reward function is calculated by the following formula:

[0111]

[0112] where r t represents the reward function, and α represents the reward coefficient.

[0113] In this embodiment, in deep reinforcement learning (DRL), the training environment is the core for the agent to learn decision-making capabilities. The design of the environment directly affects the performance of the agent, and the goal of training is to enable the agent to learn how to optimize network traffic scheduling through reinforcement learning algorithms.

[0114] The training environment is represented by env = (M, L), where it includes: M: The adjacency matrix of the network topology, representing the node connection situation and link bandwidth in the current network. L: The set of traffic matrices, representing the set of traffic requests in the network, which is input into the environment for training.

[0115] The environment is responsible for updating the network state (such as link utilization rate) according to the actions of the agent and feeding back the reward value to guide the agent to optimize decisions.

[0116] The SR path is determined by only one intermediate SR node. That is, each forwarding path is divided into two segments, namely the source node to the SR node and the SR node to the destination node, each following the OSPF path, and the complete forwarding path is represented as R st = R sk + R kt , indicating that the source node s is forwarded to the destination node t through the intermediate node k.

[0117] Specifically, the segment routing path is simplified to only include one intermediate SR node in this application. That is, each forwarding path can be divided into two segments:

[0118] The first segment: The path R sk from the source node s to the intermediate SR node k, following the shortest path determined by the OSPF (Open Shortest Path First) protocol.

[0119] The second segment: The path R kt from the intermediate SR node k to the destination node t, also following the OSPF shortest path.

[0120] The complete forwarding path is represented as Rst = R sk + R kt 。

[0121] Note: If the SR node k is the same as the source node s, then ignore the intermediate node k and directly follow the OSPF path from s to t.

[0122] If the SR node k is the same as the destination node t, then directly follow the OSPF path from s to t.

[0123] The DRL training environment is represented as env = (M, L), which includes the network topology adjacency matrix and the set of traffic matrices. The state space is represented as obs, and len(obs) = k. The action space is Action, and action ij represents the traffic allocation ratio of the i-th traffic request in the current traffic matrix on the path with the j-th node as the SR node. If the SR node is the source node, it is ignored. If the SR node is the destination node, it is forwarded along the Open Shortest Path First (OSPF) path.

[0124] The state space obs is a set of variables used to describe the current state of the network. It provides the information required for the agent to make decisions, including:

[0125] Link utilization: used to describe the current bandwidth occupancy of each link.

[0126] Traffic matrix information: the source node, destination node, and required bandwidth of each traffic request in the current traffic matrix.

[0127] Network topology: the adjacency matrix M provides the structural information of the network.

[0128] Assume that the length of the state space is len(obs) = k, which represents the number of links in the current topology, and each element in it represents the remaining bandwidth of a link.

[0129] The action space action defines all possible actions that the agent can take. In this design, each action corresponds to a path allocation method for a traffic request:

[0130] action ij represents that the i-th traffic request in the current traffic matrix selects the j-th node as the SR node and allocates a certain proportion of traffic to this path.

[0131] If the SR node is the source node, it is ignored.

[0132] If the SR node is the destination node, it is directly forwarded along the OSPF shortest path.

[0133] Through the action space, the agent can dynamically adjust the forwarding path and allocation ratio of traffic to optimize network performance.

[0134] The reward function encourages the agent to find a scheduling strategy that can balance the load and avoid congestion by weighing the maximum link utilization and the average link utilization.

[0135] The agent learns through interaction with the environment, and the specific steps are as follows:

[0136] Initialize the environment: Input the adjacency matrix M and the set of traffic matrices L.

[0137] Observe the state: Obtain the current network state from the state space obs.

[0138] Select an action: The agent selects the optimal action from the action space Action (such as selecting SR nodes and allocation ratios).

[0139] Update the environment: Update the link utilization and network state according to the action.

[0140] Calculate the reward: Calculate the reward value of the current action according to the reward function.

[0141] Optimize the policy: Adjust the policy according to the reward value, and optimize the decision-making ability of the agent through deep reinforcement learning algorithms (such as PPO, A3C).

[0142] Repeat training: Continuously repeat the above steps until the reward value converges.

[0143] In a feasible embodiment, still taking a small network topology with 5 nodes as an example:

[0144] Input the environment with the adjacency matrix M as described above and the traffic matrix L1 as above.

[0145] Then perform SR path decision: Traffic request (0, 4, 2): The agent selects node 2 as the SR node, and the path is R st = R 02 + R 24 .

[0146] Traffic request (1, 2, 8): The agent selects node 3 as the SR node, and the path is R st = R 13 + R 32 .

[0147] Assume the initial maximum link utilization U max (b1) = 0.6, and the average link utilization U avg (b1) = 0.4. After performing the action, the maximum link utilization U max (b t) = 0.5, average link utilization U avg (b t ) = 0.35.

[0148] Calculate Reward r t = e 1.37-1 ≈ 1.45.

[0149] Adjust the path selection according to the reward value to make future decisions more inclined to reduce the link utilization.

[0150] Through the above design and embodiments, the agent gradually learns how to effectively allocate traffic, optimize network resource utilization, and can be used for traffic scheduling in a real network environment after the training is completed. By introducing the SR path and combining with the self-learning ability of the DRL agent, network resource allocation can be effectively optimized, network congestion can be reduced, link utilization can be improved, and dynamic traffic demands can be met. This method can achieve an efficient and autonomous traffic scheduling strategy in a complex network environment. It can reduce the complexity of path calculation while maintaining the consistency and efficiency of the entire traffic scheduling model.

[0151] For step S102, deploy the actual network topology according to the training environment.

[0152] Specifically, deploy multiple network routers, connect them according to the topology structure of the training phase, and run the IGP routing protocol.

[0153] The routing node configures routing table entries for multiple SR paths, differentiates them with different subnet masks, and forwards a traffic request according to different forwarding paths to achieve traffic allocation.

[0154] In a real network environment, the network topology and agent decisions designed in the training phase need to be mapped to actual network hardware devices. The deployment process includes two parts: physical connection and logical configuration. The goal of this step is to ensure that the real network can operate correctly according to the design of the training phase.

[0155] According to the topology structure of the training phase, deploy multiple physical or virtual network routers and connect them according to the topology structure defined by the adjacency matrix M. The connection of each link should meet two conditions:

[0156] Link bandwidth limit: Ensure that the bandwidth configuration of the actual link matches the bandwidth in the training topology. For example, if the link bandwidth between two nodes in the training environment is 10 Gbps, the actual deployed link needs to be configured as 10 Gbps.

[0157] Two-way communication ability: Each pair of connected nodes in the network needs to support full-duplex communication, that is, the link bandwidth is symmetric in both directions.

[0158] After that, run the Interior Gateway Protocol (such as OSPF) on all routers to generate the shortest paths from each node to other nodes in the network. The OSPF protocol can calculate the shortest path from the source node to the destination node based on the weights of the links and dynamically update the routing table.

[0159] The purpose of this step is to achieve basic connectivity between nodes and ensure that the SR paths designed in the training phase can be forwarded based on the sub-paths determined by OSPF.

[0160] For steps S103 and S104, in response to the service allocation requirement, input the traffic matrix corresponding to the service into the trained agent to obtain the target action. Determine the traffic allocation ratio for each segment routing path based on the target action, and perform traffic allocation processing on the service according to the traffic allocation ratio.

[0161] In some embodiments, the determining the traffic allocation ratio for each segment routing path based on the target action and performing traffic allocation processing on the service according to the traffic allocation ratio includes: determining the traffic allocation ratio corresponding to the segment routing node and the transmission path based on the target action; configuring a subnet mask for each transmission path; and performing traffic allocation processing on the service based on the traffic allocation ratio and the transmission path.

[0162] In this embodiment, to support the path allocation requirement in SR (Segment Routing), the agent needs to be able to forward traffic through different paths. Therefore, it is necessary to configure routing table entries for multiple SR paths and distinguish different paths through subnet masks.

[0163] In a feasible embodiment, assume that multiple SR paths need to be configured for the following traffic requests:

[0164] Traffic request (0, 4, 2), that is, from node 0 to node 4, requesting a bandwidth of 2 Gbps.

[0165] Agent decision: Allocate 60% of the traffic to path R 02 +R 24 and allocate 40% of the traffic to path R 03 +R 34 .

[0166] After that, configure subnet 10.0.0.0 / 24 for path R 02 +R 24 , and configure subnet 10.1.0.0 / 24 for path R 03 +R 34 .

[0167] After that, the path control and traffic allocation rules are sent to each router node in the network, and the traffic is forwarded along the planned path through the actual routing table entries and policy deployment, so as to realize the final traffic allocation process.

[0168] As can be seen from the above embodiments, in the traffic allocation method described in the embodiments of the present application, in the constructed training environment, an agent based on a deep learning algorithm is trained according to a preset reward function; the actual network topology is deployed according to the training environment; in response to the service allocation requirement, the traffic matrix corresponding to the service is input into the trained agent to obtain a target action; the target action includes the traffic allocation ratio for each segment routing path; the traffic allocation ratio for each segment routing path is determined based on the target action, and the traffic of the service is allocated according to the traffic allocation ratio. In the embodiments of the present application, an agent capable of making quick autonomous decisions is trained through deep reinforcement learning (DRL), and reasonable path selection and traffic allocation are performed for each traffic request, thereby reducing the global network link load. It should be noted that the difference between the present application and the prior art is that the segment routing (SR) decision is directly made through reinforcement learning without determining the SR through linear programming. Therefore, compared with the prior art, the present application greatly reduces the decision-making time and effectively reduces the cost of path planning.

[0169] It should be noted that the method in the embodiments of the present application can be executed by a single device, such as a computer or a server. The method in this embodiment can also be applied to a distributed scenario and completed by the cooperation of multiple devices. In this case of a distributed scenario, one of the multiple devices can only execute one or more steps in the method in the embodiments of the present application, and these multiple devices will interact with each other to complete the described method.

[0170] It should be noted that some embodiments of the present application have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the above embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain implementations, multitasking and parallel processing are also possible or may be advantageous.

[0171] Based on the same inventive concept, corresponding to the method in any of the above embodiments, the present application also provides a traffic allocation device.

[0172] Refer to Figure 2, the traffic allocation device includes:

[0173] A training module 21, configured to train an agent based on a deep learning algorithm in a constructed training environment according to a preset reward function;

[0174] A deployment module 22, configured to deploy an actual network topology according to the training environment;

[0175] An input module 23, configured to input a traffic matrix corresponding to the service into the trained agent in response to a service allocation requirement, and obtain a target action; the target action includes a traffic allocation ratio for each segment routing path;

[0176] An allocation module 24, configured to determine a traffic allocation ratio for each segment routing path based on the target action, and perform traffic allocation processing on the service according to the traffic allocation ratio.

[0177] For the convenience of description, the above device is described by function as various modules respectively. Of course, when implementing the present application, the functions of each module can be implemented in the same or multiple software and / or hardware.

[0178] The device in the above embodiment is used to implement the corresponding traffic allocation method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0179] Based on the same inventive concept, corresponding to the method in any of the above embodiments, the present application further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the traffic allocation method described in any of the above embodiments.

[0180] Figure 3 FIG. shows a more specific schematic diagram of the hardware structure of the electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. Among them, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other inside the device through the bus 1050.

[0181] The processor 1010 can be implemented in the form of a general-purpose CPU (Central Processing Unit), a microprocessor, an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0182] The memory 1020 can be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.

[0183] The input / output interface 1030 is used to connect to the input / output module to achieve information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Among them, the input device can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device can include a display, a speaker, a vibrator, an indicator light, etc.

[0184] The communication interface 1040 is used to connect to a communication module (not shown in the figure) to achieve communication interaction between this device and other devices. Among them, the communication module can achieve communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as a mobile network, WIFI, Bluetooth, etc.).

[0185] The bus 1050 includes a path for transmitting information between various components of the device (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).

[0186] It should be noted that although only the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050 are shown in the above device, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solutions of the embodiments of this specification, and do not necessarily include all the components shown in the figure.

[0187] The electronic device of the above embodiment is used to implement the corresponding traffic allocation method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0188] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present application also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the traffic allocation method as described in any of the above embodiments.

[0189] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.

[0190] The computer instructions stored in the storage medium of the above embodiment are used to cause the computer to execute the traffic allocation method as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0191] Based on the same inventive concept, corresponding to the traffic allocation method described in any of the above embodiments, the present disclosure also provides a computer program product including computer program instructions. In some embodiments, the computer program instructions can be executed by one or more processors of the computer to cause the computer and / or the processor to execute the traffic allocation method. Corresponding to the execution subject of each step in each embodiment of the traffic allocation method, the processor executing the corresponding step can belong to the corresponding execution subject.

[0192] The computer program product of the above embodiment is used to cause the computer and / or the processor to execute the traffic allocation method as described in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0193] Those of ordinary skill in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of the present application (including the claims) is limited to these examples; within the concept of the present application, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present application as described above, and they are not provided in detail for the sake of brevity.

[0194] In addition, for the sake of simplicity of explanation and discussion, and in order not to make the embodiments of the present application difficult to understand, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the accompanying drawings. Further, the apparatus may be shown in block diagram form in order to avoid making the embodiments of the present application difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram apparatuses are highly dependent on the platform on which the embodiments of the present application are to be implemented (i.e., these details should be fully within the understanding of those skilled in the art). In cases where specific details (such as circuits) are set forth to describe exemplary embodiments of the present application, it will be apparent to those skilled in the art that the embodiments of the present application may be implemented without these specific details or with variations of these specific details. Accordingly, these descriptions should be regarded as illustrative rather than restrictive.

[0195] Although the present application has been described in connection with specific embodiments of the present application, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art based on the foregoing description. For example, other memory architectures (such as dynamic RAM (DRAM)) may be used with the embodiments discussed.

[0196] The embodiments of the present application are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present application shall be included within the protection scope of the present application.

Claims

1. A flow distribution method, characterized in that, Including: In a constructed training environment, an agent based on a deep learning algorithm is trained according to a preset reward function; Deploy an actual network topology according to the training environment; In response to a service allocation requirement, input a traffic matrix corresponding to the service into the trained agent to obtain a target action; the target action includes the traffic allocation ratio for each segment routing path; Determine the traffic allocation ratio for each segment routing path based on the target action, and perform traffic allocation processing on the service according to the traffic allocation ratio.

2. The method according to claim 1, wherein The training environment includes a first network topology and a first traffic matrix set; The training environment is constructed through the following steps: Use an adjacency matrix to represent the topological structure of the first network topology; Use a triple to represent the first traffic matrix set; the triple includes the source node, destination node, and requested bandwidth size of the traffic request.

3. The method according to claim 1, wherein The reward function is calculated through the following steps: Obtain the first maximum link utilization rate and the first average link utilization rate in the initial state of the preset state space; Obtain the second maximum link utilization rate and the second average link utilization rate in the current state of the preset state space; Determine a reward coefficient based on the corresponding action in the preset action space, the first maximum link utilization rate, the first average link utilization rate, the second maximum link utilization rate, and the second average link utilization rate; Calculate the reward function according to the state of the reward coefficient.

4. The method according to claim 3, wherein The reward coefficient is calculated by the following formula: Among them, α represents the reward coefficient, U max (b1) represents the first maximum link utilization rate, U avg (b1) represents the first average link utilization rate, U max (b t ) represents the second maximum link utilization rate, U avg (b t ) represents the second average link utilization rate.

5. The method according to claim 3 or 4, characterized in that, The reward function is calculated by the following formula: Among them, r t represents the reward function, and α represents the reward coefficient.

6. The method according to claim 1, characterized in that, The determining the traffic allocation ratio for each segment routing path based on the target action and performing traffic allocation processing on the service according to the traffic allocation ratio includes: Determine the traffic allocation ratio corresponding to the segment routing node and the transmission path based on the target action; Configure a subnet mask for each transmission path; Perform traffic allocation processing on the service based on the traffic allocation ratio and the transmission path.

7. A flow distribution device, characterized in that, Including: A training module configured to train an agent based on a deep learning algorithm in a constructed training environment according to a preset reward function; A deployment module configured to deploy an actual network topology according to the training environment; An input module configured to input a traffic matrix corresponding to the service into the trained agent in response to a service allocation requirement to obtain a target action; The target action includes the traffic allocation ratio for each segment routing path; An allocation module configured to determine the traffic allocation ratio for each segment routing path based on the target action and perform traffic allocation processing on the service according to the traffic allocation ratio.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method according to any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 6.

10. A computer program product including computer program instructions that, when running on a computer, cause the computer to execute the method according to any one of claims 1 to 6.

Citation Information

Cited By

  • Burst traffic distribution method based on deep learning and cooperative game

    CN121907771A

  • A burst traffic allocation method based on deep learning and cooperative game

    CN121907771B