Flow forwarding method and communication system based on fat tree data center network topology structure

By adopting large-stream fixed paths and small-stream multi-path forwarding methods in the fat-tree data center network, the problem of link congestion imbalance is solved, the low-latency transmission of small-streams is ensured, and efficient load balancing and real-time performance of the data center network is achieved.

CN120528871AActive Publication Date: 2025-08-22SHANGHAI XINLIJI SEMICON CO LTD

Patent Information

Application Number
CN202511029503.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-08-22
Estimated Expiration
2045-07-25

AI Technical Summary

Technical Problem

In the fat tree data center network, the traditional load balancing strategy based on five-tuple hash results in link congestion and imbalance, affecting network bandwidth utilization. Especially in high concurrency and small stream dense business scenarios, the transmission delay and packet loss rate of small streams are high. The existing random packet spraying technology affects the real-time nature of small streams when large streams are concurrent.

Method used

The traffic forwarding method based on the network topology of Fat Tree Data Center is adopted. Through the fixed path of large flow and the multi-path spray forwarding of small flows, we ensure that the large flow is forwarded along the fixed path, and the small flow is equally forwarded in the multipath to avoid interference from the large flow to the small flow, and extreme loads are handled through the waiting queue and standby path mechanism.

Benefits of technology

It improves the transmission efficiency and timeliness of small streams, ensures the real-time and agility of data center services, avoids additional computing power overhead, and improves the stability and reliability of the network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120528871A_ABST
    Figure CN120528871A_ABST
Patent Text Reader

Abstract

The invention discloses a flow forwarding method and a communication system based on a fat tree data center network topology structure, and the method comprises the main steps: a source access layer switch determines a target terminal device and an access layer switch set corresponding to the target flow when receiving the target flow sent by a source terminal device; judging whether the source access layer switch and the target access layer switch are the same access layer switch or not; if not, determining the size of the target flow, and judging whether the target flow is a large flow; if the target flow is the large flow, searching a target large flow fixed path corresponding to the access layer switching unit in a pre-configured large flow fixed path static mapping table, and configuring the target large flow fixed path as a special forwarding path of the target flow; otherwise, the data packet of the target flow is sprayed to the equivalent shortest path which is not occupied by the large flow in all the equivalent shortest paths, the large flow forwarding of the data center is reasonably controlled, and the real-time performance of small flow forwarding is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of data center networks, and in particular relates to a traffic forwarding method and a communication system based on a fat-tree data center network topology structure. Background Art

[0002] With the rapid development of cloud computing, big data, and artificial intelligence, east-west traffic within data centers has increased significantly, network structures have become flatter, and multi-level multi-path topologies, such as the fat-tree data center network topology, have been widely adopted.

[0003] To fully utilize the numerous equal-cost multi-path (ECMP) paths in fat-tree data center network topologies, the traditional approach is to use a five-tuple hash-based load balancing strategy, assigning each data flow to a fixed path. However, hash collisions or the concentration of hot traffic can cause some links to become overloaded while others are idle, leading to the so-called "link congestion imbalance" problem, which significantly limits network bandwidth utilization.

[0004] To address these issues, various complex algorithms can be used to select paths for large flows in real time and balance links. However, the more large flows there are, the greater the computing overhead and the higher the requirements for system equipment. Furthermore, large flow forwarding paths are uncertain. If large flows occur continuously, multiple equal-cost paths may still be occupied by them. This makes it unsuitable for high-concurrency, small-flow-intensive business scenarios in data centers.

[0005] Random Packet Spraying (RPS) is a packet-based forwarding strategy that independently and evenly distributes each packet within a flow onto all available equal-cost shortest paths, avoiding the hotspotting problem associated with traditional ECMP fixed-path selection. RPS enables finer-grained load balancing, thereby improving overall network throughput, reducing link utilization skew, and significantly alleviating head-of-line blocking. It is particularly well-suited for high-concurrency, small-flow-intensive business scenarios in data centers. However, RPS still has significant limitations in practical data center networks. Data center network communications exhibit a typical traffic distribution characterized by a "majority of small flows, and bandwidth-intensive large flows": the vast majority of traffic consists of short-lived, small flows (also known as "mice flows"), while only a few large flows (also known as "elephant flows") consume the majority of network bandwidth. Small flows are typically generated by applications such as control signals, small file transfers, and web page requests. Mice flows are often very latency-sensitive and require fast transmission to ensure real-time and responsive applications. Large flows are typically generated by applications such as file transfers, video streaming, and large-scale data backups. RPS technology uses a mechanism that sprays and forwards all data packets. This does help improve load balancing in scenarios with dense short flows. However, when large flows are concurrent, their packets are evenly distributed across all equal-cost shortest paths in the network. The bandwidth on these paths is occupied for extended periods, potentially leading to higher queuing delays and packet loss rates for small flows. This undermines the inherent low-latency advantage of small flows and impacts the real-time and agility of data center services. Summary of the Invention

[0006] The purpose of the present invention is to provide a traffic forwarding method and communication system based on a fat-tree data center network topology structure, in which large flow forwarding is reasonably controlled and small flow forwarding real-time performance is guaranteed.

[0007] In order to achieve the above object, the technical solution adopted by the present invention is as follows: A first aspect of the present invention provides a traffic forwarding method based on a fat-tree data center network topology structure, comprising the following steps: When a source access layer switch in a fat tree data center network topology receives target traffic sent by a source terminal device, it confirms the target terminal device and the access layer switch group corresponding to the target traffic, the access layer switch group includes the source access layer switch and the target access layer switch, the target access layer switch is the access layer switch directly connected to the target terminal device, and determines whether the source access layer switch and the target access layer switch are the same access layer switch; If they are not the same, confirming the size of the target flow and determining whether the target flow is a large flow; If it is a large flow, then search for a target large flow fixed path corresponding to the access layer switch group in a pre-configured large flow fixed path static mapping table, wherein the large flow fixed path static mapping table includes access layer switch groups and large flow fixed paths with corresponding relationships, and the large flow fixed path is one of multiple equal-cost shortest paths from the source access layer switch in the access layer switch group to the target access layer switch, and configure the target large flow fixed path as a dedicated forwarding path for the target traffic; Otherwise, the data packets of the target traffic are sprayed onto the equal-cost shortest paths that are not occupied by large flows among all equal-cost shortest paths from the source access layer switch to the target access layer switch.

[0008] The traffic forwarding method of the present invention separates large and small flows in a fat-tree data center network topology according to different forwarding strategies, based on the traffic characteristics of the network topology in actual application scenarios. By managing large and small flows separately, the forwarding path of large flows is fixed, ensuring that the bandwidth usage of large flows does not arbitrarily interfere with the forwarding of small flows. Small flows are effectively load-balanced across multiple paths, improving the timeliness of their transmission. Furthermore, static mapping of large flows to fixed paths avoids additional computing power overhead, does not significantly increase system burden, and is user-friendly to system equipment.

[0009] Furthermore, in order to avoid collision or accumulation of large flows between different large flow fixed paths as much as possible, and to minimize the impact of large flows in large flow fixed paths on the transmission efficiency of small flows in other small flow transmission paths with overlapping paths, the large flow fixed path static mapping table is pre-configured according to the following rules: Each access layer switch group has only one corresponding large flow fixed path. For any two fixed paths of large flows, there is preferably no overlapping path between them. If there must be an overlapping path between the two fixed paths of large flows, the overlapping path is preferably the shortest; and / or, Determine the large flow path that has an overlapping path with other large flow fixed paths as the target counting path, and the total number of the target counting paths is the smallest; and / or, The total length of overlapping paths in the fixed path of the control flow is minimized.

[0010] Furthermore, the target large flow fixed path is configured to exit the current dedicated forwarding path of the target flow after the target access layer switch receives all the data packets of the target flow. The exited path continues to be used to forward the data packets of the small flow, thereby ensuring the timeliness of the small flow forwarding.

[0011] Furthermore, the data packets of the target traffic are sprayed according to the following method: in the preconfigured dynamic path table, an equivalent shortest path group corresponding to the access layer switch group and not occupied by large flows is selected as the spraying range of the data packets of the target traffic, and the data packets of the target traffic are randomly or polledly distributed within the spraying range through the RPS technology.

[0012] By using a hash algorithm to randomly distribute data packets to multiple equal-cost shortest paths for load balancing, it can not only fully utilize multi-path network resources, but also avoid overloading a single path and ensure that small flows are forwarded quickly.

[0013] In order to avoid collisions between small flow packets and large flow packets, in some specific implementations, the dynamic path table is preconfigured in the following manner: The dynamic path table includes multiple sub-tables and a temporary suspension table. The sub-tables have corresponding access layer switch groups and equal-cost shortest path groups. The equal-cost shortest path group is a set of equal-cost shortest paths from the source access layer switch to the target access layer switch in the access layer switch group. The temporary suspension table includes equal-cost shortest paths occupied by large flows. The dynamic path table is configured as follows: The dynamic path table is associated with the large flow fixed path static mapping table. When the target large flow fixed path is found in the large flow fixed path static mapping table, the corresponding sub-table is found according to the access layer switch group, and the equivalent shortest path corresponding to the target large flow fixed path is moved from the equivalent shortest path group of the sub-table where the target large flow fixed path is located to the temporary suspension table. After there is no large flow on the equivalent shortest path corresponding to the target large flow fixed path, the target large flow fixed path is moved from the temporary suspension table to the equivalent shortest path group where the target large flow fixed path is located. An equivalent shortest path that overlaps with the target large flow fixed path is searched in the equivalent shortest path groups of other sub-tables and moved to the temporary suspension table. When the equivalent shortest path corresponding to the target large flow fixed path is removed from the temporary suspension table, the equivalent shortest path is moved from the temporary suspension table to the equivalent shortest path group where it is located.

[0014] Furthermore, after finding the target large flow fixed path, in response to confirming that the uplink port of the source access layer switch corresponding to the target large flow fixed path is occupied, the target traffic is added to a preset waiting queue according to a preset rule, and forwarded through the target large flow fixed path in sequence according to the sorting of the waiting queue. The waiting queue is configured as follows: when the large flow in the waiting queue is forwarded by its corresponding source access layer switch, the large flow is deleted from the waiting queue.

[0015] By placing large flows to be forwarded into a waiting queue, we ensure that even under extreme load conditions, large flows can still be processed in priority order, avoiding packet loss due to resource constraints and further improving the stability and reliability of the system.

[0016] Furthermore, the preset rule is: adding the current target traffic to the end of the queue.

[0017] In some implementations, in order to reduce the probability of large flow packets or large flows being received out of order by the target terminal, the waiting queues are sorted according to the following rules: Sort by the order in which each large flow is sent from the source terminal device. If there is target traffic sent from different source terminal devices at the same time, it is determined whether their traffic sizes are the same. If they are different, they are sorted from small to large according to the traffic size. If they are the same, they are randomly sorted.

[0018] Furthermore, considering the possibility of more extreme continuous high-volume traffic, a new dedicated path for forwarding large traffic is temporarily configured as follows: Starting from the time when the first large flow to be forwarded appears in the waiting queue, counting the total number and / or total byte count of the large flows to be forwarded in the waiting queue begins, and updating the total number and / or total byte count of the traffic to be forwarded in the waiting queue when the large flow at the head of the queue is removed and / or a new large flow is added to the tail of the queue; and / or, When each large flow joins the waiting queue, start timing to obtain the waiting time. When the total number of large flows to be forwarded in the waiting queue exceeds a preset number threshold and / or the total number of bytes exceeds a preset traffic threshold, and / or the waiting time of the large flow at the head of the waiting queue is greater than a preset time threshold, one or more equivalent-cost shortest paths that are different from the fixed path of the target large flow are selected from the equivalent-cost shortest paths as large flow backup paths, and the large flow backup paths are temporarily configured as dedicated paths for forwarding large flows.

[0019] In some specific implementations, the backup high-volume flow path is temporarily configured according to the following rules: The number of the backup high-flow paths does not exceed 1 / 4 of the total number of the equivalent-cost paths; and / or, The backup large flow path and the target large flow path are controlled to have no overlapping paths.

[0020] In some specific implementations, in response to any of the following preset conditions being met, the backup path for the large flow exits the dedicated path for forwarding the large flow: Condition 1: queuing is detected on any other equivalent shortest path different from the fixed path of the target large flow, or the queuing delay exceeds a preset threshold; Condition 2: There is no large flow in the waiting queue. The priority of the first condition is higher than that of the second condition.

[0021] The above exit condition setting can avoid sacrificing the timeliness of small flows.

[0022] Furthermore, data packets of the same flow carry the same identifier, data packets of different flows carry different identifiers, and different data packets in the same flow have different sequence numbers. In response to receiving and identifying the identifiers and sequence numbers of data packets arriving through different equivalent shortest paths, the target terminal device reassembles the data packets in the correct order according to the identifiers and the sequence numbers to obtain the target flow.

[0023] Furthermore, whether the target flow is a large flow is determined by the following method: Starting byte counting when an access layer switch directly connected to the source terminal device starts receiving the target traffic; If the accumulated bytes have reached a preset flow threshold before the target flow is completely received, it is determined to be a large flow.

[0024] Furthermore, if the source access layer switch and the target access layer switch are the same access layer switch, the packets are forwarded directly.

[0025] A second aspect of the present invention provides a communication system, which includes: terminal devices and a fat-tree data center network topology structure for connecting the terminal devices in pairs with electrical signals, a traffic judgment module and a path selection module, wherein the traffic judgment module and the path selection module are configured to forward traffic according to the above-mentioned traffic forwarding method.

[0026] The beneficial effects brought about by the technical solution provided by the present invention are as follows: The traffic forwarding method of the present invention separates large flows and small flows for differentiated processing. Large flows are forwarded using a fixed path, while small flows are forwarded using data packets distributed over multiple paths. This differentiated traffic management method avoids excessive occupation of network bandwidth by large flows, and ensures that the low-latency characteristics of short flows are not affected by large flows. The present invention ensures that large flows are always forwarded along fixed paths through a static mapping table for large flow paths, avoids bandwidth competition and path congestion caused by the randomness of path selection, ensures the transmission stability of large flows, and reduces the additional overhead caused by dynamic path selection through complex algorithms. According to the traffic forwarding method of the present invention, while effectively avoiding interference of large flows on small flows, the transmission efficiency and timeliness of small flows are improved, which is conducive to improving the real-time and agility of data center services. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0028] Figure 1 A flow chart of a traffic forwarding method provided in an embodiment; Figure 2 A schematic diagram of a standard fat-tree data center network topology structure with k=4 listed in the embodiment; Figure 3 A flowchart of the waiting queue mechanism and dynamic multi-path concurrency mechanism provided in the embodiment. DETAILED DESCRIPTION

[0029] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0030] It should be noted that the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, apparatus, product or equipment that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or are inherent to these processes, methods, products or equipment.

[0031] It should be noted that the exemplary "fat tree data center network topology" in the present invention is a standard fat tree data center network topology: The FatTree data center network topology consists of three layers of switches: core switches, aggregation switches, and access switches. The aggregation switches and access switches form multiple clusters (pods). Specifically, each switch has k ports, and the number of core switches is (k / 2)^2. There are k pods, each consisting of k switches, with k / 2 aggregation switches and k / 2 access switches. Any two pods are connected to an aggregation switch in each pod via a core switch. Each core switch connects to all pods. Each access switch can directly connect to k / 2 end devices.

[0032] Communication between different Pods: Source terminal device in Pod A → Access switch in Pod A → Aggregation switch in Pod A → Core switch → Aggregation switch in Pod B → Access switch in Pod B → Target source terminal device in Pod B, total (k / 2) 2 Equivalent shortest paths.

[0033] Communication between the same pods: source terminal device in Pod A → access layer switch A1 in Pod A → aggregation layer switch in Pod A → access layer switch A2 in Pod A → target source terminal device in Pod A. The source and target terminal devices are directly connected to different access layer switches, with a total of k / 2 equal-cost shortest paths.

[0034] Illustration 2 of communication between the same pods: source terminal device in Pod A → access layer switch A1 in Pod A → target source terminal device in Pod A. The source and target terminal devices are directly connected to the same access layer switch, with only one path.

[0035] Unless otherwise specified, the "shortest path with equal cost" in the present invention refers to the shortest path with equal cost between two different access layer switches.

[0036] In one embodiment of the present invention, a traffic forwarding method and communication system based on a fat-tree data center network topology are provided. The communication system of this embodiment includes: terminal devices, a fat-tree data center network topology for electrically connecting terminal devices in pairs, a traffic determination module, and a path selection module. The traffic determination module and the path selection module are configured to forward traffic according to the traffic forwarding method of this embodiment. In this embodiment, the traffic determination module is deployed at the access layer switch and is responsible for monitoring and classifying each communication flow passing through the switch. The path selection module includes a processor electrically connected to all switches in the fat-tree data center network topology and is responsible for controlling the specific traffic forwarding method based on the feedback from the traffic determination module.

[0037] like Figure 1 As shown, the traffic forwarding method of this embodiment includes the following main steps: When a source access layer switch in the fat tree data center network topology receives target traffic sent by a source terminal device, it confirms the target terminal device and the access layer switch group corresponding to the target traffic, where the access layer switch group includes the source access layer switch and the target access layer switch, where the target access layer switch is the access layer switch directly connected to the target terminal device, and determines whether the source access layer switch and the target access layer switch are the same access layer switch; If they are the same, forward them directly; If not, confirm the size of the target flow and determine whether the target flow is a large flow; If it is a large flow, then search for a target large flow fixed path corresponding to the access layer switch group in a pre-configured large flow fixed path static mapping table, wherein the large flow fixed path static mapping table includes access layer switch groups and large flow fixed paths with corresponding relationships, and the large flow fixed path is one of multiple equal-cost shortest paths from the source access layer switch in the access layer switch group to the target access layer switch, and configure the target large flow fixed path as a dedicated forwarding path for the target flow; If it is not a large flow, the data packets of the target flow are sprayed to the equal-cost shortest paths that are not occupied by large flows among all equal-cost shortest paths from the source access layer switch to the target access layer switch.

[0038] Specifically, this embodiment is described using a standard fat tree data center network topology structure with k=4 as an example. In other embodiments, a standard fat tree data center network topology structure with a larger k value can be used according to actual needs. The larger the k value, the more obvious the advantages of the traffic forwarding method of the present invention.

[0039] like Figure 2 As shown, the standard fat-tree data center network topology of this embodiment includes a three-layer architecture: access layer, aggregation layer, and core layer. The access layer includes eight access layer switches (labeled T1-T8), each of which is connected to two terminal devices (for example, servers S1 and S2 are connected to access layer switch T1, and servers S15 and S16 are connected to access layer switch T8). Each access layer switch is connected to two aggregation layer switches (labeled A1-A8) in its pod, and each aggregation layer switch is further connected to four core layer switches (labeled C1-C4). The connections between each layer form a fully interconnected multi-root tree structure, allowing multiple equal shortest paths between any two different access layer switches, and thus multiple equal shortest paths between any two terminal devices that are not directly connected to the same access layer switch. For example, for the source server S1 and the target server S16 located in different pods, there are four equal-cost shortest paths of the same length in the network: the outbound traffic of S1 starts from the access layer switch T1, may pass through the aggregation layer switch A1 or A2 in the source pod, and then pass through any core layer switch C1, C2, C3, C4 to reach the aggregation layer switch A7 or A8 of the target pod, and finally reach the access layer switch T8 and be delivered to S16.

[0040] The traffic forwarding method of this embodiment is further described below by taking the access layer switch T1 as the source access layer switch and the access layer switch T8 as the destination switch.

[0041] When the source access layer switch T1 receives the target traffic sent by the source terminal device (source server S1), the traffic judgment module reads the five-tuple (source IP, destination IP, source port, destination port, protocol) of the target traffic, indexes the flow, and counts the transmitted bytes of the traffic in real time. At the same time, it confirms the target terminal device (target server S16) and access layer switch group corresponding to the target traffic. The access layer switch group includes the source access layer switch T1 and the target access layer switch T8 (the access layer switch directly connected to the target terminal device), and determines whether the source access layer switch and the target access layer switch are the same access layer switch.

[0042] If they are on the same access layer switch, they are forwarded directly.

[0043] If the target traffic is not on the same access layer switch, the traffic determination module determines whether the target traffic is a large flow based on a preset large flow threshold (e.g., 500 KB in this embodiment). Specifically, if the cumulative bytes reach the preset large flow threshold before the target traffic is fully received, the flow is considered a large flow. If the cumulative bytes do not reach the preset large flow threshold after the target traffic is fully received, the flow is considered a small flow. In this embodiment, target traffic information is recorded in a table format and notified to the path selection module. The table includes the target traffic's five-tuple and traffic characteristics. Large flows are also identified, facilitating a rapid response from the path selection module.

[0044] If the target traffic is a large flow, the path selection module searches the preconfigured large flow fixed path static mapping table for the target large flow fixed path corresponding to the access layer switch group. The large flow path static mapping table includes corresponding access layer switch groups and large flow fixed paths. The large flow path is one of multiple equivalent shortest paths from the source access layer switch in the access layer switch group to the target access layer switch. The target large flow fixed path is configured as a dedicated forwarding path for the target traffic. In this embodiment, the large flow fixed path static mapping table is specifically a fixed port mapping table (Port-MapTable), which is preconfigured by the network administrator or controller during the system initialization phase and distributed to each switch. When a large flow is determined to be a large flow (the special mark of the large flow is identified), it is forwarded along the predetermined fixed path. This process is implemented through static mapping, ensuring the path stability of the large flow and effectively avoiding the overhead and uncertainty caused by dynamic path selection. In this embodiment, in order to avoid collision or accumulation of large flows between different large flow fixed paths as much as possible, and to minimize the influence of large flows in large flow fixed paths on the transmission efficiency of small flows in other small flow transmission paths with overlapping paths, a static mapping table of large flow fixed paths is pre-configured according to the following rules: (1) Each access layer switch group has only one large flow fixed path corresponding to it; (2) For any two large flow fixed paths, there is preferably no overlapping path between them. If there must be an overlapping path between the two large flow fixed paths, the shortest overlapping path is preferred; (3) The large flow path that has an overlapping path with other large flow fixed paths is determined as the target counting path, and the total number of target counting paths is minimized; (4) The total length of the overlapping paths in the large flow fixed paths is controlled to be minimized.

[0045] If the target flow is not a large flow, that is, a small flow, the data selection module will not find the special mark for large flows in the table corresponding to the target flow. The data selection module will select the equal-cost path group corresponding to the access layer switch group and not occupied by large flows from the pre-configured dynamic path table as the spray range for the target flow data packets. The module will use RPS technology to distribute the target flow data packets randomly or in a round-robin manner within the spray range. Specifically, a hash algorithm is used to randomly distribute data packets to multiple equal-cost shortest paths for load balancing. This not only fully utilizes multi-path network resources, but also avoids overloading a single path and ensures that small flows are forwarded quickly. In this embodiment, in order to avoid collision between small flow data packets and large flow as much as possible, the dynamic path table is pre-configured in the following manner: the dynamic path table includes multiple sub-tables and a temporary pause table, the sub-table has corresponding access layer switch groups and equal-cost shortest path groups, the equal-cost shortest path group is a collection of equal-cost shortest paths from the source access layer switch to the target access layer switch in the access layer switch group, the temporary pause table includes the equal-cost shortest paths occupied by the large flow, and the dynamic path table is configured as follows: (1) the dynamic path table is associated with the large flow fixed path static mapping table, when the target large flow fixed path is found in the large flow fixed path static mapping table, According to the access layer switch group, the corresponding sub-table is found, and the equivalent shortest path corresponding to the target large flow fixed path is moved from the equivalent shortest path group to the temporary suspension table. After there is no large flow on the equivalent shortest path corresponding to the target large flow fixed path, it is moved from the temporary suspension table to the equivalent shortest path group. (2) In the equivalent shortest path group of other sub-tables, an equivalent shortest path that overlaps with the target large flow fixed path is found and moved to the temporary suspension table. When the equivalent path corresponding to the target large flow fixed path is removed from the temporary suspension table, the equivalent shortest path is moved from the temporary suspension table to the equivalent shortest path group. (2) It may not be necessary.

[0046] In this embodiment, in order to avoid some extreme situations, the following steps are further added: Adding a waiting queue: After finding a fixed path for a target large flow, in response to confirming that the uplink port corresponding to the fixed path for the target large flow on the source access layer switch is occupied, the target flow is added to a preset waiting queue according to a preset rule and forwarded sequentially according to the queue's order. The waiting queue is configured such that once a large flow in the waiting queue is forwarded by its corresponding source access layer switch, it is removed from the waiting queue. By placing the large flow to be forwarded in the waiting queue, even under extreme load conditions, it is prioritized and processed in order, avoiding packet loss due to resource constraints and further improving system stability and reliability. In this embodiment, the preset rule is to add the current target flow to the end of the queue. To reduce the probability of large flow packets or large flows being received out of order by the target terminal, the waiting queue is sorted according to the following rule: each large flow is sorted in the order in which it was sent from the source terminal device. If there are target flows sent from different source terminal devices at the same time, their flow sizes are determined to be the same. If they are different, they are sorted in ascending order of flow size. If they are the same, they are randomly sorted. By adding a waiting queue mechanism, we ensure that large flows can still be processed first under extreme load conditions, avoiding packet loss or network congestion due to resource constraints, thereby improving the stability and reliability of the system.

[0047] Furthermore, a dynamic multi-path concurrency mechanism is added: When the first large flow to be forwarded appears in the waiting queue, the total number and / or total byte count of the large flows to be forwarded in the waiting queue is counted, and when the leading large flow is removed and / or a new large flow is added to the tail, the total number and / or total byte count of the traffic to be forwarded in the waiting queue is updated; and / or, a timer is started when each large flow joins the waiting queue to determine the waiting time. When the total number of large flows to be forwarded in the waiting queue exceeds a preset number threshold and / or the total byte count exceeds a preset traffic threshold, and / or the waiting time of the leading large flow in the waiting queue exceeds a preset time threshold, one or more equal-cost shortest paths different from the fixed path of the target large flow are selected from the equal-cost shortest paths as backup paths for the large flow, and the backup paths are temporarily configured as dedicated paths for forwarding the large flow. Because the k value of the fat-tree topology in this embodiment is small, there are two equal-cost shortest paths between different access layer switches in the same pod. To avoid sacrificing the timeliness of small flows, the strategy of temporarily configuring backup paths for large flows is not enabled. There are four equal-cost shortest paths between different access layer switches in different pods. You can consider enabling a strategy for temporarily configuring backup large flow paths. For example, you can choose to set the time threshold to 50ms, limit the number of backup large flow paths to one, and select an equal-cost shortest path that has no overlapping paths with the target large flow path to ensure that the new path is topologically equivalent to the original path and there is no conflict. In other embodiments, as the k value increases, the number of equal-cost shortest paths increases. Multiple backup large flow paths can be set according to the situation, preferably not more than 1 / 4 of the total number of equal-cost paths. After adding a new path for the large flow, the following two diversion strategies can be adopted: One is to perform load balancing at the packet level in the same way as RPS, and randomly or in a round-robin manner distribute the data packets of the current large flow to two or more paths. This packet-by-packet distribution method can make fine-grained use of the bandwidth of the dual paths, thereby improving concurrent throughput, but the packet disorder problem needs to be considered, and the sequence packets need to be reordered at the target terminal. The second method is to split the entire large flow currently being forwarded, dividing the large flow into two or more larger sub-flows, which are transmitted through the original path and the newly added path respectively. This method of splitting by flow or flow block ensures the order of messages within the sub-flow on each path and reduces the out-of-order overhead, but it requires dividing the traffic at the sending end and converging it at the receiving end. In practical applications, a specific solution can be selected based on the network's tolerance for out-of-order and implementation complexity. It can also be combined with the use of flow-granular dynamic flowlet technology to achieve a balance between maintaining message order and achieving load balancing. In this embodiment, when the backup large flow path is enabled, the equivalent shortest path corresponding to the backup large flow path is moved from the equivalent shortest path group of the sub-table where it is located to the temporary suspension table to avoid the small flow data packets being sprayed onto it when it is used as a backup large flow path.

[0048] In this embodiment, in order to avoid sacrificing the timeliness of small flows, in response to satisfying any of the following preset conditions, the large flow backup path exits the dedicated path for forwarding large flows: Condition 1: It is detected that queuing occurs in any other equivalent shortest path that is different from the fixed path of the target large flow, or the queuing delay exceeds a preset threshold. For example, the uplink port status of all switches on other equivalent shortest paths is monitored in real time, and it is found that data packets are queued or the data packet queuing delay exceeds the preset threshold on any uplink port; Condition 2: There is no large flow in the waiting queue, and the priority of Condition 1 is greater than that of Condition 2.

[0049] In this embodiment, data packets of the same flow carry the same identifier, data packets of different flows carry different identifiers, and different data packets in the same flow have different sequence numbers. In response to receiving and identifying the identifiers and sequence numbers of data packets arriving through different equivalent shortest paths, the target terminal device reassembles the data packets in the correct order according to the identifiers and sequence numbers to obtain the target flow.

[0050] As an illustration, the application prospects of the traffic forwarding method of the present invention include: Optimization of data center and cloud computing networks: With the widespread adoption of cloud computing and virtualization technologies, data center network traffic patterns are characterized by the coexistence of a large number of small flows and a small number of large flows. The traffic forwarding method of the present invention enables data centers to more efficiently manage network traffic, avoids excessive bandwidth usage by large flows, and ensures low-latency forwarding of small flows. Therefore, the method has broad application prospects in large-scale cloud computing environments, significantly improving the service quality of cloud platforms and ensuring the smooth operation of various applications and services.

[0051] Improving real-time performance in high-frequency trading and financial systems: In financial systems, especially in applications such as high-frequency trading (HFT), where latency requirements are extremely stringent, the timeliness of small-flow forwarding is crucial. This invention can effectively reduce small-flow congestion and latency, avoid transaction delays caused by network resource utilization, ensure efficient and timely transaction processes, and provide strong technical support for high-frequency trading in financial markets.

[0052] Traffic management in 5G and Internet of Things (IoT) networks: In 5G networks and IoT applications, a large number of devices generate different types of traffic. In IoT, small and large flows coexist significantly, placing extremely high demands on network real-time performance and stability. The traffic forwarding method of this invention enables differentiated management of large and small flows, optimizing network resource allocation, ensuring that small, real-time-sensitive flows receive priority transmission while large flows are properly controlled, thereby improving overall network efficiency and reliability.

[0053] The above is only a specific implementation method of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A traffic forwarding method based on a fat-tree data center network topology, characterized by: It includes the following steps: When a source access layer switch in a fat tree data center network topology receives target traffic sent by a source terminal device, it confirms the target terminal device and the access layer switch group corresponding to the target traffic, the access layer switch group includes the source access layer switch and the target access layer switch, the target access layer switch is the access layer switch directly connected to the target terminal device, and determines whether the source access layer switch and the target access layer switch are the same access layer switch; If they are not the same, confirming the size of the target flow and determining whether the target flow is a large flow; If it is a large flow, then search for a target large flow fixed path corresponding to the access layer switch group in a pre-configured large flow fixed path static mapping table, wherein the large flow fixed path static mapping table includes access layer switch groups and large flow fixed paths with corresponding relationships, and the large flow fixed path is one of multiple equal-cost shortest paths from the source access layer switch in the access layer switch group to the target access layer switch, and configure the target large flow fixed path as a dedicated forwarding path for the target traffic; Otherwise, the data packets of the target traffic are sprayed onto the equal-cost shortest paths that are not occupied by large flows among all equal-cost shortest paths from the source access layer switch to the target access layer switch.

2. The traffic forwarding method based on the fat-tree data center network topology structure according to claim 1 is characterized in that: Pre-configure the static mapping table for the large flow fixed path according to the following rules: Each access layer switch group has only one corresponding large flow fixed path. For any two fixed paths of large flows, there is preferably no overlapping path between them. If there must be an overlapping path between the two fixed paths of large flows, the overlapping path is preferably the shortest; and / or, Determine the large flow path that has an overlapping path with other large flow fixed paths as the target counting path, and the total number of the target counting paths is the smallest; and / or, The total length of overlapping paths in the fixed path of the control flow is minimized.

3. The traffic forwarding method based on a fat-tree data center network topology structure according to claim 1 is characterized in that: The target large flow fixed path is configured to exit the dedicated forwarding path of the target flow after the target access layer switch receives all data packets of the target flow.

4. The traffic forwarding method based on a fat-tree data center network topology structure according to claim 1, characterized in that: Spray the target traffic packets as follows: In the preconfigured dynamic path table, the equivalent shortest path group corresponding to the access layer switch group and not occupied by the large flow is selected as the spray range of the data packets of the target flow, and the data packets of the target flow are randomly or polled distributed within the spray range through the RPS technology.

5. The traffic forwarding method based on a fat-tree data center network topology structure according to claim 4 is characterized in that: Preconfigure the dynamic path table as follows: The dynamic path table includes multiple sub-tables and a temporary suspension table. The sub-tables have corresponding access layer switch groups and equal-cost shortest path groups. The equal-cost shortest path group is a set of equal-cost shortest paths from the source access layer switch to the target access layer switch in the access layer switch group. The temporary suspension table includes equal-cost shortest paths occupied by large flows. The dynamic path table is configured as follows: The dynamic path table is associated with the large flow fixed path static mapping table. When the target large flow fixed path is found in the large flow fixed path static mapping table, the corresponding sub-table is found according to the access layer switch group, and the equivalent shortest path corresponding to the target large flow fixed path is moved from the equivalent shortest path group of the sub-table where the target large flow fixed path is located to the temporary suspension table. After there is no large flow on the equivalent shortest path corresponding to the target large flow fixed path, the equivalent shortest path corresponding to the target large flow fixed path is moved from the temporary suspension table to the equivalent shortest path group where the target large flow fixed path is located. An equivalent shortest path that overlaps with the target large flow fixed path is searched in the equivalent shortest path groups of other sub-tables and moved to the temporary suspension table. When the equivalent shortest path corresponding to the target large flow fixed path is removed from the temporary suspension table, the equivalent shortest path is moved from the temporary suspension table to the equivalent shortest path group where it is located.

6. The traffic forwarding method based on a fat-tree data center network topology structure according to any one of claims 1 to 5, characterized in that: After finding the target large flow fixed path, in response to confirming that the uplink port of the source access layer switch corresponding to the target large flow fixed path is occupied, the target traffic is added to a preset waiting queue according to a preset rule, and forwarded in sequence according to the order of the waiting queue. The waiting queue is configured as follows: when the large flow in the waiting queue is forwarded by its corresponding source access layer switch, the large flow is deleted from the waiting queue.

7. The traffic forwarding method based on a fat-tree data center network topology structure according to claim 6, characterized in that: The preset rule is: add the current target flow to the end of the queue, And / or, the waiting queue is sorted according to the following rules: Sort by the order in which each large flow is sent from the source terminal device. If there is target traffic sent from different source terminal devices at the same time, it is determined whether their traffic sizes are the same. If they are different, they are sorted from small to large according to the traffic size. If they are the same, they are randomly sorted.

8. The traffic forwarding method based on a fat-tree data center network topology structure according to claim 6, characterized in that: Starting from the time when the first large flow to be forwarded appears in the waiting queue, counting the total number and / or total byte count of the large flows to be forwarded in the waiting queue begins, and updating the total number and / or total byte count of the traffic to be forwarded in the waiting queue when the large flow at the head of the queue is removed and / or a new large flow is added to the tail of the queue; and / or, When each large flow joins the waiting queue, start timing to obtain the waiting time. When the total number of large flows to be forwarded in the waiting queue exceeds a preset number threshold and / or the total number of bytes exceeds a preset traffic threshold, and / or the waiting time of the large flow at the head of the waiting queue is greater than a preset time threshold, one or more equivalent-cost shortest paths that are different from the fixed path of the target large flow are selected from the equivalent-cost shortest paths as large flow backup paths, and the large flow backup paths are temporarily configured as dedicated paths for forwarding large flows.

9. The traffic forwarding method based on a fat-tree data center network topology structure according to claim 8, characterized in that: Temporarily configure the backup high-volume flow path according to the following rules: The number of the backup high-flow paths does not exceed 1 / 4 of the total number of the equivalent-cost paths; and / or, The backup large flow path and the target large flow path are controlled to have no overlapping paths.

10. The traffic forwarding method based on a fat-tree data center network topology structure according to claim 8, characterized in that: In response to any of the following preset conditions being met, the backup path for the large flow exits the dedicated path for forwarding the large flow: Condition 1: queuing is detected on any other equivalent shortest path different from the fixed path of the target large flow, or the queuing delay exceeds a preset threshold; Condition 2: There is no large flow in the waiting queue. The priority of the first condition is higher than that of the second condition.

11. The traffic forwarding method based on a fat-tree data center network topology structure according to claim 1, characterized in that: Data packets of the same flow have the same identifier, data packets of different flows have different identifiers, and different data packets in the same flow have different sequence numbers. In response to receiving and identifying the identifiers and sequence numbers of data packets arriving through different equal-cost shortest paths, the target terminal device reassembles the data packets in the correct order according to the identifiers and the sequence numbers to obtain the target flow.

12. The traffic forwarding method based on a fat-tree data center network topology structure according to claim 1, characterized in that: Determine whether the target traffic is a large flow by the following method: Starting byte counting when an access layer switch directly connected to the source terminal device starts receiving the target traffic; If the accumulated bytes have reached a preset large flow determination threshold before the target flow is completely received, it is determined to be a large flow.

13. The traffic forwarding method based on a fat-tree data center network topology structure according to claim 1, characterized in that: If the source access layer switch and the target access layer switch are the same access layer switch, the packets are forwarded directly.

14. A communication system comprising: Terminal devices, a fat-tree data center network topology structure for connecting the terminal devices in pairs with electrical signals, a traffic judgment module and a path selection module, wherein the traffic judgment module and the path selection module are configured to forward traffic according to the traffic forwarding method according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Data center network flow balancing method and device oriented to software definition

    CN105915467A

  • Differential flow control method and device for cloud computing data center network

    CN106533970A

  • Data center network flow scheduling method for minimization of network congestion and Qos (Quality of Service) guarantee

    CN115174489A

  • Data center intelligent routing decision-making method and device based on size of zoning flow

    CN118316861A

  • Network security data processing system and switch

    CN119210904A

Cited By

  • Data traffic scheduling method and communication system based on traffic characteristics and path selection

    CN121283964A