Communication method and device based on global load, storage medium and electronic equipment
By exchanging global load information and optimizing port selection in the Spine-Leaf networking architecture, the problem that traditional dynamic load sharing solutions cannot ensure smooth transmission paths is solved, and more efficient and reliable data transmission is achieved.
Patent Information
- Application Number
- CN202510678438.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-08-26
AI Technical Summary
In the existing Spine-Leaf networking architecture, the dynamic load sharing solution can only alleviate local network congestion and cannot ensure that the entire transmission path is always smooth, resulting in a degradation of network performance and reduced data transmission reliability.
Global load information exchange is carried out between leaf switches and spinal switches. By acquiring and combining congestion information of local and remote ports, the optimal sending port is selected for data transmission, the port bandwidth is evaluated using exponential weighted moving average or periodic discount algorithm, and the congestion information is efficiently transmitted through multicast technology.
It significantly reduces communication quality problems caused by local congestion, improves the communication efficiency and reliability of the network, and ensures the stability and efficiency of the data transmission path.
Smart Images

Figure CN120547154A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of communications, and in particular to a communication method, device, storage medium, and electronic device based on global load. Background Art
[0002] like Figure 1 As shown in the figure, a spine-leaf networking architecture is commonly used in modern data center networks. This architecture consists of two switching layers: the spine layer and the leaf layer. The switches in the spine layer are called spine switches. The figure shows four spine switches, designated Spine1 through Spine4. The switches in the leaf layer are called leaf switches. The figure shows four leaf switches, designated Leaf1 through Leaf4. The main function of a leaf switch is to aggregate traffic from servers and directly connect to the spine switches in the spine layer. In a typical spine-leaf networking architecture, each spine switch typically establishes connections with all leaf switches, forming a fully interconnected network topology.
[0003] Because the Spine-Leaf networking architecture flattens the network structure by reducing network layers, it can significantly improve network performance and meet the high throughput and low latency requirements of modern applications. Specifically, under this architecture, the data transmission path between any two servers remains consistent, requiring only three hops: starting from the leaf switch (Leaf layer) where the source server is located, forwarding through the spine switch (Spine layer), and finally reaching the leaf switch (Leaf layer) where the destination server is located. Therefore, the fixed three-hop path design not only simplifies data flow management, but also ensures that the data transmission latency and path length remain stable regardless of changes in the source and destination.
[0004] Based on the above characteristics, the Spine-Leaf network architecture has become the core network design choice of major data centers. Moreover, its fully interconnected network topology further strengthens this advantage. For example, Figure 1 The data packets from any leaf switch to another leaf switch can be forwarded through any spine switch, and all possible forwarding paths have the same number of hops.
[0005] To improve overall network throughput, leaf switches often employ dynamic load balancing (DLB) during forwarding. This approach selects the port with the lowest load based on the load of each leaf switch's egress, reducing network congestion. However, in practice, this approach only resolves local network congestion and does not guarantee a consistently unobstructed transmission path. Summary of the Invention
[0006] In order to overcome at least one deficiency in the prior art, the present application provides a communication method, apparatus, storage medium, and electronic device based on global load, specifically including:
[0007] In a first aspect, the present application provides a global load-based communication method, applied to a first leaf switch in a leaf-spine topology network, wherein the leaf-spine topology network further includes multiple spine switches and a second leaf switch, the method comprising:
[0008] Obtaining congestion information for a plurality of first ports in the first leaf switch, wherein the plurality of first ports are used for communication connection with the plurality of spine switches;
[0009] Receiving status information from each of the spine switches, wherein the status information includes congestion information of a second port in the corresponding spine switch, and each of the spine switches is communicatively connected to the second leaf switch through its second port;
[0010] Based on the congestion information of each first port and the congestion information of the second port in the corresponding spine switch, the optimal sending port of the target information is determined from the multiple first ports, wherein the target information is the information sent to the second leaf switch.
[0011] In a second aspect, the present application provides a global load-based communication method, which is applied to a spine switch in a leaf-spine topology network, wherein the leaf-spine topology network further includes a first leaf switch and a second leaf switch, and the method includes:
[0012] Obtaining status information of the spine switch, wherein the status information includes congestion information of a second port in the spine switch, the second port being used for communication connection with the second leaf switch;
[0013] Sending the state information to the first leaf switch, wherein the plurality of first ports in the first leaf switch are used for communication connection with the spine switch and other spine switches in the leaf-spine topology network;
[0014] The first leaf switch obtains congestion information of the multiple first ports, and determines the optimal sending port of target information from the multiple first ports based on the congestion information of each first port and the congestion information of the second port in the corresponding spine switch, wherein the target information is information sent to the second leaf switch.
[0015] In a third aspect, the present application provides a global load-based communication device, which is applied to a spine switch in a leaf-spine topology network, wherein the leaf-spine topology network further includes a first leaf switch and a second leaf switch, and the device includes:
[0016] a first detection module, configured to obtain congestion information of a plurality of first ports in the first leaf switch, wherein the plurality of first ports are configured to be communicatively connected with the plurality of spine switches;
[0017] a status receiving module configured to receive status information from each of the spine switches, wherein the status information includes congestion information of a second port in the corresponding spine switch, and each of the spine switches is communicatively connected to the second leaf switch via its second port;
[0018] A port screening module is used to determine the optimal sending port of the target information from the multiple first ports based on the congestion information of each first port and the congestion information of the second port in the corresponding spine switch, wherein the target information is the information sent to the second leaf switch.
[0019] In a fourth aspect, the present application provides a global load-based communication device, which is applied to a spine switch in a leaf-spine topology network, wherein the leaf-spine topology network further includes a first leaf switch and a second leaf switch, and the device includes:
[0020] a second detection module configured to obtain status information of the spine switch, wherein the status information includes congestion information of a second port in the spine switch, the second port being configured to communicate with the second leaf switch;
[0021] a state sending module, configured to send the state information to the first leaf switch, wherein the plurality of first ports in the first leaf switch are used for communication connection with the spine switch and other spine switches in the leaf-spine topology network;
[0022] The first leaf switch obtains congestion information of the multiple first ports, and determines the optimal sending port of target information from the multiple first ports based on the congestion information of each first port and the congestion information of the second port in the corresponding spine switch, wherein the target information is information sent to the second leaf switch.
[0023] In a fifth aspect, the present application provides a storage medium storing a computer program, which implements the global load-based communication method when executed by a processor.
[0024] In a sixth aspect, the present application provides an electronic device, comprising a processor and a memory, wherein the memory stores a computer program, and the computer program implements the global load-based communication method when executed by the processor.
[0025] Compared with the prior art, this application has the following beneficial effects:
[0026] The present application provides a communication method, device, storage medium and electronic device based on global load. When the method is applied to the first leaf switch in a leaf-spine topology network, the leaf-spine topology network also includes multiple spine switches and a second leaf switch. In the method, congestion information of multiple first ports in the first leaf switch is obtained, wherein the multiple first ports are used to communicate with multiple spine switches; status information from each spine switch is received; wherein the status information includes congestion information of the second port in the corresponding spine switch, and each spine switch is communicated with the second leaf switch through its own second port; based on the congestion information of each first port and the congestion information of the second port in the corresponding spine switch, the optimal sending port for the target information is determined from the multiple first ports. In this way, the load information in the global range is combined when selecting the sending port, which can significantly reduce the communication quality problems caused by local congestion. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0028] Figure 1 A schematic diagram of the structure of a leaf-spine topology network provided in an embodiment of the present application;
[0029] Figure 2 A schematic diagram of a flow chart of a global load-based communication method applied to a first leaf switch according to an embodiment of the present application;
[0030] Figure 3 A schematic diagram of the multicast message encapsulation format provided in an embodiment of the present application;
[0031] Figure 4 A flowchart of a global load-based communication method applied to a spine switch provided in an embodiment of the present application;
[0032] Figure 5 A schematic diagram of the complete workflow provided in the embodiments of the present application;
[0033] Figure 6 A schematic diagram of the structure of a global load-based communication device applied to a first leaf switch provided in an embodiment of the present application;
[0034] Figure 7 A schematic diagram of the structure of a global load-based communication method applied to a spine switch provided in an embodiment of the present application;
[0035] Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0036] To make the purpose, technical solutions, and advantages of the embodiments of the present application (hereinafter referred to as the present embodiments) clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Generally, the components of the embodiments of the present application described and shown in the drawings herein can be arranged and designed in various different configurations.
[0037] Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application for protection, but merely represents selected embodiments of the present application. All other embodiments obtained by persons of ordinary skill in the art based on the embodiments in the present application without creative work are within the scope of protection of the present application.
[0038] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.
[0039] In the description of this application, it should be noted that the terms "first", "second", "third", etc. are used only to distinguish descriptions and should not be understood as indicating or implying relative importance. In addition, the terms "comprises", "comprising" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element.
[0040] Based on the above statement, as introduced in the background technology, it is found in the actual operating environment that although the DLB solution based on local load evaluation can alleviate the network congestion problem of the leaf switch itself to a certain extent, it cannot ensure that the entire transmission path remains unobstructed at all times.
[0041] For example, see Figure 1 Assume that a data packet needs to be transmitted from leaf switch Leaf1 to leaf switch Leaf3 or leaf switch Leaf4. At this time, it can be forwarded through any of the spine switches Spine1, Spine2, Spine3, or Spine4. The transmission path is shown in the following table:
[0042]
[0043] However, in traditional DLB solutions, leaf switch Leaf1 makes decisions based solely on the congestion information of its local ports when selecting a packet egress. However, this load balancing strategy based on local congestion information can lead to unexpected network congestion. For example, when Leaf1 selects spine switch Spine2 as the egress based on the congestion information of its local ports, if Spine2 is currently heavily loaded or even congested, the data packet may be discarded at the node where Spine2 is located. Therefore, traditional DLB solutions, because they fail to fully consider the actual load status of downstream spine switches, can lead to decreased network performance and reduced data transmission reliability.
[0044] Based on the discovery of the above technical problems, the following technical solutions are proposed after creative work to solve or improve the above problems. It should be noted that the defects existing in the solutions in the above prior art are the results obtained after practice and careful study. Therefore, the discovery process of the above problems and the solutions proposed in the embodiments of this application for the above problems below should be regarded as contributions to this application in the process of invention and creation, and should not be understood as technical contents known to those skilled in the art.
[0045] In view of the above technical problems, this embodiment provides a communication method based on global load. The method is applied to the first leaf switch in a leaf-spine topology network, and the leaf-spine topology network also includes multiple spine switches and a second leaf switch. Figure 2 As shown, the method includes:
[0046] S1A: Obtain congestion information of multiple first ports in the first leaf switch.
[0047] Among them, multiple first ports are used to communicate with multiple spine switches.
[0048] S2A,receives status information from each spine switch.
[0049] The status information includes congestion information of the second port in the corresponding spine switch, and each spine switch is connected to the second leaf switch through its own second port.
[0050] S3A, determining an optimal sending port for the target information from the plurality of first ports based on the congestion information of each first port and the congestion information of the second port in the corresponding spine switch.
[0051] The target information is the information sent to the second leaf switch. This allows the selection of the sending port to rely not only on the local congestion information of the first port but also to fully incorporate global load information, namely the congestion information of the second port of the spine switch. This allows for a more accurate assessment of the overall load status of each possible transmission path and selects the path with the lowest load for data transmission. This significantly reduces communication quality issues caused by local congestion, effectively improving communication efficiency and reliability across the entire network.
[0052] It should be understood that the leaf switch in this embodiment refers to a switching device deployed at the edge of the network, which is mainly used to aggregate traffic from servers and forward this traffic to the spine switch. Therefore, the leaf switch is generally located at the access layer of the data center, connecting many servers or terminal devices, and usually needs to have powerful Layer 2 and Layer 3 forwarding capabilities, while supporting virtualization technology and dynamic load balancing functions. The spine switch is a switching device deployed at the core of the network, which plays the role of interconnecting all leaf switches and ensuring that any two leaf switches can communicate in a fully interconnected manner. Therefore, the spine switch usually works at the core layer of the network and needs to have higher port density, stronger processing power and more stable equipment performance.
[0053] To make the solution provided by this embodiment clearer, Figure 2 Each step of the method shown is described in detail. However, it should be understood that the operations of the flowchart can be implemented in any order, and steps that have no logical contextual relationship can be reversed or implemented simultaneously. In addition, those skilled in the art can add one or more other operations to the flowchart or remove one or more operations from the flowchart under the guidance of the content of this application. Figure 2 , the method comprising:
[0054] S1A: Obtain congestion information of multiple first ports in the first leaf switch.
[0055] The multiple first ports are used to communicate with multiple spine switches. It should be understood that a leaf-spine topology network includes multiple leaf switches, where a first leaf switch is any leaf switch in the leaf-spine topology network, and a second leaf switch is any leaf switch in the leaf-spine topology network other than the first leaf switch. Furthermore, in this embodiment, to facilitate distinguishing between the ports on the spine switch and the ports on the first leaf switch, the ports on the first leaf switch are referred to as first ports, and the ports on the spine switch are referred to as second ports.
[0056] Based on the above statement, in some embodiments, the processor of the first leaf switch includes a first switching chip, and the first switching chip is internally designed with a timer for periodically triggering the first switching chip to use mature algorithms such as Exponentially Weighted Moving Average (EWMA) or periodic discounted rate evaluation (DRE) to perform real-time evaluation of the bandwidth of multiple local first ports; and store the evaluation results in the DsPortMetric table. It should be understood that the DsPortMetric table is used to record the real-time congestion measurement value of each first port library. Therefore, when it is necessary to select the optimal sending port later, the DsPortMetric table can be queried to obtain the congestion information of each local first port.
[0057] In this embodiment, the congestion information of the port (leaf switch or spine switch) is also called the Metric value. The Metric value occupies a total of 4 bits of storage space, so its value range is 0 to 15. Therefore, the size of the Metric value directly reflects the congestion of the port. That is, the larger the value, the higher the congestion level of the port; conversely, the smaller the value, the more unobstructed the port is. For example, when the Metric value of a port is 0, it means that the port is almost not congested and data transmission can be carried out efficiently; when the Metric value is close to 15, it means that the port is almost in a completely congested state and data transmission may be significantly restricted.
[0058] Based on the explanation of step S1A in the above embodiment, continue to refer to Figure 2 , next Figure 2 Step S2A in the following example is explained:
[0059] S2A,receives status information from each spine switch.
[0060] The status information includes congestion information for the second port of the corresponding spine switch. Each spine switch is connected to the second leaf switch via its second port. It should be noted that the status information includes congestion information for the second port of the spine switch used to communicate with the second switch. This is merely a typical implementation provided in this embodiment and does not necessarily limit the content of the status information to this. The status information may also include congestion information for the ports of the spine switch used to communicate with other leaf switches.
[0061] In some embodiments, in order to alleviate the computing pressure of the first switching chip in the spine switch, the processor of the first leaf switch also includes a first coprocessor; the status information from each spine switch is processed by the first coprocessor and summarized and stored in a local table named DsRemoteMetric.
[0062] In this embodiment, each spine switch has an internal hardware circuit that uses the EWMA or DRE algorithm to evaluate the bandwidth of each outbound port in real time, generating a quantitative value as the congestion information for each port. These evaluation results are stored in the DsPortMetric table. Therefore, the DsPortMetric table records the current congestion information for each port.
[0063] Furthermore, to periodically transmit congestion information about these ports to the first leaf switch, a timer is installed within the spine switch. This timer triggers an operation at regular intervals, reading the association between each port and the leaf switch from the DsLeaf2Port table. Therefore, the DsLeaf2Port table acts as a mapping, clearly recording which leaf switches each spine switch port is connected to. The spine switch can read the congestion information for each port from the DsPortMetric table, compose a message data in a specific format, and send it to the connected leaf switch.
[0064] In practice, we've discovered that CPUs, as general-purpose processors, are designed to execute complex instruction sets and handle diverse tasks, rather than being optimized for fast network packet processing. When a spine switch needs to send state information to the first leaf switch, using the CPU to accomplish this task requires a complex series of steps, including reading the state information from memory, converting its format, encapsulating it into a packet, and finally sending it to the network interface. These operations not only involve numerous context switches but also potentially consume valuable CPU computing resources, increasing processing latency.
[0065] To this end, the processor in the spine switch includes a second switching chip and a second coprocessor that assists this chip. When the spine switch needs to send status information to a leaf switch, the second switching chip and the coprocessor work together to complete this operation. Specifically, after receiving the status information, the second coprocessor converts it into an initial message in a pre-defined format. This initial message in the pre-defined format is a customized message structure in this embodiment, used to carry congestion information for each port in the spine switch.
[0066] For example, the following is combined Figure 3 The structure of the initial message is explained intuitively. Assuming there are 128 leaf switches, it means there are 128 second ports. The congestion information of each port is quantified by a 4-bit value. Therefore, Figure 3 The initial message shown can be organized in the form of 128*4 bits, where each 4 bits represents congestion information sent from the spine switch to a leaf switch. Figure 3 The first 4-bit field in the initial message records the congestion information of the port of the spine switch used to connect to the leaf switch numbered 0 (ie, Leaf0).
[0067] The initial message generated by the second coprocessor is then passed to the second switching chip. As the core module for network data processing, the second switching chip is responsible for further encapsulating the initial message into a multicast message. This utilizes the second switching chip's inherent multicast replication and IP tunneling encapsulation capabilities to efficiently send status information to all directly connected leaf switches. Because multicast technology enables spine switches to send messages to all leaf switches in a leaf-spine topology at once, eliminating the need for unicast transmissions, the speed and efficiency of message synchronization are significantly improved.
[0068] For example, see Figure 3 , based on the initial message, the initial message is further encapsulated through the second switching chip. The encapsulated information includes IP GRE (Generic Routing Encapsulation) tunnel information, IP information, and MAC address information (MacDa, MacSa). IP GRE tunnel is a network encapsulation technology used to encapsulate data packets of one protocol into another protocol for transmission. In this embodiment, the main function of the IP GRE tunnel is to provide a reliable and standardized channel for message transmission between the spine switch and the leaf switch. The IPSA in the figure is the source address field in the IP GRE encapsulation, which is used to identify the sender of the message, that is, the device identifier of the spine switch; IPDA is the destination address field in the IP GRE encapsulation, which is used to identify the target device of the message, that is, the device identifier of the leaf switch.
[0069] Based on the explanation of step S2A in the above implementation, step S3A will be explained next:
[0070] S3A, determining an optimal sending port for the target information from the plurality of first ports based on the congestion information of each first port and the congestion information of the second port in the corresponding spine switch.
[0071] Since there are certain limitations when relying solely on the congestion information of each local port or the congestion information of each port of the remote spine switch, this embodiment shifts from evaluating the congestion information of the port to evaluating the congestion information of the transmission line. Therefore, as an optional implementation, step S3A may include:
[0072] S3A-1, based on the congestion information of each first port and the congestion information of the second port in the corresponding spine switch, obtain comprehensive congestion information of the transmission path corresponding to each first port.
[0073] It can be understood that in this embodiment, the first leaf switch first obtains the congestion information of multiple local first ports, and at the same time receives status information from each spine switch, where this status information includes the congestion information of the second port in the corresponding spine switch. By combining the data of these two dimensions, namely the congestion information of the first port and the congestion information of the second port in the corresponding spine switch, the comprehensive congestion information of the transmission path corresponding to each first port can be further calculated. Therefore, the comprehensive congestion information in this embodiment is no longer limited to the status of a single port, but reflects the overall congestion situation on the entire transmission path, thereby realizing the transition from port congestion information evaluation to transmission line congestion information evaluation.
[0074] During the research process, it was discovered that when a leaf switch needs to select the optimal sending port from multiple first ports, it must consider the congestion information of the first port itself and the congestion information of the second port of the associated spine switch. In traditional weighted fusion methods, these two factors require complex mathematical operations. For example, the congestion information of the first port and the congestion information of the second port are multiplied according to certain weights and then accumulated to obtain the comprehensive congestion information. However, in large-scale network environments, leaf switches may need to simultaneously evaluate the comprehensive congestion information of dozens or even hundreds of ports. In this case, repeated multiplication operations not only consume a lot of time, but also occupy hardware resources, thus affecting the real-time performance and efficiency of the entire system.
[0075] In view of this, in this embodiment, the first leaf switch converts the congestion information of each first port and the congestion information of the second port in the corresponding spine switch into an index coefficient of each first port; according to the index coefficient of each first port, the comprehensive congestion information of the transmission path corresponding to each first port is indexed from the preset congestion evaluation table.
[0076] In the above implementation, to reduce the complexity of constructing the index coefficient, the congestion information for each first port is directly concatenated with the congestion information for the second port of the corresponding spine switch. This approach simply combines the two values into a new value according to a predetermined rule, eliminating the need for complex mathematical operations. Because the concatenation operation itself is a simple data processing method, its computational efficiency is significantly improved compared to traditional weighted averaging or multiplication operations.
[0077] For example, as described in the above embodiment, the first leaf switch records the current congestion information of each first port in real time through the DsPortMetric table, so as to provide accurate basic information for subsequent comprehensive evaluation. Therefore, the first leaf switch reads the data stored in the DsPortMetric table to obtain the congestion information of the local first port; then, according to the device identifier (Spine ID) of the spine switch that is connected to each first port, the congestion metric value of the corresponding spine switch second port is read from the DsRemoteMetric table. Among them, the first leaf switch records the congestion status of the second port in all spine switches through the DsRemoteMetric table. It is assumed here that there are 128 spine switches in total, and the congestion information of each port in the spine switch occupies 4 bits of space. At this time, the structure of the DsRemoteMetric table is shown in the following table:
[0078] Leaf0 Metric Leaf1 Metric … Leaf127 Metric Leaf0 Metric Leaf1 Metric … Leaf127 Metric Leaf0 Metric Leaf1 Metric … Leaf127 Metric Leaf0 Metric Leaf1 Metric … Leaf127 Metric
[0079] Each row in the DsRemoteMetric table corresponds to a spine switch and is indexed by the device ID of the spine switch. Each column in the DsRemoteMetric table records the congestion information of each spine switch's port connecting to the same leaf switch.
[0080] The congestion information for the first port obtained in the above embodiment is concatenated with the congestion information for the second port of the corresponding spine switch to generate an index coefficient. Since the congestion information for the first port occupies 4 bits of storage space, and the congestion information for the second port of the corresponding spine switch also occupies 4 bits of storage space, the index coefficient obtained by concatenating the two occupies 8 bits of storage space.
[0081] In this embodiment, the congestion evaluation table in the first leaf switch is referred to as DsGlbMetricMap. The concatenated index coefficients are matched against the DsGlbMetricMap table, and comprehensive congestion information corresponding to the index coefficients is returned based on the rules defined in the table. In this embodiment, comprehensive congestion information is referred to as FinalMetric, which represents the overall congestion status of the transmission path corresponding to the current first port.
[0082] Since the index coefficient occupies 8 bits, its value range is 0 to 255. By dividing this range into multiple sub-ranges and defining different mapping rules for each sub-range, the system can dynamically adjust the evaluation strategy based on actual congestion conditions. This design is based on the fact that index coefficients in different ranges may correspond to different congestion characteristics or network conditions, so using differentiated mapping rules can more accurately reflect the actual situation.
[0083] For example, assume the range of 8-bit index coefficients is divided into three subranges: 0-63, 64-191, and 192-255. The first range (0-63) likely corresponds to paths in a low-congestion state in the network. In this case, a looser mapping rule can be used to reduce the possibility of misjudgment. The second range (64-191) likely corresponds to paths in a moderately congested state. In this case, a more refined mapping rule can be used to accurately distinguish different congestion levels. The third range (192-255) likely corresponds to paths in a highly congested state. In this case, a stricter mapping rule can be used to ensure timely identification of potential network bottlenecks.
[0084] By pre-establishing a mapping between different index values and comprehensive congestion information, the system can quickly complete congestion assessments during actual operation, eliminating the need for time-consuming real-time calculations. This approach not only simplifies implementation but also significantly improves response speed, especially in large-scale leaf-spine networks, significantly reducing latency issues associated with congestion assessments.
[0085] Based on the description of the comprehensive congestion information in the above implementation, step S3A further includes:
[0086] S3A-2, determining an optimal sending port from the plurality of first ports based on the comprehensive congestion information of the transmission path corresponding to each first port.
[0087] In this embodiment, the first leaf switch may select the first port with the smallest comprehensive congestion information as the optimal sending port, and send the target message to the second leaf switch through the port.
[0088] In the above embodiment, the communication method based on global load provided by this embodiment is described from the perspective of a leaf switch. Next, the communication method based on global load is further explained from the perspective of a spine switch. Therefore, the method is applied to a spine switch in a leaf-spine topology network, and the leaf-spine topology network also includes a first leaf switch and a second leaf switch. Figure 4 As shown, the method includes:
[0089] S1B, obtains the status information of the spine switch.
[0090] The status information includes congestion information of a second port in the spine switch, where the second port is used for communication connection with the second leaf switch.
[0091] The spine switch can be any spine switch in a leaf-spine topology network, and the first and second leaf switches can be any two switches in the leaf-spine topology network. As an alternative to the above method, each spine switch has an internal hardware circuit that uses the EWMA or DRE algorithm to evaluate the bandwidth of each outbound port in real time, generating a quantitative value as port congestion information. These evaluation results are stored in the DsPortMetric table. Therefore, the DsPortMetric table records the current congestion information for each port.
[0092] S2B sends the status information to the first leaf switch.
[0093] As an optional embodiment, the spine switch includes a processor, which includes a second switching chip and a second coprocessor that assists the second switching chip. The spine switch can convert state information into an initial message in a preset format via the second coprocessor; encapsulate the initial message into a multicast message via the second switching chip; and send the multicast message to the first leaf switch.
[0094] For example, a spine switch can periodically send status information to leaf switches in a leaf-spine topology. To enable the spine switch to periodically transmit port congestion information to the first leaf switch, a timer is internally configured in the spine switch. This timer triggers at a fixed interval, reads the port-to-leaf switch association from the mapping table DsLeaf2Port, and obtains congestion information for each port from the DsPortMetric table. This information is then formatted into message data and sent to the leaf switch.
[0095] However, if this task is performed by the CPU, since its design is not optimized for fast processing of network packets, it needs to go through complex steps such as reading status information, format conversion, message encapsulation and sending to the network interface, which will cause a large number of context switches and resource occupation, increasing latency. To this end, the spine switch adopts a solution in which the second switching chip and the second coprocessor work together. The second coprocessor is responsible for receiving status information and converting it into an initial message in a predefined format. For example, in the case of 128 leaf switches, the congestion information of each port is represented by 4 bits, and the initial message is organized in the form of 128×4 bits. The first 4-bit field records the congestion information of the port with the interface identifier 0.
[0096] The initial message generated by the second coprocessor is passed to the second switching chip, and the second switching chip uses its own multicast replication capability and IP tunnel encapsulation capability to further encapsulate the initial message into a multicast message. Through multicast technology, the spine switch can efficiently synchronize messages to all directly connected first leaf switches at one time, avoiding one-by-one unicast transmission, significantly improving synchronization speed and efficiency. It should be noted that information such as IP GRE (Generic Routing Encapsulation), IP and MAC addresses are also added during the encapsulation process. The IP GRE tunnel provides a reliable standardized communication channel, in which the ipsa field identifies the message sender (spine switch) and the ipda field identifies the receiver (leaf switch). In this way, the reliable delivery of messages in the leaf-spine topology network is ensured.
[0097] Because the multiple first ports of the first leaf switch are used to communicate with the spine switch and other spine switches in the leaf-spine topology network, the first leaf switch obtains congestion information of the multiple first ports and, based on the congestion information of each first port and the congestion information of the second port of the corresponding spine switch, determines the optimal port for sending target information from the multiple first ports, where the target information is information sent to the second leaf switch.
[0098] In the above embodiment, each step of the communication method based on global load balancing has been described separately. In order to enable those skilled in the art to more clearly understand the overall technical solution of the present invention, the complete workflow of the method will be described in detail in conjunction with a specific embodiment.
[0099] like Figure 5 As shown, for any spine switch, the second switching chip inside it will use the industry's mature EWMA or DRE algorithm to evaluate the bandwidth quantization value of each egress port in real time, representing the congestion information of the corresponding port, and will be stored in a table called DsPortMetric. Subsequently, the spine switch will set a hardware timer, which triggers an operation at regular intervals. During this process, the association between the egress port and the leaf switch is read from the DsLeaf2Port table, and the corresponding congestion information is extracted from the DsPortMetric table. These congestion information will be combined into an initial message with a fixed format. Each 4-bit data segment in the initial message represents the congestion information between different leaf switches. For example, the value in the first 4 bits represents the congestion information between the leaf switch Leaf0, the value in the second 4 bits represents the congestion information between the leaf switch Leaf1, and so on.
[0100] Next, the initial messages are passed to the second coprocessor inside the spine switch. Upon receiving these initial messages, the second coprocessor organizes them into custom-formatted multicast messages. These multicast messages not only contain the aforementioned congestion information but also utilize the switch chip's inherent multicast replication and IP tunnel encapsulation capabilities for further processing. Finally, these encapsulated messages are sent as multicasts to all directly connected leaf switches.
[0101] Next, two leaf switches are selected from the plurality of leaf switches as the first leaf switch and the second leaf switch, and the processing flow in each leaf switch is exemplarily described. The first leaf switch needs to send target information to the second leaf switch.
[0102] Continue to see Figure 5 , when the first leaf switch receives a message sent by any spine switch, the first coprocessor inside it will take over the subsequent operations. The first leaf switch parses the source address field (ipsa) in the message through the first coprocessor to clarify which spine switch the message data comes from; then, the first coprocessor directly writes the valid data into the DsRemoteMetric table. The DsRemoteMetric table is designed in a 64*128*4bit format, where 64 represents the number of spine switches, 128 represents the number of leaf switches, and 4 represents the space required to record congestion information. Therefore, each row in the DsRemoteMetric table corresponds to a spine switch, and each column records the congestion information between the spine switch and each leaf switch.
[0103] At the same time, the first chip inside the first leaf switch is also designed with an independent timer. This timer will start periodically and use the same EWMA or DRE algorithm as the spine switch to evaluate the bandwidth of the local egress port, which is stored as the port congestion information in the DsPortMetric table. In addition, in order to calculate the final comprehensive congestion information, the first leaf switch will perform a series of complex evaluation operations. Specifically, when the timer is triggered, the spine switch associated with each egress port is first searched from the Port2Spine table; and based on the device identification of the spine switch, a list of all leaf switches associated with the spine switch is obtained from the Spine2LeafBmp table, and each leaf switch is processed in turn.
[0104] For each local port, the first leaf switch will read the congestion information of the local outbound port (denoted as LocalMetric) through the DsPortMetric table, and read the congestion information from the corresponding spine switch (denoted as RemoteMetric) through the DsRemoteMetric table. The first leaf switch will splice the LocalMetric and RemoteMetric into an 8-bit index value and use it to index the DsGlbMetricMap table to obtain the final comprehensive congestion information (FinalMetric); finally, the FinalMetric will be written to the DsPortFinalMetric table, which uses the device identifier of the leaf switch as the index and records the comprehensive congestion information evaluated by the chip and passed through each port to reach the leaf switch. Therefore, the first leaf switch can query the comprehensive congestion information of each transmission path through each spine switch to send the target information to the second leaf switch through the device identifier of the second leaf switch, and thus select the optimal sending port to forward the target information to the second leaf switch.
[0105] Based on the same inventive concept as the global load-based communication method applied to the first leaf switch in the leaf-spine topology network provided in this embodiment, this embodiment also provides a global load-based communication device applied to the first leaf switch in the leaf-spine topology network, and the leaf-spine topology network also includes multiple spine switches and a second leaf switch. The device includes at least one software function module that can be stored in the memory in software form or solidified in the first leaf switch. The processor in the first leaf switch is used to execute the executable module stored in the memory. For example, the software function modules and computer programs included in the device. Please refer to Figure 6 , functionally speaking, the device can include:
[0106] A first detection module 11A is configured to obtain congestion information of a plurality of first ports in a first leaf switch, wherein the plurality of first ports are configured to communicate with a plurality of spine switches;
[0107] a status receiving module 12A configured to receive status information from each spine switch, wherein the status information includes congestion information of a second port in the corresponding spine switch, and each spine switch is communicatively connected to the second leaf switch via its second port;
[0108] The port screening module 13A is used to determine the optimal sending port of the target information from multiple first ports based on the congestion information of each first port and the congestion information of the second port in the corresponding spine switch, wherein the target information is the information sent to the second leaf switch.
[0109] In this embodiment, the first detection module 11A is used to implement Figure 1In step S1A, the status receiving module 12A is used to implement Figure 1 In step S2A, the port screening module 13A is used to implement Figure 1 Therefore, for a detailed description of each of the above modules, please refer to the specific implementation of the corresponding step.
[0110] Optionally, the port screening module 13A is further specifically configured to:
[0111] Obtaining comprehensive congestion information of a transmission path corresponding to each first port based on the congestion information of each first port and the congestion information of the second port in the corresponding spine switch;
[0112] An optimal sending port is determined from the plurality of first ports according to comprehensive congestion information of a transmission path corresponding to each first port.
[0113] Optionally, the port screening module 13A is further specifically configured to:
[0114] converting the congestion information of each first port and the congestion information of the second port in the corresponding spine switch into an index coefficient for each first port;
[0115] According to the index coefficient of each first port, the comprehensive congestion information of the transmission path corresponding to each first port is indexed from the preset congestion evaluation table.
[0116] Optionally, the port screening module 13A is further specifically configured to:
[0117] The congestion information of each first port is concatenated with the congestion information of the second port in the corresponding spine switch to obtain an index coefficient of each first port.
[0118] Based on the same inventive concept as the global load-based communication method for the spine switch in the leaf-spine topology network provided in this embodiment, this embodiment also provides a global load-based communication device, which is applied to the spine switch in the leaf-spine topology network. The leaf-spine topology network also includes a first leaf switch and a second leaf switch. The device includes at least one software function module that can be stored in the memory in software form or solidified in the first leaf switch. The processor in the first leaf switch is used to execute the executable module stored in the memory. For example, the software function modules and computer programs included in the device. Please refer to Figure 7 , functionally speaking, the device can include:
[0119] A second detection module 11B is configured to obtain status information of the spine switch, wherein the status information includes congestion information of a second port in the spine switch, where the second port is used for communication with the second leaf switch;
[0120] A state sending module 12B is configured to send state information to the first leaf switch, wherein the plurality of first ports in the first leaf switch are configured to communicate with the spine switch and other spine switches in the leaf-spine topology network;
[0121] The first leaf switch obtains congestion information of multiple first ports, and determines the optimal sending port of target information from the multiple first ports based on the congestion information of each first port and the congestion information of the second port in the corresponding spine switch, wherein the target information is information sent to the second leaf switch.
[0122] Optionally, the spine switch includes a processor, the processor includes a second switching chip and a second coprocessor assisting the second switching chip, and the status sending module 12B is further specifically configured to:
[0123] converting the state information into an initial message in a preset format by a second coprocessor;
[0124] The initial message is encapsulated into a multicast message through the second switching chip, and the multicast message is sent to the first leaf switch.
[0125] In addition, the functional modules in each embodiment of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0126] It should also be understood that if the above embodiments are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application.
[0127] Therefore, this embodiment further provides a storage medium, which is a computer-readable storage medium. The storage medium stores a computer program, and when the computer program is executed by a processor, the global load-based communication method provided in this embodiment for a spine switch or a first leaf switch is implemented. The storage medium can be a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, among other media that can store program code.
[0128] This embodiment provides an electronic device. Figure 8As shown, the electronic device includes a processor 22 and a memory 21, wherein the memory 21 stores a computer program. When the electronic device is a spine switch, the processor implements the global load-based communication method for the spine switch provided in this embodiment by reading and executing the computer program corresponding to the above embodiments in the memory 21. When the electronic device is a first leaf switch, the processor implements the global load-based communication method for the first leaf switch provided in this embodiment by reading and executing the computer program corresponding to the above embodiments in the memory 21.
[0129] Continue to see Figure 8 The electronic device further includes a communication unit 23. The memory 21, the processor 22 and the communication unit 23 are electrically connected to each other directly or indirectly via a system bus 24 to achieve data transmission or interaction.
[0130] The memory 21 may be an information recording device based on any electronic, magnetic, optical or other physical principles, for recording execution instructions, data, etc. In some embodiments, the memory 21 may be, but is not limited to, a volatile memory, a non-volatile memory, a storage drive, etc.
[0131] In some embodiments, the volatile memory may be a random access memory (RAM); in some embodiments, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, etc.; in some embodiments, the storage drive may be a magnetic disk drive, a solid-state drive, any type of storage disk (such as a CD, DVD, etc.), or a similar storage medium, or a combination thereof.
[0132] The communication unit 23 is used to send and receive data through a network. In some embodiments, the network may include a wired network, a wireless network, a fiber optic network, a telecommunications network, an intranet, the Internet, a local area network (LAN), a wide area network (WAN), a wireless local area network (WLAN), a metropolitan area network (MAN), a wide area network (WAN), a public switched telephone network (PSTN), a Bluetooth network, a ZigBee network, or a near field communication (NFC) network, or any combination thereof. In some embodiments, the network may include one or more network access points. For example, the network may include a wired or wireless network access point, such as a base station and / or a network switching node, through which one or more components of the service request processing system can connect to the network to exchange data and / or information.
[0133] The processor 22 may be an integrated circuit chip having signal processing capabilities, and the processor may include one or more processing cores (e.g., a single-core processor or a multi-core processor). By way of example only, the processor may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), an application-specific instruction set processor (ASIP), a graphics processing unit (GPU), a physical processing unit (PPU), a digital signal processor (DSP), a field programmable gate array (FPGA), a programmable logic device (PLD), a controller, a microcontroller unit, a reduced instruction set computer (RISC), or a microprocessor, or any combination thereof.
[0134] I understand. Figure 8The structure shown is for reference only. Figure 8 More or fewer components than shown, or with Figure 8 Different configurations shown. Figure 8 The components shown may be implemented in hardware, software, or a combination thereof.
[0135] It should be understood that the devices and methods disclosed in the above embodiments may also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of the devices, methods, and computer program products according to multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram may represent a module, a program segment, or a portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box may also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes may actually be executed substantially in parallel, or they may sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, may be implemented using a dedicated hardware-based system that performs a specified function or action, or may be implemented using a combination of dedicated hardware and computer instructions.
[0136] The above descriptions are merely examples of various embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any modifications or substitutions that can be readily conceived by a person skilled in the art within the technical scope disclosed in the present application should be included within the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A communication method based on global load, characterized in that: Applied to a first leaf switch in a leaf-spine topology network, the leaf-spine topology network further comprising a plurality of spine switches and a second leaf switch, the method comprising: Obtaining congestion information for a plurality of first ports in the first leaf switch, wherein the plurality of first ports are used for communication connection with the plurality of spine switches; Receiving status information from each of the spine switches, wherein the status information includes congestion information of a second port in the corresponding spine switch, and each of the spine switches is communicatively connected to the second leaf switch through its second port; Based on the congestion information of each first port and the congestion information of the second port in the corresponding spine switch, the optimal sending port of the target information is determined from the multiple first ports, wherein the target information is the information sent to the second leaf switch.
2. The communication method based on global load according to claim 1, characterized in that: Determining an optimal port for sending target information from the plurality of first ports based on congestion information of each first port and congestion information of a second port in a corresponding spine switch includes: Obtaining comprehensive congestion information of a transmission path corresponding to each of the first ports based on the congestion information of each of the first ports and the congestion information of the second port in the corresponding spine switch; The optimal sending port is determined from the plurality of first ports according to comprehensive congestion information of the transmission path corresponding to each first port.
3. The communication method based on global load according to claim 2, characterized in that: Obtaining comprehensive congestion information of a transmission path corresponding to each first port based on the congestion information of each first port and the congestion information of the second port in the corresponding spine switch, including: converting the congestion information of each of the first ports and the congestion information of the second port in the corresponding spine switch into an index coefficient of each of the first ports; According to the index coefficient of each first port, comprehensive congestion information of the transmission path corresponding to each first port is indexed from a preset congestion evaluation table.
4. The communication method based on global load according to claim 2, characterized in that: Converting the congestion information of each of the first ports and the congestion information of the second port in the corresponding spine switch into an index coefficient of each of the first ports includes: The congestion information of each of the first ports is concatenated with the congestion information of the second port in the corresponding spine switch to obtain an index coefficient of each of the first ports.
5. A communication method based on global load, characterized in that: A spine switch applied to a leaf-spine topology network, wherein the leaf-spine topology network further includes a first leaf switch and a second leaf switch, wherein the method includes: Obtaining status information of the spine switch, wherein the status information includes congestion information of a second port in the spine switch, the second port being used for communication connection with the second leaf switch; Sending the state information to the first leaf switch, wherein the plurality of first ports in the first leaf switch are used for communication connection with the spine switch and other spine switches in the leaf-spine topology network; The first leaf switch obtains congestion information of the multiple first ports, and determines the optimal sending port of target information from the multiple first ports based on the congestion information of each first port and the congestion information of the second port in the corresponding spine switch, wherein the target information is information sent to the second leaf switch.
6. The communication method based on global load according to claim 5, characterized in that: The spine switch includes a processor, the processor includes a second switching chip and a second coprocessor assisting the second switching chip, and sending the status information to the first leaf switch includes: converting the state information into an initial message in a preset format by a second coprocessor; The initial message is encapsulated into a multicast message by the second switching chip, and the multicast message is sent to the first leaf switch.
7. A communication device based on global load, characterized in that: A spine switch applied to a leaf-spine topology network, wherein the leaf-spine topology network further includes a first leaf switch and a second leaf switch, wherein the apparatus includes: a first detection module, configured to obtain congestion information of a plurality of first ports in the first leaf switch, wherein the plurality of first ports are configured to communicate with the spine switch; a status receiving module configured to receive status information from each of the spine switches, wherein the status information includes congestion information of a second port in the corresponding spine switch, and each of the spine switches is communicatively connected to the second leaf switch via its second port; A port screening module is used to determine the optimal sending port of the target information from the multiple first ports based on the congestion information of each first port and the congestion information of the second port in the corresponding spine switch, wherein the target information is the information sent to the second leaf switch.
8. A communication device based on global load, characterized in that: A spine switch applied to a leaf-spine topology network, wherein the leaf-spine topology network further includes a first leaf switch and a second leaf switch, wherein the apparatus includes: a second detection module configured to obtain status information of the spine switch, wherein the status information includes congestion information of a second port in the spine switch, the second port being configured to communicate with the second leaf switch; a state sending module, configured to send the state information to the first leaf switch, wherein the plurality of first ports in the first leaf switch are used for communication connection with the spine switch and other spine switches in the leaf-spine topology network; The first leaf switch obtains congestion information of the multiple first ports, and determines the optimal sending port of target information from the multiple first ports based on the congestion information of each first port and the congestion information of the second port in the corresponding spine switch, wherein the target information is information sent to the second leaf switch.
9. A storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the computer program implements the global load-based communication method according to any one of claims 1 to 4 or 4 to 6.
10. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory stores a computer program, and the computer program, when executed by the processor, implements the global load-based communication method according to any one of claims 1-4 or 5-6.