Bandwidth allocation method and device, electronic device and storage medium

By adopting the rounding strategy and quadratic calibration method in the bandwidth allocation of computing nodes and switches, the problem of non-integer multiples in bandwidth allocation is solved, and efficient bandwidth allocation and data transmission stability are achieved. It is suitable for high-performance computing clusters and super-node AI servers.

CN120547067BActive Publication Date: 2025-10-03INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511037430.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-10-03
Estimated Expiration
2045-07-25

AI Technical Summary

Technical Problem

In the existing technology, in a fully interconnected topology of computing nodes and switches, since the processor bandwidth and the switching chip bandwidth are fixed values, non-rate integer situations occur during bandwidth allocation, making it impossible to achieve effective design.

Method used

By determining the initial bandwidth allocated by each processor to each switch, and adopting a rounding strategy when the initial bandwidth is a non-integer multiple, adjustments are made in combination with the preset processor bandwidth to ensure that the bandwidth allocation is an integer multiple. At the same time, secondary calibration is performed based on the total receive bandwidth of the switch to optimize the target bandwidth allocation.

Benefits of technology

It achieves flexible adaptation under different processor numbers and switching chip bandwidth conditions, ensures the accuracy and efficiency of bandwidth allocation, avoids data transmission errors or blockages, and is suitable for the stable operation of high-performance computing clusters and super-node AI servers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120547067B_ABST
    Figure CN120547067B_ABST
Patent Text Reader

Abstract

The present application provides a bandwidth allocation method and device, an electronic device, and a storage medium, which determine the initial bandwidth allocated by each processor to each switch, wherein each processor is any one of the multiple processors of each computing node, each switch is any one of the multiple switches of a predetermined integer target number of switches, and multiple processors are connected to each switch respectively; when the initial bandwidth is a non-integer multiple of the preset transmission channel bandwidth, the initial bandwidth is rounded by a rounding strategy to determine the target bandwidth allocated by each processor to each switch; based on the target bandwidth, the total receiving bandwidth of each switch is determined, and based on the total receiving bandwidth, the target bandwidth is adjusted to control each processor to allocate the adjusted target bandwidth to each switch. Compared with related technologies, the present application can avoid the problem of difficult design of processor bandwidth allocation due to bandwidth granularity. At the same time, it significantly improves the efficiency of bandwidth allocation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a bandwidth allocation method and device, an electronic device, and a storage medium. Background Art

[0002] Commonly used servers in the industry consist of compute nodes that include multiple processors (such as Graphics Processing Units (GPUs)) and multiple switches interconnected with them. Any processor in a compute node can achieve high-speed signal transmission with the switching chip of any switch through a high-speed internet network. With the help of switch relays, it can complete data transmission with processors in other compute nodes.

[0003] Currently, when implementing a fully interconnected topology of compute nodes and switches, the bandwidth of each processor in a compute node is evenly distributed across each switch chip in each switch node, achieving full processor bandwidth interconnection. However, because the bandwidth of processor and switch chips is fixed, this can lead to non-integer bandwidth when evenly distributing processor bandwidth, making the bandwidth allocation design incomplete. Summary of the Invention

[0004] This application provides a bandwidth allocation method and device, an electronic device, and a storage medium, the main purpose of which is to solve the problem in related technologies that processor bandwidth allocation cannot be designed due to bandwidth granularity.

[0005] According to a first aspect of the present application, a bandwidth allocation method is provided, comprising:

[0006] Determine an initial bandwidth allocated by each processor to each switch, where each processor is any one of multiple processors included in each computing node among the multiple computing nodes, each switch is any one of multiple switches corresponding to a predetermined target number of switches, the multiple processors included in each computing node are connected to each switch, and the target number of switches is an integer;

[0007] When the initial bandwidth is a non-integer multiple of the preset transmission channel bandwidth, the initial bandwidth is rounded up using a rounding strategy to determine the target bandwidth allocated by each processor to each switch;

[0008] Based on the target bandwidth, the total receiving bandwidth of each switch is determined, and the target bandwidth is adjusted based on the total receiving bandwidth to control each processor to allocate the adjusted target bandwidth to each switch.

[0009] In some embodiments, determining the initial bandwidth allocated by each processor to each switch includes: obtaining the preset processor bandwidth of each processor, the preset switching chip bandwidth of each switching chip, the number of computing nodes of multiple computing nodes, and the number of processors of multiple processors, each switch including a preset number of switching chips; determining the total node bandwidth of multiple computing nodes based on the product of the preset processor bandwidth, the number of computing nodes, and the number of processors; determining the number of target switches connected to the multiple computing nodes based on the total node bandwidth, the preset switching chip bandwidth, and the preset number of switching chips; and determining the initial bandwidth allocated by each processor to each switch based on the target number of switches.

[0010] In some embodiments, determining a target number of switches connected to a plurality of computing nodes based on the total node bandwidth, the preset switching chip bandwidth, and the preset number of switching chips includes: determining the total number of switching chips of the switching chip based on a ratio of the total node bandwidth to the preset switching chip bandwidth; determining the total number of switching chips as the initial number of switches based on the preset number of switching chips; when the initial number of switches is an integer, determining the initial number of switches as the target number of switches; and when the initial number of switches is a non-integer, rounding up the initial number of switches to obtain the target number of switches.

[0011] In some embodiments, determining the initial bandwidth allocated by each processor to each switch based on the target number of switches includes: determining a first theoretical bandwidth allocated by each computing node to each switch based on a preset processor bandwidth, the number of processors, and the target number of switches; determining a second theoretical bandwidth allocated by each processor to each switch based on the first theoretical bandwidth and the number of processors, and determining the second theoretical bandwidth as the initial bandwidth.

[0012] In some embodiments, when the initial bandwidth is a non-integer multiple of the preset transmission channel bandwidth, rounding the initial bandwidth using a rounding strategy to determine the target bandwidth allocated by each processor to each switch includes: when the initial bandwidth is a non-integer multiple of the preset transmission channel bandwidth, rounding the initial bandwidth using a rounding strategy, and combining it with the preset processor bandwidth to determine the target bandwidth allocated by each processor to each switch.

[0013] In some embodiments, the target bandwidth includes a high-range bandwidth and a low-range bandwidth. The initial bandwidth is rounded up using a rounding strategy, and combined with the preset processor bandwidth, the target bandwidth allocated by each processor to each switch is determined, including: rounding up or down the initial bandwidth to obtain the high-range bandwidth and low-range bandwidth corresponding to the initial bandwidth; based on the high-range bandwidth, the low-range bandwidth and the preset processor bandwidth, determining the high-range bandwidth or low-range bandwidth allocated by each processor to each switch.

[0014] In some embodiments, rounding the initial bandwidth up or down to obtain the high-range bandwidth and low-range bandwidth corresponding to the initial bandwidth includes: determining the initial bandwidth as the rounding center value; based on the rounding center value and taking the preset transmission channel bandwidth as the reference bandwidth unit, rounding the initial bandwidth up or down to obtain the high-range bandwidth and low-range bandwidth corresponding to the initial bandwidth, the high-range bandwidth and the low-range bandwidth being integer multiples of the reference bandwidth unit.

[0015] In some embodiments, determining the high-range bandwidth or low-range bandwidth allocated to each switch by each processor based on the high-range bandwidth, the low-range bandwidth, and the preset processor bandwidth includes: determining, based on the high-range bandwidth, the low-range bandwidth, and the preset processor bandwidth, in combination with a preset bandwidth constraint algorithm, the number of high-range switches in the target number of switches to which the high-range bandwidth needs to be allocated and the number of low-range switches to which the low-range bandwidth needs to be allocated; and determining the high-range bandwidth or low-range bandwidth allocated to each switch by each processor based on the number of high-range switches and the number of low-range switches.

[0016] In some embodiments, based on the target bandwidth, determining the total receiving bandwidth of each switch, and adjusting the target bandwidth based on the total receiving bandwidth includes: determining the total receiving bandwidth of each switch based on the target bandwidth allocated by each processor to each switch; comparing the total receiving bandwidth of each switch with the preset switching chip bandwidth; when the total receiving bandwidth is greater than the preset switching chip bandwidth, adjusting the target bandwidth allocated by each processor to each switch, the total receiving bandwidth being greater than the preset switching chip bandwidth indicating that each processor in multiple computing nodes allocates the same target bandwidth to each switch.

[0017] In some embodiments, determining the total receive bandwidth of each switch based on the target bandwidth allocated by each processor to each switch includes determining the total receive bandwidth of each switch based on a product of the target bandwidth, the number of processors, and the number of computing nodes.

[0018] In some embodiments, adjusting the target bandwidth allocated by each processor to each switch includes adjusting the target bandwidth allocated by each processor to each switch based on the low-range bandwidth, the high-range bandwidth, the number of low-range switches, the number of high-range switches, the number of processors, and the number of computing nodes.

[0019] In some embodiments, adjusting the target bandwidth allocated by each processor to each switch based on the low-range bandwidth, the high-range bandwidth, the number of low-range switches, the number of high-range switches, the number of processors, and the number of computing nodes includes: adjusting the target bandwidth allocated by each processor to each switch based on the low-range bandwidth, the high-range bandwidth, the number of processors, and the number of computing nodes, to obtain multiple second target bandwidth combinations allocated by each computing node to each switch, where the second target bandwidth combination is a combination of multiple first target bandwidth combinations that meets a bandwidth restriction condition, each first target bandwidth combination is composed of the number of processors of each computing node that allocate low-range bandwidth and high-range bandwidth to each switch, and the total number of processors of each computing node that allocate low-range bandwidth and high-range bandwidth to each switch is the same as a preset number of processors; and determining the adjusted target bandwidth allocated by each processor to each switch based on the number of low-range switches, the number of high-range switches, and the multiple second target bandwidth combinations.

[0020] In some embodiments, adjusting the target bandwidth allocated by each processor to each switch based on the low-range bandwidth, the high-range bandwidth, the number of processors, and the number of computing nodes to obtain multiple second target bandwidth combinations allocated by each computing node to each switch includes: adjusting the low-range bandwidth or the high-range bandwidth allocated by each processor to each switch based on the low-range bandwidth, the high-range bandwidth, the number of processors, and the number of computing nodes to obtain multiple first target bandwidth combinations allocated by each computing node to each switch; determining a first total receive bandwidth of each switch corresponding to each first target bandwidth combination; determining multiple combinations that meet a bandwidth restriction condition from the multiple first target bandwidth combinations, and determining the multiple combinations as multiple second target bandwidth combinations, where the bandwidth restriction condition is that the first total receive bandwidth is less than or equal to a preset switching chip bandwidth.

[0021] In some embodiments, determining the adjusted target bandwidth allocated by each processor to each switch based on the number of low-range switches, the number of high-range switches, and multiple second target bandwidth combinations includes: sorting the second total received bandwidths based on the second total received bandwidths of each switch corresponding to each second bandwidth combination, and determining the second bandwidth combination corresponding to the second total received bandwidth with a sorted position as a target position as a third target bandwidth combination; and determining the adjusted target bandwidth allocated by each processor to each switch based on the number of low-range switches, the number of high-range switches, and the third target bandwidth combination.

[0022] In some embodiments, determining the adjusted target bandwidth allocated by each processor to each switch based on the number of low-range switches, the number of high-range switches, and the third target bandwidth combination includes: determining the number of processors in each computing node in the third target bandwidth combination that allocate low-range bandwidth and high-range bandwidth to each switch; and determining the adjusted target bandwidth allocated by each processor to each switch from the third target bandwidth combination based on the number of processors in each computing node in the third target bandwidth combination that allocate low-range bandwidth and high-range bandwidth to each switch, in combination with constraints on the number of low-range switches and the number of high-range switches, so that a difference between multiple third total receive bandwidths of the multiple switches is less than or equal to a preset error range, and the third total receive bandwidth is obtained by each switch based on the adjusted target bandwidth of each processor.

[0023] In some embodiments, multiple computing nodes and multiple switches are arranged in a front-to-back orientation, and multiple processors included in each computing node are orthogonally connected to each switch via an orthogonal connector.

[0024] According to a second aspect of the present application, a bandwidth allocation device is provided, including:

[0025] a determining unit, configured to determine an initial bandwidth allocated by each processor to each switch, wherein each processor is any one of a plurality of processors included in each of the plurality of computing nodes, each switch is any one of a plurality of switches corresponding to a predetermined target number of switches, the plurality of processors included in each computing node are connected to each switch, and the target number of switches is an integer;

[0026] a first adjusting unit, configured to, when the initial bandwidth is a non-integer multiple of the preset transmission channel bandwidth, round the initial bandwidth using a rounding strategy to determine a target bandwidth allocated by each processor to each switch;

[0027] The second adjustment unit is configured to determine the total receiving bandwidth of each switch based on the target bandwidth, and adjust the target bandwidth based on the total receiving bandwidth to control each processor to distribute the adjusted target bandwidth to each switch.

[0028] In some embodiments, the determination unit is used to: obtain the preset processor bandwidth of each processor, the preset switching chip bandwidth of each switching chip, the number of computing nodes of multiple computing nodes, and the number of processors of multiple processors, each switch including a preset number of switching chips; determine the total node bandwidth of multiple computing nodes based on the product of the preset processor bandwidth, the number of computing nodes, and the number of processors; determine the number of target switches connected to the multiple computing nodes based on the total node bandwidth, the preset switching chip bandwidth, and the preset number of switching chips; and determine the initial bandwidth allocated to each switch by each processor based on the target number of switches.

[0029] In some embodiments, the determining unit is configured to: determine the total number of switch chips of the switch chip based on a ratio of the total node bandwidth to the preset switch chip bandwidth; determine the total number of switch chips as the initial number of switches based on the preset number of switch chips; when the initial number of switches is an integer, determine the initial number of switches as the target number of switches; and when the initial number of switches is a non-integer, round up the initial number of switches to obtain the target number of switches.

[0030] In some embodiments, the determination unit is used to: determine a first theoretical bandwidth allocated by each computing node to each switch based on a preset processor bandwidth, the number of processors, and the target number of switches; determine a second theoretical bandwidth allocated by each processor to each switch based on the first theoretical bandwidth and the number of processors, and determine the second theoretical bandwidth as the initial bandwidth.

[0031] In some embodiments, the first adjustment unit is configured to: when the initial bandwidth is a non-integer multiple of the preset transmission channel bandwidth, round the initial bandwidth using a rounding strategy, and determine the target bandwidth allocated by each processor to each switch in combination with the preset processor bandwidth.

[0032] In some embodiments, the target bandwidth includes a high-range bandwidth and a low-range bandwidth, and the first adjustment unit is used to: round up or down the initial bandwidth to obtain the high-range bandwidth and low-range bandwidth corresponding to the initial bandwidth; and determine the high-range bandwidth or low-range bandwidth allocated by each processor to each switch based on the high-range bandwidth, the low-range bandwidth, and the preset processor bandwidth.

[0033] In some embodiments, the first adjustment unit is used to: determine the initial bandwidth as a rounded center value; based on the rounded center value and with the preset transmission channel bandwidth as the benchmark bandwidth unit, round the initial bandwidth up or down to obtain a high-end bandwidth and a low-end bandwidth corresponding to the initial bandwidth, and the high-end bandwidth and the low-end bandwidth are integer multiples of the benchmark bandwidth unit.

[0034] In some embodiments, the first adjustment unit is configured to: determine, based on the high-range bandwidth, the low-range bandwidth, and the preset processor bandwidth, in combination with a preset bandwidth constraint algorithm, the number of high-range switches in the target number of switches to which the high-range bandwidth needs to be allocated and the number of low-range switches to which the low-range bandwidth needs to be allocated; and determine, based on the number of high-range switches and the number of low-range switches, the high-range bandwidth or the low-range bandwidth to be allocated to each switch by each processor.

[0035] In some embodiments, the second adjustment unit is used to: determine the total receiving bandwidth of each switch based on the target bandwidth allocated by each processor to each switch; compare the total receiving bandwidth of each switch with the preset switching chip bandwidth; when the total receiving bandwidth is greater than the preset switching chip bandwidth, adjust the target bandwidth allocated by each processor to each switch, and the total receiving bandwidth being greater than the preset switching chip bandwidth indicates that each processor in multiple computing nodes allocates the same target bandwidth to each switch.

[0036] In some embodiments, the second adjustment unit is configured to determine the total receiving bandwidth of each switch based on the product of the target bandwidth, the number of processors, and the number of computing nodes.

[0037] In some embodiments, the second adjustment unit is configured to adjust the target bandwidth allocated by each processor to each switch based on the low-range bandwidth, the high-range bandwidth, the number of low-range switches, the number of high-range switches, the number of processors, and the number of computing nodes.

[0038] In some embodiments, the second adjustment unit is configured to: adjust the target bandwidth allocated by each processor to each switch based on the low-range bandwidth, the high-range bandwidth, the number of processors, and the number of computing nodes, to obtain multiple second target bandwidth combinations allocated by each computing node to each switch, where the second target bandwidth combination is a combination of multiple first target bandwidth combinations that meets the bandwidth restriction condition, each first target bandwidth combination is composed of the number of processors of each computing node that allocate low-range bandwidth and high-range bandwidth to each switch, and the sum of the number of processors of each computing node that allocate low-range bandwidth and high-range bandwidth to each switch is the same as the preset number of processors; and determine the adjusted target bandwidth allocated by each processor to each switch based on the number of low-range switches, the number of high-range switches, and the multiple second target bandwidth combinations.

[0039] In some embodiments, the second adjustment unit is configured to: adjust the low-range bandwidth or high-range bandwidth allocated by each processor to each switch based on the low-range bandwidth, the high-range bandwidth, the number of processors, and the number of computing nodes, to obtain multiple first target bandwidth combinations allocated by each computing node to each switch; determine the first total receive bandwidth of each switch corresponding to each first target bandwidth combination; determine multiple combinations that meet a bandwidth restriction condition from the multiple first target bandwidth combinations, and determine the multiple combinations as multiple second target bandwidth combinations, where the bandwidth restriction condition is that the first total receive bandwidth is less than or equal to a preset switching chip bandwidth.

[0040] In some embodiments, the second adjustment unit is configured to: sort the second total receive bandwidths based on the second total receive bandwidths of each switch corresponding to each second bandwidth combination, and determine the second bandwidth combination corresponding to the second total receive bandwidths at the target sorting position as a third target bandwidth combination; and determine the adjusted target bandwidth allocated by each processor to each switch based on the number of low-range switches, the number of high-range switches, and the third target bandwidth combination.

[0041] In some embodiments, the second adjustment unit is configured to: determine the number of processors in each computing node in the third target bandwidth combination that allocate low-range bandwidth and high-range bandwidth to each switch; determine, based on the number of processors in each computing node in the third target bandwidth combination that allocate low-range bandwidth and high-range bandwidth to each switch, in combination with constraints on the number of low-range switches and the number of high-range switches, an adjusted target bandwidth allocated by each processor to each switch from the third target bandwidth combination, so that a difference between multiple third total receive bandwidths of multiple switches is less than or equal to a preset error range, and the third total receive bandwidth is obtained by each switch based on the adjusted target bandwidth of each processor.

[0042] In some embodiments, multiple computing nodes and multiple switches are arranged in a front-to-back orientation, and multiple processors included in each computing node are orthogonally connected to each switch via an orthogonal connector.

[0043] According to a third aspect of the present application, an electronic device is provided, including:

[0044] at least one processor; and

[0045] a memory communicatively connected to at least one processor; wherein,

[0046] The memory stores instructions that can be executed by at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the bandwidth allocation method of the first aspect.

[0047] According to a fourth aspect of the present application, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the bandwidth allocation method of the first aspect.

[0048] According to a fifth aspect of the present application, a computer program product is provided, comprising a computer program, which implements the bandwidth allocation method of the first aspect when executed by a processor.

[0049] The present application provides a bandwidth allocation method and device, an electronic device, and a storage medium, which relate to the field of distributed computing technology. Compared with related technologies, the present application first accurately determines the initial bandwidth allocated to each processor to each switch when the number of switches is an integer; and when the initial bandwidth is not an integer multiple of the preset transmission channel bandwidth, a rounding strategy is used for optimization and adjustment, which fundamentally avoids the problem that the processor bandwidth allocation is difficult to design due to the bandwidth granularity. In addition, the target bandwidth is calibrated twice according to the total receiving bandwidth of each switch, which greatly improves the bandwidth allocation efficiency. Therefore, in scenarios where the processor bandwidth and the switching chip bandwidth are complex and changeable, the present application can flexibly adapt to the combination of different numbers of processors and computing nodes, efficiently meet the full interconnection bandwidth allocation needs, and provide reliable guarantees for the stable operation and performance improvement of the server.

[0050] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present application.

[0052] Figure 1 A schematic diagram of a computing node and a switch performing orthogonal interconnection according to an embodiment of the present application;

[0053] Figure 2 A schematic diagram of a related technology provided in an embodiment of the present application to achieve full interconnection between computing nodes and switches;

[0054] Figure 3 A flowchart of the first bandwidth allocation method provided in an embodiment of the present application;

[0055] Figure 4 A flowchart of a second bandwidth allocation method provided in an embodiment of the present application;

[0056] Figure 5 A schematic diagram of a flow chart of a third bandwidth allocation method provided in an embodiment of the present application;

[0057] Figure 6 A schematic diagram of an embodiment of the present application providing multiple processors in each computing node allocating target bandwidth to multiple switches;

[0058] Figure 7 A schematic diagram of an embodiment of the present application providing a method for allocating low-level bandwidth and high-level bandwidth to multiple switches by multiple processors in each computing node;

[0059] Figure 8 A flowchart of the fourth bandwidth allocation method provided in an embodiment of the present application;

[0060] Figure 9 A schematic diagram of a preliminary target bandwidth allocated by each processor in a single computing node to each switch provided in an embodiment of the present application;

[0061] Figure 10 A schematic diagram of a target bandwidth allocated by each processor in a single computing node to each switch after adjusting the target bandwidth provided in an embodiment of the present application;

[0062] Figure 11 A schematic diagram of the structure of a bandwidth allocation device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0063] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0064] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0065] The internal structure of a commonly used supernode artificial intelligence (AI) server in the industry currently consists of multiple compute nodes equipped with processors and multiple switches interconnecting the compute nodes. In this architecture, any processor in any compute node can communicate with the switching chip in any switch via a high-speed interconnection network. Furthermore, the relay function of the switch allows data transmission between processors in any two compute nodes.

[0066] There are two main physical forms of implementing this type of interconnection in related technologies: cable tray interconnection and orthogonal interconnection. Cable tray interconnection is to set up a cable backplane on the back of the cabinet of the super-node artificial intelligence server, and the computing nodes are connected to the switch through cables to achieve data transmission; orthogonal interconnection is to arrange the computing nodes and the switch in front and back, and use the orthogonal connectors set on them to achieve 90-degree orthogonal direct plug-in. Figure 1 , which is a schematic diagram of a computing node and a switch performing orthogonal interconnection provided by the present application.

[0067] In the related art, as shown in FIG2 , the present application provides a schematic diagram of a related art for realizing full interconnection between computing nodes and switches. Figure 2 In order to achieve a fully interconnected topology, the relevant technology will evenly distribute the bandwidth of each processor in each computing node to each orthogonal connector, and then further evenly distribute it to each switching chip in each switching node, so as to achieve full interconnection of the processor bandwidth in the computing node.

[0068] Therefore, to achieve full interconnection of processor bandwidth in a compute node, the total external bandwidth of all processors must be equal to the total external bandwidth of all switches. During the design process, the number of compute nodes and switches can be determined by design specifications.

[0069] However, since the bandwidth of processors and switch chips is fixed, this creates a bandwidth granularity problem. Figure 1 The orthogonal architecture shown here is unique in that the orthogonal connectors of each compute node are paired one-to-one with the orthogonal connectors of each switch, and the connector bandwidths are identical. Therefore, when evenly allocating processor bandwidth, it's possible that the processor bandwidth might not be an integer. For example, if the minimum unit of processor bandwidth is 100G, evenly allocating it might result in a value of 211.8G. In this case, bandwidth allocation will not work as designed.

[0070] In addition, because the number of switches must be an integer, according to the design of evenly distributing bandwidth, the final calculation result may include a decimal number of switches, which will also make it impossible to design and distribute bandwidth normally.

[0071] To address the issues presented in related solutions, the present invention proposes a fully interconnected bandwidth allocation method for orthogonal architecture cabinets. This method can flexibly adapt to various combinations of processor numbers and node numbers under varying processor bandwidth and switch chip bandwidth conditions. To address the potential non-integer multiple issues that may arise during bandwidth allocation, a rounding strategy is employed to ensure effective bandwidth allocation across physical links. Furthermore, a secondary, refined calibration of the target bandwidth is performed based on the total receive bandwidth of each switch, further improving bandwidth allocation efficiency.

[0072] It should be noted that the bandwidth allocation method of this application is suitable for scenarios such as super-node AI servers and high-performance computing clusters that require full interconnection of multiple processors, and has significant advantages in orthogonal architecture cabinets.

[0073] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0074] Figure 3 This is a flow chart of the first bandwidth allocation method provided in an embodiment of the present application.

[0075] like Figure 3 As shown, the method comprises the following steps:

[0076] Step 101: Determine the initial bandwidth allocated by each processor to each switch.

[0077] In the present application, each processor is any one of the multiple processors included in each computing node in the multiple computing nodes, each switch is any one of the multiple switches corresponding to a predetermined target number of switches, the multiple processors included in each computing node are respectively connected to each switch, and the target number of switches is an integer.

[0078] In some embodiments, in a fully interconnected bandwidth allocation scenario of an orthogonal architecture cabinet, in order to achieve full interconnection, it is necessary to determine the initial bandwidth allocated by each processor to each switch, so as to adjust and optimize the bandwidth based on the initial bandwidth to ensure efficient bandwidth allocation.

[0079] The initial bandwidth is the bandwidth value allocated by each processor to each switch. It is the starting value of bandwidth allocation and may not meet the requirements of physical link implementation.

[0080] Each compute node does not refer to a specific compute node, but rather to an arbitrary selection from the multiple compute nodes included in the server, representing each of these multiple compute nodes. In other words, the operations and features described below for each compute node apply to all compute nodes in the server.

[0081] Each processor is selected from the multiple processors included in each computing node and can represent every processor in each computing node. In addition, each computing node in the system is equipped with the same number of processors to ensure the uniformity of the system architecture and the balance of data processing.

[0082] Each switch is randomly selected from a plurality of switches corresponding to a predetermined target number of switches and may represent each of the plurality of switches. The number of target switches is an integer. It is understood that when the number of target switches is not an integer, it is necessary to round the number of target switches to ensure that it is an integer and meets the bandwidth allocation design requirements.

[0083] It is important to emphasize that this application focuses on fully interconnected scenarios under an orthogonal architecture. In this orthogonal architecture, multiple computing nodes and multiple switches are arranged in a front-to-back orientation. At the same time, multiple processors are orthogonally connected to each switch through orthogonal connectors, thereby building an efficient and stable data transmission channel.

[0084] Step 102: When the initial bandwidth is a non-integer multiple of the preset transmission channel bandwidth, the initial bandwidth is rounded up using a rounding strategy to determine a target bandwidth allocated by each processor to each switch.

[0085] In some embodiments, in actual bandwidth allocation, since the bandwidth specifications of processors and switches are fixed, the initial bandwidth may not be an integer multiple of the preset transmission channel bandwidth. If transmission is performed according to a bandwidth that is not an integer multiple, the physical link cannot be accurately implemented, and data transmission may be erroneous or blocked. Therefore, this application adopts a rounding strategy to process the initial bandwidth, adjusting it to an integer multiple of the preset transmission channel bandwidth to ensure the feasibility of the physical link and the stability of data transmission.

[0086] The preset transmission channel bandwidth refers to the minimum unit of bandwidth for the processor (GPU), namely, the fixed bandwidth value of each serializer / deserializer (SerDes) channel in the GPU. Since the GPU bandwidth is composed of several ports, each port contains multiple serializer / deserializer (SerDes) channels, and the bandwidth of each channel is a fixed value (such as the common 50G channel and 100G channel). Taking a 100G channel as an example, its bandwidth is 100G and the transmission rate is 100G / s. This value is the minimum indivisible unit of bandwidth allocation, namely, the preset transmission channel bandwidth of this application.

[0087] The rounding strategy of this application includes rounding up and rounding down, which are used to adjust the initial bandwidth to an integer multiple of the preset transmission channel bandwidth. The specific rounding method can be determined according to the actual needs and design requirements of the system.

[0088] The target bandwidth refers to a bandwidth value that meets the physical link implementation requirement and is allocated by the first processor to the first switch after the rounding policy is processed.

[0089] Step 103 : determining the total receiving bandwidth of each switch based on the target bandwidth, and adjusting the target bandwidth based on the total receiving bandwidth to control each processor to distribute the adjusted target bandwidth to each switch.

[0090] In some embodiments, after determining the target bandwidth, the total bandwidth received by each switch from all connected processors needs to be considered. If the total bandwidth received by each switch exceeds its own processing capacity, the target bandwidth needs to be adjusted. By adjusting the target bandwidth, the present application can ensure that each switch can operate within a reasonable load range, avoiding data congestion or loss, while ensuring unobstructed data transmission in the fully interconnected topology of the entire server.

[0091] The total bandwidth received by each switch under the target bandwidth is the total bandwidth received by each switch from all connected processors, assuming each processor allocates bandwidth to each switch according to the target bandwidth. This total bandwidth reflects the load of each switch under the current bandwidth allocation.

[0092] The adjusted target bandwidth is the new bandwidth value obtained by adjusting the original target bandwidth based on the total receive bandwidth of each switch.

[0093] In summary, this application achieves a reasonable allocation of bandwidth between processors and switches in an orthogonal architecture cabinet. By determining the initial bandwidth, rounding it up, and adjusting the target bandwidth, it solves the non-integer multiple problem and switch load imbalance that may arise in bandwidth allocation, ensuring the feasibility of physical links and non-blocking data transmission on servers. It can adapt to different processor bandwidth specifications, switch chip bandwidth, and various combinations of processor numbers and node numbers.

[0094] Figure 4 The flowchart of the second bandwidth allocation method proposed in this application is further shown. Figure 4 The embodiment shown in FIG. 1 further explains step 101. Figure 4 The following steps may be included.

[0095] Step 201 : Obtain a preset processor bandwidth of each processor, a preset switch chip bandwidth of each switch chip, a number of computing nodes of a plurality of computing nodes, and a number of processors of a plurality of processors, wherein each switch includes a preset number of switch chips.

[0096] In some embodiments, when determining the total node bandwidth of multiple computing nodes, it is necessary to first obtain the preset processor bandwidth of each processor, the preset switching chip bandwidth of each switching chip, the number of computing nodes of the multiple computing nodes, and the number of processors of the multiple processors.

[0097] The preset processor bandwidth represents the data transmission capability of each processor in each computing node, and in this application, it can be set to a.

[0098] Each switch includes a preset number of switch chips, where the preset number of switch chips can be 1, that is, each switch includes 1 switch chip. The preset switch chip bandwidth of each switch chip reflects the data processing capability of the switch chip, which can be set to b in this application.

[0099] The number of computing nodes of multiple computing nodes is the total number of computing nodes in the server, which can be set to Y in this application.

[0100] The processor quantity of the multiple processors is the number of processors included in each computing node, which can be set to X in this application.

[0101] Step 202 : Determine the total node bandwidth of the plurality of computing nodes based on the product of the preset processor bandwidth, the number of computing nodes, and the number of processors.

[0102] In some embodiments, the present application may calculate the total node bandwidth P of multiple computing nodes according to the formula P=a×X×Y based on the preset processor bandwidth, the number of computing nodes, and the number of processors.

[0103] The principle of the above formula is that the bandwidth of each processor is a, and each computing node has X processors, so the bandwidth of a computing node is a×X; there are Y computing nodes in the server, so the total node bandwidth of multiple computing nodes is a×X×Y.

[0104] For example, if the bandwidth of each GPU a = 100G, the number of GPUs in each computing node X = 8, and the number of computing nodes Y = 4, then the total node bandwidth of the computing nodes P = 100 × 8 × 4 = 3200G.

[0105] Step 203 : Determine the number of target switches connected to the plurality of computing nodes based on the total node bandwidth, the preset switch chip bandwidth, and the preset number of switch chips.

[0106] In some embodiments, after obtaining the total node bandwidth through step 202 above, the present application can determine the total number of switching chips based on the total node bandwidth and the preset switching chip bandwidth, and determine the target number of switches based on the total number of switching chips.

[0107] In other words, the present application can determine the total number of switch chips of the switch chip based on the ratio of the total node bandwidth to the preset switch chip bandwidth; based on the preset number of switch chips, the total number of switch chips is determined as the initial number of switches; when the initial number of switches is an integer, the initial number of switches is determined as the target number of switches; when the initial number of switches is a non-integer, the initial number of switches is rounded up to obtain the target number of switches.

[0108] Specifically, in a fully interconnected non-blocking topology, to ensure that data from all computing nodes can be transmitted smoothly through the switch, each switch is equipped with a preset number of switch chips. Here, the preset number of switch chips is 1, so the total bandwidth of the nodes should be equal to the total bandwidth of the switch chips. , that is, the total number of switching chips is a×X× .

[0109] For example, if the bandwidth of each switching chip b = 800G, and the total node bandwidth P calculated above = 3200G, then the total number of switching chips is 3200 ÷ 800 = 4.

[0110] Since each switch is equipped with one switching chip, the initial number of switches Z is equal to the total number of switching chips, that is, Z = a × X × .

[0111] However, in real-world applications, processor bandwidth and switch chip bandwidth are determined based on actual chip conditions, resulting in a certain degree of granularity. This means that the initial number of switches (total number of switch chips) calculated using the above formula may be a non-integer value. In reality, however, both the number of switch chips and the number of switches can only be integers, as there are no half-switch chips or half-switches. Therefore, when the initial number of switches Z is an integer, the target number of switches Z′ is directly determined as Z. When the initial number of switches Z is a non-integer, the initial number of switches Z must be rounded up to obtain the target number of switches Z′.

[0112] The purpose of rounding up is to ensure that the total bandwidth of the switch can meet the data transmission requirements of all computing nodes and avoid data congestion caused by insufficient bandwidth.

[0113] Rounding up symbol Indicates that the target number of switches Z′= , which means taking a value greater than a×X× The smallest integer.

[0114] For example, if the calculated initial number of switches Z = 3.2, then the target number of switches Z′ = 4 after rounding up.

[0115] Step 204 : Determine the initial bandwidth allocated by each processor to each switch based on the target number of switches.

[0116] In some embodiments, in an orthogonal architecture fully interconnected scenario, after determining the target number of switches, the present application needs to further clarify the initial bandwidth allocated by each processor to each switch.

[0117] Overall, the data generated by all processors in each compute node needs to be distributed to each target switch for transmission. To ensure that data is evenly and reasonably distributed to each switch, this application can calculate the first theoretical bandwidth allocated by each compute node to each switch based on the preset processor bandwidth, the number of processors, and the number of target switches.

[0118] Calculation formula and explanation: According to the formula Q=a× To calculate the bandwidth Q that each computing node provides to each switch, that is, the first theoretical bandwidth.

[0119] Where a×X represents the total bandwidth capacity of all processors in each computing node. By evenly distributing this total bandwidth to Z′ target switches, we can obtain the first theoretical bandwidth Q allocated by each computing node to each switch.

[0120] For example, if the preset bandwidth a of each processor is 100G, the number of processors in each computing node is X=8, and the number of target switches is Z′=4, then the first theoretical bandwidth Q allocated by each computing node to each switch is Q=100×8÷4=200G.

[0121] After determining that each computing node allocates the first theoretical bandwidth Q to each switch, since each computing node has X processors, in order to further clarify the bandwidth allocated by each processor to each switch, this application needs to evenly distribute the first theoretical bandwidth Q to each processor.

[0122] This application can be based on the formula Q′= = To calculate the bandwidth Q' that each processor provides to each switch, that is, the second theoretical bandwidth, and determine it as the initial bandwidth. Among them, Q' reflects the bandwidth allocated by a single processor to each switch. From another perspective, because Q = a × , substitute it into Q′= In this case, we can get Q′= For example, if the preset bandwidth a of each processor is 100G, the target number of switches Z′ is 4, and the first theoretical bandwidth Q is 200G, then the initial bandwidth Q′ allocated by each processor to each switch is 200÷8=100÷4=25G.

[0123] In summary, this application accurately determines the target number of switches based on parameters such as the processor bandwidth and number of computing nodes and the switching chip bandwidth, and rounds up to avoid the problem of bandwidth allocation not being able to be performed normally due to a decimal number of switches; on this basis, based on the target number of switches, the initial bandwidth allocated by computing nodes and processors to the switches is reasonably allocated, achieving a scientific distribution of bandwidth among various components, and building a stable, efficient and non-blocking data transmission foundation for the fully interconnected servers with orthogonal architecture, ensuring that the servers can run smoothly and perform at their best.

[0124] Figure 5 The flowchart of the third bandwidth allocation method proposed in this application is further shown. Figure 5 The embodiment shown in FIG. 1 further explains step 102. Figure 5 The following steps may be included.

[0125] Step 301 : When the initial bandwidth is a non-integer multiple of the preset transmission channel bandwidth, the initial bandwidth is rounded up using a rounding strategy to determine a target bandwidth allocated by each processor to each switch.

[0126] In some embodiments, in a bandwidth allocation scenario of a fully interconnected orthogonal architecture, when the initial bandwidth obtained above is not a non-integer multiple of the preset transmission channel bandwidth (using the serializer / deserializer (SerDes) channel bandwidth as the minimum unit, such as a 50G channel, a 100G channel, etc., and 100G is used as an example here, that is, the minimum bandwidth unit of the signal rate is 100G), since the serializer / deserializer (SerDes) channel cannot actually separate and allocate bandwidth using non-integer multiples of the minimum bandwidth unit, it is necessary to round the initial bandwidth and combine it with the preset processor bandwidth to finally determine the target bandwidth allocated by each processor to each switch. The target bandwidth of this application includes high-end bandwidth and low-end bandwidth.

[0127] First, this application requires the calculated initial bandwidth to be used as the center value for rounding. This initial bandwidth is a theoretical value based on the overall server parameters and the bandwidth allocation algorithm. However, in actual allocation, due to the limitation of the minimum bandwidth unit, it is necessary to round around this center value.

[0128] Based on the determined rounding center value and using the preset transmission channel bandwidth as the base bandwidth unit, the initial bandwidth is rounded up and down. This minimizes the difference in bandwidth load allocated by the processor to each switch and ensures that the target bandwidth is an integer multiple of the base bandwidth unit.

[0129] The result obtained by rounding down is recorded as the low-range bandwidth, and is represented by the symbol The result obtained by rounding up is recorded as the high-bit bandwidth, and the symbol Indicates. Among them, To round down, take the smallest integer less than this value. To round up the value, take the smallest integer greater than this value.

[0130] For example, if the initial bandwidth is 230G and the minimum bandwidth unit is 100G (i.e. the benchmark bandwidth unit in this application), then the low-end bandwidth =200G, high-end bandwidth =300G. Through this rounding operation, two possible bandwidth values ​​corresponding to the initial bandwidth are obtained, which provides a basis for the subsequent target bandwidth determination.

[0131] Since the computing nodes and switches are orthogonally connected, the bandwidth allocated by the processor on the computing node to the corresponding switch is equivalent to the bandwidth allocated to the corresponding orthogonal connector on the computing node, e.g. Figure 6 The diagram shows a schematic diagram of multiple processors in each computing node allocating target bandwidth to multiple switches respectively.

[0132] Furthermore, in order to meet the conditions of the preset processor bandwidth (assuming the preset bandwidth of each processor is a) and the target number of switches (assuming Z′), the present application can reasonably allocate high-end bandwidth and low-end bandwidth through a preset constraint algorithm (i.e., a linear equation of two variables), and determine, for a single computing node, the number of high-end switches that need to be allocated high-end bandwidth and the number of low-end switches that need to be allocated low-end bandwidth in the target number of switches.

[0133] Specifically, in this application, the number of switches with forward-rounded bandwidth (low-end bandwidth) can be set as m, and the number of switches with backward-rounded bandwidth (high-end bandwidth) can be set as n. Based on the total amount of bandwidth allocated and the total number of switches, the following set of linear equations can be established:

[0134] m× +n× =a, m+n=Z′.

[0135] The above equation states that the total bandwidth allocated to all switches must equal the target processor bandwidth. Specifically, the total bandwidth allocated by m switches with low-range bandwidth plus the total bandwidth allocated by n switches with high-range bandwidth equals the target processor bandwidth a for a single processor. m + n = Z′, meaning the sum of the number of switches with low-range bandwidth and the number of switches with high-range bandwidth equals the target number of switches, Z′.

[0136] By solving the above two-variable linear equations, we can get integer solutions for m and n. We need to ensure that the solution is an integer because the number of switches must be an integer. For example, if the low-range bandwidth =200G, high-end bandwidth =300G, the preset processor bandwidth a = 2600G, and the target number of switches Z′ = 10. Substituting into the equations yields: 200m + 300n = 2600, where m + n = 10. Solving this equation, transforming the second equation to m = 10 - n, and substituting into the first equation yields: 200 × (10 - n) + 300n = 2600, or 2000 - 200n + 300n = 2600, where 100n = 600. Solving for n = 6, we obtain m = 10 - 6 = 4.

[0137] Based on the values ​​of m and n obtained, the target bandwidth allocated by each processor to each switch in a single computing node is preliminarily determined. On the computing node, each processor allocates a smaller bandwidth to m orthogonal connectors (i.e., the corresponding switches), i.e. ; After multiplying by the number of processors X on the computing node, the bandwidth allocated to m switches by the computing node is X× At the same time, the processor on the computing node allocates a larger bandwidth to n orthogonal connectors (i.e., the corresponding switches), i.e. ; After multiplying by the number of processors X on the computing node, the bandwidth allocated to n switches by the computing node is X× .like Figure 7 The diagram shows that multiple processors in each computing node allocate low-range bandwidth and high-range bandwidth to multiple switches respectively.

[0138] In summary, when the initial bandwidth is not an integer multiple of the minimum bandwidth unit, this application reasonably and scientifically determines the target bandwidth allocated by each processor to each switch, ensuring the accuracy and effectiveness of bandwidth allocation, minimizing the difference in bandwidth load allocated by the processor to each switch, and thus optimizing the data transmission performance of the server.

[0139] Figure 8 The flowchart of the fourth bandwidth allocation method proposed in this application is further shown. Figure 8 The embodiment shown in FIG. 1 further explains step 103. Figure 8The following steps may be included.

[0140] Step 401 : Determine the total receiving bandwidth of each switch based on the target bandwidth allocated by each processor to each switch.

[0141] In some embodiments, in a bandwidth allocation scenario of an orthogonal architecture fully interconnected scenario, after completing the preliminary determination of the target bandwidth, the present application needs to further determine the total receiving bandwidth of each switch under the determined target bandwidth.

[0142] Specifically, in some embodiments, the total receiving bandwidth of each switch is determined based on the product of the target bandwidth, the number of processors, and the number of computing nodes. In the orthogonal cabinet scenario, assuming that the bandwidth allocated to each computing node has a larger bandwidth value (i.e., a high-end bandwidth), ), and since the processor on the computing node allocates a larger bandwidth to some switches, the total bandwidth received by the switch is equal to the maximum bandwidth allocated to each computing node multiplied by the number of computing nodes, which can be expressed as X×Y× Where X represents the number of processors (i.e., the number of processors in each compute node), and Y represents the number of compute nodes. This calculation method accurately calculates the total bandwidth data received by each switch from all relevant compute nodes under the target bandwidth setting.

[0143] Step 402 : Compare the total receiving bandwidth of each switch with the preset switching chip bandwidth. When the total receiving bandwidth is greater than the preset switching chip bandwidth, adjust the target bandwidth allocated by each processor to each switch.

[0144] In the present application, the total received bandwidth being greater than the preset switching chip bandwidth indicates that each processor in the plurality of computing nodes allocates the same target bandwidth to each switch.

[0145] In some embodiments, after obtaining the total bandwidth received by each switch, the present application needs to compare it with the preset switching chip bandwidth. In an orthogonal cabinet environment, it is necessary to confirm the maximum bandwidth received by the switch is X×Y× Whether the bandwidth limit of the switching chip is exceeded, you can ×X×Yb≤0, where b represents the preset switching chip bandwidth.

[0146] like If ×X×Yb≤0 holds, it means that the bandwidth allocated to the processors of all computing nodes on each switch chip does not exceed the bandwidth limit of the switch chip itself, and the current target bandwidth setting is feasible. If the equation does not hold, it means that the bandwidth of the processors allocated to each switch chip has exceeded the bandwidth limit of the switch chip itself. In this case, the target bandwidth needs to be adjusted to ensure normal operation of the system and avoid performance problems or failures caused by bandwidth excess.

[0147] When the target bandwidth needs to be adjusted, in some embodiments, the adjustment is made based on multiple parameters such as low-range bandwidth, high-range bandwidth, number of low-range switches, number of high-range switches, number of processors, and number of computing nodes.

[0148] First, because the bandwidth overlimit problem occurs under the initial target bandwidth setting, it is necessary to try different bandwidth allocation combinations to find a feasible solution that meets the bandwidth limitation conditions. The present application can first adjust the low-range bandwidth or high-range bandwidth allocated by each processor to each switch based on the low-range bandwidth, high-range bandwidth, number of processors, and number of computing nodes, and obtain multiple first target bandwidth combinations allocated by each computing node to each switch. Each first target bandwidth combination is composed of the number of processors that allocate low-range bandwidth and high-range bandwidth to each switch by multiple processors of each computing node, and the sum of the number of processors that allocate low-range bandwidth and high-range bandwidth to each switch by multiple processors of each computing node is the same as the preset number of processors.

[0149] After obtaining multiple first target bandwidth combinations, a first total receive bandwidth of each switch corresponding to each first target bandwidth combination can be further determined. Multiple combinations that meet a bandwidth restriction condition are determined from the multiple first target bandwidth combinations, and the multiple combinations are determined as multiple second target bandwidth combinations, where the bandwidth restriction condition is that the first total receive bandwidth is less than or equal to a preset switch chip bandwidth. The second target bandwidth combinations are combinations that meet the bandwidth restriction condition among the multiple first target bandwidth combinations.

[0150] Finally, after obtaining the second target bandwidth combination, the second total receive bandwidths are sorted based on the second total receive bandwidth of each switch corresponding to each second bandwidth combination, and the second bandwidth combination corresponding to the second total receive bandwidth with the target sorting position is determined as the third target bandwidth combination; based on the number of low-range switches, the number of high-range switches, and the third target bandwidth combination, the adjusted target bandwidth allocated by each processor to each switch is determined.

[0151] Among them, after obtaining the final third target bandwidth combination, the present application can specifically first determine the number of processors in each computing node in the third target bandwidth combination that allocate low-range bandwidth and high-range bandwidth to each switch; based on the number of processors in each computing node in each third target bandwidth combination that allocate low-range bandwidth and high-range bandwidth to each switch, combined with the constraints of the number of low-range switches and the number of high-range switches, determine the adjusted target bandwidth allocated by each processor to each switch from the third target bandwidth combination, so that the difference between multiple third receiving total bandwidths of multiple switches is less than or equal to a preset error range, and the third receiving total bandwidth is obtained by each switch based on the adjusted target bandwidth of each processor.

[0152] This application comprehensively considers factors such as the overall bandwidth allocation balance and performance optimization of the system through the above screening process to ensure that the screened target bandwidth can not only meet the bandwidth limitation of the switching chip, but also maximize the data transmission efficiency and stability of the system.

[0153] Specifically, in an optional embodiment of the present application, in the above-mentioned initial bandwidth allocation process, for each orthogonal connector (or corresponding switch), the original allocation method of each processor in the computing node is the same, for example, each processor allocates the same bandwidth to a certain orthogonal connector Z'. Bandwidth, which results in some orthogonal connectors having fixed bandwidth greater than others, easily leading to bandwidth overruns. To address this issue, the way each processor in a compute node is assigned to each orthogonal connector is adjusted so that the bandwidth values ​​from different processors on the orthogonal connector are different.

[0154] It is known that each processor has two bandwidth values ​​allocated, namely the smaller bandwidth (low-range bandwidth) and larger bandwidth (High-bit bandwidth). Therefore, each processor can be allocated a smaller bandwidth on each orthogonal connector. Or allocate larger bandwidth Since each computing node consists of X processors, there are multiple possible total bandwidths on each orthogonal connector. The minimum value is the smaller bandwidth Multiply by the total number of processors, that is X; maximum value is the larger bandwidth Multiply by the total number of processors, that is X; intermediate values ​​are smaller bandwidths Multiply by the number of partial processors plus the maximum bandwidth Multiply by several intermediate values ​​of the number of partial processors, e.g. + (X-1), 2+ (X-2) and so on. In this way, we get several suitable combinations of the number of pre-rounded processors and the number of post-rounded processors, and the total bandwidth of each orthogonal connector belongs to the set { X; + (X-1); 2+ (X-2);...; X}.

[0155] Because the original allocation method, each processor in each computing node allocates the same target bandwidth to each orthogonal connector (switch), which results in the processor of each computing node possibly allocating a larger bandwidth to a certain orthogonal connector. , then the total bandwidth allocated to the switch by all computing nodes will exceed the capacity limit of the switch chip.

[0156] Therefore, this application can allocate different target bandwidths to different processors of a computing node on a certain orthogonal connector (switch). That is, a processor in a computing node can allocate low-range bandwidth to a certain switch, while another processor in the computing node can allocate high-range bandwidth to the same switch. Therefore, on this orthogonal connector (switch), there are multiple combinations of X×low-range bandwidth, X-1×low-range bandwidth + 1×high-range bandwidth…X×high-range bandwidth. That is, the target bandwidth on each orthogonal connector is X; + (X-1); 2+ (X-2);...; (k)+ (Xk) There are k target bandwidth combination schemes in total. In this application, the multiple target bandwidth combinations initially obtained are used as multiple first target bandwidth combinations.

[0157] After obtaining multiple first target bandwidth combinations, the present application can calculate the first total receiving bandwidth (i.e., k+ (Xk)); Based on the first received total bandwidth, a first target bandwidth combination (i.e., a first received total bandwidth less than or equal to a preset switching chip bandwidth) is obtained from the first target bandwidth combination. (k)+ (Xk) < b) and determining the first target bandwidth combination obtained by screening as the second bandwidth combination of the present application. At this time, it can be determined that the second total receive bandwidth of each switch corresponding to the second bandwidth combination does not exceed the bandwidth limit of the switching chip.

[0158] Then, based on the second total receive bandwidth of each switch corresponding to each second bandwidth combination, the second total receive bandwidths are sorted, and the second bandwidth combination corresponding to the largest second total receive bandwidth (i.e., the first second total receive bandwidth in descending order or the last second total receive bandwidth in ascending order) is selected to be determined as the third target bandwidth combination of the present application; the number of processors in each computing node in the third target bandwidth combination that allocate low-range bandwidth and high-range bandwidth to each switch is determined; based on the number of processors in each computing node in the third target bandwidth combination that allocate low-range bandwidth and high-range bandwidth to each switch, combined with constraints on the number of low-range switches and the number of high-range switches, an adjusted target bandwidth allocated by each processor to each switch is determined from the third target bandwidth combination, so that the difference between multiple third total receive bandwidths of multiple switches is less than or equal to a preset error range (so that the target bandwidth allocated by each switch is as consistent as possible), and the third total receive bandwidth is obtained by each switch based on the adjusted target bandwidth of each processor.

[0159] In summary, this application calculates the total receiving bandwidth of each switch based on the target bandwidth, the number of processors, and the number of computing nodes, providing a key data benchmark for evaluating whether the system bandwidth allocation is reasonable; and compares the total receiving bandwidth with the preset switching chip bandwidth. If it exceeds the limit, the target bandwidth allocation method is innovatively adjusted based on multiple parameters to screen out the target bandwidth that meets the bandwidth limit, ultimately ensuring that the total receiving bandwidth of the first switch does not exceed the switching chip capacity, ensuring the rationality and stability of the server's bandwidth allocation, avoiding performance problems caused by bandwidth excess, and optimizing the overall data transmission performance.

[0160] For ease of understanding, this application is further explained with a specific example, as follows:

[0161] Assume that each compute node has 4 GPUs (4 processors), each GPU has a bandwidth of 4.8T (the default processor bandwidth is 4.8T), there are 18 compute nodes (the total number of compute nodes is 18), and each switch chip has a bandwidth of 25.6T (the default switch chip bandwidth is 25.6T).

[0162] The initial number of switches is calculated by the number of processors, the preset processor bandwidth, the number of computing nodes, and the switching chip bandwidth. =13.5, the initial number of switches is 13.5.

[0163] At this point, we determine that 13.5 is not an integer, so we round up the initial number of switches to get the target number of switches, which is ⌈13.5⌉, resulting in a target number of 14 switches.

[0164] Afterwards, the first theoretical bandwidth allocated by each computing node to each switch is determined by presetting the processor bandwidth, the number of processors, and the target number of switches, i.e. =1.37T. It can be seen that the first theoretical bandwidth is a non-integer multiple of the preset transmission channel bandwidth (0.1T (100G)).

[0165] Based on the first theoretical bandwidth and the number of processors, the second theoretical bandwidth provided by each GPU to each switch is determined, and the second theoretical bandwidth is used as the initial bandwidth. =0.34T, which shows that the initial bandwidth is also a non-integer multiple of the preset transmission channel bandwidth (0.1T (100G)).

[0166] According to the second theoretical bandwidth obtained above, with 0.34T as the center point and 0.1T as the reference bandwidth unit, we can get =0.3T, =0.4T. That is, a single GPU can allocate bandwidth to some switches at 0.3T and to other switches at 0.4T.

[0167] By establishing the quadratic equation m×0.3+n×0.4=4.8, m+n=14, we obtain m=8, n=6. This means that in each compute node, each GPU allocates 0.3T to 8 orthogonal connectors (i.e., 8 switches), and each GPU allocates 0.4T to 6 orthogonal connectors (i.e., 6 switches). Considering that the allocation method for each processor in each compute node is the same in the original case, each of the four processors in each compute node allocates the same 0.4T or 0.3T to a specific orthogonal connector, we can obtain the following: Figure 9 The figure shows a preliminary diagram of the target bandwidth allocated by each processor in a single computing node to each switch.

[0168] Reference Figure 9 In a single compute node, there are four GPUs (labeled A, B, C, and D). These GPUs will allocate target bandwidth to 14 orthogonal connectors (i.e., 18 switches) according to the following allocation rules:

[0169] Each GPU will allocate bandwidth in two ways: 1. Allocate 0.3T to 8 orthogonal connectors (corresponding to 8 switches), and 2. Allocate 0.4T to 6 orthogonal connectors (corresponding to 6 switches). Since the allocation method of each GPU in the computing node is consistent, the four GPUs will allocate bandwidth to the orthogonal connectors according to the same standard - either 0.4T or 0.3T. Based on this rule, we get Figure 9The figure shows a schematic diagram of the target bandwidth (0.4T or 0.3T) allocated by each GPU to each switch in a single compute node. Multiple GPUs (A, B, C, and D) are shown repeatedly in the figure to clearly illustrate the bandwidth allocation details of each GPU in the compute node.

[0170] Reference Figure 9 When the total receiving bandwidth of the switch is 0.4×4×18-25.6>0, it is determined that the bandwidth of the GPU allocated to the switch chip has exceeded the bandwidth limit of the switch chip itself.

[0171] Each orthogonal connector may have a combination of GPU allocation of 0.3T and GPU allocation of 0.4T. The bandwidth of each orthogonal connector ranges from 1.2T to 1.6T, namely, multiple first target bandwidth combinations of (0.3×4); (0.3×3+0.4); (0.3×2+0.4×2); (0.3+0.4×3); and (0.4×4). Among them, 1.2×18 = 21.6 < 25.6; 1.3×18 = 23.4 < 25.6; 1.4×18 = 25.2 < 25.6; 1.5×18 = 27 > 25.6; and 1.6×18 = 28.8 > 25.6.

[0172] Therefore, each orthogonal connector can only have the following situations: (0.3×4); (0.3×3+0.4); (0.3×2+0.4×2), that is, multiple second target bandwidth combinations are (0.3×4); (0.3×3+0.4); (0.3×2+0.4×2).

[0173] Further sorting is performed as (0.3×4); (0.3×3+0.4); (0.3×2+0.4×2), and (0.3×2+0.4×2) corresponding to the maximum value 25.2 is selected as the third target bandwidth combination.

[0174] According to the above m=8,n=6, in the computing node, each GPU allocates 0.3T to 8 orthogonal connectors (i.e., 8 switches); each GPU allocates 0.4T to 6 orthogonal connectors (i.e., 6 switches), and combined with the third target bandwidth combination of (0.3×2+0.4×2), adjust the target bandwidth allocated by each processor to each switch, thus obtaining the following: Figure 10 Figure 2 shows the target bandwidth allocated by each processor to each switch in a single compute node after target bandwidth adjustment. When adjusting the target bandwidth allocated by each processor to each switch, the difference between the multiple total receive bandwidths of the multiple switches must be less than or equal to a preset error range. This means that the target bandwidth allocated to each switch is as consistent as possible.

[0175] Reference Figure 10The figure clearly shows the target bandwidth (0.3T or 0.4T) allocated to each of the 14 orthogonal connectors (switches) for the four GPUs (A, B, C, and D) within a single compute node. This demonstrates the more balanced bandwidth allocation after adjustment, meeting the target combination and error requirements. Understandably, the repeated display of multiple GPUs (A, B, C, and D) in the figure represents the four GPUs in the same compute node. This is to clearly demonstrate the bandwidth allocation details for each GPU within that compute node.

[0176] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0177] The embodiment of the present application further provides a bandwidth allocation device 1100, Figure 11 A schematic diagram of a bandwidth allocation device provided in an embodiment of the present application is shown in FIG. Figure 11 Shown, including:

[0178] A determining unit 1110 is configured to determine an initial bandwidth allocated by each processor to each switch, where each processor is any one of multiple processors included in each of the multiple computing nodes, each switch is any one of multiple switches corresponding to a predetermined target number of switches, the multiple processors included in each computing node are connected to each switch, and the target number of switches is an integer.

[0179] A first adjusting unit 1120 is configured to round the initial bandwidth using a rounding strategy when the initial bandwidth is a non-integer multiple of the preset transmission channel bandwidth, and determine a target bandwidth allocated by each processor to each switch;

[0180] The second adjusting unit 1130 is configured to determine the total receiving bandwidth of each switch based on the target bandwidth, and adjust the target bandwidth based on the total receiving bandwidth to control each processor to allocate the adjusted target bandwidth to each switch.

[0181] Furthermore, in a possible implementation of an embodiment of the present application, the determination unit 1110 is used to: obtain the preset processor bandwidth of each processor, the preset switching chip bandwidth of each switching chip, the number of computing nodes of multiple computing nodes, and the number of processors of multiple processors, each switch including a preset number of switching chips; determine the total node bandwidth of multiple computing nodes based on the product of the preset processor bandwidth, the number of computing nodes, and the number of processors; determine the number of target switches connected to the multiple computing nodes based on the total node bandwidth, the preset switching chip bandwidth, and the preset number of switching chips; and determine the initial bandwidth allocated to each switch by each processor based on the target number of switches.

[0182] Furthermore, in a possible implementation of an embodiment of the present application, the determining unit 1110 is configured to: determine the total number of switch chips of the switch chip based on a ratio of the total node bandwidth to the preset switch chip bandwidth; determine the total number of switch chips as the initial number of switches based on the preset number of switch chips; when the initial number of switches is an integer, determine the initial number of switches as the target number of switches; and when the initial number of switches is a non-integer, round up the initial number of switches to obtain the target number of switches.

[0183] Furthermore, in a possible implementation of an embodiment of the present application, the determination unit 1110 is used to: determine a first theoretical bandwidth allocated by each computing node to each switch based on a preset processor bandwidth, the number of processors, and the target number of switches; determine a second theoretical bandwidth allocated by each processor to each switch based on the first theoretical bandwidth and the number of processors, and determine the second theoretical bandwidth as the initial bandwidth.

[0184] Furthermore, in a possible implementation of an embodiment of the present application, the first adjustment unit 1120 is configured to: when the initial bandwidth is a non-integer multiple of the preset transmission channel bandwidth, round the initial bandwidth using a rounding strategy, and determine the target bandwidth allocated by each processor to each switch in combination with the preset processor bandwidth.

[0185] Furthermore, in a possible implementation of an embodiment of the present application, the target bandwidth includes a high-range bandwidth and a low-range bandwidth, and the first adjustment unit 1120 is used to: round up or down the initial bandwidth to obtain a high-range bandwidth and a low-range bandwidth corresponding to the initial bandwidth; based on the high-range bandwidth, the low-range bandwidth and the preset processor bandwidth, determine the high-range bandwidth or low-range bandwidth allocated by each processor to each switch.

[0186] Furthermore, in a possible implementation of an embodiment of the present application, the first adjustment unit 1120 is used to: determine the initial bandwidth as a rounded center value; based on the rounded center value and with the preset transmission channel bandwidth as the reference bandwidth unit, round the initial bandwidth up or down to obtain a high-level bandwidth and a low-level bandwidth corresponding to the initial bandwidth, where the high-level bandwidth and the low-level bandwidth are integer multiples of the reference bandwidth unit.

[0187] Furthermore, in a possible implementation of the embodiment of the present application, the first adjustment unit 1120 is configured to: determine, based on the high-range bandwidth, the low-range bandwidth, and the preset processor bandwidth, in combination with a preset bandwidth constraint algorithm, the number of high-range switches in the target number of switches to which the high-range bandwidth needs to be allocated and the number of low-range switches to which the low-range bandwidth needs to be allocated; and determine, based on the number of high-range switches and the number of low-range switches, the high-range bandwidth or the low-range bandwidth to be allocated to each switch by each processor.

[0188] Furthermore, in a possible implementation of an embodiment of the present application, the second adjustment unit 1130 is configured to: determine the total receive bandwidth of each switch based on the target bandwidth allocated by each processor to each switch; compare the total receive bandwidth of each switch with the preset switching chip bandwidth; and adjust the target bandwidth allocated by each processor to each switch when the total receive bandwidth is greater than the preset switching chip bandwidth. The fact that the total receive bandwidth is greater than the preset switching chip bandwidth indicates that the target bandwidth allocated by each processor in multiple computing nodes to each switch is the same.

[0189] Furthermore, in a possible implementation of the embodiment of the present application, the second adjustment unit 1130 is configured to determine the total receiving bandwidth of each switch based on the product of the target bandwidth, the number of processors, and the number of computing nodes.

[0190] Furthermore, in a possible implementation of an embodiment of the present application, the second adjustment unit 1130 is used to adjust the target bandwidth allocated by each processor to each switch based on the low-range bandwidth, the high-range bandwidth, the number of low-range switches, the number of high-range switches, the number of processors, and the number of computing nodes.

[0191] Furthermore, in a possible implementation of an embodiment of the present application, the second adjustment unit 1130 is configured to: adjust the target bandwidth allocated by each processor to each switch based on the low-range bandwidth, the high-range bandwidth, the number of processors, and the number of computing nodes, to obtain multiple second target bandwidth combinations allocated by each computing node to each switch, where the second target bandwidth combination is a combination of multiple first target bandwidth combinations that meets the bandwidth restriction condition, each first target bandwidth combination is composed of the number of processors of each computing node that allocate low-range bandwidth and high-range bandwidth to each switch, and the sum of the number of processors of each computing node that allocate low-range bandwidth and high-range bandwidth to each switch is the same as the preset number of processors; and determine the adjusted target bandwidth allocated by each processor to each switch based on the number of low-range switches, the number of high-range switches, and the multiple second target bandwidth combinations.

[0192] Furthermore, in a possible implementation of an embodiment of the present application, the second adjustment unit 1130 is configured to: adjust the low-range bandwidth or high-range bandwidth allocated by each processor to each switch based on the low-range bandwidth, the high-range bandwidth, the number of processors, and the number of computing nodes, to obtain multiple first target bandwidth combinations allocated by each computing node to each switch; determine the first total receive bandwidth of each switch corresponding to each first target bandwidth combination; determine multiple combinations that meet a bandwidth restriction condition from the multiple first target bandwidth combinations, and determine the multiple combinations as multiple second target bandwidth combinations, where the bandwidth restriction condition is that the first total receive bandwidth is less than or equal to a preset switch chip bandwidth.

[0193] Furthermore, in a possible implementation of the embodiment of the present application, the second adjustment unit 1130 is configured to: sort the second total receive bandwidths based on the second total receive bandwidths of each switch corresponding to each second bandwidth combination, and determine the second bandwidth combination corresponding to the second total receive bandwidth with the sorting position being the target position as a third target bandwidth combination; and determine the adjusted target bandwidth allocated by each processor to each switch based on the number of low-range switches, the number of high-range switches, and the third target bandwidth combination.

[0194] Furthermore, in a possible implementation of the embodiment of the present application, the second adjustment unit 1130 is configured to: determine the number of processors in each computing node in the third target bandwidth combination that allocate the low-range bandwidth and the high-range bandwidth to each switch; and determine, based on the number of processors in each computing node in the third target bandwidth combination that allocate the low-range bandwidth and the high-range bandwidth to each switch, in combination with constraints on the number of low-range switches and the number of high-range switches, from the third target bandwidth combination, an adjusted target bandwidth to be allocated by each processor to each switch, such that a difference between multiple third total receive bandwidths of the multiple switches is less than or equal to a preset error range, and the third total receive bandwidth is obtained by each switch based on the adjusted target bandwidth of each processor.

[0195] Furthermore, in a possible implementation of the embodiment of the present application, multiple computing nodes and multiple switches are arranged in a front-to-back orientation, and multiple processors included in each computing node are orthogonally connected to each switch through an orthogonal connector.

[0196] For the description of the features in the embodiment corresponding to the bandwidth allocation device, reference can be made to the relevant description of the embodiment corresponding to the bandwidth allocation method, which will not be described in detail here.

[0197] An embodiment of the present application further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned bandwidth allocation method embodiments.

[0198] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any of the above-mentioned bandwidth allocation method embodiments when running.

[0199] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0200] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above-mentioned bandwidth allocation method embodiments are implemented.

[0201] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above-mentioned bandwidth allocation method embodiments are implemented.

[0202] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0203] The above describes in detail the bandwidth allocation method, electronic device, storage medium, and product provided by this application. This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is intended only to facilitate understanding of the method and core concepts of this application. It should be noted that those skilled in the art may make various improvements and modifications to this application without departing from the principles of this application, and such improvements and modifications fall within the scope of protection of the claims of this application.

Claims

1. A bandwidth allocation method, characterized in that: include: Determining an initial bandwidth allocated by each processor to each switch, where each processor is any one of multiple processors included in each computing node among the multiple computing nodes, and each switch is any one of multiple switches corresponding to a predetermined target number of switches, and the multiple processors included in each computing node are connected to each switch, respectively, where the target number of switches is an integer; When the initial bandwidth is a non-integer multiple of the preset transmission channel bandwidth, the initial bandwidth is rounded using a rounding strategy to determine a target bandwidth allocated by each processor to each switch; wherein the preset transmission channel bandwidth is the minimum unit of bandwidth of the processor; and the target bandwidth is the bandwidth value allocated by the processor to the switch that meets the physical link implementation requirements after the rounding strategy is processed; Based on the target bandwidth, a total receiving bandwidth of each switch is determined, and the target bandwidth is adjusted based on the total receiving bandwidth, so as to control each processor to allocate the adjusted target bandwidth to each switch.

2. The method according to claim 1, characterized in that Determining the initial bandwidth allocated by each processor to each switch includes: Obtaining a preset processor bandwidth of each processor, a preset switch chip bandwidth of each switch chip, a number of computing nodes of the plurality of computing nodes, and a number of processors of the plurality of processors, wherein each switch includes a preset number of switch chips; Determining a total node bandwidth of the plurality of computing nodes based on a product of the preset processor bandwidth, the number of computing nodes, and the number of processors; Determining the number of target switches connected to the plurality of computing nodes based on the total node bandwidth, the preset switch chip bandwidth, and the preset number of switch chips; Based on the target number of switches, an initial bandwidth allocated by each processor to each switch is determined.

3. The method according to claim 2, characterized in that The determining, based on the total node bandwidth, the preset switching chip bandwidth, and the preset number of switching chips, the target number of switches connected to the plurality of computing nodes includes: Determining the total number of switch chips of the switch chip based on the ratio of the total node bandwidth to the preset switch chip bandwidth; Based on the preset number of switch chips, determining the total number of switch chips as the initial number of switches; When the initial number of switches is an integer, determining the initial number of switches as the target number of switches; When the initial number of switches is not an integer, the initial number of switches is rounded up to obtain the target number of switches.

4. The method according to claim 2, characterized in that The determining, based on the target number of switches, the initial bandwidth allocated by each processor to each switch includes: Determining a first theoretical bandwidth allocated by each computing node to each switch based on the preset processor bandwidth, the number of processors, and the target number of switches; Based on the first theoretical bandwidth and the number of processors, a second theoretical bandwidth allocated by each processor to each switch is determined, and the second theoretical bandwidth is determined as the initial bandwidth.

5. The method according to claim 4, characterized in that When the initial bandwidth is a non-integer multiple of the preset transmission channel bandwidth, rounding the initial bandwidth using a rounding strategy to determine the target bandwidth allocated by each processor to each switch includes: When the initial bandwidth is a non-integer multiple of the preset transmission channel bandwidth, the initial bandwidth is rounded up using a rounding strategy, and the target bandwidth allocated by each processor to each switch is determined in combination with the preset processor bandwidth.

6. The method according to claim 5, characterized in that The target bandwidth includes a high-range bandwidth and a low-range bandwidth. The rounding-up process is performed on the initial bandwidth by a rounding-up strategy, and combined with the preset processor bandwidth, to determine the target bandwidth allocated by each processor to each switch. The process includes: Rounding up or down the initial bandwidth to obtain a high-range bandwidth and a low-range bandwidth corresponding to the initial bandwidth; Based on the high-range bandwidth, the low-range bandwidth, and the preset processor bandwidth, a high-range bandwidth or a low-range bandwidth allocated by each processor to each switch is determined.

7. The method according to claim 6, characterized in that The rounding up or down of the initial bandwidth to obtain the high-range bandwidth and the low-range bandwidth corresponding to the initial bandwidth includes: Determining the initial bandwidth as a rounded center value; According to the rounded center value and taking the preset transmission channel bandwidth as a reference bandwidth unit, the initial bandwidth is rounded up or down to obtain a high-level bandwidth and a low-level bandwidth corresponding to the initial bandwidth, and the high-level bandwidth and the low-level bandwidth are integer multiples of the reference bandwidth unit.

8. The method according to claim 6, characterized in that The determining, based on the high bandwidth, the low bandwidth, and the preset processor bandwidth, the high bandwidth or the low bandwidth allocated by each processor to each switch includes: Based on the high-range bandwidth, the low-range bandwidth, and the preset processor bandwidth, and in combination with a preset bandwidth constraint algorithm, determine the number of high-range switches to be allocated the high-range bandwidth and the number of low-range switches to be allocated the low-range bandwidth among the target number of switches; Based on the number of high-range switches and the number of low-range switches, a high-range bandwidth or a low-range bandwidth allocated by each processor to each switch is determined.

9. The method according to claim 8, characterized in that Determining the total receiving bandwidth of each switch based on the target bandwidth, and adjusting the target bandwidth based on the total receiving bandwidth includes: determining a total receive bandwidth of each switch based on a target bandwidth allocated by each processor to each switch; Comparing the total receiving bandwidth of each switch with the preset switching chip bandwidth respectively; When the total received bandwidth is greater than the preset switching chip bandwidth, the target bandwidth allocated by each processor to each switch is adjusted. The total received bandwidth being greater than the preset switching chip bandwidth indicates that each processor in the multiple computing nodes allocates the same target bandwidth to each switch.

10. The method according to claim 9, characterized in that Determining the total receive bandwidth of each switch based on the target bandwidth allocated by each processor to each switch includes: The total receiving bandwidth of each switch is determined based on the product of the target bandwidth, the number of processors, and the number of computing nodes.

11. The method according to claim 9, characterized in that The adjusting the target bandwidth allocated by each processor to each switch comprises: The target bandwidth allocated by each processor to each switch is adjusted based on the low-range bandwidth, the high-range bandwidth, the number of low-range switches, the number of high-range switches, the number of processors, and the number of computing nodes.

12. The method according to claim 11, characterized in that The adjusting the target bandwidth allocated by each processor to each switch based on the low-range bandwidth, the high-range bandwidth, the number of low-range switches, the number of high-range switches, the number of processors, and the number of computing nodes includes: Based on the low-range bandwidth, the high-range bandwidth, the number of processors, and the number of computing nodes, adjust the target bandwidth allocated by each processor to each switch to obtain multiple second target bandwidth combinations allocated by each computing node to each switch, where the second target bandwidth combinations are combinations that meet the bandwidth restriction condition among the multiple first target bandwidth combinations, each first target bandwidth combination is composed of the number of processors of each computing node that allocate the low-range bandwidth and the high-range bandwidth to each switch, and the total number of processors of each computing node that allocate the low-range bandwidth and the high-range bandwidth to each switch is the same as the preset number of processors; An adjusted target bandwidth allocated by each processor to each switch is determined based on the number of low-range switches, the number of high-range switches, and the plurality of second target bandwidth combinations.

13. The method according to claim 12, characterized in that The step of adjusting the target bandwidth allocated by each processor to each switch based on the low-range bandwidth, the high-range bandwidth, the number of processors, and the number of computing nodes to obtain multiple second target bandwidth combinations allocated by each computing node to each switch includes: Adjusting the low-range bandwidth or the high-range bandwidth allocated by each processor to each switch based on the low-range bandwidth, the high-range bandwidth, the number of processors, and the number of computing nodes, to obtain multiple first target bandwidth combinations allocated by each computing node to each switch; Determine a first total receive bandwidth of each switch corresponding to each first target bandwidth combination; A plurality of combinations satisfying the bandwidth restriction condition are determined from the plurality of first target bandwidth combinations, and the plurality of combinations are determined as the plurality of second target bandwidth combinations, wherein the bandwidth restriction condition is that the first receiving total bandwidth is less than or equal to the preset switching chip bandwidth.

14. The method according to claim 12, characterized in that The determining, based on the number of the low-range switches, the number of the high-range switches, and the plurality of second target bandwidth combinations, of the adjusted target bandwidth allocated by each processor to each switch comprises: sorting the second total received bandwidths based on the second total received bandwidths of each switch corresponding to each second bandwidth combination, and determining the second bandwidth combination corresponding to the second total received bandwidth having a sorting position as a target position as a third target bandwidth combination; An adjusted target bandwidth allocated by each processor to each switch is determined based on the number of low-range switches, the number of high-range switches, and the third target bandwidth combination.

15. The method according to claim 14, characterized in that The determining, based on the number of low-range switches, the number of high-range switches, and the third target bandwidth combination, the adjusted target bandwidth allocated by each processor to each switch includes: Determining the number of processors in each computing node in the third target bandwidth combination that allocate low-range bandwidth and high-range bandwidth to each switch; Based on the number of processors in each computing node in the third target bandwidth combination that allocate low-range bandwidth and high-range bandwidth to each switch, and in combination with constraints on the number of low-range switches and the number of high-range switches, an adjusted target bandwidth allocated by each processor to each switch is determined from the third target bandwidth combination, such that a difference between multiple third total receive bandwidths of the multiple switches is less than or equal to a preset error range, the third total receive bandwidth being obtained by each switch based on the adjusted target bandwidth of each processor.

16. The method according to claim 1, wherein The multiple computing nodes and the multiple switches are arranged in a front-to-back orientation, and the multiple processors included in each computing node are orthogonally connected to each switch via an orthogonal connector.

17. A bandwidth allocation device, characterized in that: include: a determining unit, configured to determine an initial bandwidth allocated by each processor to each switch, wherein each processor is any one of a plurality of processors included in each computing node among the plurality of computing nodes, and each switch is any one of a plurality of switches corresponding to a predetermined target number of switches, wherein the plurality of processors included in each computing node are connected to each switch, respectively, and the target number of switches is an integer; a first adjustment unit configured to, when the initial bandwidth is a non-integer multiple of a preset transmission channel bandwidth, round the initial bandwidth using a rounding strategy to determine a target bandwidth allocated by each processor to each switch; wherein the preset transmission channel bandwidth is a minimum unit of bandwidth for the processor; and the target bandwidth is a bandwidth value allocated by the processor to the switch that meets physical link implementation requirements after the rounding strategy is applied; The second adjustment unit is configured to determine a total receiving bandwidth of each switch based on the target bandwidth, and adjust the target bandwidth based on the total receiving bandwidth to control each processor to allocate the adjusted target bandwidth to each switch.

18. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the bandwidth allocation method according to any one of claims 1 to 16.

19. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable the computer to execute the bandwidth allocation method according to any one of claims 1-16.

20. A computer program product, characterized in that The method comprises a computer program, which, when executed by a processor, implements the bandwidth allocation method according to any one of claims 1 to 16.

Citation Information

Patent Citations

  • Communication method and device

    CN113543229A

  • Bandwidth resource allocation method and device and nonvolatile storage medium

    CN118474801A