Bandwidth scheduling method and device
By setting up a token bucket on the outgoing port of the AI training network device, scheduling traffic with residual tokens and preset thresholds, dynamically adjusting bandwidth, solving the congestion problem caused by insufficient bandwidth, improving data transmission efficiency, and meeting the low latency and low packet loss requirements of AI training networks.
Patent Information
- Application Number
- CN202410178206.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-08
- Publication Date
- 2025-08-08
AI Technical Summary
In AI training networks, traffic congestion caused by insufficient bandwidth affects the data transmission efficiency between network devices.
By setting the token bucket on the outgoing port of the network device, scheduling traffic using the remaining tokens and preset thresholds in the token bucket, dynamically adjusting the bandwidth of the incoming port to avoid traffic congestion on the outgoing port.
It effectively avoids traffic congestion on the outgoing port, improves the data transmission efficiency between network devices, and meets the low latency and low packet loss requirements of AI-trained networks.
Smart Images

Figure CN120455377A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of communications, and in particular to a bandwidth scheduling method and device. Background Art
[0002] With the development of technology, more and more data is transmitted on the network, and accordingly, the demand for network bandwidth is also increasing. In some scenarios, traffic congestion occurs due to insufficient bandwidth.
[0003] As a specific example, in an artificial intelligence (AI) training network, multiple network devices can be trained in parallel to train the AI model. In this scenario, AI training data needs to be transmitted between network devices in the AI training network. For example, network device A and network device B each send their own training data to network device C, so that network device C can send the received training data to the AI training device for training. In this scenario, if the bandwidth of the egress port through which network device C sends training data to the AI training device is insufficient, traffic congestion will occur on that egress port.
[0004] Therefore, there is an urgent need for a solution that can solve or partially solve the above problems. Summary of the Invention
[0005] The embodiments of the present application provide a bandwidth scheduling method and apparatus, which can effectively avoid traffic congestion.
[0006] In a first aspect, an embodiment of the present application provides a bandwidth scheduling method that can be applied to a first network device, the first network device including an egress port. The traffic sent by the egress port to an AI training device connected to the egress port includes AI training data. The egress port has a corresponding token bucket, which is used to schedule traffic on the egress port. The first network device can determine a first remaining token in the token bucket, where the first remaining token is an unused token in the token bucket. Taking into account the unused tokens in the token bucket, the remaining bandwidth of the egress port can be represented. Accordingly, the quantitative relationship between the first remaining tokens and a preset threshold can represent the degree of traffic congestion at the egress port. Therefore, the first network device can determine a first bandwidth adjustment value corresponding to the preset threshold based on the first remaining tokens and the preset threshold. The first bandwidth adjustment value is used to adjust the traffic on an ingress port of the first network device. The ingress port is connected to a second network device. After determining the first bandwidth adjustment value, the first network device can send the first bandwidth adjustment value to the second network device, so that the source network device can determine the bandwidth for sending traffic to the ingress port based on the first bandwidth adjustment value. It can be seen that in the embodiment of the present application, the first network device can adjust the bandwidth of the traffic sent by the second network device to the ingress port according to the traffic congestion level of the egress port, thereby avoiding traffic congestion of the egress port as much as possible.
[0007] In one possible implementation, the second network device is the aforementioned source network device. In this case, after the second network device receives the first bandwidth adjustment value sent by the first network device, the second network device can adjust the bandwidth of the traffic sent to the first network device based on the first bandwidth adjustment value.
[0008] In one possible implementation, a third network device can send traffic to a first network device via a second network device. In this case, the third network device is the source network device. After receiving the first bandwidth adjustment value sent by the first network device, the second network device can further send the first bandwidth adjustment value to the third network device, so that the third network device can adjust the bandwidth of the traffic it sends to the first network device based on the first bandwidth adjustment value. It is not difficult to understand that when the third network device adjusts its bandwidth to the first network device, the bandwidth of the traffic sent by the second network device to the ingress port of the first network device is also adjusted accordingly.
[0009] In one possible implementation, the first network device determines a first bandwidth adjustment value corresponding to the preset threshold based on the first remaining tokens and the preset threshold. In a specific implementation, the first bandwidth adjustment weight corresponding to the preset threshold can be first determined based on the first remaining tokens and the preset threshold. Furthermore, the first bandwidth adjustment value can be determined based on the first bandwidth adjustment weight. In this manner, the first network device can obtain the first bandwidth adjustment value based on the first remaining tokens and the preset threshold, and then send the first bandwidth adjustment value to the second network device, so that the source network device can determine the bandwidth for sending traffic to the ingress port based on the first bandwidth adjustment value.
[0010] In one possible implementation, when the first network device determines the first bandwidth adjustment value based on the first bandwidth adjustment weight, it may directly determine the first bandwidth adjustment value using the first bandwidth adjustment weight. In other words, the first network device may directly send the first bandwidth adjustment weight to the second network device, so that the source network device can determine the bandwidth for sending traffic to the ingress port based on the first bandwidth adjustment weight.
[0011] In one possible implementation, the first network device may determine the first bandwidth adjustment value based on the first bandwidth adjustment weight. In specific implementations, the first bandwidth adjustment value may be determined based on the bandwidth of the egress port and the first bandwidth adjustment weight. As an example, the first network device may allocate an initial bandwidth to the third network device based on the bandwidth of the egress port, and then multiply the initial bandwidth by the first bandwidth adjustment weight to obtain the first bandwidth adjustment value. In this case, the first bandwidth adjustment value may be the target bandwidth allocated by the first network device to the third network device. In this case, the source network device may send traffic to the ingress port based on the target bandwidth.
[0012] In one possible implementation, the aforementioned preset threshold may include at least one threshold, and the at least one threshold includes a first threshold. Each of the at least one threshold corresponds to one or more bandwidth adjustment weights. In other words, the first threshold corresponds to one or more bandwidth adjustment weights. The first network device may save a correspondence between each threshold and the bandwidth adjustment weight corresponding to each threshold. For this case, in one example, the first network device determines the first bandwidth adjustment weight corresponding to the preset threshold based on the first remaining token and the preset threshold. In a specific implementation, when the first remaining token is less than or equal to the first threshold, the bandwidth adjustment weight corresponding to the first threshold is obtained as the first bandwidth adjustment weight. That is: when the congestion level of the egress port reaches the congestion level indicated by the first threshold, the bandwidth adjustment weight corresponding to the first threshold is obtained as the first bandwidth adjustment weight.
[0013] In a possible implementation, when the first threshold has a unique bandwidth adjustment weight, the first network device may obtain the bandwidth adjustment weight uniquely corresponding to the first threshold as the first bandwidth adjustment weight.
[0014] In one possible implementation, when the first threshold has multiple bandwidth adjustment weights, each of the multiple bandwidth adjustment weights corresponds to a priority. In this case, the first network device may, when the first remaining token is less than or equal to the first threshold, determine the bandwidth adjustment weight corresponding to the first threshold and having the same priority as the third network device as the first bandwidth adjustment weight based on the priority of the third network device. In this case, for the source network device (i.e., the third network device) that sends traffic to the first network device, its corresponding bandwidth adjustment weight matches the priority of the source network device, and the bandwidth adjustment weights corresponding to source network devices of different priorities may be different. This enables differentiated scheduling of bandwidth for source network devices of different priorities to send traffic to the first network device.
[0015] In one possible implementation, the first network device may maintain one or more thresholds corresponding to each of a plurality of priorities, and each of the one or more thresholds corresponds to a bandwidth adjustment weight. In this case, for the convenience of description, the priority of the third network device is referred to as the target priority, and the first threshold may be one of the one or more thresholds corresponding to the target priority. In other words, in this scenario, the aforementioned preset threshold may be one or more thresholds corresponding to the target priority. In this case, since the first threshold uniquely corresponds to one bandwidth adjustment weight, the first network device may obtain the bandwidth adjustment weight uniquely corresponding to the first threshold as the first bandwidth adjustment weight.
[0016] In one possible implementation, in a scenario where a first network device maintains one or more thresholds corresponding to each of a plurality of priorities, in order to ensure as much as possible that a high-priority source network device has sufficient bandwidth for sending traffic to the first network device, thereby ensuring as much as possible that the traffic sent by the high-priority source network device has a lower transmission delay, assuming that the aforementioned multiple priorities include a first priority and a second priority, and that the first priority is higher than the second priority, for a first target threshold corresponding to the first priority and a second target threshold corresponding to the second priority, if the first target threshold is greater than the second target threshold, then the bandwidth adjustment weight corresponding to the first target threshold is greater than or equal to the bandwidth adjustment weight corresponding to the second target threshold.
[0017] In one possible implementation, at least one of the preset thresholds includes a second threshold, the second threshold being less than the first threshold, and the first network device may further determine a second remaining token in the token bucket; when the second remaining token is less than or equal to the second threshold, the first network device obtains a second bandwidth adjustment weight corresponding to the second threshold, the second threshold being less than the first threshold, and the second bandwidth adjustment weight being less than the first bandwidth adjustment weight; the first network device determines a second bandwidth adjustment value based on the second bandwidth adjustment weight, and sends the second bandwidth adjustment value to the second network device. In this manner, when the congestion level of the egress port tends to increase, the bandwidth of the traffic sent by the source network device to the first network device can be further reduced, thereby effectively alleviating the congestion level of the egress port.
[0018] In one possible implementation, at least one of the preset thresholds includes a third threshold, where the third threshold is greater than the first threshold. The first network device may also determine a third remaining token in the token bucket. If the third remaining token is greater than the first threshold and less than or equal to the third threshold, the first network device obtains a third bandwidth adjustment weight corresponding to the third threshold, where the third bandwidth adjustment weight is greater than the first bandwidth adjustment weight. The first network device determines a third bandwidth adjustment value based on the third bandwidth adjustment weight and sends the third bandwidth adjustment value to the second network device. In this manner, when the congestion level of the egress port is alleviated, the bandwidth of the traffic sent from the source network device to the first network device can be further increased, thereby rationally utilizing the bandwidth resources of the egress port.
[0019] In one possible implementation, the egress port mentioned in the embodiments of the present application can be a separate egress port, for example, a separate physical port. That is, the aforementioned token bucket can be a physical port-specific token bucket. In other words, using the solution of the embodiments of the present application, a physical port-specific token bucket can be used to schedule traffic at the ingress port, thereby minimizing traffic congestion at the egress port.
[0020] In one possible implementation, the egress ports mentioned in the embodiments of this application may be egress port groups. The egress port groups mentioned herein may include at least one individual egress port, i.e., an egress port group is a group consisting of at least one individual egress port. In some scenarios, an egress port group may also be referred to as a "plane." In other words, using the solutions of the embodiments of this application, a token bucket in the plane dimension can be used to schedule traffic on the ingress ports, thereby minimizing congestion in the plane.
[0021] In one possible implementation, the egress port mentioned in the embodiments of this application can be a chip comprising at least one egress port group. In other words, the aforementioned token bucket can be a chip-specific token bucket. In other words, using the solution of the embodiments of this application, a chip-specific token bucket can be used to schedule traffic on the ingress port, thereby minimizing traffic congestion on the chip.
[0022] In a possible implementation, the AI training device may be a device that is independent of the first network device and has AI training capabilities.
[0023] In one possible implementation, the AI training device may be an NPU or GPU included in the first network device. For example, in an AI training network, if the first network device corresponds to a leaf device, the AI training device may be an NPU or GPU included in the leaf device.
[0024] In a second aspect, an embodiment of the present application provides a bandwidth scheduling device, which is applied to a first network device, and the device includes: a processing unit, used to determine a first remaining token in a token bucket, the token bucket is used to schedule the traffic of the egress port of the first network device, the first remaining token is an unused token in the token bucket, and the traffic sent by the egress port to the artificial intelligence AI training device connected to the egress port includes AI training data; based on the first remaining token and a preset threshold, a first bandwidth adjustment value corresponding to the preset threshold is determined, and the first bandwidth adjustment value is used to adjust the traffic of the ingress port of the first network device; a sending unit, used to send the first bandwidth adjustment value to a second network device, and the second network device is the network device connected to the ingress port.
[0025] In a possible implementation, the processing unit is configured to: determine a first bandwidth adjustment weight corresponding to the preset threshold according to the first remaining tokens and the preset threshold; and determine the first bandwidth adjustment value according to the first bandwidth adjustment weight.
[0026] In a possible implementation manner, determining the first bandwidth adjustment value according to the first bandwidth adjustment weight includes: determining the first bandwidth adjustment weight as the first bandwidth adjustment value.
[0027] In a possible implementation, determining the first bandwidth adjustment value according to the first bandwidth adjustment weight includes: determining the first bandwidth adjustment value according to the bandwidth of the egress port and the first bandwidth adjustment weight.
[0028] In one possible implementation, the preset threshold includes at least one threshold, the at least one threshold includes a first threshold, and determining the first bandwidth adjustment weight corresponding to the preset threshold based on the first remaining token and the preset threshold includes: when the first remaining token is less than or equal to the first threshold, obtaining the bandwidth adjustment weight corresponding to the first threshold as the first bandwidth adjustment weight.
[0029] In a possible implementation, the first threshold corresponds to a bandwidth adjustment weight.
[0030] In one possible implementation, the first threshold corresponds to multiple bandwidth adjustment weights, each of the multiple bandwidth adjustment weights corresponds to a priority, and obtaining the bandwidth adjustment weight corresponding to the first threshold as the first bandwidth adjustment weight includes: according to the priority of a third network device, determining a bandwidth adjustment weight among the multiple bandwidth adjustment weights that has the same priority as the priority of the third network device as the first bandwidth adjustment weight, wherein the third network device is the second network device, or the third network device sends traffic to the first network device via the second network device.
[0031] In one possible implementation, the first network device maintains one or more thresholds corresponding to each of multiple priorities, each of the one or more thresholds corresponds to a bandwidth adjustment weight, and the first threshold is one of the one or more thresholds corresponding to the target priority among the multiple priorities, and the target priority is the priority of the third network device.
[0032] In one possible implementation, the multiple priorities include a first priority and a second priority, and the first priority is higher than the second priority. Then: if the first target threshold corresponding to the first priority is greater than or equal to the second target threshold corresponding to the second priority, then the bandwidth adjustment weight corresponding to the first target threshold is greater than or equal to the bandwidth adjustment weight corresponding to the second target threshold.
[0033] In one possible implementation, the multiple thresholds also include a second threshold, which is smaller than the first threshold; the processing unit is further used to determine a second remaining token in the token bucket; when the second remaining token is smaller than or equal to the second threshold, obtain a second bandwidth adjustment weight corresponding to the second threshold, the second threshold is smaller than the first threshold, and the second bandwidth adjustment weight is smaller than the first bandwidth adjustment weight; determine a second bandwidth adjustment value based on the second bandwidth adjustment weight; the sending unit is further used to send the second bandwidth adjustment value to the second network device.
[0034] In one possible implementation, the multiple thresholds also include a third threshold, and the third threshold is greater than the first threshold; the processing unit is further used to determine a third remaining token in the token bucket; when the third remaining token is greater than the first threshold and less than or equal to the third threshold, obtain a third bandwidth adjustment weight corresponding to the third threshold, and the third bandwidth adjustment weight is greater than the first bandwidth adjustment weight; determine a third bandwidth adjustment value based on the third bandwidth adjustment weight; and the sending unit is further used to send the third bandwidth adjustment value to the second network device.
[0035] In a possible implementation, the egress port is: a single egress port, an egress port group, or a chip including at least one egress port group, wherein: the egress port group includes at least one single egress port.
[0036] In a possible implementation, the AI training device connected to the output port includes: a graphics processor GPU of the first network device, or a neural network processor NPU of the first network device.
[0037] In a third aspect, an embodiment of the present application provides a device comprising: a processor and a memory; the memory is used to store instructions or computer programs; the processor is used to execute the instructions or computer programs and execute the method described in the first aspect and any one of the above first aspects.
[0038] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, including instructions or a computer program, which, when executed on a computer, enables the computer to execute the method described in the first aspect and any one of the above first aspects.
[0039] In a fifth aspect, an embodiment of the present application provides a computer program product comprising instructions or a computer program, which, when executed on a computer, enables the computer to execute the method described in the first aspect and any one of the above aspects. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0041] Figure 1a A schematic diagram of the structure of an AI training network provided in an embodiment of the present application;
[0042] Figure 1b A schematic diagram of a bandwidth allocation result provided in an embodiment of the present application;
[0043] Figure 2 A schematic diagram of a bandwidth scheduling method provided in an embodiment of the present application;
[0044] Figure 3 A flowchart of another bandwidth scheduling method provided in an embodiment of the present application;
[0045] Figure 4Flowcharts of two further bandwidth scheduling methods provided in embodiments of the present application;
[0046] Figure 5 A schematic structural diagram of a first network device provided in this application;
[0047] Figure 6 A schematic diagram of the structure of a bandwidth scheduling device provided in an embodiment of the present application;
[0048] Figure 7 A schematic diagram of the structure of a device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0049] The embodiments of the present application provide a bandwidth scheduling method and apparatus, which can effectively avoid traffic congestion.
[0050] Currently, AI technology has been widely used in various fields, and the key to AI technology is AI training. With the rise of AI technology and its increasing scale, the demand for large-scale AI training networks with high bandwidth, low latency, low jitter, and low packet loss has been raised. Specifically:
[0051] AI technology involves training large models, which is computationally intensive and requires thousands or even tens of thousands of graphics processing units (GPUs) or neural network processing units (NPUs) to perform parallel computations. Furthermore, data exchange between different GPUs / NPUs is necessary, creating the need for ultra-large-scale networking.
[0052] During large-scale model training, AI models are typically trained in parallel, which creates a large demand for communication between network devices. For example, for training large models with hundreds of billions of parameters, the amount of communication data generated by parallel training can reach gigabytes (GB). Consequently, AI training networks require high bandwidth.
[0053] Furthermore, the training process for large models involves serial communication and computation. For example, network devices A and B each send their own training data to network device C. After receiving the training data from network devices A and B, network device C sends the training data from network devices A and B to the AI training device, which then performs the next computation.
[0054] If the delay between network device A and the AI training device is large, the AI training device will not be able to receive the training data from network device A in time, which will reduce the efficiency of AI training. Similarly, if the network jitter is large, the aforementioned AI training device will not be able to receive the training data from a certain network device (such as network device A) in time, which will reduce the efficiency of AI training. In addition, if the data received by the AI training device is lost, the training data involved in the AI training will be incomplete, which will cause the AI training to be unable to be completed quickly, which will also reduce the efficiency of AI training. Therefore, the AI training network has the demand for low latency, low jitter and low packet loss.
[0055] Next, combine Figure 1a The AI training network shown here explains the relevant content of AI training. Figure 1a A schematic diagram of the structure of an AI training network provided in an embodiment of the present application.
[0056] like Figure 1a As shown, the AI training network includes multiple spine devices and multiple leaf devices. Leaf devices are devices involved in AI training. Specifically, leaf devices include GPUs and / or NPUs that can process AI training data.
[0057] At present, in order to improve the efficiency of AI training, multiple leaf devices can use parallel training to perform AI training. In a specific example, in a certain training stage, each leaf device in the multiple leaf devices can be trained based on their corresponding training data, and the training results obtained by training are sent to the target leaf device. Accordingly, the training results obtained by the aforementioned multiple leaf devices can be used as training data for the target leaf device, and the target leaf device can further perform AI training based on the training data. Specifically, the target leaf device can send the training results received from the aforementioned multiple leaf devices to its own GPU and / or NPU, so that the GPU and / or NPU can further perform AI training based on the training results.
[0058] Currently, data transmission between leaf devices can be relayed through the spine device. For example, if leaf device 1 wants to send its training results to leaf device 2, leaf device 1 can first send the training results to the spine device, and then the spine device can further send the training results to leaf device 2.
[0059] As mentioned earlier, AI training networks require low latency and low packet loss. However, congestion in AI training networks can lead to high latency and high packet loss. Therefore, preventing congestion in AI training networks is crucial.
[0060] Load balancing can be used to prevent traffic congestion on spine devices. This section does not provide a detailed description of load balancing technology.
[0061] As mentioned above, in an AI training network, there are scenarios where multiple leaf devices send their training results to the same target leaf device. In this scenario, traffic is easily congested at the egress port of the target leaf device. For this situation, the inventors of this application have discovered that a push-on-pull (POP) technique can be used to avoid traffic congestion at the egress port of the target leaf device.
[0062] POP is a scheduling algorithm that dynamically allocates bandwidth to the upstream based on the downstream egress bandwidth. Specifically, the upstream first pushes a request message to the downstream. The request message carries the upstream bandwidth scheduling request amount. The downstream distributes bandwidth resources to the upstream based on the bandwidth of the egress port and the upstream request information, thereby minimizing traffic congestion at the egress port. In a scenario where leaf device 1 and leaf device 2 both send traffic to the target leaf device, the upstream corresponds to leaf device 1 and leaf device 2, and the downstream corresponds to the target leaf device. That is, leaf device 1 and leaf device 2 can each send a request message to the target leaf device. Accordingly, the target leaf device allocates bandwidth resources to leaf device 1 and leaf device 2 based on the bandwidth of the egress port.
[0063] However, using POP technology cannot completely avoid congestion. This is because the round-trip time (RTT) between the downstream and different upstreams is different. Due to the influence of different RTTs of different source ends, there is still a certain amount of burst in the actual network. Now let's take the scenario where leaf device 1 and leaf device 2 both send traffic to the target leaf device as an example. Assuming that the bandwidth of the outbound port of the target leaf device is 10G, you can refer to Figure 1b To understand, Figure 1b A schematic diagram of a bandwidth allocation result provided by an embodiment of the present application. The bandwidth allocated by the target leaf device to leaf device 1 and leaf device 2 at different times is as follows: Figure 1b As shown in the upper half of the figure, the shaded part is the bandwidth allocated by the target leaf device to leaf device 1 at different times, and the unshaded part is the bandwidth allocated by the target leaf device to leaf device 2 at different times.
[0064] Since the RTT between leaf device 1 and the target leaf device is greater than the RTT between leaf device 2 and the target leaf device, the bandwidth of the traffic sent by leaf device 1 and leaf device 2 to the target leaf device at each time can be expressed as follows: Figure 1b As shown in the lower half of Figure 1b As shown, at time 2, time 3, and time 5, the total bandwidth of the traffic sent by leaf device 1 and leaf device 2 to the target leaf device is greater than the bandwidth of the egress port of the target leaf device, causing traffic congestion at the egress port of the target leaf device.
[0065] In order to solve the above problems, an embodiment of the present application provides a bandwidth scheduling method. Next, the bandwidth scheduling method provided by the embodiment of the present application is introduced in conjunction with the accompanying drawings.
[0066] See also Figure 2 , which is a flow chart of a bandwidth scheduling method provided in an embodiment of the present application.
[0067] Figure 2 The method shown can be applied to a first network device. In one example, the first network device can be a network device in an AI training network. In a specific example, when the architecture of the AI training network is Figure 1a In the network architecture shown, the first network device can be, for example, a leaf device in the AI training network.
[0068] In the introduction Figure 2 Before describing the method, a brief description of the scenario of an embodiment of the present application is first given. In an embodiment of the present application, a source network device may send traffic to a first network device. To avoid traffic congestion at the egress port of the first network device, the source network device may send a request message to the first network device requesting bandwidth resources before sending traffic to the first network device. After receiving the request message, the first network device may allocate bandwidth to the source network device and send the allocated bandwidth to the source network device, so that the source network device can send traffic to the first network device based on the bandwidth.
[0069] In one example, the source network device may directly send traffic to the first network device. In this case, the source network device may also directly send the aforementioned request information to the first network device.
[0070] In yet another example, the source network device may send traffic to the first network device through the second network device. In this case, the source network device may also send the aforementioned request information to the first network device through the second network device.
[0071] Figure 2 The method shown may include the following S101-S103.
[0072] S101: A first network device determines a first remaining token in a token bucket, where the token bucket is used to schedule traffic at an egress port of the first network device. The first remaining token is an unused token in the token bucket, and traffic sent from the egress port to an artificial intelligence (AI) training device connected to the egress port includes AI training data.
[0073] In an embodiment of the present application, the first network device includes an egress port, and the first network device is provided with a token bucket for the egress port, and the token bucket is used to schedule the traffic of the egress port. As a specific example, when traffic arrives at the egress port, if there are remaining tokens in the token bucket, the tokens in the token bucket can be deducted based on the arriving traffic, and forwarding resources can be allocated to the traffic, for example, the traffic is arranged to enter a corresponding queue. In addition, the first network device periodically performs a filling operation for the token bucket based on the bandwidth of the egress port and a preset acceleration ratio. The preset acceleration ratio is greater than 1, for example, the preset acceleration ratio can be a value between 1.01 and 1.10.
[0074] In one example, the egress port mentioned in the embodiment of the present application may be a separate egress port, for example, a separate physical port. In another example, the egress port mentioned in the embodiment of the present application may be an egress port group, and the egress port group mentioned here may include at least one separate egress port, that is, the egress port group is a combination of at least one separate egress port. In some scenarios, the egress port group may also be referred to as a "plane". In another example, the egress port mentioned in the embodiment of the present application may be a chip including at least one egress port group. That is to say, the aforementioned token bucket may be a token bucket of the physical port dimension, a token bucket of the plane dimension, or a token bucket of the chip dimension.
[0075] The first network device can obtain the first remaining token in the token bucket, where the first remaining token refers to the unused token in the token bucket. It is easy to understand that the unused token in the token bucket (e.g., the first remaining token) can be used to indicate the remaining bandwidth of the egress port.
[0076] In an embodiment of the present application, the egress port of the first network device is connected to the AI training device, and the traffic sent by the egress port to the AI training device includes AI training data. Accordingly, after receiving the AI training data, the AI training device can further perform AI training based on the AI training data.
[0077] The embodiment of the present application does not specifically limit the AI training device. In one example, the AI training device may be a device independent of the first network device and having AI training capabilities. In another example, the AI training device may be an NPU or GPU included in the first network device. For example, Figure 1a In the application scenario shown, the first network device corresponds to a leaf device, and the AI training device can be the NPU or GPU included in the leaf device.
[0078] S102: The first network device determines a first bandwidth adjustment value corresponding to the preset threshold according to the first remaining tokens and the preset threshold, where the first bandwidth adjustment value is used to adjust traffic at an ingress port of the first network device.
[0079] Because the first remaining token can represent the remaining bandwidth of the egress port, the remaining bandwidth of the egress port can, to a certain extent, represent the degree of traffic congestion at the egress port. For example, a large remaining bandwidth of the egress port indicates that the traffic at the egress port is not congested, or that the traffic at the egress port is slightly congested. For another example, a small remaining bandwidth of the egress port indicates that the traffic at the egress port is severely congested.
[0080] Therefore, based on the first remaining tokens and the preset threshold, the congestion level of traffic at the egress port can be determined. Specifically, based on the quantitative relationship between the first remaining tokens and the preset threshold, the congestion level of traffic at the egress port can be determined. The preset threshold, for example, may include at least one threshold, for example, the preset threshold may include the value of remaining tokens in the token bucket at different congestion levels. In the embodiment of the present application, the preset threshold may, for example, be a value set based on experience.
[0081] In one example, during the specific implementation of S102, the first network device may pre-store a correspondence between a preset threshold and the first bandwidth adjustment value. Therefore, the first network device may determine the first bandwidth adjustment value based on the preset threshold and the correspondence.
[0082] In another example, S102 may include the following steps A1-A2 during specific implementation.
[0083] Step A1: The first network device determines a first bandwidth adjustment weight corresponding to the preset threshold according to the first remaining tokens and the preset threshold.
[0084] In one example, the first network device may determine the first bandwidth adjustment weight corresponding to the preset threshold when the first remaining token and the preset threshold satisfy a certain quantitative relationship. As a specific example, the first network device may determine the first bandwidth adjustment weight corresponding to the preset threshold when the first remaining token is less than or equal to the preset threshold. The first network device may pre-store the correspondence between the preset threshold and the first bandwidth adjustment weight, and when the first remaining token and the preset threshold satisfy a certain quantitative relationship, the first bandwidth adjustment weight may be obtained based on the correspondence between the preset threshold and the first bandwidth adjustment weight and the preset threshold.
[0085] In another example, the preset threshold may include at least one threshold, and the at least one threshold includes a first threshold. Each of the at least one threshold corresponds to one or more bandwidth adjustment weights. In other words, the first threshold corresponds to one or more bandwidth adjustment weights. The first network device may save a correspondence between each threshold and the bandwidth adjustment weight corresponding to each threshold. In this case, when step A1 is specifically implemented, when the first remaining token is less than or equal to the first threshold, the bandwidth adjustment weight corresponding to the first threshold can be obtained as the first bandwidth adjustment weight.
[0086] As mentioned above, the first threshold corresponds to one or more bandwidth adjustment weights. In one example, when the first threshold has only one bandwidth adjustment weight, the first network device may obtain the bandwidth adjustment weight uniquely corresponding to the first threshold as the first bandwidth adjustment weight.
[0087] In another example, when the first threshold has multiple bandwidth adjustment weights, each of the multiple bandwidth adjustment weights corresponds to a priority. For example, the first threshold corresponds to bandwidth adjustment weight 1 and bandwidth adjustment weight 2, bandwidth adjustment weight 1 corresponds to the first priority, and bandwidth adjustment weight 2 corresponds to the second priority. For this case, when step A1 is specifically implemented, when the first remaining token is less than or equal to the first threshold, the bandwidth adjustment weight with the same priority as the third network device among the multiple bandwidth adjustment weights corresponding to the first threshold can be determined as the first bandwidth adjustment weight based on the priority of the third network device. For this case, for the source network device (i.e., the third network device) that sends traffic to the first network device, its corresponding bandwidth adjustment weight matches the priority of the source network device, and the bandwidth adjustment weights corresponding to source network devices of different priorities may be different. This achieves differentiated scheduling of bandwidth for source network devices of different priorities to send traffic to the first network device. In one example, if the first priority is higher than the second priority, the bandwidth adjustment weight corresponding to the first priority is greater than the bandwidth adjustment weight corresponding to the second priority. In this way, even if the output port of the first network device is congested, it is possible to ensure that the high-priority source network device has sufficient bandwidth for sending traffic to the first network device, thereby ensuring that the traffic sent by the high-priority source network device has a lower transmission delay.
[0088] Regarding the first threshold corresponding to multiple bandwidth adjustment weights, please refer to the following Table 1 for understanding.
[0089] Table 1
[0090]
[0091] The third network device mentioned here refers to the source network device that sends traffic to the first network device.
[0092] In a scenario where the source network device directly sends traffic to the first network device, the third network device may be the second network device. In this case, in one example, the first network device and the second network device may both be leaf devices.
[0093] In a scenario where a source network device sends traffic to a first network device via a second network device, the third network device may be a different device from the second network device. In this case, in one example, the first network device and the third network device may be two different leaf devices, and the second network device may be a spine device.
[0094] Regarding the priority of the source network device, it should be noted that, in one example, the first network device may be provided with request information storage queues corresponding to different priorities. In another example, the first network device may determine the priority of the queue storing request information sent by the third network device as the priority of the third network device.
[0095] In another example, the first network device may maintain one or more thresholds corresponding to each of a plurality of priorities, and each of the one or more thresholds corresponds to a bandwidth adjustment weight. In this case, for the convenience of description, the priority of the third network device is referred to as the target priority, and the first threshold may be one of the one or more thresholds corresponding to the target priority. In this scenario, the aforementioned preset threshold may be one or more thresholds corresponding to the target priority. In this case, since the first threshold is unique to one bandwidth adjustment weight, the first network device may obtain the bandwidth adjustment weight uniquely corresponding to the first threshold as the first bandwidth adjustment weight.
[0096] In a scenario where a first network device maintains one or more thresholds corresponding to each of multiple priorities, in order to ensure that a high-priority source network device has sufficient bandwidth to send traffic to the first network device, thereby ensuring that the traffic sent by the high-priority source network device has a low transmission latency, assuming that the multiple priorities include a first priority and a second priority, and that the first priority is higher than the second priority, for a first target threshold corresponding to the first priority and a second target threshold corresponding to the second priority, if the first target threshold is greater than the second target threshold, then the bandwidth adjustment weight corresponding to the first target threshold is greater than or equal to the bandwidth adjustment weight corresponding to the second target threshold. In a specific example, if the first target threshold is greater than the second target threshold, then the bandwidth adjustment weight corresponding to the first target threshold is greater than the bandwidth adjustment weight corresponding to the second target threshold. If the first target threshold is equal to the second target threshold, then the bandwidth adjustment weight corresponding to the first target threshold is greater than or equal to the bandwidth adjustment weight corresponding to the second target threshold. This is because the first congestion level of the egress port when the number of tokens in the token bucket drops to the first target threshold is lower than the second congestion level of the egress port when the number of tokens in the token bucket drops to the second target threshold. When the congestion level at the egress port of the first network device is low (corresponding to the first congestion level and the first target threshold), the bandwidth allocated to the source network device with a higher priority (first priority) should not be lower than the bandwidth allocated to the source network device with a lower priority (second priority) when the congestion level at the egress port is high (corresponding to the second congestion level and the second target threshold). In addition, if the first target threshold is lower than the second target threshold, the quantitative relationship between the bandwidth adjustment weight corresponding to the first target threshold and the bandwidth adjustment weight corresponding to the second target threshold is not specifically limited in this embodiment of the application.
[0097] The first target threshold is one of the multiple thresholds corresponding to the first priority, and the second target threshold is one of the multiple thresholds corresponding to the second priority. In one example, the first threshold may be the first target threshold; in another example, the first threshold may be the second target threshold.
[0098] Regarding the first target threshold, the second target threshold, the bandwidth adjustment weight corresponding to the first target threshold, and the bandwidth adjustment weight corresponding to the second target threshold, please first understand them in conjunction with Table 2 below.
[0099] Table 2
[0100]
[0101] As shown in Table 2:
[0102] When the first target threshold is 192KB, the second target threshold can be 128KB, 64KB, 0KB or -64KB. The first target threshold is greater than the second target threshold, and the bandwidth adjustment weight corresponding to the first target threshold is greater than the bandwidth weight corresponding to the second target threshold.
[0103] Similarly, when the first target threshold is 128KB, the second target threshold can be 128KB, 64KB, 0KB, or -64KB, the first target threshold is greater than the second target threshold, and the bandwidth adjustment weight corresponding to the first target threshold is greater than the bandwidth weight corresponding to the aforementioned second target threshold. When the first target threshold is 64KB, the second target threshold can be 64KB, 0KB, or -64KB, the first target threshold is greater than the second target threshold, and the bandwidth adjustment weight corresponding to the first target threshold is greater than the bandwidth weight corresponding to the aforementioned second target threshold. When the first target threshold is 0KB, the second target threshold can be 0KB or -64KB, the first target threshold is greater than the second target threshold, and the bandwidth adjustment weight corresponding to the first target threshold is greater than the bandwidth weight corresponding to the aforementioned second target threshold.
[0104] When one or more thresholds corresponding to the first priority are exactly the same as one or more thresholds corresponding to the second priority, it is equivalent to the case where the aforementioned one threshold corresponds to multiple bandwidth adjustment weights. Please refer to the description of Table 1 above and will not be repeated here.
[0105] In addition, in the embodiments of the present application, the number of thresholds corresponding to different priorities can be the same or different, and the embodiments of the present application do not specifically limit this. For example, the first priority corresponds to 4 thresholds, while the second priority corresponds to 3 thresholds; or for another example, the first priority and the second priority each correspond to 4 thresholds.
[0106] The embodiments of the present application do not strictly control the specific values of the thresholds corresponding to different priorities. For example, one or more thresholds corresponding to the first priority and one or more thresholds corresponding to the second priority can be exactly the same. For another example, one or more thresholds corresponding to the first priority and one or more thresholds corresponding to the second priority can be completely different. For another example, one or more thresholds corresponding to the first priority and one or more thresholds corresponding to the second priority can be different.
[0107] In addition, the aforementioned multiple priorities are not limited to the first priority and the second priority, and may also include a third priority or even a fourth priority. Here, the first priority and the second priority are taken as examples.
[0108] Step A2: The first network device determines the first bandwidth adjustment value according to the first bandwidth adjustment weight.
[0109] After determining the first bandwidth adjustment weight, the first network device may determine the first bandwidth adjustment value based on the first bandwidth adjustment weight. There are multiple implementations of the first network device determining the first bandwidth adjustment value based on the first bandwidth adjustment weight. Two possible implementations are described below.
[0110] In an example, the first network device may directly determine the first bandwidth adjustment weight as the first bandwidth adjustment value. That is, the first network device may directly determine the first bandwidth adjustment weight as the first bandwidth adjustment value.
[0111] In another example, the first network device may determine the first bandwidth adjustment value based on the bandwidth of the egress port and the first bandwidth adjustment weight. As an example, the first network device may allocate an initial bandwidth to the third network device based on the bandwidth of the egress port, and then multiply the initial bandwidth by the first bandwidth adjustment weight to obtain the first bandwidth adjustment value. In this case, the first bandwidth adjustment value may be the target bandwidth allocated by the first network device to the third network device. The method in which the first network device allocates the initial bandwidth to the third network device based on the bandwidth of the egress port may follow the method in which the first network device allocates bandwidth to the third network device in traditional technology, and will not be repeated here. In an embodiment of the present application, the first bandwidth adjustment value is used to adjust the traffic of the ingress port of the first network device. The ingress port mentioned here may be the ingress port corresponding to the aforementioned egress port. The traffic received by the first network device from the ingress port is forwarded to the AI training device through the egress port.
[0112] S103: The first network device sends the first bandwidth adjustment value to a second network device, where the second network device is the network device connected to the ingress port.
[0113] After determining the first bandwidth adjustment value, the first network device may send the first bandwidth adjustment value to the second network device.
[0114] As mentioned above, in one example, the second network device is a source network device. In this case, after the second network device receives the first bandwidth adjustment value sent by the first network device, the second network device can adjust the bandwidth of the traffic it sends to the first network device based on the first bandwidth adjustment value.
[0115] In another example, a third network device may send traffic to a first network device via a second network device. In this case, after receiving a first bandwidth adjustment value from the first network device, the second network device may further send the first bandwidth adjustment value to the third network device, so that the third network device can adjust the bandwidth of traffic sent to the first network device based on the first bandwidth adjustment value.
[0116] As mentioned above, the first bandwidth adjustment value may be the first bandwidth adjustment weight, or may be the target bandwidth allocated by the first network device to the third network device.
[0117] If the first bandwidth adjustment value is the first bandwidth adjustment weight, then the source network device (the second network device or the third network device) can obtain the bandwidth for sending traffic to the first network device based on the first bandwidth adjustment weight and its own required bandwidth. For example, the source network device can adjust the bandwidth for sending traffic to the first network device to the product of the first bandwidth adjustment weight and its own required bandwidth. If the first bandwidth adjustment value is the target bandwidth, the source network device can adjust the bandwidth for sending traffic to the first network device to the target bandwidth.
[0118] From the above description, it can be seen that, using the solution of the embodiment of the present application, the first network device can adjust the bandwidth of the traffic sent by the second network device to the ingress port according to the traffic congestion level of the egress port, thereby avoiding congestion of the egress port as much as possible.
[0119] In the embodiment of the present application, the first network device may periodically obtain the remaining tokens in the token bucket. In one example, after the first network device executes S103, it may also execute Figure 3 S201-S203 shown or Figure 4 S301-S303 shown. Figure 3 and Figure 4 Schematic diagram of the flow of two other bandwidth scheduling methods provided in the embodiments of the present application.
[0120] Figure 3 The method shown may include the following S201-S203.
[0121] S201: The first network device determines second remaining tokens in the token bucket.
[0122] In the embodiment of the present application, the second remaining token is a token that has not been used in the token bucket. In the embodiment of the present application, the second remaining token is smaller than the first remaining token.
[0123] S202: When the second remaining token is less than or equal to the second threshold, the first network device obtains a second bandwidth adjustment weight corresponding to the second threshold, the second threshold is less than the first threshold, and the second bandwidth adjustment weight is less than the first bandwidth adjustment weight.
[0124] In the embodiment of the present application, the preset threshold includes not only the first threshold but also a second threshold, the second threshold is smaller than the first threshold, and the first remaining token is smaller than or equal to the first threshold and larger than the second threshold.
[0125] In the embodiment of the present application, after executing S103 , the remaining tokens in the token bucket become fewer, which indicates that the congestion level of the egress port has not been alleviated and even has a tendency to worsen.
[0126] In this case, the first network device may obtain the second bandwidth adjustment weight corresponding to the second threshold when the second remaining token is less than or equal to the second threshold.
[0127] The implementation principle of the first network device obtaining the second bandwidth adjustment weight corresponding to the second threshold is the same as the implementation principle of the first network device obtaining the first bandwidth adjustment weight corresponding to the first threshold. Therefore, for the specific implementation of "the first network device obtaining the second bandwidth adjustment weight corresponding to the second threshold", please refer to the description of the first network device obtaining the first bandwidth adjustment weight in the previous article, and will not be repeated here.
[0128] In this embodiment of the present application, the second bandwidth adjustment weight is less than the first bandwidth adjustment weight. In this way, when the congestion level of the egress port tends to increase, the bandwidth of the traffic sent by the source network device to the first network device can be further reduced, thereby effectively alleviating the congestion level of the egress port.
[0129] S203: The first network device determines a second bandwidth adjustment value according to the second bandwidth adjustment weight, and sends the second bandwidth adjustment value to the second network device.
[0130] Similar to the function of the first bandwidth adjustment value, the second bandwidth adjustment value is also used to adjust the traffic of the ingress port of the first network device.
[0131] The implementation principle of the first network device determining the second bandwidth adjustment value based on the second bandwidth adjustment weight is the same as the implementation principle of the first network device determining the first bandwidth adjustment value based on the first bandwidth adjustment weight. Therefore, for the specific implementation of "the first network device determining the second bandwidth adjustment value based on the second bandwidth adjustment weight", please refer to the description of "the first network device determining the first bandwidth adjustment value based on the first bandwidth adjustment weight" above, and will not be repeated here.
[0132] In one example, after the first network device sends the second bandwidth adjustment value to the second network device, the second network device may adjust the bandwidth of traffic sent to the first network device based on the second bandwidth adjustment value. The specific implementation of "the second network device adjusting the bandwidth of traffic sent to the first network device based on the second bandwidth adjustment value" can be found in the description of "the second network device adjusting the bandwidth of traffic sent to the first network device based on the first bandwidth adjustment value" above, and is not repeated here.
[0133] In one example, after the first network device sends the second bandwidth adjustment value to the second network device, the second network device may send the second bandwidth adjustment value to a third network device, so that the third network device adjusts the bandwidth of traffic sent to the first network device based on the second bandwidth adjustment value. For the specific implementation of "the third network device adjusting the bandwidth of traffic sent to the first network device based on the second bandwidth adjustment value," please refer to the description of "the third network device adjusting the bandwidth of traffic sent to the first network device based on the first bandwidth adjustment value" above, and will not be repeated here.
[0134] Figure 4 The method shown may include the following S301-S303.
[0135] S301: The first network device determines a third remaining token in the token bucket.
[0136] In the embodiment of the present application, the third remaining token is a token that has not been used in the token bucket. In the embodiment of the present application, the third remaining token is greater than the first remaining token.
[0137] S302: When the third remaining token is greater than the first threshold and less than or equal to the third threshold, the first network device obtains a third bandwidth adjustment weight corresponding to the third threshold, where the third bandwidth adjustment weight is greater than the first bandwidth adjustment weight.
[0138] In the embodiment of the present application, the aforementioned preset threshold value includes not only the first threshold value but also a third threshold value, and the third threshold value is greater than the first threshold value.
[0139] In the embodiment of the present application, after executing S103 , the number of remaining tokens in the token bucket becomes more, which indicates that the congestion level of the egress port is alleviated.
[0140] In this case, the first network device may obtain a third bandwidth adjustment weight corresponding to the third threshold when the third remaining token is less than or equal to the third threshold and greater than the first threshold.
[0141] The implementation principle of the first network device obtaining the third bandwidth adjustment weight corresponding to the third threshold is the same as the implementation principle of the first network device obtaining the first bandwidth adjustment weight corresponding to the first threshold. Therefore, for the specific implementation of "the first network device obtaining the third bandwidth adjustment weight corresponding to the third threshold", please refer to the description of the first network device obtaining the first bandwidth adjustment weight in the previous section, and will not be repeated here.
[0142] In this embodiment of the present application, the third bandwidth adjustment weight is greater than the first bandwidth adjustment weight. In this way, when the congestion level of the egress port is alleviated, the bandwidth of the traffic sent from the source network device to the first network device can be further increased, thereby rationally utilizing the bandwidth resources of the egress port.
[0143] S303: The first network device determines a third bandwidth adjustment value according to the third bandwidth adjustment weight, and sends the third bandwidth adjustment value to the second network device.
[0144] Similar to the function of the first bandwidth adjustment value, the third bandwidth adjustment value is also used to adjust the traffic of the ingress port of the first network device.
[0145] The implementation principle of the first network device determining the third bandwidth adjustment value based on the third bandwidth adjustment weight is the same as the implementation principle of the first network device determining the first bandwidth adjustment value based on the first bandwidth adjustment weight. Therefore, for the specific implementation of "the first network device determining the third bandwidth adjustment value based on the third bandwidth adjustment weight", please refer to the description of "the first network device determining the first bandwidth adjustment value based on the first bandwidth adjustment weight" above, and will not be repeated here.
[0146] In one example, after the first network device sends the third bandwidth adjustment value to the second network device, the second network device may adjust the bandwidth of traffic sent to the first network device based on the third bandwidth adjustment value. The specific implementation of "the second network device adjusting the bandwidth of traffic sent to the first network device based on the third bandwidth adjustment value" can be found in the description of "the second network device adjusting the bandwidth of traffic sent to the first network device based on the first bandwidth adjustment value" above, and is not repeated here.
[0147] In one example, after the first network device sends the third bandwidth adjustment value to the second network device, the second network device may send the third bandwidth adjustment value to the third network device, so that the third network device adjusts the bandwidth of traffic sent to the first network device based on the third bandwidth adjustment value. For the specific implementation of "the third network device adjusting the bandwidth of traffic sent to the first network device based on the third bandwidth adjustment value," please refer to the description of "the third network device adjusting the bandwidth of traffic sent to the first network device based on the first bandwidth adjustment value" above, and will not be repeated here.
[0148] The above introduces the bandwidth scheduling method provided in the embodiment of the present application. Next, the solution provided in the embodiment of the present application is described in combination with specific scenarios.
[0149] See also Figure 5 , this figure is a structural diagram of a first network device provided in this application.
[0150] like Figure 5 As shown, the first network device includes multiple independent egress ports, multiple egress ports correspond to one egress port group, and multiple egress port groups ( Figure 5 Two output port groups are shown) corresponding to one chip.
[0151] In one example, a token bucket may be set for a separate egress port, for example, a token bucket may be set for one or more separate egress ports, that is, a token bucket of a separate egress port dimension may be set.
[0152] In another example, a token bucket may be set for an egress port group, for example, a token bucket may be set for each egress port group in at least one egress port group, that is, a token bucket of an egress port group dimension may be set.
[0153] In another example, a token bucket may be set for a chip, that is, a token bucket of chip dimension may be set.
[0154] Regardless of whether the token bucket is a token bucket of a single outbound port dimension, a token bucket of an outbound port group dimension, or a token bucket of a chip dimension, the bandwidth scheduling operation principle performed by the first communication device based on the token bucket is the same. The difference lies in the different threshold values corresponding to the token buckets of different dimensions. For example: Figure 5 The three values of TH3′, TH3 and TH3′ shown in the figure may be different; the three values of TH2′, TH2 and TH2″ may be different; the three values of TH1′, TH1 and TH1″ may be different; the three values of TH0′, TH0 and TH0″ may be different.
[0155] Next, taking the token bucket as the token bucket of the egress port group dimension as an example, the bandwidth scheduling method provided in the embodiment of the present application is introduced.
[0156] In one example, the preset threshold includes three thresholds, namely, threshold TH3, threshold TH2, threshold TH1, and threshold TH0. The embodiment of the present application does not specifically limit the specific values of TH3, TH2, TH1, and TH0. In one example, TH0 can be -64 kilobytes (KB), TH1 can be 0 KB, TH2 can be 64 KB, and TH3 can be 128 KB.
[0157] In a specific example, the egress port group has 16 egress ports, and the bandwidth of each egress port is 100G, so the bandwidth of the egress port group is 1600G.
[0158] Then, for the egress port group, a token bucket may be set for the egress port group, the bucket depth of the token bucket may be 512KB, and the filling rate of the token bucket may be 1600G multiplied by a preset acceleration ratio.
[0159] In some embodiments, the first network device stores the corresponding relationship shown in Table 3 below:
[0160] Table 3
[0161] Threshold Bandwidth adjustment weight TH3 (e.g. 128KB) 0.9 TH2 (e.g. 64KB) 0.8 TH1 (e.g. 0KB) 0.5 TH0 (e.g. - 64KB) 0.2
[0162] When traffic arrives at the egress port group, tokens in the token bucket may be deducted based on the length of the traffic.
[0163] In one example, if a traffic burst occurs in the egress port group, more tokens in the token bucket will be deducted. At this time, the remaining tokens in the token bucket are less than or equal to TH3. The first network device can then obtain a bandwidth adjustment weight of 0.9 based on the corresponding relationship shown in Table 3, and determine a bandwidth adjustment value for the source network device based on the bandwidth adjustment weight of 0.9.
[0164] After determining a bandwidth adjustment value for the source network device based on the bandwidth adjustment weight of 0.9, if the traffic burst of the outbound port group cannot be controlled, for example, if the remaining tokens in the token bucket are less than or equal to TH2, the first network device may obtain a bandwidth adjustment weight of 0.8 based on the correspondence shown in Table 3, and determine a bandwidth adjustment value for the source network device based on the bandwidth adjustment weight of 0.8.
[0165] After determining a bandwidth adjustment value for the source network device based on the bandwidth adjustment weight of 0.8, if the traffic burst of the outbound port group cannot be controlled, for example, if the remaining tokens in the token bucket are less than or equal to TH1, the first network device may obtain a bandwidth adjustment weight of 0.5 based on the correspondence shown in Table 3, and determine a bandwidth adjustment value for the source network device based on the bandwidth adjustment weight of 0.5.
[0166] After determining a bandwidth adjustment value for the source network device based on the bandwidth adjustment weight of 0.5, if the traffic burst of the outbound port group cannot be controlled, for example, if the remaining tokens in the token bucket are less than or equal to TH0, the first network device may obtain a bandwidth adjustment weight of 0.2 based on the correspondence shown in Table 3, and determine a bandwidth adjustment value for the source network device based on the bandwidth adjustment weight of 0.2.
[0167] After determining a bandwidth adjustment value for the source network device based on a bandwidth adjustment weight of 0.2, if traffic bursts of the outbound port group are controlled, for example, if the remaining tokens in the token bucket are less than or equal to TH1 and greater than TH0, the first network device may obtain a bandwidth adjustment weight of 0.5 based on the correspondence shown in Table 3, and determine a bandwidth adjustment value for the source network device based on the bandwidth adjustment weight of 0.5.
[0168] After determining a bandwidth adjustment value for the source network device based on a bandwidth adjustment weight of 0.5, if traffic bursts of the outbound port group are controlled, for example, if the remaining tokens in the token bucket are less than or equal to TH2 and greater than TH1, the first network device may obtain a bandwidth adjustment weight of 0.8 based on the correspondence shown in Table 3, and determine a bandwidth adjustment value for the source network device based on the bandwidth adjustment weight of 0.8.
[0169] After determining a bandwidth adjustment value for the source network device based on a bandwidth adjustment weight of 0.8, if traffic bursts of the outbound port group are controlled, for example, if the remaining tokens in the token bucket are less than or equal to TH3 and greater than TH2, the first network device may obtain a bandwidth adjustment weight of 0.9 based on the correspondence shown in Table 3, and determine a bandwidth adjustment value for the source network device based on the bandwidth adjustment weight of 0.9.
[0170] After determining the bandwidth adjustment value for the source network device based on the bandwidth adjustment weight of 0.9, if the traffic burst of the outbound port group is controlled, for example, if the remaining tokens in the token bucket are greater than TH3, the first network device may determine a bandwidth adjustment value for the source network device based on the bandwidth adjustment weight of 1.0.
[0171] In other embodiments, the first device maintains multiple thresholds corresponding to high priorities, and each of the multiple thresholds corresponding to high priorities corresponds to a bandwidth adjustment weight. In addition, the first device also maintains multiple thresholds corresponding to low priorities, and each of the multiple thresholds corresponding to low priorities corresponds to a bandwidth adjustment weight, and the multiple thresholds corresponding to high priorities are exactly the same as the multiple thresholds corresponding to low priorities. Then, the first network device stores the corresponding relationship shown in Table 4 below:
[0172] Table 4
[0173]
[0174] Bandwidth adjustment weight 1 is the bandwidth adjustment weight corresponding to the high priority, and bandwidth adjustment weight 2 is the bandwidth adjustment weight corresponding to the low priority. The high priority mentioned here can be equivalent to the first priority in the above embodiment, and the second priority mentioned here can be equivalent to the second priority in the above embodiment.
[0175] When traffic arrives at the egress port group, tokens in the token bucket may be deducted based on the length of the traffic.
[0176] In one example, if a traffic burst occurs in the outbound port group, more tokens in the token bucket will be deducted. At this time, the remaining tokens in the token bucket are less than or equal to TH3. The first network device can then obtain bandwidth adjustment weights of 1.0 and 0.9 based on the corresponding relationship shown in Table 4, and determine a bandwidth adjustment value for the source network device based on the bandwidth adjustment weights of 1.0 and 0.9. Specifically, a bandwidth adjustment value can be determined for a high-priority source network device based on a bandwidth adjustment weight of 1.0, and a bandwidth adjustment value can be determined for a low-priority source network device based on a bandwidth adjustment weight of 0.9.
[0177] After determining a bandwidth adjustment value for the source network device based on bandwidth adjustment weights 1.0 and 0.9, if the traffic burst of the outbound port group cannot be controlled, for example, if the remaining tokens in the token bucket are less than or equal to TH2, the first network device may obtain bandwidth adjustment weights 0.8 and 0.7 based on the corresponding relationship shown in Table 4, and determine a bandwidth adjustment value for the source network device based on the bandwidth adjustment weights 0.8 and 0.7. Specifically, a bandwidth adjustment value may be determined for a high-priority source network device based on a bandwidth adjustment weight of 0.8, and a bandwidth adjustment value may be determined for a low-priority source network device based on a bandwidth adjustment weight of 0.7.
[0178] After determining a bandwidth adjustment value for the source network device based on bandwidth adjustment weights of 0.8 and 0.7, if the traffic burst of the outbound port group cannot be controlled, for example, if the remaining tokens in the token bucket are less than or equal to TH1, the first network device can obtain bandwidth adjustment weights of 0.5 and 0.4 based on the corresponding relationship shown in Table 4, and determine a bandwidth adjustment value for the source network device based on the bandwidth adjustment weights of 0.5 and 0.4. Specifically, a bandwidth adjustment value can be determined for a high-priority source network device based on a bandwidth adjustment weight of 0.5, and a bandwidth adjustment value can be determined for a low-priority source network device based on a bandwidth adjustment weight of 0.4.
[0179] After determining a bandwidth adjustment value for the source network device based on bandwidth adjustment weights of 0.5 and 0.4, if the traffic burst of the outbound port group cannot be controlled, for example, if the remaining tokens in the token bucket are less than or equal to TH0, the first network device can obtain bandwidth adjustment weights of 0.2 and 0.1 based on the corresponding relationship shown in Table 4, and determine a bandwidth adjustment value for the source network device based on the bandwidth adjustment weights of 0.2 and 0.1. Specifically, a bandwidth adjustment value can be determined for a high-priority source network device based on a bandwidth adjustment weight of 0.2, and a bandwidth adjustment value can be determined for a low-priority source network device based on a bandwidth adjustment weight of 0.1.
[0180] After determining a bandwidth adjustment value for the source network device based on bandwidth adjustment weights 0.2 and 0.1, if traffic bursts of the outbound port group are controlled, for example, if the remaining tokens in the token bucket are less than or equal to TH1 and greater than TH0, the first network device can obtain bandwidth adjustment weights 0.5 and 0.4 based on the correspondence shown in Table 4, and determine a bandwidth adjustment value for the source network device based on the bandwidth adjustment weights 0.5 and 0.4.
[0181] After determining a bandwidth adjustment value for the source network device based on bandwidth adjustment weights 0.5 and 0.4, if traffic bursts of the outbound port group are controlled, for example, if the remaining tokens in the token bucket are less than or equal to TH2 and greater than TH1, the first network device can obtain bandwidth adjustment weights 0.8 and 0.7 based on the correspondence shown in Table 4, and determine a bandwidth adjustment value for the source network device based on the bandwidth adjustment weights 0.8 and 0.7.
[0182] After determining a bandwidth adjustment value for the source network device based on bandwidth adjustment weights 0.8 and 0.7, if traffic bursts of the outbound port group are controlled, for example, if the remaining tokens in the token bucket are less than or equal to TH3 and greater than TH2, the first network device can obtain bandwidth adjustment weights 1.0 and 0.9 based on the correspondence shown in Table 4, and determine a bandwidth adjustment value for the source network device based on the bandwidth adjustment weights 1.0 and 0.9.
[0183] After determining a bandwidth adjustment value for the source network device based on bandwidth adjustment weights of 1.0 and 0.9, if traffic bursts in the outbound port group are controlled, for example, if the remaining tokens in the token bucket are greater than TH3, the first network device can determine a bandwidth adjustment value for the source network device based on the bandwidth adjustment weight of 1.0. In other words, the bandwidth adjustment value is determined based on the bandwidth adjustment weight of 1.0 for both high-priority and low-priority source network devices.
[0184] It should be noted that in the above example, the multiple thresholds corresponding to the high priority and the multiple thresholds corresponding to the low priority are exactly the same. When the multiple thresholds corresponding to the high priority and the multiple thresholds corresponding to the low priority are not exactly the same (or completely different), the bandwidth scheduling principle of the first network device is similar, so it will not be repeated here.
[0185] It can be seen from the above examples that, by adopting the solution of the embodiment of the present application, when a burst occurs at the egress port, resulting in a certain degree of congestion at the egress port, the traffic from a certain source network device is not directly shut off. This is because once the traffic from a certain source network device is shut off, the AI training device will not be able to receive the traffic (training data) of the source network device in a timely manner, which will reduce the efficiency of AI training accordingly. Therefore, in an embodiment of the present application, the bandwidth of the source network device is reduced according to the bandwidth adjustment weight. In this way, the problem of directly shutting off the traffic from a certain source network device and affecting the efficiency of AI training in the AI training scenario can be avoided. That is, by adopting this solution, the efficiency of AI training can be guaranteed.
[0186] Based on the bandwidth scheduling method provided in the above embodiment, the embodiment of the present application also provides a corresponding device, which is described below in conjunction with the accompanying drawings.
[0187] See also Figure 6 , which is a structural diagram of a bandwidth scheduling device provided in an embodiment of the present application. Figure 6 The bandwidth scheduling apparatus can be applied to the first network device provided in the above embodiment, and is used to execute the bandwidth scheduling method provided in the above method embodiment and executed by the first network device.
[0188] like Figure 6 As shown, the bandwidth scheduling device 600 includes: a processing unit 601 and a sending unit 602.
[0189] Processing unit 601 is configured to determine a first remaining token in a token bucket, the token bucket being used to schedule traffic on an egress port of the first network device, the first remaining token being an unused token in the token bucket, and the traffic sent by the egress port to an artificial intelligence (AI) training device connected to the egress port including AI training data; and determine, based on the first remaining token and a preset threshold, a first bandwidth adjustment value corresponding to the preset threshold, the first bandwidth adjustment value being used to adjust traffic on an ingress port of the first network device;
[0190] The sending unit 602 is configured to send the first bandwidth adjustment value to a second network device, where the second network device is the network device connected to the ingress port.
[0191] In a possible implementation, the processing unit 601 is configured to: determine a first bandwidth adjustment weight corresponding to the preset threshold according to the first remaining tokens and the preset threshold; and determine the first bandwidth adjustment value according to the first bandwidth adjustment weight.
[0192] In a possible implementation manner, determining the first bandwidth adjustment value according to the first bandwidth adjustment weight includes: determining the first bandwidth adjustment weight as the first bandwidth adjustment value.
[0193] In a possible implementation, determining the first bandwidth adjustment value according to the first bandwidth adjustment weight includes: determining the first bandwidth adjustment value according to the bandwidth of the egress port and the first bandwidth adjustment weight.
[0194] In one possible implementation, the preset threshold includes at least one threshold, the at least one threshold includes a first threshold, and determining the first bandwidth adjustment weight corresponding to the preset threshold based on the first remaining token and the preset threshold includes: when the first remaining token is less than or equal to the first threshold, obtaining the bandwidth adjustment weight corresponding to the first threshold as the first bandwidth adjustment weight.
[0195] In a possible implementation, the first threshold corresponds to a bandwidth adjustment weight.
[0196] In one possible implementation, the first threshold corresponds to multiple bandwidth adjustment weights, each of the multiple bandwidth adjustment weights corresponds to a priority, and obtaining the bandwidth adjustment weight corresponding to the first threshold as the first bandwidth adjustment weight includes: according to the priority of a third network device, determining a bandwidth adjustment weight among the multiple bandwidth adjustment weights that has the same priority as the priority of the third network device as the first bandwidth adjustment weight, wherein the third network device is the second network device, or the third network device sends traffic to the first network device via the second network device.
[0197] In one possible implementation, the first network device maintains one or more thresholds corresponding to each of multiple priorities, each of the one or more thresholds corresponds to a bandwidth adjustment weight, and the first threshold is one of the one or more thresholds corresponding to the target priority among the multiple priorities, and the target priority is the priority of the third network device.
[0198] In one possible implementation, the multiple priorities include a first priority and a second priority, and the first priority is higher than the second priority. Then: if the first target threshold corresponding to the first priority is greater than or equal to the second target threshold corresponding to the second priority, then the bandwidth adjustment weight corresponding to the first target threshold is greater than or equal to the bandwidth adjustment weight corresponding to the second target threshold.
[0199] In one possible implementation, the multiple thresholds also include a second threshold, which is smaller than the first threshold; the processing unit 601 is further used to determine a second remaining token in the token bucket; when the second remaining token is smaller than or equal to the second threshold, obtain a second bandwidth adjustment weight corresponding to the second threshold, the second threshold is smaller than the first threshold, and the second bandwidth adjustment weight is smaller than the first bandwidth adjustment weight; determine a second bandwidth adjustment value based on the second bandwidth adjustment weight; the sending unit 602 is further used to send the second bandwidth adjustment value to the second network device.
[0200] In one possible implementation, the multiple thresholds also include a third threshold, and the third threshold is greater than the first threshold; the processing unit 601 is further used to determine a third remaining token in the token bucket; when the third remaining token is greater than the first threshold and less than or equal to the third threshold, obtain a third bandwidth adjustment weight corresponding to the third threshold, and the third bandwidth adjustment weight is greater than the first bandwidth adjustment weight; determine a third bandwidth adjustment value based on the third bandwidth adjustment weight; the sending unit 602 is further used to send the third bandwidth adjustment value to the second network device.
[0201] In a possible implementation, the egress port is: a single egress port, an egress port group, or a chip including at least one egress port group, wherein: the egress port group includes at least one single egress port.
[0202] In a possible implementation, the AI training device connected to the output port includes: a graphics processor GPU of the first network device, or a neural network processor NPU of the first network device.
[0203] It should be noted that the hardware structure of the bandwidth scheduling device 600 mentioned above can be as follows: Figure 7 The structure shown, Figure 7 A schematic diagram of the structure of a device provided in an embodiment of the present application.
[0204] See also Figure 7 As shown, the device 700 includes: a processor 710, a communication interface 720 and a memory 730. The number of the processor 710 in the device 700 can be one or more. Figure 7 In the embodiment of the present application, the processor 710, the communication interface 720 and the memory 730 may be connected via a bus system or other means, wherein: Figure 7 The connection via bus system 740 is taken as an example.
[0205] Processor 710 may be a central processing unit (CPU), an NP, or a combination of a CPU and an NP. Processor 710 may further include a hardware chip. The hardware chip may be an ASIC, a programmable logic device (PLD), or a combination thereof. The PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0206] Memory 730 may include volatile memory, such as random-access memory (RAM); non-volatile memory, such as flash memory, a hard disk drive (HDD), or a solid-state drive (SSD); or a combination of these types of memory. Memory 730 may, for example, store the correspondence between the aforementioned first threshold and the bandwidth adjustment weight corresponding to the first threshold.
[0207] Optionally, the memory 730 stores an operating system and programs, executable modules, or data structures, or subsets thereof, or extended sets thereof, wherein the programs may include various operating instructions for implementing various operations. The operating system may include various system programs for implementing various basic services and processing hardware-based tasks. The processor 710 may read the programs in the memory 730 to implement the bandwidth scheduling method provided in the embodiments of the present application.
[0208] The bus system 740 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus system 740 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0209] An embodiment of the present application provides a computer-readable storage medium, including instructions or a computer program, which, when executed on a computer, enables the computer to execute the method described in the above method embodiment.
[0210] An embodiment of the present application provides a computer program product comprising instructions or a computer program, which, when executed on a computer, enables the computer to execute the method described in the above method embodiment.
[0211] The terms "first," "second," "third," "fourth," and the like (if any) in the specification and claims of this application and in the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or sequential sequence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" and "having," and any variations thereof, are intended to cover non-exclusive inclusions, e.g., a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0212] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0213] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical business division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.
[0214] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0215] In addition, each business unit in each embodiment of the present application can be integrated into a processing unit, each unit can exist physically separately, or two or more units can be integrated into a single unit. The above-mentioned integrated units can be implemented in the form of hardware or software business units.
[0216] If the integrated unit is implemented in the form of a software business unit and sold or used as a separate product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0217] Those skilled in the art will appreciate that, in one or more of the above examples, the services described herein can be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these services can be stored on a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. Computer-readable media include computer storage media and communication media, including any medium that facilitates the transmission of computer programs from one location to another. Storage media can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0218] The above specific implementation methods further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above are only specific implementation methods of the present invention.
[0219] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A bandwidth scheduling method, characterized in that: The method comprises: The first network device determines a first remaining token in a token bucket, where the token bucket is used to schedule traffic on an egress port of the first network device, where the first remaining token is an unused token in the token bucket, and traffic sent by the egress port to an artificial intelligence (AI) training device connected to the egress port includes AI training data; The first network device determines, based on the first remaining token and a preset threshold, a first bandwidth adjustment value corresponding to the preset threshold, where the first bandwidth adjustment value is used to adjust traffic on an ingress port of the first network device; The first network device sends the first bandwidth adjustment value to a second network device, where the second network device is the network device connected to the ingress port.
2. The method according to claim 1, characterized in that The first network device determines, according to the first remaining token and a preset threshold, a first bandwidth adjustment value corresponding to the preset threshold, including: The first network device determines, according to the first remaining token and the preset threshold, a first bandwidth adjustment weight corresponding to the preset threshold; The first network device determines the first bandwidth adjustment value according to the first bandwidth adjustment weight.
3. The method according to claim 2, characterized in that The first network device determining the first bandwidth adjustment value according to the first bandwidth adjustment weight includes: The first network device determines the first bandwidth adjustment weight as the first bandwidth adjustment value.
4. The method according to claim 2, characterized in that The first network device determining the first bandwidth adjustment value according to the first bandwidth adjustment weight includes: The first network device determines the first bandwidth adjustment value according to the bandwidth of the egress port and the first bandwidth adjustment weight.
5. The method according to any one of claims 2 to 4, characterized in that: The preset threshold includes at least one threshold, the at least one threshold includes a first threshold, and the first network device determines a first bandwidth adjustment weight corresponding to the preset threshold according to the first remaining token and the preset threshold, including: When the first remaining token is less than or equal to the first threshold, the first network device obtains a bandwidth adjustment weight corresponding to the first threshold as the first bandwidth adjustment weight.
6. The method according to claim 5, characterized in that The first threshold corresponds to a bandwidth adjustment weight.
7. The method according to claim 5, characterized in that The first threshold corresponds to a plurality of bandwidth adjustment weights, each of the plurality of bandwidth adjustment weights corresponds to a priority, and obtaining the bandwidth adjustment weight corresponding to the first threshold as the first bandwidth adjustment weight includes: According to the priority of a third network device, a bandwidth adjustment weight among the multiple bandwidth adjustment weights whose priority is the same as the priority of the third network device is determined as the first bandwidth adjustment weight, wherein the third network device is the second network device, or the third network device sends traffic to the first network device via the second network device.
8. The method according to claim 5 or 6, characterized in that The first network device maintains one or more thresholds corresponding to each of multiple priorities, each of the one or more thresholds corresponds to a bandwidth adjustment weight, the first threshold is one of the one or more thresholds corresponding to the target priority among the multiple priorities, and the target priority is the priority of the third network device.
9. The method according to claim 8, characterized in that The multiple priorities include a first priority and a second priority, and the first priority is higher than the second priority. Then: if the first target threshold corresponding to the first priority is greater than or equal to the second target threshold corresponding to the second priority, then the bandwidth adjustment weight corresponding to the first target threshold is greater than or equal to the bandwidth adjustment weight corresponding to the second target threshold.
10. The method according to any one of claims 5 to 9, characterized in that: The at least one threshold further includes a second threshold, the second threshold being smaller than the first threshold, the method further including: The first network device determines a second remaining token in the token bucket; When the second remaining token is less than or equal to the second threshold, the first network device obtains a second bandwidth adjustment weight corresponding to the second threshold, the second threshold is less than the first threshold, and the second bandwidth adjustment weight is less than the first bandwidth adjustment weight; The first network device determines a second bandwidth adjustment value according to the second bandwidth adjustment weight, and sends the second bandwidth adjustment value to the second network device.
11. The method according to any one of claims 5 to 9, characterized in that: The at least one threshold further includes a third threshold, the third threshold being greater than the first threshold, and the method further includes: The first network device determines a third remaining token in the token bucket; When the third remaining token is greater than the first threshold and less than or equal to the third threshold, the first network device obtains a third bandwidth adjustment weight corresponding to the third threshold, where the third bandwidth adjustment weight is greater than the first bandwidth adjustment weight. The first network device determines a third bandwidth adjustment value according to the third bandwidth adjustment weight, and sends the third bandwidth adjustment value to the second network device.
12. The method according to any one of claims 1 to 11, characterized in that The output ports are: A single egress port, a group of egress ports, or a chip comprising at least one egress port group, wherein the egress port group comprises at least one single egress port.
13. The method according to any one of claims 1 to 11, characterized in that The AI training device connected to the output port includes: The graphics processor GPU of the first network device, or the neural network processor NPU of the first network device.
14. A bandwidth scheduling device, characterized in that: Applied to a first network device, the apparatus includes: a processing unit, configured to determine a first remaining token in a token bucket, the token bucket being used to schedule traffic on an egress port of the first network device, the first remaining token being an unused token in the token bucket, and the traffic sent by the egress port to an artificial intelligence (AI) training device connected to the egress port including AI training data; and determining, based on the first remaining token and a preset threshold, a first bandwidth adjustment value corresponding to the preset threshold, the first bandwidth adjustment value being used to adjust traffic on an ingress port of the first network device; A sending unit is configured to send the first bandwidth adjustment value to a second network device, where the second network device is the network device connected to the ingress port.
15. The device according to claim 14, characterized in that The processing unit is configured to: Determining a first bandwidth adjustment weight corresponding to the preset threshold according to the first remaining tokens and the preset threshold; The first bandwidth adjustment value is determined according to the first bandwidth adjustment weight.
16. The device according to claim 15, characterized in that The determining the first bandwidth adjustment value according to the first bandwidth adjustment weight includes: The first bandwidth adjustment weight is determined as the first bandwidth adjustment value.
17. The device according to claim 15, characterized in that The determining the first bandwidth adjustment value according to the first bandwidth adjustment weight includes: The first bandwidth adjustment value is determined according to the bandwidth of the egress port and the first bandwidth adjustment weight.
18. The device according to any one of claims 15 to 17, characterized in that The preset threshold includes at least one threshold, the at least one threshold includes a first threshold, and determining a first bandwidth adjustment weight corresponding to the preset threshold according to the first remaining token and the preset threshold includes: When the first remaining token is less than or equal to the first threshold, a bandwidth adjustment weight corresponding to the first threshold is obtained as the first bandwidth adjustment weight.
19. A device, characterized in that include: processor and memory; The memory is used to store instructions or computer programs; The processor is configured to execute the instructions or computer program and perform the method according to any one of claims 1 to 13.
20. A computer-readable storage medium, characterized in that The method comprises instructions or computer programs, which, when executed on a computer, cause the computer to execute the method according to any one of claims 1 to 13.