Flow control method and apparatus
Patent Information
- Application Number
- PCT/CN2025/147510
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-10
- Filing Date
- 2025-12-30
- Publication Date
- 2026-09-17
Smart Images

Figure CN2025147510_17092026_PF_FP_ABST
Abstract
Description
A flow control method and device
[0001] This application claims priority to Chinese Patent Application No. 202510287911.5, filed on March 10, 2025, entitled "A Flow Control Method and Apparatus", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of computer networks, and more particularly to a flow control method and apparatus. Background Technology
[0003] As network bandwidth demands increase, it is becoming increasingly common to implement high-bandwidth network devices using bare chips with limited bandwidth. Examples include network devices with multiple bare chips (network chips).
[0004] Low-bandwidth bare chips can have the same or different functions, but they all include a cache for packet store-and-forward. In a multi-bare-chip architecture, the caches in each bare chip are used and operate independently. Therefore, each bare chip's physical cache corresponds to a cache resource management (Admission Control) mechanism (ADM) to manage the physical cache resources within that bare chip, ensuring reasonable and full utilization of the cache and guaranteeing the Quality of Service (QoS) of the network in which the multi-bare-chip devices reside. When a bare chip receives a packet, the ADM checks the cache resources. If there are remaining resources in the cache, the packet is received and enqueued, increasing the cache usage; otherwise, packet loss or flow control operations are performed. This flow control operation involves sending flow control signals to upstream devices, causing them to stop transmitting traffic under certain conditions, thus alleviating the cache pressure on the bare chip sending the flow control signal.
[0005] However, current flow control operations may result in some bare chip caches not being effectively utilized, which will inevitably lead to low cache utilization of the entire network device or network chip. Summary of the Invention
[0006] This application provides a flow control method and apparatus to achieve flow control operations at the bare chip level, thereby improving the cache utilization of the entire device and increasing the distance of connection between devices.
[0007] Firstly, a flow control method is provided, applicable to a network device including a chip comprising multiple bare chips, each with a cache deployed within it. Specifically, the method includes: determining that the cache occupancy of a first queue on a first bare chip satisfies a first condition; and performing flow control on the first queue using a preceding bare chip. Here, the first bare chip is one of multiple bare chips, and the first condition indicates that the first queue is congested on the first bare chip. If the first bare chip is a first-level bare chip of the network device, the preceding bare chip is a preceding device of the network device.
[0008] The solution provided in the first aspect of this application implements flow control between bare chips within a network device by initiating flow control to the preceding bare chip after congestion is determined in a certain bare chip. This avoids situations where some bare chip caches become idle due to flow control by any bare chip within the network device on the preceding device, effectively improving the overall device's cache utilization and increasing the distance of connections between devices.
[0009] One possible implementation of the method may further include: if the network device satisfies a third condition, performing the above-described determination that the buffer occupancy of the first queue for the first bare chip satisfies the first condition, and performing flow control on the first queue by the preceding bare chip of the first bare chip. The third condition is used to indicate that the network device and the preceding device are in a short-distance scenario. Because the distance between devices is short in a short-distance scenario, per-bare-chip flow control can quickly alleviate congestion and has a good flow control effect.
[0010] Another possible implementation method may further include: determining that the second queue's cache occupancy for multiple bare chips in the network device satisfies the second condition if the network device does not meet the third condition; then, the preceding device of the network device performs flow control on the second queue; during the flow control process, if the reserved space in the second bare chip's cache is unavailable, the preceding bare chip of the second bare chip performs flow control on the second queue. The second condition is used to indicate that the second queue is congested in the network device. The second bare chip is one of multiple bare chips. The reserved space is the space configured in the cache for flow control. From a global perspective, it is determined that the queue initiates flow control on the preceding device after multiple bare chips in the entire network device become congested, and during the flow control process, if the space in the cache of a certain bare chip for flow control is unavailable, flow control is performed on the preceding bare chip. Combining the aforementioned per-bare-chip-level flow control, it can more effectively avoid the situation where some bare chip caches are idle due to flow control of any bare chip in the network device on the preceding device, effectively improving the cache utilization of the entire device and increasing the distance of inter-device connections.
[0011] Secondly, another flow control method is provided, which can be applied to network devices including chips, which comprise multiple bare chips with caches deployed within them. Specifically, the method includes: determining that the cache occupancy of the second queue across the multiple bare chips in the network device satisfies a second condition; then, the preceding stage device of the network device performs flow control on the second queue; during the flow control process, if the reserved space in the second bare chip's cache is unavailable, the preceding stage bare chip performs flow control on the second queue. The second condition indicates that the second queue is congested within the network device. The second bare chip is one of multiple bare chips. The reserved space is the space configured in the cache for flow control.
[0012] The solution provided in the second aspect of this application determines, from a global perspective, that when multiple bare chips in the entire network device become congested, flow control is initiated on the front-end device. During the flow control process, if the space in the cache of a certain bare chip for flow control is unavailable, flow control is performed on the next-level bare chip. In this way, the situation where some bare chip caches are idle due to flow control of the front-end device by any bare chip in the network device can be avoided, which effectively improves the cache utilization of the entire device and increases the distance of the connection between devices.
[0013] One possible implementation of the method may further include: if the network device does not meet the third condition, performing the above-described determination that the cache occupancy of the second queue for multiple bare chips meets the second condition, and the preceding device of the network device performing flow control on the second queue. The third condition is used to indicate a short-distance scenario between the network device and the preceding device. Because in long-distance scenarios, the distance between devices is long, globally determining that the queue becomes congested in the current network device before the preceding device performs flow control can quickly alleviate congestion and achieve good flow control results.
[0014] Another possible implementation method may further include: determining that the cache occupancy of the first queue on the first bare chip satisfies a first condition; and performing flow control on the first queue by the bare chip preceding the first bare chip. Here, the first bare chip is one of a plurality of bare chips, and the first condition indicates that the first queue is congested on the first bare chip. If the first bare chip is the first-level bare chip of the network device, the bare chip preceding the first bare chip is the preceding device of the network device. Based on global flow control, combined with the aforementioned per-bare-chip-level flow control, it can more effectively avoid situations where partial bare chip cache idleness is caused by flow control of any bare chip within the network device on the preceding device, effectively improving the overall device cache utilization and increasing the distance of inter-device connections.
[0015] Another possible implementation involves the network device's cache including reserved space. This reserved space is used to store traffic during flow control, or it can be controlled by the front-end device or the bare chip. The third condition mentioned above can include: the network device's reserved space is less than the cache size of the network device's first-level bare chip divided by Z. Z is less than or equal to m, where m is the number of queues corresponding to one cache of the first-level bare chip. Since the size of the reserved space is strongly correlated with the distance between devices, this implementation method, which determines the scenario between network devices based on the size of the reserved space, is highly feasible.
[0016] Another possible implementation, the third condition mentioned above, includes: the physical distance between the network device and the front-end device is less than a distance threshold. Determining the scenario based on the physical distance between devices is simple to implement.
[0017] Another possible implementation involves the network device's cache including a shared space, which is shared by multiple queues. Accordingly, the first condition includes: the space corresponding to the first queue within the remaining space of the shared space in the first bare chip is greater than or equal to this remaining space. Correspondingly, the preceding stage bare chip performs flow control on the first queue, including sending flow control information to the preceding stage bare chip, which instructs to stop sending traffic. This application employs a passive flow control approach to implement the flow control.
[0018] Another possible implementation may further include: sending first reference information to the preceding stage bare chip, the first reference information indicating the remaining space in the cache of the first bare chip in the controllable portion of the preceding stage bare chip. Correspondingly, the first condition includes the first reference information indicating no remaining space. The flow control of this application is implemented using an active flow control approach.
[0019] Another possible implementation involves the network device's cache further including shared space. The second condition may include: space corresponding to the second queue within the remaining space of the network device's shared space, which is greater than or equal to that of the second queue. Accordingly, the preceding device of the network device performs flow control on the second queue, including sending flow control information to the preceding device, the flow control information indicating a halt to traffic transmission. This application employs a passive flow control approach to implement flow control.
[0020] Another possible implementation may further include: sending second reference information to the preceding device, the second reference information indicating the remaining space in the network device's cache within the portion controllable by the preceding device. Correspondingly, the second condition may include: determining, based on the second reference information, that there is no remaining space in the network device's cache within the portion controllable by the preceding device. This application employs active flow control to implement the flow control.
[0021] Another possible implementation involves using the second reference information as the received cumulative value plus the available buffer space; the received cumulative value indicates the cumulative number of received packets. Determining no remaining space based on the second reference information includes: if the transmitted cumulative value is greater than the second reference information, then it is determined that there is no remaining space in the network device's buffer within the controllable portion of the upstream device. The transmitted cumulative value indicates the cumulative number of transmitted packets.
[0022] Another possible implementation involves configuring a sharing coefficient for the queues. The space corresponding to the third queue in the target remaining space includes: the target remaining space multiplied by the sharing coefficient of the third queue. The third queue can be any queue entering the network device for forwarding. The sharing coefficient allows for control over the proportion of shared space allocated to different queues.
[0023] Another possible implementation method may further include: accumulating the cache occupancy (cnt) of the first queue on the first bare chip when a packet from the first queue enters the cache queue of the first bare chip; and decrementing cnt when a packet from the first queue is dequeued from the cache queue of the first bare chip. Real-time statistics of the cache occupancy of the queue on the bare chip are used to accurately implement flow control.
[0024] Another possible implementation is that the second queue covers the cache occupancy of multiple bare chips in the network device, including the sum of the cache occupancy of each bare chip in the network device.
[0025] Thirdly, a flow control device is provided, which is applied to a network device including a chip, the chip including multiple bare chips, and a cache deployed in the bare chips. The device may include: a first determining unit and a processing unit. Wherein:
[0026] The first determining unit is used to determine that the cache occupancy of the first queue for the first bare chip meets a first condition. The first bare chip is one of a plurality of bare chips. The first condition is used to indicate that the first queue is congested on the first bare chip.
[0027] The processing unit is configured to determine, in the first determining unit, that the buffer occupancy of the first queue for the first bare chip meets a first condition, and control the preceding bare chip of the first bare chip to perform flow control on the first queue. Wherein, if the first bare chip is the first-level bare chip of a network device, the preceding bare chip of the first bare chip is the preceding device of that network device.
[0028] It should be noted that the flow control device provided in the third aspect of this application is used to execute the flow control method provided in the first aspect. The specific implementation can refer to the first aspect or any possible implementation method, which will not be repeated here.
[0029] Fourthly, another flow control device is provided, which is applied to a network device including a chip, the chip comprising multiple bare chips, in which caches are deployed. The device includes: a second determining unit, a processing unit, and a third determining unit. Wherein:
[0030] The second determining unit is used to determine whether the buffer usage of the second queue for multiple bare chips meets a second condition. The second condition indicates that the second queue is congested in the network device. The second bare chip is one of multiple bare chips.
[0031] The processing unit is used to determine, in the second determining unit, that the cache occupancy of the second queue for multiple bare chips meets the second condition, and to control the front-end device of the network device to perform flow control on the second queue.
[0032] The third determining unit is used to: determine that there is no available space in the reserved space of the second bare chip cache during the flow control process of the second queue.
[0033] The processing unit is also configured to: after the third determining unit determines that there is no available space in the reserved space of the second bare chip cache, control the preceding bare chip of the second bare chip to perform flow control on the second queue. The second bare chip is one of a plurality of bare chips. The reserved space is the space configured in the cache for flow control.
[0034] It should be noted that the flow control device provided in the fourth aspect of this application is used to execute the flow control method provided in the second aspect above. The specific implementation can refer to the first aspect above or any possible implementation method, which will not be repeated here.
[0035] Fifthly, a network device is provided, including a processor and a memory. The processor is configured to execute instructions stored in the memory to cause the network device to perform operational steps as described in the first aspect, the second aspect, or any possible implementation thereof.
[0036] A sixth aspect provides a computer-readable storage medium comprising: computer software instructions; which, when executed in a processor, cause the processor to perform operational steps of the method as described in the first aspect, the second aspect, or any of the possible implementations above.
[0037] In a seventh aspect, a computer program product is provided that, when run on a computer, causes the computer to perform the operational steps of the method as described in the first aspect, the second aspect, or any of the possible implementations above.
[0038] Eighthly, a chip is provided, comprising: a processor and a power supply circuit; wherein the power supply circuit is used to supply power to the processor; the processor is used to perform operational steps of the method in the first aspect or the second aspect or any possible implementation of the first aspect.
[0039] The technical effects of any of the design methods in aspects three through eight can be found in the technical effects of different design methods in aspects one or two, and will not be repeated here.
[0040] Based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods. Attached Figure Description
[0041] Figure 1 is a schematic diagram of a network scenario;
[0042] Figure 2 is a schematic diagram of the architecture of a network transmission system provided in an embodiment of this application;
[0043] Figure 3 is a schematic diagram of the structure of a network device provided in an embodiment of this application;
[0044] Figure 4 is a flowchart illustrating a flow control method provided in an embodiment of this application;
[0045] Figure 5 is a flowchart illustrating another flow control method provided in an embodiment of this application;
[0046] Figure 6 is a schematic diagram of a network architecture scenario provided in an embodiment of this application;
[0047] Figure 7 is a schematic diagram of another network architecture scenario provided in the embodiments of this application;
[0048] Figure 8 is a timing diagram of the occupancy of physical cache in a network device according to an embodiment of this application;
[0049] Figure 9 is a timing diagram of physical cache occupancy in another network device provided in an embodiment of this application;
[0050] Figure 10 is a timing diagram of physical cache occupancy in another network device provided in an embodiment of this application;
[0051] Figure 11 is a schematic diagram of another network architecture scenario provided in an embodiment of this application;
[0052] Figure 12 is a schematic diagram of the timing of physical cache occupancy in another network device provided in an embodiment of this application;
[0053] Figure 13 is a schematic diagram of the timing of physical cache occupancy in another network device provided in an embodiment of this application;
[0054] Figure 14 is a schematic diagram of the timing of physical cache occupancy in another network device provided in an embodiment of this application;
[0055] Figure 15 is a schematic diagram of the timing of physical cache occupancy in another network device provided in an embodiment of this application;
[0056] Figure 16 is a schematic diagram of the timing of physical cache occupancy in another network device provided in an embodiment of this application;
[0057] Figure 17 is a timing diagram of physical cache occupancy in another network device provided in an embodiment of this application;
[0058] Figure 18 is a timing diagram of physical cache occupancy in another network device provided in an embodiment of this application;
[0059] Figure 19 is a schematic diagram of the timing of physical cache occupancy in another network device provided in an embodiment of this application;
[0060] Figure 20 is a schematic diagram of the timing of physical cache occupancy in another network device provided in an embodiment of this application;
[0061] Figure 21 is a schematic diagram of a flow control device provided in an embodiment of this application;
[0062] Figure 22 is a schematic diagram of another flow control device provided in an embodiment of this application;
[0063] Figure 23 is a schematic diagram of the structure of a network device provided in an embodiment of this application. Detailed Implementation
[0064] The technical solutions of the embodiments of this application will be described below with reference to the accompanying drawings. In the description of this application, unless otherwise stated, " / " indicates that the objects before and after are in an "or" relationship. For example, A / B can represent A or B. "And / or" in this application is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone, where A and B can be singular or plural. Furthermore, in the description of this application, unless otherwise stated, "multiple" refers to two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple. Furthermore, to facilitate a clear description of the technical solutions in the embodiments of this application, the terms "first" and "second" are used in the embodiments of this application to distinguish identical or similar items with substantially the same function and effect. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and the terms "first" and "second" are not necessarily different.
[0065] To facilitate understanding, the main terms used in this application will be explained first.
[0066] Network device: A physical device that acts as a single node in a network. Network devices can take the form of chips or other forms.
[0067] Flow control refers to the process of controlling network traffic (such as pausing transmission) through relevant strategies and methods.
[0068] As network bandwidth demands increase, the area of network chips also grows. Due to process engineering constraints, high-bandwidth network chips are typically implemented using a multi-bare-chip approach. A bare chip can be the smallest unit of granularity within a chip, such as a chiplet or even smaller. Multiple bare chips within a network chip can have the same or different functions, but they all include buffers for packet store-and-forward, and queue management and scheduling chips for packet enqueueing and dequeueing. For the entire network chip, physical buffers are distributed across various bare chips, which operate independently. A bare chip can also be divided into multiple identical or different smaller units, each with its own buffer for packet store-and-forward. From the perspective of the entire chip, the chip's physical buffers are distributed across the smallest granularity bare chips, which operate independently. To effectively, rationally, and fully utilize the buffers, buffer resources need to be managed to ensure the QoS performance of the network in which the network chip operates.
[0069] In a multi-chip network architecture, each chip's physical buffer operates independently. Therefore, each chip's physical buffer corresponds to an ADM (Application Management Controller), which manages the physical buffer resources within its own chip. When a chip receives a packet, the ADM corresponding to the chip's physical buffer checks the buffer resources. If there are remaining resources, the packet is received and enqueued, increasing the buffer usage; otherwise, packet loss or flow control operations are initiated. When a packet is dequeued, the ADM decrements the buffer usage. If flow control operations have been performed, the corresponding operations may be revoked.
[0070] To fully and rationally utilize cache space, flow control is a crucial aspect of cache management. Flow control mechanisms are divided into active flow control and passive flow control. A common active flow control method is the credit flow control mode, while a common passive flow control method is the PFC flow control mode. The cache management actions differ depending on the flow control mechanism.
[0071] In PFC flow control mode, the physical buffer is divided into an overshoot protection space (headroom) and a shared space for different packet buffering functions. The shared space is used by all queues to maximize the use of buffer space to absorb traffic bursts. The overshoot protection space (headroom) is used to absorb all incoming traffic before the flow stops (the preceding device stops flow after receiving the flow control signal) when this device generates flow control (XOFF) for a preceding device.
[0072] In Credit flow control mode, the physical buffer is typically divided into static reserved space and shared space for different packet caching functions. The shared space is used by all queues to maximize the use of buffer space to absorb traffic bursts. The static reserved space is controlled and managed by the front-end device. During network data forwarding, devices exchange information in real time about the usage of the static reserved space. When the static reserved space is exhausted, the front-end device proactively stops sending traffic.
[0073] In practical applications, the distributed physical cache in network chips can be partitioned and managed based on PFC flow control mode or credit flow control mode. From the perspective of the entire chip, when some caches are in use, other caches may be in an idle and unoccupied state, resulting in the overall chip's cache not being fully utilized.
[0074] Figure 1 illustrates a network scenario. As shown in Figure 1, this device includes bare chip A and bare chip B. The buffers in bare chips A and B are used for packet storage and forwarding. In PFC flow control mode, each buffer is divided into headroom and shared space. When the shared space of any buffer in any bare chip is exhausted, XOFF flow control information is generated and sent to the front-end device. Assuming that the buffer capacity of bare chip B is 100 units, and the buffer capacity of each bare chip A is 10 units, and n is 32, the total buffer capacity is 32 * 10 = 320 units. When traffic passes through bare chip A and enters bare chip B, at a certain moment, bare chip B becomes congested (for example, the 100 units of buffer capacity are exhausted), and traffic accumulates in the buffer of bare chip B. Then, bare chip B generates XOFF flow control information for the front-end device n (indicated by the dashed arrow in Figure 1). The front-end device receives the XOFF flow control information and stops sending traffic. At this time, the buffer of bare chip A may be in an idle state, and the entire chip's buffer may have 320 units in an idle state. Therefore, from the perspective of the entire chip of this device, all physical caches are not fully utilized before the flow control signal from the front-end device is generated, resulting in very low cache utilization. Similarly, the credit flow control mode has the same problem.
[0075] Based on this, this application provides a flow control scheme. By initiating flow control on the previous-level bare chip after determining that the queue is congested in a certain bare chip, flow control between bare chips within the network device can be achieved. Alternatively, from a global perspective, flow control can be initiated on the previous-level device after determining that the queue is congested on multiple bare chips in the entire network device. In subsequent transmission, when a certain bare chip has no available space, flow control can be initiated on the previous-level bare chip. In this way, the situation where some bare chip buffers are idle due to flow control of the previous-level device by any bare chip in the network device can be avoided, which effectively improves the buffer utilization of the entire device and increases the distance of the connection between devices.
[0076] One possible implementation method is to add interaction between bare chips, such as exchanging flow control information or cache occupancy information, to achieve global cache resource management across bare chips, manage all cache resources of the network chip, and realize full utilization of the entire chip cache.
[0077] The solution provided in this application can be applied to the network transmission system illustrated in Figure 2. As shown in Figure 2, the network transmission system includes multiple network devices 201.
[0078] Network device 201 is used to receive and forward traffic packets according to service requirements. For example, network device 201 can implement the functions of a switch, router, or other network forwarding devices. The number of network devices 201 in the network transmission system can be deployed according to actual needs; Figure 2 is only an example and does not constitute a limitation.
[0079] The network device 201 can be a physical server or a chip. This application embodiment does not limit the product form of the network device 201.
[0080] Furthermore, network device 201 deploys chips, which consist of multiple small-granularity bare chips, such as chiplets and dies.
[0081] For example, one possible structure of network device 201 is shown in Figure 3. As shown in Figure 3, network device 201 includes multiple bare chips 2011 and a global cache resource management module 2012.
[0082] A bare chip 2011 is managed and scheduled by one or more queues. A bare chip 2011 includes one or more caches 20111 and a corresponding cache resource management module 20112. The cache resource management modules 20112 in each bare chip 2011 are connected to the global cache resource management module 2012. The bare chips 2011 interact with each other; the interaction can include flow control information and cache occupancy information. The flow control information can be implemented based on PFC flow control mode, credit flow control mode, or other flow control modes.
[0083] The physical cache of network device 201 is divided into reserved space and shared space from a holistic perspective. This division can be performed by the global cache resource management module 2012. For example, headroom space and shared space in PFC flow control mode, or static reserved space and shared space in Credit flow control mode.
[0084] The cache resource management module 20112 is used to manage cache resources for a single cache 20111. The cache resource management module 20112 reports cache management-related information to the global cache resource management module 2012, which then manages the physical cache of the network device 201 from a global perspective.
[0085] The Global Resource Cache Management Module 2012 can be implemented in software or hardware.
[0086] In one possible implementation, the global cache resource management module 2012 can be deployed inside the network device 201, at the same level as the bare chip 2011 (as shown in Figure 3).
[0087] In another possible implementation, the global cache resource management module 2012 can also be deployed inside any bare chip 2011 (not shown in Figure 3).
[0088] The queue described in this application can be a port-level queue or a priority-based queue. The queue can be implemented using any method, such as FIFO or linked list.
[0089] The cache 20111 described in this application can be implemented based on any form, such as first-in-first-out (FIFO) or static random-access memory (SRAM). The inter-queue scheduling corresponding to each cache can be any scheduling type, such as round-robin (RR), strict priority (SP), or weighted round-robin (WRR).
[0090] The cache resource management described in this application can be based on static reservation, dynamic sharing, or a combination of both of port queues, or it can be based on static reservation, dynamic sharing, or a combination of both of priority queues.
[0091] The Credit flow control mode described in this application supports the implementation of any protocol, including but not limited to credit-based flow control (CBFC) under the InfiniBand (IB) protocol, CBFC under the Ultra Ethernet Consortium (UEC) protocol, centralized firewall control (CFC) under the Universal Bus Protocol (UB) protocol, and CFC under the EET protocol.
[0092] The solution provided in this application will now be described in detail with reference to the accompanying drawings.
[0093] This application provides a flow control method that can be applied to a network device including chips, where the chips include multiple bare chips. For example, the network device can be network device 201 illustrated in Figure 2 or Figure 3. This method is specifically applied in a computer network including multiple network devices during the data forwarding process.
[0094] Specifically, each bare chip in a network device deploys one or more caches. Each cache can correspond to a cache resource management module, used to manage the cache's resources. Each cache also corresponds to one or more queues, used to store and forward packets within those queues.
[0095] Before performing data forwarding and flow control, the physical cache of network devices can be allocated to reserve space for storing traffic during flow control. The name and type of the reserved space may differ under different flow control modes, and this application embodiment does not limit this. All space used for storing traffic during flow control falls under the category of the reserved space described in this application. In practical applications, reserved space can be allocated in the physical cache according to actual needs, and this application embodiment does not limit the specific scheme for allocating reserved space.
[0096] For example, the size of the reserved space can be positively correlated with the one-way time obtained by dividing the physical distance between the two devices by the transmission rate of the fiber optic link. That is, the greater the physical distance between the two devices, the larger the reserved space.
[0097] In one possible implementation, the reserved space described in this application can be the headroom space in the PFC flow control mode, used to absorb the flight traffic during the flow control process (i.e., the traffic received after the flow control is initiated).
[0098] In another possible implementation, the reserved space described in this application can be the reserved space in the Credit flow control mode, used for flow control by the front-end device or the previous bare chip, and the transmission of traffic is stopped when there is no remaining reserved space.
[0099] Specifically, in this application, the global cache management module allocates reserved space and configures it into the cache management modules in each bare chip, synchronizing the size of the reserved space and the shared space within the cache management modules in the bare chip.
[0100] In one possible implementation, the reserved space can be distributed starting from the first-level bare chip of the network device until the reserved space size is reached. For example, if the reserved space size is less than or equal to the cache size of the first-level bare chip, the reserved space is only deployed in the cache of the first-level bare chip. If the reserved space size is greater than the cache size of the first-level bare chip, the entire cache of the first-level bare chip is used as reserved space, and the remaining reserved space is deployed in the cache of the second-level bare chip.
[0101] In another possible implementation, the reserved space can be distributed across the bare chips at various levels of the network device.
[0102] On one hand, embodiments of this application provide a flow control method, as shown in FIG4, which may include:
[0103] S401, Record queue cache usage on bare chip.
[0104] For example, S401 can be executed by the cache resource management module corresponding to the cache, which is deployed within the bare chip.
[0105] Among them, cache usage is used to indicate the cache occupancy status of the queue, which can be reflected through cache usage information. Cache usage information can include queue length, the cache usage amount corresponding to the queue, changes in queue length, changes in the amount of cache usage in the queue, message length, and any other information that reflects cache usage.
[0106] Optionally, the cache usage information may also include additional information required by the flow control mode.
[0107] Specifically, S401 can be implemented as follows: when a packet enters a queue in a bare chip, the buffer occupancy of the queue is incremented. When a packet is dequeued from the bare chip, the buffer occupancy of the bare chip is decremented.
[0108] The increment and decrement steps can be configured according to actual needs. For example, the increment and decrement steps can be 1.
[0109] For example, the buffer occupancy cnt of the first queue on the first bare chip is used to indicate the occupancy status of the first queue's packets stored in the first bare chip's buffer. cnt is incremented when a packet enters the buffer queue of the first bare chip; cnt is decremented when a packet is dequeued from the buffer queue of the first bare chip. Here, the first bare chip is any bare chip in the network device, and the first queue is any queue within the first bare chip.
[0110] It should be understood that each bare chip in the network device performs the S401 operation to record the cache usage of each queue on each bare chip.
[0111] Furthermore, the cache occupancy of the second queue for multiple bare chips includes: the sum of the cache occupancy of each bare chip in the multiple bare chips of the network device in the second queue.
[0112] For example, suppose a network device includes bare chip A and bare chip B. Aq_cnt represents the cache usage of queue q for bare chip A, Bq_cnt represents the cache usage of queue q for bare chip B, Aq_cnt+Bq_cnt represents the cache usage of queue q for the entire network device, or, represents the cache usage of queue q for all bare chips in the network device.
[0113] When messages are enqueued or dequeued, the queue's cache usage can be sent to the global cache resource management module for global resource management and flow control decisions.
[0114] S402. Determine that the cache occupancy of the first queue for the first bare chip meets the first condition.
[0115] In this context, the first bare chip is one of a plurality of bare chips included in the network device. If the first bare chip is the first-level bare chip of the network device, the preceding bare chip is the preceding device of the network device. The first queue is any queue among the first bare chips.
[0116] Specifically, the first condition is used to indicate that the first queue is congested on the first bare chip. The content of the first condition can be configured according to actual needs.
[0117] The following examples illustrate the specific implementation of the first condition, but do not constitute a specific limitation:
[0118] Implementation 1: The network device's cache includes not only reserved space but also shared space. This shared space is shared by multiple queues. The first condition includes: the space corresponding to the first queue within the remaining space of the shared space in the first bare chip, which is greater than or equal to that of the first queue.
[0119] When a cache is shared by multiple queues, the degree to which each queue owns the cache, i.e. the size of the cache occupied by the queue, is the space corresponding to that queue.
[0120] For example, a sharing coefficient for the queue can be configured to indicate the degree to which the queue occupies the shared space. The space corresponding to the third queue in the target remaining space includes: the target remaining space multiplied by the sharing coefficient of the third queue. The third queue can be any queue in the network device.
[0121] For example, the network device uses PFC flow control mode between bare chips, and the shared space is called the share space. Assuming the network device includes bare chip A and bare chip B, the cache occupancy of queue q for bare chip B is determined to satisfy the first condition when the cache occupancy of queue q for bare chip B is Bq_cnt ≥ the remaining share space in the cache of bare chip B multiplied by the dynamic coefficient a of queue q. Here, the remaining share space in the cache of bare chip B is: the total size of the share space in the cache of bare chip B minus the already occupied size.
[0122] It is understandable that implementing 1 can be a specific implementation of the first condition in a passive flow control scenario.
[0123] Implementation 2: The bare chip feeds back first reference information to the previous bare chip. The first reference information is used to indicate the remaining space in the cache of this bare chip in the controllable part of the previous bare chip. The first condition includes the first reference information indicating that there is no remaining space.
[0124] The first reference information can be either the remaining capacity or the used capacity. The first reference information indicating no remaining space can be understood as: based on the first reference information, it is determined that there is no remaining space in the controllable portion of the preceding stage bare chip's cache.
[0125] For example, the remaining space in the cache of this bare chip, which is controlled by the previous bare chip, can be the reserved space in the Credit flow control mode.
[0126] In one possible implementation, the sending bare chip (TX) and the receiving bare chip (RX) maintain flow control-related information and interact with each other, and the information fed back from the RX bare chip to the TX bare chip is used as the aforementioned first reference information.
[0127] For example, the TX bare chip maintains a transmit cumulative value (FCTBS), which is incremented for each packet transmitted. The RX bare chip maintains a receive cumulative value (ABR), which is incremented for each packet received. Simultaneously, the available buffer capacity of the RX bare chip is recorded as RAVAIL. The bare chip calculates ABR + RAVAIL as first reference information and periodically sends it to the preceding bare chip. When the preceding bare chip determines that the FTBS of the first queue is greater than or equal to ABR + RAVAIL, it determines that the first condition is met and stops transmitting traffic.
[0128] In implementation 2, the flow control method provided in this application embodiment may further include: sending first reference information to the preceding stage bare chip of the first bare chip, the first reference information being used to indicate the remaining space in the cache of the first bare chip of the controllable portion of the preceding stage bare chip. The first condition includes the first reference information indicating no remaining space.
[0129] In one possible implementation, the first reference information can be sent in real time to the previous stage bare chip of the first bare chip.
[0130] In another possible implementation, the first reference information can be sent to the previous stage bare chip when the remaining space in the cache of the first bare chip starts to decrease. Of course, the timing of sending the first reference information can also be configured according to actual needs, and this application embodiment does not limit this.
[0131] In one possible implementation, the cache resource management module (used to manage the cache corresponding to the first queue) in the first bare chip can send the first reference information to the previous bare chip.
[0132] In another possible implementation, the global cache resource management module can send the first reference information to the previous stage bare chip of the first bare chip.
[0133] In another possible implementation, the first bare chip is the first-level bare chip of the network device, the first condition adopts the implementation method 2, and the first reference information can be the following second reference information, which is used to indicate the remaining space of the controllable part of the front-end device in this network device.
[0134] In implementation 2, the preceding bare chip of the first bare chip stops sending traffic to the first queue when it determines that the first reference information indicates that there is no remaining space, and resumes sending traffic to the first queue when the first reference information indicates that there is remaining space.
[0135] It is understandable that implementation 2 can be a specific implementation of the first condition in the active flow control scenario.
[0136] To achieve step 3, the first condition is that the threshold value is greater than or equal to the first threshold value.
[0137] The first threshold is a pre-configured upper limit for the cache usage of the first queue on the first bare chip. The value of the first threshold can be configured according to actual needs.
[0138] It is understandable that implementation 3 can be a specific implementation of the first condition in a passive flow control scenario.
[0139] S403, the first-level bare chip preceding the first bare chip performs flow control on the first queue.
[0140] In this context, "the preceding stage of the first bare chip performs flow control on the first queue" means that the preceding stage of the first bare chip stops sending traffic to the first queue.
[0141] In one possible implementation, corresponding to implementation 1 or implementation 3 in S402, S403 can specifically include: sending flow control information to the preceding stage bare chip, the flow control information being used to instruct the cessation of traffic transmission. Upon receiving the flow control information, the preceding stage bare chip stops transmitting traffic in the first queue. For example, the flow control information can be an XOFF flow control signal, but is not limited to this.
[0142] In one possible implementation, the cache resource management module (used to manage the cache corresponding to the first queue) in the first bare chip can send flow control information to the previous bare chip.
[0143] In another possible implementation, the global cache resource management module can send flow control information to the previous stage bare chip of the first bare chip.
[0144] Furthermore, in implementation 1 or implementation 3, when the buffer occupancy of the first queue for the first bare chip no longer meets the first condition, a release message needs to be sent to the previous bare chip to restore the transmission flow of the first queue. For example, the release message can be an XON signal.
[0145] In another possible implementation, corresponding to implementation 2 in S402, S403 can be specifically implemented as follows: after determining that the buffer occupancy of the first queue on the first bare chip meets the first condition, the first-level bare chip before the first bare chip actively stops sending the traffic of the first queue.
[0146] The solution provided in this application provides that, after determining that congestion has occurred in a certain bare chip, flow control is initiated on the previous bare chip, thereby achieving flow control between bare chips within the network device. In this way, the situation where some bare chip caches are idle due to flow control of the previous device by any bare chip in the network device can be avoided, which effectively improves the cache utilization of the entire device and increases the distance of the connection between devices.
[0147] On the other hand, this application provides another flow control method, as shown in FIG5, which may include:
[0148] S501 records the cache usage of the raw chip by the record queue.
[0149] It should be noted that the implementation of S501 can refer to the aforementioned S401, and will not be repeated here.
[0150] S502. Determine that the cache occupancy of multiple bare chips in the network device for the second queue meets the second condition.
[0151] The second queue can be any queue in the network device. The second bare chip is one of a plurality of bare chips.
[0152] Specifically, the second condition is used to indicate that the second queue is congested on the network device. The content of the second condition can be configured according to actual needs.
[0153] The following examples illustrate the specific implementation of the second condition, but do not constitute a specific limitation:
[0154] To achieve this, the network device's cache includes not only reserved space but also shared space. This shared space is shared by multiple queues. The second condition includes: the space corresponding to the second queue within the remaining space of the network device's shared space, which is greater than or equal to that of the second queue.
[0155] For example, the network devices use PFC flow control mode between bare chips, and the shared space is called the share space. Assuming the network device includes bare chip A and bare chip B, the second condition is met when the cache usage of queue q for the network device is Aq_cnt + Bq_cnt ≥ the remaining space of the first space * the dynamic coefficient a of queue q. Here, the first space is the total space of the share space of the network device, and the remaining space of the first space is the total size of the first space minus the occupied size.
[0156] It is understandable that this involves the specific implementation of 'a' as the first condition in a passive flow control scenario.
[0157] Implementation b: The network device feeds back the second reference information to the previous level device. The second reference information is used to indicate the remaining space in the buffer of this device that is controllable by the previous level device. The second condition includes: based on the second reference information, it is determined that there is no remaining space in the buffer of the network device that is controllable by the previous level device.
[0158] The second reference information can be the remaining capacity or the used capacity.
[0159] For example, the remaining space in the cache of a network device that is controllable by the preceding device can be the reserved space in the Credit flow control mode.
[0160] In one possible implementation, the transmitting device (TX) and the receiving device (RX) maintain flow control-related information and interact with each other, and the information fed back by the RX device to the TX device serves as the aforementioned second reference information.
[0161] For example, the second reference information is the received cumulative value plus the available buffer space; the received cumulative value is used to indicate the cumulative number of received packets. Determining that there is no remaining space based on the second reference information includes: if the transmitted cumulative value is greater than or equal to the second reference information, then it is determined that there is no remaining space in the controllable portion of the network device's buffer; the transmitted cumulative value is used to indicate the cumulative number of transmitted packets.
[0162] For example, the TX device maintains a transmit cumulative quantity FCTBS, which is incremented for each packet transmitted, and the RX device maintains a receive cumulative quantity ABR, which is incremented for each packet received. At the same time, the available buffer quantity of the RX device is recorded as RAVAIL.
[0163] The network device calculates ABR+RAVAIL as the second reference information and periodically sends it to the preceding device. The preceding device determines that the second condition is met when the FCTBS of the second queue is greater than or equal to ABR+RAVAIL.
[0164] In implementation b, the flow control method provided in this application embodiment may further include: sending second reference information to the front-end device of the network device, wherein the second reference information is used to indicate the remaining space in the cache of the network device that is controllable by the front-end device.
[0165] In one possible implementation, second reference information can be sent to the preceding device of this network device in real time.
[0166] In another possible implementation, the second reference information can be sent to the preceding device when the remaining space in the cache of this network device, which is controllable by the preceding device, begins to decrease. Of course, the timing of sending the second reference information can also be configured according to actual needs, and this application embodiment does not limit this.
[0167] It should be understood that the second reference information can be determined by any bare chip in this network device (the target bare chip) and sent to the front-end device, while other bare chips will synchronize their respective cache usage to the target bare chip.
[0168] In one possible implementation, the global cache resource management module can determine whether the second condition is met, and if the second condition is met, send the second reference information to the upstream device of this network device.
[0169] It is understandable that this involves the specific implementation of 'b' as the first condition in a proactive flow control scenario.
[0170] To achieve c, the second condition is that it is greater than or equal to the second threshold.
[0171] The second threshold is the pre-configured upper limit for the second queue's cache usage on this network device. The value of the second threshold can be configured according to actual needs.
[0172] S503, the front-end equipment of the network device performs flow control on the second queue.
[0173] In this context, when the network device's front-end device performs flow control on the second queue, it means that the network device's front-end device stops sending traffic to the second queue.
[0174] In one possible implementation, corresponding to implementation a or c in S502, S503 can specifically include: sending flow control information to the preceding device of the network device, the flow control information being used to instruct the cessation of traffic transmission. Upon receiving the flow control information, the preceding device of the network device stops transmitting traffic in the second queue. For example, the flow control information can be an XOFF flow control signal, but is not limited to this.
[0175] In one possible implementation, the global cache resource management module can determine whether the cache occupancy of all bare chips of the network device by the second queue meets the second condition, and if the second condition is met, send flow control information to the front-end device of the network device.
[0176] Furthermore, in the implementation of scheme a or scheme c, when the second queue's cache occupancy on the network device no longer meets the second condition, it is necessary to send a release message to the preceding device of this network device in order to restore the traffic of sending the second queue.
[0177] In another possible implementation, corresponding to implementation b in S502, S503 can be specifically implemented as follows: when the preceding device of this network device determines, according to the second reference information, that there is no remaining space in the controllable part of the preceding device in the buffer of this network device, it actively stops sending traffic in the second queue; when it determines, according to the second reference information, that there is remaining space in the controllable part of the preceding device in the buffer of this network device, it resumes sending traffic in the second queue.
[0178] S504. During the flow control process for the second queue, if there is no available space in the reserved space of the second bare chip cache, the first-level bare chip preceding the second bare chip performs flow control on the second queue.
[0179] One possible implementation is that the cache management module inside the bare chip determines whether there is available space in the internal reserved space. If there is no available space, the first-level bare chip before the second bare chip performs flow control on the second queue.
[0180] The process of flow control of the second queue by the first-level bare chip preceding the second bare chip can be referred to the implementation of S403 mentioned above, and will not be repeated here.
[0181] The solution provided in this application, from a global perspective, determines that when multiple bare chips in the entire network device become congested, flow control is initiated on the front-end device. During the flow control process, when the space in the buffer of a certain bare chip used for flow control is unavailable, the front-end bare chip stops sending traffic. In this way, the situation where some bare chip buffers are idle due to flow control of the front-end device by any bare chip in the network device can be avoided, which effectively improves the buffer utilization of the entire device and increases the distance of the connection between devices.
[0182] In the solution provided in this application, the flow control between bare chips and / or with the front-end device can be PFC flow control mode, credit flow control mode, a mode in which PFC flow control mode and credit flow control mode coexist and can be switched between each other, or any variant mode, or other flow control modes (packet loss, explicit congestion notification (ECN), congestion-aware queue management (CAQM), etc.). The embodiments of this application do not limit this.
[0183] In one possible implementation, the flow control method illustrated in Figure 4 or Figure 5 of this application can be used alone or in combination (i.e., the schemes illustrated in Figure 4 and Figure 5 are both deployed in network devices for execution).
[0184] In another possible implementation, depending on the network scenario, one can choose to execute either the scheme shown in Figure 4 or the scheme shown in Figure 5.
[0185] For example, if the network device meets the third condition, the scheme shown in Figure 4 is adopted; if the network device does not meet the third condition, the scheme shown in Figure 5 is adopted.
[0186] For example, the global resource management module in the network device can determine whether the network device meets the third condition. Of course, other units can also determine whether the network device meets the third condition, and this application embodiment does not limit this.
[0187] In one possible implementation, the third condition can be used to indicate a short-distance scenario between the network device and the front-end device. If the network device meets the third condition, it can be understood as a short-distance scenario between the network device and the front-end device. If the network device does not meet the third condition, it can be understood as a long-distance scenario between the network device and the front-end device. The content of the third condition can be configured according to actual needs, and this embodiment of the application does not limit it in this regard.
[0188] In another possible implementation, the third condition can also be used to indicate the deployment of reserved space within a single bare chip. Failure to meet the third condition means that the reserved space is deployed across modules.
[0189] For example, the third condition may include: the reserved space of the network device is less than or equal to the cache size of the first-level bare chip of the network device divided by Z. Z is less than or equal to m, where m is the number of queues corresponding to one cache of the first-level bare chip.
[0190] For example, the third condition includes: the physical distance between the network device and the front-end device is less than or equal to a distance threshold.
[0191] Figure 6 illustrates a network architecture scenario. As shown in Figure 6, the network device includes two bare chips: chip A and chip B. Chip A deploys n caches, from cache 1 to cache n, and n cache resource management modules (shown as cache resource management modules 1 to n in Figure 6) corresponding to each cache. Queues 11 to 1m (m queues) of chip A correspond to cache 1, and queues n1 to nm correspond to cache n. Chip B deploys one cache, which is shared by all queues. Chip B also deploys a cache resource management module to manage this single cache. A global cache resource management module is deployed throughout the network device.
[0192] The following description, with reference to Figure 6, uses PFC flow control mode between bare chips and between network devices as an example to illustrate the solution provided in this application. The process includes the following steps:
[0193] Step 11: Configure the headroom space of the network device, the remaining cache is shared space, and configure the sharing coefficient 'a' for each queue.
[0194] Step 12: Record the queue's usage of the cache.
[0195] For each queue q, it is assumed that it uses the space of buffer L in A die (buffer L is one of buffer 1 to buffer n). When a message enters queue q of die A, the buffer occupancy Aq_cnt of queue q is accumulated, and when a message is dequeued from die A, Aq_cnt is decremented. When a message enters and leaves die B, the buffer occupancy Bq_cnt of queue q is accumulated and decremented respectively.
[0196] When a message is enqueued and dequeued, both Aq_cnt and Bq_cnt are transmitted to the global buffer resource management module, so as to perform global resource management and flow control decision-making.
[0197] Two flow control processes are provided, one flow control process is step 13a to step 13c, and the other flow control process is step 14a to step 14c.
[0198] Step 13a: when the buffer occupancy Bq_cnt of queue q in die B is greater than the remaining space of the share space of die B multiplied by the dynamic coefficient a of queue q, die B initiates a flow control signal to die A, so that die A stops sending the traffic of queue q to BDIE / die A.
[0199] Step 13b: after die A receives the flow control signal, it stops sending the traffic of queue q. Thereafter, the buffer of die A accumulates, and when the buffer occupancy Aq_cnt of queue q in ASDIE / die A is greater than the remaining space of the share space of die A multiplied by the dynamic coefficient a of queue q, die A initiates flow control to the upstream device, and the headroom space absorbs the in-flight traffic before the upstream device stops the flow.
[0200] Step 13c: after the upstream device receives the flow control instruction, it stops sending the traffic of queue q. Thereafter, the buffer occupancy of die A gradually decreases, and when Aq_cnt is less than the remaining space of the share space of die A multiplied by the dynamic coefficient a of queue q, die A cancels the flow control to the upstream device. The buffer occupancy of die B gradually decreases, and when Bq_cnt is less than the remaining space of the share space of die B multiplied by the dynamic coefficient a of queue q, die B cancels the flow control to die A. This process is repeated, which improves the buffer utilization of the network device.
[0201] Step 14a: when the sum of the buffer occupancy Bq_cnt and Aq_cnt of queue q on each die of the network device is greater than the remaining space of the share space of the network device multiplied by the dynamic coefficient a of queue q, flow control is initiated to the upstream device, and the headroom space of the network device absorbs the in-flight traffic when the upstream device stops the flow.
[0202] Step 14b: When the headroom in B is fully occupied, that is, the B bare die only accommodates part of the in-flight traffic, and there is no available space in the reserved space of the B bare die, the B bare die generates flow control for the A bare die. After receiving the flow control information, the A bare die stops sending the traffic of queue q to the B bare die. At this time, the A bare die continues to accommodate in-flight traffic until the previous-stage device stops sending traffic.
[0203] Step 14c: After the previous-stage device stops traffic, the cache of the A bare die gradually decreases, and the cache of the BDIE / A bare die also gradually decreases. When there is free space in the headroom of the B bare die, the B bare die cancels the flow control on the A bare die. Then, when Bq_cnt + Aq_cnt < the remaining space of the share space of the network device multiplied by the dynamic coefficient a of queue q, the flow control on the previous-stage device is canceled. This process is repeated, which improves the cache utilization of the network device.
[0204] Further, the scenario of the network device can also be determined, and the solution of steps 13a to 13c or the solution of steps 14a to 14c is selected according to the scenario. For example, it is determined whether the network device and the previous-stage device are in a short-distance scenario or a long-distance scenario. In the short-distance scenario, the solution of steps 13a to 13c is selected for execution, and in the long-distance scenario, the solution of steps 14a to 14c is selected for execution.
[0205] For example, when the headroom space < the size of cache L in the A bare die / m, it is represented as a short-distance scenario; otherwise, it is a long-distance scenario.
[0206] With reference to FIG. 6 below, the solution provided by the present application is described by taking an example where a PFC flow control mode is adopted between bare dies and a Credit flow control mode is adopted between network devices. The process includes the following steps:
[0207] Step 21: Configure the reserved space of the network device, the remaining cache is the share space, and configure the sharing coefficient a of each queue.
[0208] Step 22: Record the cache occupation of queues.
[0209] When messages enter and exit the queue, both Aq_cnt and Bq_cnt are transmitted to the global cache resource management module for global resource management and flow control decision-making.
[0210] Step 23: Count the usage of the reserved space and feed it back to the previous-stage device regularly.
[0211] Two flow control processes are provided: one flow control process is steps 24a to 24c, and the other flow control process is steps 25a to 25c.
[0212] Step 24a: When the buffer occupation Bq_cnt of queue q on the B bare die exceeds the remaining space of the share space of the B bare die multiplied by the dynamic coefficient a of queue q, the B bare die sends a flow control signal to the A bare die, so that the A bare die stops sending the traffic of queue q to the B bare die.
[0213] Step 24b: After receiving the flow control signal, the A bare die stops sending the traffic of queue q. Thereafter, the buffer of the A bare die accumulates. When the buffer occupation Aq_cnt of queue q in ASDIE / A bare die exceeds the remaining space of the share space of the A bare die multiplied by the dynamic coefficient a of queue q, the A bare die generates flow control on the upstream device. The usage of the reserved space is continuously transmitted in step 23, which reflects the actual used space or remaining space of the reserved space to the upstream device. When the buffer occupation Aq_cnt of queue q in ASDIE / A bare die is less than the remaining space of the share space of the A bare die multiplied by the dynamic coefficient a of queue q, the used space of reserved transmitted to the upstream device is 0 or it always indicates that there is remaining space; after the buffer occupation Aq_cnt of queue q in ASDIE / A bare die exceeds the remaining space of the share space of the A bare die multiplied by the dynamic coefficient a of queue q, the used space of reserved transmitted to the upstream device gradually increases, or the remaining space gradually decreases.
[0214] Step 24c: After the upstream device determines that there is no available space in the reserved space, it will stop sending traffic. Thereafter, the buffer occupation of the A bare die gradually decreases, and when the upstream device determines that there is available space in the reserved space, it will resume sending traffic. When the buffer occupation of the B bare die gradually decreases until Bq_cnt is less than the remaining space of the share space of the B bare die multiplied by the dynamic coefficient a of queue q, the B bare die cancels the flow control on the A bare die. This process repeats.
[0215] Step 25a: After queue q is congested on the current network device, the used space of reserved transmitted to the upstream device in step 23 gradually increases, or the remaining space gradually decreases.
[0216] Step 25b: During the process where the remaining space of the reserved space gradually decreases, the reserved space in the B bare die is exhausted first. When the reserved space in the B bare die is exhausted, the B bare die generates flow control on the A bare die, so that the A bare die stops sending traffic to the B bare die. The buffer of the A bare die accumulates, the available reserved space of the current network device continues to decrease, and the information about the continuously decreasing available reserved space is continuously fed back to the upstream device through step 23.
[0217] Step 25c: After the upstream device determines that there is no available space in the reserved of the current network device, it will stop sending traffic. After the upstream device stops traffic, the buffer occupancy of the A bare die gradually decreases, and the buffer occupancy of the B bare die gradually decreases. When there is available space in the reserved space of the B bare die, the B bare die cancels the flow control on the A bare die. After the upstream device determines that there is available space in the reserved of the current network device, it will continue to send traffic. This repeated process improves the buffer utilization of the network device.
[0218] Further, the scenario of the network device may also be determined, and a solution among steps 24a to 24c or a solution among steps 25a to 25c is selected for execution according to the scenario. For example, determining whether the network device and the upstream device are in a short-distance scenario or a long-distance scenario. In the short-distance scenario, the solution of steps 24a to 24c is selected for execution, and in the long-distance scenario, the solution of steps 25a to 25c is selected for execution.
[0219] For example, when the reserved space < the size of buffer L in A bare die / m, it is represented as a short-distance scenario; otherwise, it is a long-distance scenario.
[0220] Further, the solution provided by the embodiments of the present application can be applied to a scenario where a bare die is further divided into small-granularity modules, so as to implement flow control per small-granularity module. Specifically, when the solution provided by the embodiments of the present application is applied to a scenario where a bare die is further divided into small-granularity modules, the operations performed by the bare die in the above description of the solution can be replaced with those performed by the small-granularity module, which will not be repeated herein.
[0221] The solution provided by the present application is described below through several examples.
[0222] Example 1:
[0223] Figure 7 schematically shows a network architecture scenario. As shown in Figure 7, the network device includes two bare dies, A bare die and B bare die. A buffer is deployed in each bare die, and each buffer corresponds to one queue. The buffer capacity of the A bare die is 2K, and the buffer capacity of the B bare die is 100K + bias. bias represents the space reserved for flow control between two DIEs / bare dies, which requires very little buffer space. For convenience of description, bias is separately reserved in Example 1 and is not used as a part of buffer space division. The buffer space of the entire network device chip is divided into headroom space and share space. The global buffer resource management module is implemented in the B bare die. PFC flow control mode is adopted between bare dies and between devices.
[0224] Based on the scenario illustrated in Figure 7, assuming the reserved (headroom) space of this network device is 0.95K, and the headroom space is only distributed in bare chip A, the remaining 1.05K space in bare chip ADEI / A, excluding the headroom space, is shared space, and the entire cache space of bare chip B is shared space. Since the headroom space (0.95K) < 2K, this network device is determined to be in a short-range scenario. The timing of physical cache occupancy in the network device under the short-range scenario is shown in Figure 8, and the flow control process under the short-range scenario is as follows:
[0225] During the transmission and forwarding of data traffic by the network device, congestion occurs in bare chip B, and the cache of bare chip B accumulates. At time t1, the cache of bare chip B occupies 100K of cache space. Bare chip B generates flow control on bare chip A. At this time, the cache distribution of the network device is shown in Figure 8(a).
[0226] Subsequently, B absorbs the excess traffic from A's bare chip, occupying the bias cache space. At time t2, the bias cache space is full, and the cache distribution of the network device is shown in Figure 8(b).
[0227] Raw chip A receives flow control from raw chip B and stops sending traffic to raw chip B. Raw chip A's buffer gradually accumulates, reaching 1.05K of buffer space at time t3. When raw chip A occupies 1.05K of buffer space, the queue reaches the buffer congestion condition in ADIE / mod A, and raw chip A sends XOFF flow control to the upstream device. The buffer distribution of network devices at time t3 is shown in Figure 8(c).
[0228] Afterwards, bare chip A absorbs the excess traffic from the preceding device until the preceding device stops sending traffic. At time t4, bare chip A absorbs 0.95K of excess traffic, and the buffer distribution of the network device is shown in Figure 8(d).
[0229] After the preceding device stops sending traffic, the buffer usage of bare chip A and bare chip B decreases sequentially. After a period of time or after the buffer usage decreases to a fixed value, bare chip A sends XON flow control to the preceding device, and the preceding device resends traffic.
[0230] Based on the scenario illustrated in Figure 7, assume that the reserved headroom space of this network device is 22K, distributed across bare chips A and B. 2K of headroom space is deployed in the ADEI / A bare chip, 20K in the BDIE / A bare chip, and the remaining 80K in the BDEI / B bare chip is shared space, excluding the headroom space. Since the headroom space (22K) > 2K, this network device is determined to be a long-haul scenario. Assuming that the ADEI / A bare chip does not pre-occupy headroom space, this is referred to as a simple long-haul scenario. The timing of physical cache occupancy in the network device under the simple long-haul scenario is shown in Figure 9. The flow control process under the simple long-haul scenario is as follows:
[0231] During the transmission and forwarding of data traffic from the network device, congestion occurs in bare chip B. At time t1, the buffer of bare chip B occupies 80K of buffer space. Bare chip B sends XOFF flow control to the upstream device. At this time, the buffer distribution of the network device is shown in Figure 9(a). At all times, bare chip A needs to continuously output the buffer occupancy value to bare chip B (even if the buffer occupancy value is 0).
[0232] Subsequently, B absorbs excess traffic, occupying the headroom space in B's bare chip. At time t2, the 20K headroom space in B's bare chip is full, and B's bare chip generates flow control over A's bare chip. At this time, the cache distribution of the network device is shown in Figure 9(b).
[0233] Subsequently, B absorbs the excess traffic from A's bare chip, occupying the bias buffer space. A's bare chip receives flow control from B's bare chip and stops sending traffic to B's bare chip. At time t3, the bias buffer space is full, and the buffer distribution of the network devices is shown in Figure 9(c).
[0234] After chip A stops sending traffic to chip B, chip A continues to absorb excess traffic from the preceding device until the preceding device stops sending traffic. At time t4, chip A absorbs 2K of excess traffic, and the network device's buffer distribution is shown in Figure 9(d).
[0235] After the preceding device stops sending traffic, the buffer usage of bare chips A and B decreases sequentially. After a period of time, or when the buffer usage decreases to a fixed value, XON flow control is sent to the preceding device, and the preceding device resends traffic. When there is available space in the headroom of bare chip B, bare chip B cancels flow control over bare chip A.
[0236] Based on the scenario illustrated in Figure 7, assume that the reserved headroom space of this network device is 22K, distributed across bare chips A and B. 2K of the headroom space is deployed in the ADEI / A bare chip, 20K in the BDIE / A bare chip, and the remaining 80K in the BDEI / B bare chip is shared space, excluding the headroom space. Since the headroom space (22K) > 2K, this network device is determined to be a long-haul scenario. Assuming that the ADEI / A bare chip prematurely occupies the headroom space for some reason, this is called a complex long-haul scenario. The timing of physical cache occupancy in the network device under the complex long-haul scenario is shown in Figure 10. The flow control process under the complex long-haul scenario is as follows:
[0237] Throughout all time periods, bare chip A needs to continuously output its buffer usage value to bare chip B. Due to various reasons (small packet bursts or long packet processing times, etc.), bare chip A's buffer accumulates, eventually occupying 1K of buffer space (equivalent to prematurely occupying headroom space). When bare chip B experiences congestion, its buffer accumulates as well. At time t1, bare chip B's buffer occupies 79K of buffer space, and the buffer distribution of the network device at this time is shown in Figure 10(a).
[0238] When the B bare chip's cache occupies 79K of cache space, the B bare chip comprehensively judges that the total chip cache occupies 80K (1K+79K). (The 1K cache space occupied by the A bare chip is regarded as part of the shared space). When the shared space of the network device is exhausted, B generates XOFF flow control for the front-end device.
[0239] Subsequently, B absorbs excess traffic, occupying the headroom space in B's bare chip. At time t2, the 20K headroom space in B's bare chip is full, and B's bare chip generates flow control over A's bare chip. At this time, the cache distribution of the network device is shown in Figure 10(b).
[0240] Subsequently, B absorbs the excess traffic from A's bare chip, occupying the bias buffer space. A's bare chip receives flow control from B's bare chip and stops sending traffic to B's bare chip. At time t3, the bias buffer space is full, and the buffer distribution of the network device is shown in Figure 10(c).
[0241] After chip A stops sending traffic to chip B, chip A continues to absorb excess traffic from the preceding device until the preceding device stops sending traffic. At time t4, the network device absorbs 2K of excess traffic, of which chip A absorbs 1K and chip B absorbs 1K. The buffer distribution of the network device is shown in Figure 10(d).
[0242] After the preceding device stops sending traffic, the buffer usage of bare chips A and B decreases sequentially. After a period of time, or when the buffer usage decreases to a fixed value, XON flow control is sent to the preceding device, and the preceding device resends traffic. When there is available space in the headroom of bare chip B, bare chip B cancels flow control over bare chip A.
[0243] Example 2:
[0244] Figure 11 illustrates a network architecture scenario. As shown in Figure 11, the network device includes two bare chips, A and B. Bare chip A has two caches, each with a capacity of 2K, and each cache corresponds to one queue. Bare chip B has one cache, corresponding to two queues. The cache capacity of bare chip B is 100K + 2 * bias, where bias represents the space reserved for flow control between the two DIEs / bare chips. The two caches and queues of bare chip A each correspond to one bias in bare chip B. The cache space required for bias is very small. For ease of description, bias is reserved separately in Example 2 and is not included in the cache space allocation. The cache space of the entire network device chip is divided into reserved (headroom) space and shared (share) space. The global cache resource management module is implemented in bare chip B. PFC flow control mode is used between bare chips and between devices.
[0245] Based on the scenario illustrated in Figure 11, assuming the reserved headroom space of this network device is 0.95K*2, and the headroom space is only distributed in bare chip A, the remaining 1.05K*2 space in bare chip ADEI / A, besides the headroom space, is shared space, and the space of bare chip B's cache is entirely shared space. Since the headroom space of a single cache (0.95K) < 2K, this network device is determined to be a short-range scenario. The timing of physical cache occupancy in the network device under the short-range scenario is shown in Figure 12, and the flow control process under the short-range scenario is as follows:
[0246] During the transmission and forwarding of data traffic from the network device, two queues in bare chip B become congested, leading to buffer accumulation in bare chip B. At time t1, each of the two queues in bare chip B occupies 50K of buffer space, and both queues in bare chip B generate flow control over bare chip A. At this time, the buffer distribution of the network device is shown in Figure 12(a). Subsequently, B absorbs the excess traffic from bare chip A, and each of the two queues occupies one bias buffer space. At time t2, both bias buffer spaces are full, and the buffer distribution of the network device is shown in Figure 12(b).
[0247] When raw chip A receives flow control from two queues of raw chip B, the two queues of raw chip A stop sending traffic to raw chip B. The buffers of the two queues of raw chip A gradually accumulate, and at time t3, each queue of raw chip A occupies 1.05K of buffer space. When each queue of raw chip A occupies 1.05K of buffer space, the buffer congestion condition in ADIE / mod A is reached. The two queues of raw chip A then send XOFF flow control to their respective front-end devices (the front-end devices can be different queues of the same device or queues of different devices). The buffer distribution of the network devices at time t3 is shown in Figure 12(c).
[0248] Subsequently, the two queues of bare chip A absorb the excess traffic from the preceding devices until their respective preceding devices stop sending traffic. At time t4, the two queues of bare chip A each absorb 0.95K of excess traffic, and the buffer distribution of the network devices is shown in Figure 12(d).
[0249] When the front-end devices of the two queues of bare chip A stop sending traffic, the buffer usage of bare chip A and bare chip B decreases sequentially. After a period of time or when the buffer usage decreases to a fixed value, the two queues of bare chip A send XON flow control to their respective front-end devices, and the front-end devices of the two queues resend traffic.
[0250] Based on the scenario illustrated in Figure 11, assume that the reserved headroom space of this network device is 22K*2, distributed across bare chips A and B. 2K*2 headroom spaces are deployed in the ADEI / A bare chip, 20K*2 in the BDIE / A bare chip, and the remaining 30K*2 space in the BDEI / B bare chip is shared space. Since the headroom space of a single cache (22K) > 2K, this network device is determined to be a long-haul scenario. Assuming that the ADEI / A bare chip does not pre-occupy headroom space, this is referred to as a simple long-haul scenario. The timing of physical cache occupancy in the network device under the simple long-haul scenario is shown in Figure 13. The flow control process under the simple long-haul scenario is as follows:
[0251] During the transmission and forwarding of data traffic from the network device, congestion occurs in the two queues of bare chip B. At time t1, each queue of bare chip B occupies 30K of buffer space. The two queues of bare chip B send XOFF flow control to their respective upstream devices. At this time, the buffer distribution of the network device is shown in Figure 13(a). At all times, the two queues of bare chip A need to continuously output the buffer occupancy value to bare chip B (even if the buffer occupancy value is 0).
[0252] Subsequently, bare chip B absorbs the excess traffic from two queues, occupying the headroom space in bare chip B. At time t2, the two 20K headroom spaces in bare chip B are full, and bare chip B generates flow control over bare chip A. At this time, the cache distribution of the network device is shown in Figure 13(b).
[0253] Subsequently, the two queues of bare chip B absorb the excess traffic from the two queues of bare chip A, each occupying one bias buffer space. Bare chip A receives flow control from the two queues of bare chip B and stops sending traffic from the two queues to bare chip B. At time t3, both bias buffer spaces are full, and the buffer distribution of the network device is shown in Figure 13(c).
[0254] After chip A stops sending traffic to chip B in two queues, the buffers of chip A's two queues accumulate. Chips A's two queues continue to absorb excess traffic from their respective predecessor devices until they stop sending traffic. At time t4, chip A's two queues each absorb 2K of excess traffic. The buffer distribution of the network devices is shown in Figure 13(d).
[0255] When the upstream devices of both queues stop sending traffic, the buffer usage of bare chips A and B decreases sequentially. After a period of time, or when the buffer usage decreases to a fixed value, they send XON flow control to their respective upstream devices, which then resend traffic. When there is available space in the headroom of bare chip B, bare chip B cancels the flow control over bare chip A.
[0256] Based on the scenario illustrated in Figure 11, assume that the reserved headroom space of this network device is 22K*2, distributed across bare chips A and B. 2K*2 headroom spaces are deployed in the ADEI / A bare chip, 20K*2 in the BDIE / A bare chip, and the remaining 30K*2 space in the BDEI / B bare chip is shared space. Since the headroom space of a single cache (22K) > 2K, this network device is determined to be a long-haul scenario. Assuming that the ADEI / A bare chip prematurely occupies headroom space for some reason, this is called a complex long-haul scenario. The timing of physical cache occupancy in the network device under the complex long-haul scenario is shown in Figure 14, and the flow control process under the complex long-haul scenario is as follows:
[0257] Throughout all time periods, the two queues of bare chip A need to continuously output their buffer usage values to bare chip B. Due to various reasons (small packet bursts or long packet processing times, etc.), the buffer of bare chip A accumulates, eventually resulting in each of the two queues occupying 1K of buffer space (equivalent to pre-occupying headroom space). When the two queues of bare chip B become congested, the buffer of bare chip B also accumulates. At time t1, the two queues of bare chip B each occupy 29K of buffer space, and the buffer distribution of the network device at this time is shown in Figure 14(a).
[0258] When the cache of bare chip B occupies 29K of cache space, bare chip B comprehensively determines that the cache of the two queues of the whole chip occupies 30K*2((1K+29K)*2), (the 1K cache space occupied by bare chip A is regarded as part of the shared space). When the shared space is used up, the two queues of B generate XOFF flow control for their respective front-end devices.
[0259] Subsequently, the two queues of B absorb excess traffic, occupying the headroom space in the B bare chip. At time t2, the 20K headroom space of each of the two queues in the B bare chip is full, and the B bare chip generates flow control over the A bare chip. At this time, the cache distribution of the network device is shown in Figure 14(b).
[0260] Subsequently, the two queues of bare chip B absorb the excess traffic from bare chip A, each occupying one bias buffer space. Bare chip A receives flow control from the two queues of bare chip B, and both queues stop sending traffic to bare chip B. At time t3, both bias buffer spaces of bare chip B are full, and the buffer distribution of the network device is shown in Figure 14(c).
[0261] After chip A stops sending traffic to chip B, the buffers of chip A's two queues continue to accumulate, and chip A's two queues continue to absorb excess traffic from the preceding device until the preceding device stops sending traffic. At time t4, the two queues of the network device each absorb 2K of excess traffic, with chip A's two queues each absorbing 1K of excess traffic and chip B's two queues each absorbing 1K of excess traffic. The buffer distribution of the network device is shown in Figure 14(d).
[0262] After their respective upstream devices stop sending traffic, the buffer usage of bare chips A and B decreases sequentially. After a period of time, or after the buffer usage decreases to a fixed value, they send XON flow control to their respective upstream devices, and the upstream devices resend traffic. When there is available space in the headroom of bare chip B, bare chip B cancels the flow control over bare chip A.
[0263] The following example illustrates the use of Credit flow control mode between devices.
[0264] In Credit-based flow control mode, the sending device (TX) and receiving device (RX) need to maintain credit-related information and interact to determine whether the TX device should stop or send traffic. One representation is as follows: the TX device maintains a transmission cumulative value (FCTBS), which increments with each packet sent; the RX device maintains a reception cumulative value (ABR), which increments with each packet received. Simultaneously, the RX's buffer availability is recorded as RAVAIL. The flow control reference information for the devices is FCCL = ABR + RAVAIL. The RX device needs to periodically send FCCL to the TX device. When the TX device determines that FCBPS > FCCL, it stops sending traffic. The following embodiments are based on this representation.
[0265] Example 3:
[0266] Based on the network architecture shown in Figure 7, the cache space of the entire network device chip is divided into only static reserved space, without deploying shared space. Specifically, the ADEI / A bare chip cache is reserved space, and the BDIE / A bare chip cache is also reserved space. The global cache resource management module is implemented in the B bare chip. PFC flow control mode is used between bare chips, and Credit flow control mode is used between devices.
[0267] In the simplified scenario of Example 3, the process described in S403 is used. The timing of physical buffer occupancy in the network device is shown in Figure 15. At all times, raw chip A needs to continuously output its buffer occupancy value to raw chip B (even if raw chip A's buffer occupancy value is 0) for raw chip B to calculate the total available buffer space RAVAIL. Simultaneously, the entire chip needs to calculate the flow control reference information FCCL based on the received packet accumulation amount ABR and the total available buffer space RAVAIL, and periodically send it to the front-end device. The specific flow control process in this simplified scenario is as follows:
[0268] During the transmission and forwarding of data traffic by the network device, congestion occurs in bare chip B, and the cache of bare chip B accumulates. At time t1, the cache of bare chip B occupies 100K of cache space. Bare chip B generates flow control on bare chip A. At this time, the cache distribution of the network device is shown in Figure 15(a).
[0269] Subsequently, B absorbs the excess traffic from A's bare chip, occupying the bias cache space. At time t2, the bias cache space is full, and the cache distribution of the network device is shown in Figure 15(b).
[0270] The bare chip A receives the flow control from bare chip B and stops sending traffic to bare chip B. The buffer of bare chip A gradually accumulates. At time t3, bare chip A occupies 2K of buffer space, and the buffer distribution of the network device is shown in (c) of FIG. 15. When bare chip A occupies 2K of buffer space, the upstream device determines that FCTBS>FCCL and stops sending traffic.
[0271] After the upstream device stops sending traffic, the buffer occupancy of bare chip A and bare chip B decreases sequentially. After a period of time, B determines that there is available buffer space and cancels the flow control on A. The upstream device determines that FCTBS<FCCL and resumes sending traffic.
[0272] In the complex scenario of Embodiment 3, the process of the above S403 is adopted. Assuming that the ADEI / bare chip A occupies the reserved space in advance for some reasons, the occupancy sequence of physical buffers in the network device under the complex scenario is shown in FIG. 16, and the flow control process under the complex scenario is as follows:
[0273] At all times, bare chip A needs to continuously output its buffer occupancy value to bare chip B (even if the buffer occupancy value of bare chip A is 0), which is used to calculate the available buffer space RAVAIL of the entire chip. Meanwhile, the chip needs to calculate the flow control reference information FCCL according to the accumulated received packet amount ABR and the available buffer space RAVAIL of the entire chip, and periodically send it to the upstream device.
[0274] Due to some reasons (such as burst of small packets or long packet processing time, etc.), the buffer of bare chip A accumulates, and finally occupies 1K of buffer space (which is equivalent to occupying the reserved space in advance). When congestion occurs on bare chip B, the buffer of bare chip B accumulates. At time t1, the buffer of bare chip B occupies 100K of buffer space, and the buffer distribution of the network device at this time is shown in (a) of FIG. 16.
[0275] At time t1, the buffer of bare chip B occupies 100K, bare chip B generates flow control for bare chip A, absorbs the overcharged traffic of bare chip A, and occupies the bias buffer space. At time t2, the bias buffer space is fully occupied, and the buffer distribution of the network device is shown in (b) of FIG. 16.
[0276] Bare chip A receives the flow control from bare chip B and stops sending traffic to bare chip B. The buffer of bare chip A continues to accumulate. At time t3, the buffer of bare chip A accumulates another 1K, and the buffer distribution of the network device is shown in (c) of FIG. 16. At this time, the upstream device determines that FCTBS>FCCL and stops sending traffic.
[0277] After the current-stage device stops sending traffic, the buffer occupancy of the A die and the B die decreases sequentially. After a period of time, B determines that there is available buffer space and cancels the flow control for A. The previous-stage device determines that FCTBS<FCCL, and the previous-stage device restarts sending traffic.
[0278] Example 4
[0279] Based on the network architecture illustrated in FIG. 11, the buffer space of the entire network device chip is only divided into static reserved space, and no shared space is deployed. That is, the buffer of ADEI / A die is reserved space, and the buffer of BDIE / B die is reserved space. Global buffer resource management is implemented in the B die. A PFC flow control mode is adopted between the dies, and a Credit flow control mode is adopted between devices.
[0280] In the simple scenario of Embodiment 4, when the process of the foregoing S403 is adopted, the occupancy timing of physical buffers in the network device is shown in FIG. 17. At all times, the A die needs to continuously output the buffer occupancy value to the B die (even if the buffer occupancy value of the A die is 0), which is used by the B die to calculate the available buffer space RAVAIL of the entire chip. Meanwhile, the two queues of the entire chip need to calculate the flow control reference information FCCL according to their respective received message accumulation ABR and the available buffer space RAVAIL of the two queues of the entire chip, and periodically send the flow control reference information FCCL to the previous-stage device. The specific flow control process in the simple scenario is as follows:
[0281] During the transmission and forwarding of data traffic by the network device, congestion occurs in the two queues of the B die, and the buffer of the B die accumulates. At time t1, the two queues of the B die occupy 50K buffer space respectively, and the two queues of the B die respectively generate flow control for the A die. At this time, the buffer distribution of the network device is shown in (a) in FIG. 17.
[0282] Thereafter, the two queues of the B die respectively absorb the overcharged traffic from the A die, and the two queues occupy bias buffer spaces respectively. At time t2, the two bias buffer spaces are fully occupied, and the buffer distribution of the network device is shown in (b) in FIG. 17.
[0283] After the A die receives the flow control from the two queues of the B die, the two queues of the A die stop sending traffic to the B die respectively, and the buffers of the two queues of the A die gradually accumulate. At time t3, the two queues of the A die occupy 2K buffer space respectively, and the buffer distribution of the network device is shown in (c) in FIG. 17. When the two queues of the A die occupy 2K buffer space respectively, the respective previous-stage devices of the two queues respectively determine that FCTBS>FCCL and stop sending traffic.
[0284] After the current-level device stops sending traffic, the buffer occupation of the A bare die and the B bare die decreases sequentially. After a period of time, B determines that there is available buffer space and cancels the flow control on A. The respective upstream devices of the two queues determine that FCTBS<FCCL, and the respective upstream devices of the two queues restart sending traffic.
[0285] In the complex scenario of Embodiment 4, the process of the above S403 is adopted. Assuming that the two queues of ADEI / A bare die occupy the reserved space in advance for some reasons, the occupation timing of physical buffers in the network device in the complex scenario is shown in Figure 18. At all times, the A bare die needs to continuously output the buffer occupation value to the B bare die (even if the buffer occupation value of the A bare die is 0), for the B bare die to calculate the available buffer space RAVAIL of the entire die. Meanwhile, the two queues of the entire die need to calculate the flow control reference information FCCL according to their respective accumulated received message quantity ABR and the available buffer space RAVAIL of the two queues of the entire die, and periodically send it to the upstream devices. The flow control process in the complex scenario is as follows:
[0286] For the two queues of the A bare die, due to some reasons (such as small packet burst or long packet processing time), the buffer of the A bare die accumulates, and the two queues finally occupy 1K of buffer space respectively (which is equivalent to occupying the reserved space in advance). When congestion occurs in the two queues of the B bare die respectively, the buffer of the B bare die accumulates. At time t1, the buffer occupation of the two queues of the B bare die is 50K buffer space respectively, and the buffer distribution of the network device at this time is shown in (a) in Figure 18.
[0287] At time t1, the two queues of the B bare die each occupy 50K of buffer, the two queues of the B bare die respectively generate flow control on the A bare die to absorb the overcharged traffic of the A bare die, and the two queues occupy the bias buffer space respectively. At time t2, the two bias buffer spaces of the B bare die are fully occupied, and the buffer distribution of the network device is shown in (b) in Figure 18.
[0288] After the A bare die receives the flow control from the two queues of the B bare die, the two queues of the A bare die stop sending traffic to the B bare die respectively. The buffers corresponding to the two queues of the A bare die continue to accumulate. At time t3, the buffers of the two queues of the A bare die continue to accumulate by 1K respectively, and the buffer distribution of the network device at this time is shown in (c) in Figure 18. At this time, the respective upstream devices of the two queues respectively determine that FCTBS>FCCL and stop sending traffic.
[0289] After the respective upstream devices of the two queues stop sending traffic, the buffer occupation of the A bare die and the B bare die decreases sequentially. After a period of time, B determines that there is available buffer space and cancels the flow control on A. The respective upstream devices of the two queues determine that FCTBS<FCCL, and the respective upstream devices restart sending traffic.
[0290] Example 5
[0291] In the network architecture shown in Figure 7, the entire chip's cache space is divided into static reserved space and shared space. Global cache resource management is implemented in the B bare chip. PFC flow control mode is used between bare chips, and Credit flow control mode is used between devices.
[0292] Based on the scenario illustrated in Figure 7, assuming the static reserved space of this network device is 0.95K, and this reserved space is only distributed in bare chip A, the remaining 1.05K space in bare chip ADEI / A (excluding the reserved space) is shared space, and the entire buffer space of bare chip B is shared space. Since the reserved space (0.95K) < 2K, this network device is determined to be in a short-range scenario. During message transmission and reception, each bare chip maintains the received message accumulation amount (ABR), i.e., the ABR is incremented for each received message. Bare chip A calculates the flow control reference information (FCCL) of the network device based on the ABR and RAVAIL, and periodically sends the FCCL to the preceding device. The timing of physical buffer occupancy in the network device in the short-range scenario is shown in Figure 8, and the flow control process in the short-range scenario is as follows:
[0293] During the transmission and forwarding of data traffic by the network device, congestion occurs in bare chip B, causing buffer accumulation in bare chip B. At time t1, the buffer space occupied by bare chip B is 100K. Bare chip B generates flow control over bare chip A. At this time, the buffer distribution of the network device is shown in Figure 8(a). During this period, the network device needs to maintain the received packet accumulation amount ABR. That is, for each received packet, ABR is incremented, but FCCL is not calculated, and it is not sent to the upstream device.
[0294] Subsequently, bare chip B absorbs the excess traffic from bare chip A, occupying the bias cache space. At time t2, the bias cache space is full, and the cache distribution of the network device is shown in Figure 8(b).
[0295] Upon receiving flow control from bare chip B, bare chip A stops sending traffic to bare chip B. The buffer of bare chip A gradually accumulates. At time t3, bare chip A occupies 1.05K of buffer space. The buffer distribution of the network devices at this time is shown in Figure 8(c). When bare chip A occupies 1.05K of buffer space, it calculates that the available buffer space RAVAIL of the entire chip is 0.95K remaining. At this time, bare chip A calculates its respective flow control reference information FCCL based on ABR and RAVAIL, and periodically sends the FCCL to its respective upstream device. Afterwards, the buffer of bare chip A continues to accumulate.
[0296] At time t4, the reserved space of 0.95K in bare die A is fully occupied, and the cache distribution of the network device at this time is shown in (d) of Figure 8. At this time, the upstream device determines that FCTBS>FCCL, activates Credit-based flow control, and stops transmitting traffic.
[0297] After the upstream device stops transmitting traffic, the cache occupancy of bare die A and bare die B decreases sequentially. After a period of time, bare die B withdraws the flow control on bare die A. The upstream device determines that FCTBS<FCCL, and the upstream device restarts transmitting traffic.
[0298] Based on the scenario illustrated in Figure 7, it is assumed that the static reserved space of this network device is 22K, and the reserved space is distributed on bare die A and bare die B. The reserved space is deployed as 2K in ADEI / bare die A, 20K in BDIE / bare die A, and in BDEI / bare die B, except for the reserved space, the remaining 80K space is the shared space. Since the headroom space (22K) > 2K, this network device is determined to be in a long-haul scenario. Assuming that ADEI / bare die A does not occupy the headroom space in advance, this scenario is referred to as the simple long-haul scenario. During the transmission and reception of packets, each bare die maintains the accumulated received packet amount ABR, that is, ABR is accumulated every time a packet is received. Bare die B calculates the flow control reference information FCCL of the network device based on ABR and RAVAIL, and periodically sends FCCL to the upstream device. The occupancy sequence of physical cache in the network device under the simple long-haul scenario is shown in Figure 9, and the flow control process under the simple long-haul scenario is as follows:
[0299] During the process of transmitting and forwarding data traffic by the network device, congestion occurs in bare die B. At time t1, the cache of bare die B occupies 80K of cache space, and the cache distribution of the network device at this time is shown in (a) of Figure 9. At time t1, 22K of available cache space RAVAIL remains in the entire network device chip. Before time t1, the network device needs to maintain the accumulated received packet amount ABR, that is, ABR is accumulated every time a packet is received, but FCCL is not calculated and is not sent to the upstream device. After time t1, bare die A calculates its respective flow control reference information FCCL based on ABR and RAVAIL, and periodically sends FCCL to its respective upstream device.
[0300] After time t1, the cache accumulation of bare die B continues. At time t2, the cache of bare die B continues to accumulate by 20K, and the cache distribution of the network device at this time is shown in (b) of Figure 9. At time t2, bare die B initiates flow control on bare die A, absorbs the overcharged traffic from bare die A, and occupies the bias cache space. At time t3, the bias cache space is fully occupied, and the cache distribution of the network device is shown in (c) of Figure 9.
[0301] After the A bare die stops sending traffic to the B bare die, the A bare die continues to absorb overcharge traffic from the upstream device until the upstream device stops sending traffic. At time t4, the A bare die absorbs 2K of overcharge traffic, and the buffer distribution of the network device is shown in (d) of FIG. 9. At this time, the upstream device determines that FCTBS>FCCL and stops sending traffic.
[0302] After the upstream device stops sending traffic, the buffer occupancy of the A bare die and the B bare die decreases sequentially. After a period of time, the B bare die withdraws the flow control for the A bare die. After determining that FCTBS<FCCL, the upstream device resumes sending traffic.
[0303] Based on the scenario illustrated in FIG. 7, assume that the static reserved space of the network device is 22K, and the reserved space is distributed on the A bare die and the B bare die. 2K of the reserved space is deployed in ADEI / A bare die, 20K is deployed in BDIE / A bare die, and apart from the reserved space in BDEI / B bare die, the remaining 80K space is the share space. Since headroom space (22K) > 2K, the network device is determined to be in a long-haul scenario. Assume that ADEI / A bare die occupies the reserved space in advance for some reasons, which is referred to as a complex long-haul scenario. During the transmission and reception of packets, each bare die maintains an accumulated received packet count ABR, that is, every time a packet is received, ABR is accumulated. The B bare die calculates the flow control reference information FCCL of the network device based on ABR and RAVAIL, and periodically sends the FCCL to the upstream device. The occupation timing sequence of physical buffers in the network device in a simple long-haul scenario is shown in FIG. 19, and the flow control process in the complex long-haul scenario is as follows:
[0304] During all time periods, the A bare die needs to continuously output the buffer occupancy value to the B bare die for calculating the available buffer space RAVAIL of the whole die. Due to some reasons (such as burst of small packets or long packet processing time, etc.), the buffer of the A bare die accumulates, and finally occupies 1K buffer space (equivalent to occupying the reserved space in advance). When congestion occurs on the B bare die, the buffer of the B bare die accumulates. At time t1, the buffer occupancy of the B bare die is 80K, and the buffer distribution of the network device is shown in (a) of FIG. 19. Before time t1, the network device needs to maintain the accumulated received packet count ABR, that is, every time a packet is received, ABR is accumulated, but FCCL is not calculated nor sent to the upstream device. After time t1, the A bare die calculates respective flow control reference information FCCL based on ABR and RAVAIL, and periodically sends the FCCL to respective upstream devices.
[0305] When the cache of B bare chip occupies 80K cache space, B bare chip comprehensively determines that the shared cache space of the whole chip is used up, and at this time, the remaining available space RAVAIL of the whole chip is 21K. B bare chip calculates flow control reference information FCCL based on ABR and RAVAIL, and periodically sends FCCL to the upstream device.
[0306] After time t1, the cache records of B bare chip accumulate. By time t2, the cache of B bare chip continues to accumulate by 20K. At this time, the cache distribution of the network device is shown as (b) in Figure 19. At time t2, B bare chip generates flow control on A bare chip to absorb the overcharged traffic from A bare chip and occupies the bias cache space. At time t3, the bias cache space is fully occupied, and the cache distribution of the network device is shown as (c) in Figure 19.
[0307] After A bare chip stops sending traffic to B bare chip, A bare chip continues to absorb the overcharged traffic from the upstream device until the upstream device stops sending traffic. At time t4, A bare chip absorbs 1K of overcharged traffic, and the cache distribution of the network device is shown as (d) in Figure 19. At this time, the upstream device determines that FCTBS>FCCL and stops sending traffic.
[0308] After the upstream device stops sending traffic, the cache occupation of A bare chip and B bare chip decreases sequentially. After a period of time, B bare chip cancels the flow control on A bare chip. After determining that FCTBS<FCCL, the upstream device resumes sending traffic.
[0309] Embodiment 6
[0310] In Embodiment 6, based on the network architecture illustrated in Figure 11, the cache space of the whole chip is divided into static reserved space (reserved) and shared space (share space). Global cache resource management is implemented in B bare chip. A PFC flow control mode is adopted between bare chips, and a Credit flow control mode is adopted between devices.
[0311] Based on the scenario illustrated in Figure 11, assuming the static reserved space of this network device is 0.95K*2, and the reserved space is only distributed in bare chip A, the remaining 1.05K*2 space in bare chip ADEI / A, besides the reserved space, is shared space, and the space of bare chip B's cache is entirely shared space. Since the single cache reserved space (0.95K) < 2K, this network device is determined to be a short-range scenario. During message transmission and reception, each queue of each bare chip maintains the received message accumulation amount ABR, that is, the ABR is accumulated for each received message. Bare chip B calculates the flow control reference information FCCL of this network device based on ABR and RAVAIL, and periodically sends the FCCL to the front-end device. The timing of physical cache occupancy in the network device in the short-range scenario is shown in Figure 12, and the flow control process in the short-range scenario is as follows:
[0312] During the transmission and forwarding of data traffic by the network device, congestion occurs in both queues of bare chip B, leading to buffer accumulation in the buffer of bare chip B. At time t1, each queue of bare chip B occupies 50K of buffer space. Both queues of bare chip B generate flow control for bare chip A. The buffer distribution of the network device at this time is shown in Figure 12(a). Before time t1, the network device needs to maintain the received packet accumulation amount (ABR), that is, the ABR is incremented for each received packet, but the FCCL is not calculated or sent to the preceding device. After time t1, the two queues of bare chip A calculate their respective flow control reference information (FCCL) based on their respective ABR and RAVAIL, and periodically send the FCCL to their respective preceding devices.
[0313] Subsequently, the two queues of bare chip B absorb the excess traffic from bare chip A, each occupying a bias buffer space. At time t2, the two bias buffer spaces of bare chip B are full, and the buffer distribution of the network device is shown in Figure 12(b).
[0314] Raw chip A receives flow control from two queues of raw chip B. Both queues of raw chip A stop sending traffic to raw chip B, and the buffer of raw chip A gradually accumulates. At time t3, each queue of raw chip A occupies 1.05K of buffer space. The buffer distribution of the network device at this time is shown in Figure 12(c). When each queue of raw chip A occupies 1.05K of buffer space, raw chip A calculates that the remaining available buffer space RAVAIL is 0.95K*2. At this time, the two queues of raw chip A calculate their respective flow control reference information FCCL based on ABR and RAVAIL, and periodically send the FCCL to their respective upstream devices. Afterwards, the buffer of raw chip A continues to accumulate.
[0315] At time t4, the 0.95K reserved spaces corresponding to the 2 queues of the A bare die are fully occupied respectively, and the buffer distribution of the network device is shown in (d) of FIG. 12. At this time, the respective upstream devices of the 2 queues respectively determine that FCTBS>FCCL, and stop transmitting traffic.
[0316] After the respective upstream devices stop transmitting traffic, the buffer occupancy of the A bare die and the B bare die decreases sequentially. After a period of time, the B bare die cancels the flow control for the A bare die. The respective upstream devices of the 2 queues determine that FCTBS<FCCL, and the upstream devices restart transmitting traffic.
[0317] Based on the scenario illustrated in FIG. 11, it is assumed that the static reserved space of the network device is 22K*2, and the reserved space is distributed on the A bare die and the B bare die. The reserved space is deployed as 2K*2 in ADEI / A bare die, and 20K in BDIE / A bare die. Except for the reserved space in BDEI / B bare die, the remaining 30K*2 space is the shared space. Since the reserved space of a single buffer (22K) > 2K, the network device is determined to be in a long-distance scenario. Assuming that ADEI / A bare die does not occupy the reserved space in advance, this scenario is referred to as a simple long-distance scenario. During the transmission and reception of messages, each bare die maintains the accumulated received message amount ABR, that is, every time a message is received, ABR is accumulated. The B bare die calculates the flow control reference information FCCL of the network device according to ABR and RAVAIL, and periodically sends FCCL to the upstream device. The occupancy sequence of physical buffers in the network device in the simple long-distance scenario is shown in FIG. 13, and the flow control process in the simple long-distance scenario is as follows:
[0318] During the process of transmitting and forwarding data traffic by the network device, congestion occurs respectively in the 2 queues of the B bare die. At time t1, the 2 queues of the B bare die occupy 30K buffer space respectively, and the buffer distribution of the network device at this time is shown in (a) of FIG. 13. At time t1, the remaining available buffer space RAVAIL of the entire network device chip is 22K*2. Before time t1, the network device needs to maintain the accumulated received message amount ABR, that is, every time a message is received, ABR is accumulated, but FCCL is not calculated, nor is it sent to the upstream device. After time t1, the B bare die calculates the respective flow control reference information FCCL according to the respective ABR and RAVAIL of the 2 queues, and periodically sends FCCL to the respective upstream devices.
[0319] After time t1, the buffer records of the 2 queues of the B bare die accumulate. By time t2, the buffer occupancy of the 2 queues of the B bare die continues to accumulate by 20K respectively, and at this time, the buffer distribution of the network device is shown in (b) of Figure 13. At time t2, the 2 queues of the B bare die respectively generate flow control for the A bare die to absorb the overcharged traffic of the A bare die, and the 2 queues each occupy one bias buffer space. At time t3, the 2 bias buffer spaces of the B bare die are fully occupied, and the buffer distribution of the network device is shown in (c) of Figure 13.
[0320] After the 2 queues of the A bare die respectively stop sending traffic to the B bare die, the A bare die continues to absorb the overcharged traffic from the upstream device until the upstream device stops sending traffic. At time t4, the 2 queues of the A bare die respectively absorb 2K of overcharged traffic, and the buffer distribution of the network device is shown in (d) of Figure 13. At this time, the respective upstream devices of the 2 queues determine that FCTBS > FCCL and stop sending traffic.
[0321] After the respective upstream devices of the 2 queues stop sending traffic, the buffer occupancy of the A bare die and the B bare die decreases sequentially. After a period of time, the 2 queues of the B bare die respectively cancel the flow control for the A bare die. The upstream devices of the 2 queues determine that FCTBS < FCCL, and the respective upstream devices resume sending traffic.
[0322] Based on the scenario illustrated in Figure 11, it is assumed that the static reserved space of the network device is 22K*2, and the reserved space is distributed on the A bare die and the B bare die. 2K*2 of the reserved space is deployed in ADEI / A bare die, 20K is deployed in BDIE / A bare die, and excluding the reserved space in BDEI / B bare die, the remaining 30K*2 space is the share space. Since the reserved space of a single buffer (22K) > 2K, the network device is determined to be a long-haul scenario. Assuming that ADEI / A bare die occupies the reserved space in advance for some reasons, this is called a complex long-haul scenario. During the transmission and reception of packets, each bare die maintains the accumulated received packet amount ABR, that is, every time a packet is received, ABR is accumulated. The B bare die calculates the flow control reference information FCCL of the network device based on ABR and RAVAIL, and periodically sends FCCL to the upstream device. The occupation timing of physical cache in the network device under the complex long-haul scenario is shown in Figure 20, and the flow control process under the complex long-haul scenario is as follows:
[0323] During all time periods, the A bare die needs to continuously output the buffer occupancy value to the B bare die, which is used to calculate the available buffer space RAVAIL of the entire chip. Due to certain reasons (such as burst of small packets or long packet processing time, etc.), the buffers of the 2 queues of the A bare die accumulate, and eventually the 2 queues respectively occupy 1K of buffer space (equivalent to occupying the reserved space in advance). When congestion occurs respectively in the 2 queues of the B bare die, the buffers of the B bare die accumulate. At time t1, the 2 queues of the B bare die respectively occupy 30K of buffer space, and the buffer distribution of the network device at this time is shown in (a) of Figure 20. Before time t1, the network device needs to maintain the accumulated received packet amount ABR, that is, ABR is accumulated every time a packet is received, but FCCL is not calculated, nor is it sent to the upstream device. After time t1, the B bare die respectively calculates the respective flow control reference information FCCL according to the respective ABR and RAVAIL of the 2 queues, and periodically sends FCCL to the respective upstream devices.
[0324] When the 2 queues of the B bare die respectively occupy 30K of buffer space, the B bare die comprehensively determines that the shared buffer space of the 2 queues of the entire chip is used up, and at this time the available space RAVAIL of the 2 queues of the entire chip respectively remains 21K. At this time, the 2 queues of the B bare die respectively calculate the flow control reference information FCCL according to the respective ABR and RAVAIL, and periodically send FCCL to the upstream device.
[0325] After time t1, the buffer record of the B bare die accumulates. By time t2, the buffers of the 2 queues of the B bare die continue to accumulate 20K respectively, and the buffer distribution of the network device at this time is shown in (b) of Figure 20. At time t2, the 2 queues of the B bare die respectively generate flow control for the A bare die to absorb the overcharged traffic from the A bare die and occupy the bias buffer space. At time t3, the 2 bias buffer spaces of the B bare die are fully occupied, and the buffer distribution of the network device is shown in (c) of Figure 20.
[0326] After the 2 queues of the A bare die respectively stop sending traffic to the B bare die, the 2 queues of the A bare die continue to absorb the overcharged traffic from the upstream devices until the upstream devices stop sending traffic. At time t4, the 2 queues of the A bare die respectively absorb 1K of overcharged traffic, and the buffer distribution of the network device is shown in (d) of Figure 20. At this time, the upstream devices of the 2 queues determine that FCTBS>FCCL and stop sending traffic.
[0327] After the respective upstream devices of the 2 queues stop sending traffic, the buffer occupancies of the A bare die and the B bare die decrease sequentially. After a period of time, the 2 queues of the B bare die respectively cancel the flow control for the A bare die. The upstream devices of the 2 queues determine that FCTBS<FCCL, and the respective upstream devices resume sending traffic.
[0328] The foregoing mainly describes the solution provided in this application. Accordingly, this application also provides a flow control device for implementing various functions in the above method embodiments.
[0329] In some embodiments, the flow control device includes hardware structures and / or software modules corresponding to the execution of each function in order to achieve the above-described functions. Those skilled in the art will readily recognize that, based on the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0330] This application embodiment can divide the flow control device into functional modules according to the above method embodiment. For example, each function can be divided into a separate functional module, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware or as a software functional module. It should be noted that the module division in this application embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.
[0331] On one hand, this application provides a flow control device 210, which is used to implement the scheme illustrated in FIG4 of the above method embodiment. As shown in FIG21, the flow control device 210 may include a first determining unit 2101 and a processing unit 2102.
[0332] The first determining unit 2101 is used to execute operation S402 in the method illustrated in FIG4. The processing unit 2102 is used to execute operation S403 in the method illustrated in FIG4.
[0333] On the other hand, this application embodiment provides another flow control device 220, which is used to implement the scheme illustrated in FIG5 of the above method embodiment. As shown in FIG22, the flow control device 220 may include a second determining unit 2201, a processing unit 2202 and a third determining unit 2203.
[0334] The second determining unit 2201 is used to execute operation S502 in the method illustrated in FIG5. The processing unit 2202 is used to execute operation S503 or S504 in the method illustrated in FIG5. The third determining unit 2203 is used to determine that there is no available space in the reserved space of the second bare chip cache during the flow control process of the second queue.
[0335] In another aspect, embodiments of this application provide a flow control system, which includes a flow control device 210 or a flow control device 220.
[0336] Furthermore, embodiments of this application provide a network device 230, which can be a chip or other form. This network device 230 can be used to perform any of the operations illustrated in FIG4 or FIG5 above.
[0337] As shown in Figure 23, the network device 230 provided in this embodiment may include a processor 2301, a bus 2302, a communication interface 2303, and a memory 2304. The processor 2301, the memory 2304, and the communication interface 2303 communicate with each other via the bus 2302. It should be understood that this application does not limit the number of processors and memories in the network device 230.
[0338] Bus 2302 can be a PCI bus, an Extended Industry Standard Architecture (EISA) bus, or a UB bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, only one line is used in Figure 23, but this does not imply that there is only one bus or one type of bus. Bus 2302 can include pathways for transmitting information between various components of network device 230 (e.g., memory 2304, processor 2301, communication interface 2303).
[0339] Processor 2301 may include any one or more processors such as CPU, graphics processing unit (GPU), microprocessor (MP), or digital signal processor (DSP).
[0340] The memory 2304 may include volatile memory, such as random access memory (RAM). The processor 2301 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0341] The communication interface 2303 uses transceiver modules such as, but not limited to, network interface cards and transceivers to enable communication between the network device 230 and other devices or communication networks.
[0342] The memory 2304 stores executable program code, and the processor 2301 executes the executable program code to implement the functions described in the aforementioned method embodiments. That is, the memory 2304 stores instructions for executing the above-described flow control method.
[0343] Furthermore, embodiments of this application also provide a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product is run on at least one computing device, it causes the at least one computing device to perform the operational steps of the method in the above-described method embodiments.
[0344] Furthermore, embodiments of this application also provide a computer-readable storage medium. The computer-readable storage medium can be any available medium capable of being stored by a computing device, or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to perform the operational steps of the method described in the above method embodiments.
[0345] The method steps in this embodiment can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, portable hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and storage medium can reside in an ASIC. Alternatively, the ASIC can reside in a computing device. Of course, the processor and storage medium can also exist as discrete components in the computing device.
[0346] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are performed entirely or partially. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user equipment, or other programmable device. The computer program or instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. For example, the computer program or instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, hard disk, or magnetic tape; it can also be an optical medium, such as a digital video disc (DVD); or it can be a semiconductor medium, such as a solid-state drive (SSD). The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A flow control method, characterized in that, Applied to a network device including a chip, said chip comprising a plurality of bare chips in which cache is deployed; the method includes: The cache occupancy of the first queue for the first bare chip is determined to meet a first condition; the first bare chip is one of the plurality of bare chips, and the first condition is used to indicate that the first queue is congested on the first bare chip; The preceding bare chip of the first bare chip performs flow control on the first queue; wherein, if the first bare chip is the first-level bare chip of the network device, the preceding bare chip is the preceding device of the network device.
2. The method according to claim 1, characterized in that, The method further includes: If the network device meets the third condition, the process of determining that the cache occupancy of the first queue for the first bare chip meets the first condition is executed, and the first queue of the first bare chip in the previous stage of the first bare chip is subjected to flow control; the third condition is used to indicate that the network device and the previous stage device are in a short-distance scenario.
3. The method according to claim 2, characterized in that, The method further includes: If the network device does not meet the third condition, it is determined that the second queue's cache occupancy for the plurality of bare chips meets the second condition; the second condition is used to indicate that the second queue is congested in the network device. The preceding equipment of the network device performs flow control on the second queue; During the flow control process for the second queue, if the reserved space of the second bare chip cache is unavailable, the first-level bare chip preceding the second bare chip performs flow control on the second queue; the second bare chip is one of the plurality of bare chips; the reserved space is the space configured in the cache for flow control.
4. A flow control method, characterized in that, Applied to a network device including a chip, said chip comprising a plurality of bare chips in which cache is deployed; the method includes: The cache occupancy of the second queue for the plurality of bare chips is determined to meet a second condition; the second condition is used to indicate that the second queue is congested in the network device; the second bare chip is one of the plurality of bare chips. The preceding equipment of the network device performs flow control on the second queue; During the flow control process for the second queue, if there is no available space in the reserved space of the second bare chip cache, the first-level bare chip preceding the second bare chip will perform flow control on the second queue; the reserved space is the space configured in the cache for flow control.
5. The method according to claim 4, characterized in that, The method further includes: If the network device does not meet the third condition, the process of determining that the cache occupancy of the second queue for the multiple bare chips meets the second condition is executed, and the front-end device of the network device performs flow control on the second queue; the third condition is used to indicate that the network device and the front-end device are in a short-distance scenario.
6. The method according to claim 5, characterized in that, The method further includes: If the network device meets the third condition, it is determined that the cache occupancy of the first queue for the first bare chip meets the first condition, and the first queue of the first bare chip preceding the first bare chip performs flow control.
7. The method according to claim 2, 3, 5, or 6, characterized in that, The network device's cache includes reserved space, which is used to store traffic during flow control or is controlled by the front-end device or bare chip. The third condition includes: the reserved space of the network device is less than the cache size of the first-level bare chip of the network device divided by Z, where Z is less than or equal to m, and m is the number of queues corresponding to one cache of the first-level bare chip. or, The third condition includes: the physical distance between the network device and the front-end device is less than a distance threshold.
8. The method according to any one of claims 1-3 or 6 or 7, characterized in that, The network device's cache includes a shared space, which is shared by multiple queues; the first condition includes: the space corresponding to the first queue in the remaining space of the shared space within the first bare chip, which is greater than or equal to the space in the first bare chip. The preceding stage bare chip of the first bare chip performs flow control on the first queue, including: sending flow control information to the preceding stage bare chip, the flow control information being used to indicate to stop sending traffic.
9. The method according to any one of claims 1-3 or 6 or 7, characterized in that, The method further includes: sending first reference information to the previous stage bare chip, the first reference information being used to indicate the remaining space in the cache of the first bare chip in the controllable portion of the previous stage bare chip; The first condition includes the first reference information indicating that there is no remaining space.
10. The method according to any one of claims 3-6, characterized in that, The network device's cache also includes a shared space, and the second condition includes: the space corresponding to the second queue in the remaining space of the network device's shared space that is greater than or equal to the space of the network device's shared space; The preceding device of the network device performs flow control on the second queue, including sending flow control information to the preceding device, wherein the flow control information is used to instruct the cessation of traffic transmission.
11. The method according to any one of claims 3-6, characterized in that, The method further includes: sending second reference information to the front-end device, the second reference information being used to indicate the remaining space in the cache of the network device that is controllable by the front-end device; The second condition includes: determining, based on the second reference information, that there is no remaining space in the controllable portion of the front-end device in the cache of the network device.
12. The method according to claim 11, characterized in that, The second reference information is the received cumulative value plus the available buffer amount; The received cumulative value is used to indicate the cumulative number of received messages; The step of determining that there is no remaining space based on the second reference information includes: if the cumulative value of transmission is greater than the second reference information, then it is determined that there is no remaining space in the controllable portion of the front-end device in the cache of the network device; The cumulative transmission value is used to indicate the cumulative number of messages transmitted.
13. The method according to claim 8 or 10, characterized in that, The sharing coefficient of the queue is configured. The space corresponding to the third queue in the target remaining space includes: the target remaining space multiplied by the sharing coefficient of the third queue; the third queue is any queue that enters the network device for forwarding.
14. The method according to any one of claims 1-13, characterized in that, The method further includes: When a message from the first queue enters the cache queue of the first bare chip, the first queue increments the cache occupancy cnt within the first bare chip; when a message from the first queue is dequeued from the cache queue of the first bare chip, cnt is decremented.
15. The method according to any one of claims 3-6 or 10-12, characterized in that, The cache usage of the second queue for the plurality of bare chips includes: the sum of the cache usage of each bare chip in the second queue for the plurality of bare chips.
16. A flow control device, characterized in that, An apparatus for use in network devices comprising a chip including a plurality of bare chips in which cache is deployed; the apparatus includes: The first determining unit is configured to determine that the cache occupancy of the first queue for the first bare chip meets a first condition; the first bare chip is one of the plurality of bare chips, and the first condition is configured to indicate that the first queue is congested on the first bare chip; The processing unit is configured to, after the first determining unit determines that the cache occupancy of the first queue on the first bare chip meets the first condition, control the preceding bare chip of the first bare chip to perform flow control on the first queue; wherein, if the first bare chip is the first-level bare chip of the network device, the preceding bare chip is the preceding device of the network device.
17. The apparatus according to claim 16, characterized in that, The first determining unit is specifically used for: If the network device meets the third condition, the process of determining that the cache occupancy of the first queue for the first bare chip meets the first condition is executed; the third condition is used to indicate that the network device and the front-end device are in a short-distance scenario.
18. The apparatus according to claim 17, characterized in that, The apparatus further includes a second determining unit, configured to: determine that the cache occupancy of the second queue for the plurality of bare chips satisfies the second condition when the network device does not meet the third condition; the second condition is used to indicate that the second queue is congested in the network device; The processing unit is further configured to: control the front-end device of the network device to perform flow control on the second queue; The device further includes a third determining unit, used to: determine that there is no available space in the reserved space of the second bare chip cache during the flow control process of the second queue; The processing unit is further configured to: when the third determining unit determines that there is no available space in the reserved space of the second bare chip cache, control the first-level bare chip preceding the second bare chip to perform flow control on the second queue; the second bare chip is one of the plurality of bare chips; the reserved space is the space configured in the cache for flow control.
19. A flow control device, characterized in that, An apparatus for use in network devices comprising a chip including a plurality of bare chips in which cache is deployed; the apparatus includes: The second determining unit is used to determine that the cache occupancy of the second queue for the plurality of bare chips meets a second condition; the second condition is used to indicate that the second queue is congested in the network device; the second bare chip is one of the plurality of bare chips. The processing unit is configured to, after the second determining unit determines that the cache occupancy of the second queue for the plurality of bare chips meets the second condition, control the front-end device of the network device to perform flow control on the second queue; The third determining unit is used to: determine that there is no available space in the reserved space of the second bare chip cache during the flow control process of the second queue; The processing unit is further configured to: when the third determining unit determines that there is no available space in the reserved space of the second bare chip cache, control the first-level bare chip preceding the second bare chip to perform flow control on the second queue; the second bare chip is one of the plurality of bare chips; the reserved space is the space configured in the cache for flow control.
20. The apparatus according to claim 19, characterized in that, The second determining unit is specifically used for: If the network device does not meet the third condition, it is determined that the cache occupancy of the second queue for the multiple bare chips meets the second condition, and the front-end device of the network device performs flow control on the second queue; the third condition is used to indicate that the network device and the front-end device are in a short-distance scenario.
21. The apparatus according to claim 20, characterized in that, The device further includes: The first determining unit is configured to: determine, when the network device satisfies the third condition, that the buffer occupancy of the first queue on the first bare chip satisfies the first condition; the first bare chip is one of the plurality of bare chips, and the first condition is used to indicate that the first queue is congested on the first bare chip; The processing unit is further configured to: after the first determining unit determines that the cache occupancy of the first queue on the first bare chip meets the first condition, control the first-level bare chip preceding the first bare chip to perform flow control on the first queue; wherein, if the first bare chip is the first-level bare chip of the network device, the preceding-level bare chip is the preceding-level device of the network device.
22. The apparatus according to claim 17, 18, 20, or 21, characterized in that, The network device's cache includes reserved space, which is used to store traffic during flow control or is controlled by the front-end device or bare chip. The third condition includes: the reserved space of the network device is less than the cache size of the first-level bare chip of the network device divided by Z, where Z is less than or equal to m, and m is the number of queues corresponding to one cache of the first-level bare chip. or, The third condition includes: the physical distance between the network device and the front-end device is less than a distance threshold.
23. The apparatus according to any one of claims 16-18 or 21 or 22, characterized in that, The network device's cache includes a shared space, which is shared by multiple queues; the first condition includes: the space corresponding to the first queue in the remaining space of the shared space within the first bare chip, which is greater than or equal to the space in the first bare chip. The processing unit is specifically used to: send flow control information to the previous stage bare chip, wherein the flow control information is used to indicate to stop sending traffic.
24. The apparatus according to any one of claims 16-18 or 21 or 22, characterized in that, The device further includes: a first transmitting unit, configured to transmit first reference information to the preceding bare chip, the first reference information being used to indicate the remaining space in the cache of the first bare chip of the controllable portion of the preceding bare chip; The first condition includes the first reference information indicating that there is no remaining space.
25. The apparatus according to any one of claims 18-21, characterized in that, The network device's cache also includes a shared space, and the second condition includes: the space corresponding to the second queue in the remaining space of the network device's shared space that is greater than or equal to the space of the network device's shared space; The processing unit is specifically used to: send flow control information to the front-end device, wherein the flow control information is used to instruct the cessation of traffic transmission.
26. The apparatus according to any one of claims 18-21, characterized in that, The device further includes: a second sending unit, configured to send second reference information to the front-end device, the second reference information being used to indicate the remaining space in the cache of the network device that is controllable by the front-end device; The second condition includes: determining, based on the second reference information, that there is no remaining space in the controllable portion of the front-end device in the cache of the network device.
27. The apparatus according to claim 26, characterized in that, The second reference information is the received cumulative value plus the available buffer amount; The received cumulative value is used to indicate the cumulative number of received messages; The step of determining that there is no remaining space based on the second reference information includes: if the cumulative value of transmission is greater than the second reference information, then it is determined that there is no remaining space in the controllable portion of the front-end device in the cache of the network device; The cumulative transmission value is used to indicate the cumulative number of messages transmitted.
28. The apparatus according to claim 23 or 25, characterized in that, The sharing coefficient of the queue is configured. The space corresponding to the third queue in the target remaining space includes: the target remaining space multiplied by the sharing coefficient of the third queue; the third queue is any queue that enters the network device for forwarding.
29. The apparatus according to any one of claims 16-28, characterized in that, The device further includes a statistical unit for: When a message from the first queue enters the cache queue of the first bare chip, the first queue increments the cache occupancy cnt within the first bare chip; when a message from the first queue is dequeued from the cache queue of the first bare chip, cnt is decremented.
30. The apparatus according to any one of claims 18-21 or 25-27, characterized in that, The cache usage of the second queue for the plurality of bare chips includes: the sum of the cache usage of each bare chip in the second queue for the plurality of bare chips.
31. A network device, characterized in that, Including processor and memory; The processor is configured to execute instructions stored in the memory to cause the network device to perform the operational steps of the method as described in any one of claims 1 to 15.
32. A computer-readable storage medium, characterized in that, include: Computer software instructions; When the computer software instructions are executed on a computing device, the computing device causes the computing device to perform the method as described in any one of claims 1 to 15.
33. A computer program product, characterized in that, When the computer program product is run on a computing device, the computing device performs the method as described in any one of claims 1 to 15.