Preemptive cache management system oriented to on-chip shared cache exchange chip

By introducing a preemptive cache management system in the data center, the problems of high cache cost and low utilization in shared cache switching chips are solved, efficient utilization and fairness of caches are achieved, and the performance and stability of the data center network are improved.

CN120111017APending Publication Date: 2025-06-06XI AN JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510269147.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The on-chip cache cost of shared cache switching chips in data centers is high and has low utilization. The existing non-preemptive cache management strategy cannot effectively solve the cache overload problem.

Method used

A preemptive cache management system is designed, including five modules: admission decision-making, head-loss selector, arbitrator, head-loss executor and demultiplexer. Through active packet loss and dynamic adjustment of cache resource allocation, efficient utilization and fairness of cache is achieved.

Benefits of technology

It significantly improves the efficiency and fairness of cache utilization, enhances the ability of the switch to absorb burst traffic, avoids cache blockage and traffic bottlenecks, and improves the stability and efficiency of network data transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120111017A_ABST
    Figure CN120111017A_ABST
Patent Text Reader

Abstract

The invention discloses a preemptive cache management system oriented to an on-chip shared cache exchange chip. The preemptive cache management system comprises five modules, namely an admission control module, a Head-drop selector module, an arbiter module, a Head Drop Executor module and a demultiplexer module. The preemptive cache management system comprises five modules, namely an admission control module, a Head-drop Selector module, an arbiter module, a Head Drop Executor module and a demultiplexer module. A cache management (Buffer Management) strategy dynamically allocates shared caches to queues in order to achieve high efficiency in the aspects of burst absorption and high throughput and achieve relatively fair performance isolation. Through testing, the method can achieve better effects in the aspects of burst absorption, performance isolation and cache blocking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of shared cache switching chips, and in order to solve the problem of extremely high on-chip cache cost and limited utilization of shared cache switching chips in data centers, specifically relates to a preemptive cache management strategy for shared cache of on-chip cache switching chips. Background Art

[0002] In recent years, the speed of data center switches has been doubling every two years. International mainstream switch chip manufacturers including Broadcom and Cisco have increased the speed of switch chips to 51.2Tbps. In order to achieve high-speed data packet access, modern high-speed data center switches often integrate cache inside the switch chip, that is, on-chip cache structure. However, since the cache size is constrained by factors such as chip size and power consumption, it cannot grow with the increase in switch chip speed, and the cache capacity is increasingly insufficient.

[0003] To improve cache utilization efficiency, modern switching chips usually share the cache among all queues to improve cache utilization. The buffer management strategy dynamically allocates shared cache to queues, aiming to achieve high efficiency in burst absorption and high throughput, and to achieve fairer performance isolation. Currently, most switching chips use non-preemptive cache management strategies (such as dynamic threshold strategy DT), that is, packets will not be discarded after entering the cache.

[0004] A typical shared cache switching chip includes three parts: ingress packet processing, traffic manager and egress packet processing. Among them, the traffic manager is responsible for accommodating data packets and dynamically allocating cache for data packets, which is the core part of the cache management strategy. Therefore, the present invention mainly focuses on the design of the traffic manager in the existing solution.

[0005] The traffic manager accommodates packets and dynamically allocates buffers between queues. For each arriving packet, the traffic manager first decides whether to allow the packet into the buffer. Once approved, the packet is written to the packet buffer, which is a centralized, globally shared on-chip SRAM. On the egress side, the scheduler selects a queue to read the packet based on the scheduling algorithm.

[0006] In the existing technical solution, the traffic manager includes the following five parts: switch shared buffer structure (PacketBuffer Structure), queuing module (Packet Admission), statistics module (Statistics), output scheduler (Output Scheduler) and dequeue module (Packet Dequeue). The output scheduler is the scheduling module in the traditional traffic controller.

[0007] 1. Switch shared cache structure

[0008] The switch shared cache structure is responsible for accommodating data packets and maintaining a multi-queue structure that shares a cache, including the following three modules:

[0009] (1) Cell Data Memory: Data packets are divided into cells of equal size and stored in the Cell Data Memory.

[0010] (2) Cell Pointer Memory: The Cell Pointer Memory is composed of Cell pointers. Each Cell pointer stores the address of a Cell data in the Cell data memory. For a data packet divided into multiple Cells, its Cell pointers are linked together in the form of a linked list. The Cell Pointer Memory also maintains free Cell pointer units through a linked list, namely the Free Cell Pointer List. The Free Cell Pointer List contains all free Cell pointers (corresponding to the addresses of free Cells in the Cell data memory);

[0011] (3) Packet Descriptor Memory: Each packet has a packet descriptor (PD), which contains the packet metadata and the head of the cell pointer list. The queue in the switch chip is maintained in the form of a packet descriptor list.

[0012] The shared cache structure can receive the signal of the enqueue module to implement the enqueue operation of the data packet; receive the signal of the dequeue module to implement the dequeue operation of the data packet; and generate a signal to the statistics module to update the queue length information of each queue.

[0013] 2. Enqueue module

[0014] The enqueue module is responsible for sending signals to the shared cache structure to implement the enqueue operation of the data packet. When the data packet arrives, the working process is as follows:

[0015] When a data packet arrives, the switch stores the data packet descriptor in the data packet descriptor memory and links the data packet descriptor to the tail of the linked list corresponding to the queue; then the switch takes out a certain number of Cell pointers from the free Cell pointer linked list to allocate storage space for the data packet, the specific number cell_num = data packet length / Cell length; then, the switching chip links the cell_num Cell pointers obtained in the previous step in the form of a linked list, and stores the head of the linked list in the data packet descriptor; finally, the switching chip uses the values ​​stored in these Cell pointers as addresses to write the data packet into the Cell data memory.

[0016] At the same time, the queue module can also implement some cache management strategies. Taking DT as an example, the queue module maintains a threshold T through the queue length information of the statistical module. If the queue length is greater than the threshold, the queued data packet is discarded. The calculation method of the threshold T is as follows:

[0017]

[0018] Where α is a control parameter, B is the shared cache size, and q i (t) is the length of queue i at time t. In practical configurations, α is usually a power of 2, so the threshold can be simply calculated by shifting the free buffer size.

[0019] 3. Statistics module

[0020] The statistics module is mainly responsible for receiving signals from the shared cache structure, counting the lengths of each queue, and outputting the current lengths of each queue to the enqueue module.

[0021] 4. Output Scheduler

[0022] The output scheduler schedules each queue using a certain algorithm based on whether there are data packets in each queue. Its output is the next queue index for dequeuing. The index number is used by the output scheduler to retrieve data from the corresponding queue linked list in the shared cache structure.

[0023] 5. Dequeue module

[0024] The dequeue module dequeues the data packet in the queue in the shared cache structure according to the queue index output by the output scheduler, and reads the data packet from the shared cache structure. During the dequeue process, the interaction between the dequeue module and the shared cache structure is as follows:

[0025] (1) Find the queue: find the header of the data packet descriptor list corresponding to the index number;

[0026] (2) Data packet dequeue: According to the linked list header, take out the linked list header data packet descriptor from the data packet descriptor linked list, and update the linked list header of the data packet descriptor;

[0027] (3) Get Cell pointer: According to the Cell pointer list stored in the data packet descriptor, get the Cell pointer corresponding to the data packet;

[0028] (4) Release the Cell storage space: that is, move the Cell pointer to the free Cell pointer list;

[0029] (5) Reading Cell data: According to the obtained Cell pointer, read the corresponding Cell data in the Cell data memory to form a data packet out queue.

[0030] Furthermore, since a data packet may contain multiple cells, (3), (4) and (5) need to be executed multiple times. At the same time, the data packet descriptor memory, the cell pointer memory and the cell data memory may be three different memories, so the access to them can be parallelized. Summary of the invention

[0031] In order to overcome the above-mentioned shortcomings of the prior art, the present invention provides a preemptive cache management system for an on-chip shared cache switching chip. The present invention is based on the improvement of the cache management system in the prior art solution, and the present invention includes five parts: admission control, head-drop selector, arbiter, head-drop executor and demultiplexer.

[0032] The system works in coordination through multiple modules. The admission decision serves as the entrance of the system, receives the output signal from the entrance queue, and receives the queue length signal provided by the statistics module, makes admission decisions according to the queue length signal, and outputs the data packet descriptor signal, the Cell pointer signal and the Cell data signal to the data packet descriptor memory, the Cell pointer memory and the Cell data memory respectively; the head loss selector receives the queue length signal of the statistics module and generates signals with the output scheduler respectively, and outputs them to the arbitrator for decision making; the arbitrator integrates the signals of the head loss selector and the scheduler, and sends the arbitration results to the data packet descriptor memory and the demultiplexer respectively; the data packet descriptor memory and the Cell pointer memory output the stored data to the demultiplexer, and the demultiplexer demultiplexes the data packet descriptor signal and the Cell pointer signal respectively according to the output signal of the arbitrator, and sends the results to the head loss executor and the dequeue module respectively; the dequeue module serves as the exit of the system, receives the output signal of the demultiplexer and the Cell data signal of the Cell data memory, and finally outputs the processed Cell data to the exit queue.

[0033] Admission Decision

[0034] The admission decision module is located on the input side. When a data packet arrives at the switch, it needs to pass the admission decision before it can enter the cache. The admission decision module uses a threshold T to determine whether to allow the data packet to enter the cache. When the length of the target queue to which the data packet is to be queued exceeds the threshold T, the data packet is discarded; otherwise, the data packet is written to the cache. The threshold T is maintained by the statistical module of the switch and can change dynamically. If a dynamic threshold algorithm is used, the threshold can be obtained by the following formula:

[0035] T=α·F (2)

[0036] Wherein, F is the currently unused cache size, α is a control parameter, and the present invention recommends setting the parameter to 8.

[0037] Head loss selector

[0038] The head drop selector determines the queue that needs to be dropped based on the threshold T used by the admission decision module and the queue length information maintained by the statistics module. Its overall structure is as follows: Figure 2 As shown in the figure, the head loss selector consists of a bitmap, multiple comparators, and a round-robin arbiter.

[0039] The bitmap is a register composed of multiple bits, and its length is the same as the number of queues. The value of each bit in the bitmap represents whether the queue needs to be head-dropped: if the value is 1, it means that too much buffer is allocated to the queue, and the queue needs to be head-dropped; otherwise, the queue does not need to be head-dropped, and the value of the corresponding bit in the bitmap is 0.

[0040] The value of the bitmap is obtained by the output of multiple comparators. Each comparator compares the length of a queue with the queue leader threshold (i.e., T in the admission control module). If the queue leader exceeds the threshold, it is considered that the cache size allocated to the queue is too large and outputs 1; otherwise, it outputs 0.

[0041] The round-robin arbiter polls all bits in the bitmap that have a value of 1. Each poll outputs the index of the bit corresponding to the value 1 in the bitmap, indicating the queue that needs to be head dropped. When the head of the queue drops a packet, it continues to search for the next bit corresponding to the value 1 in the bitmap.

[0042] Arbitrator

[0043] The arbiter needs to select a signal output between two input signals, the two input signals are:

[0044] (3) Dequeue signal from the output scheduler: If the output scheduler needs to schedule a queue to be dequeued, it outputs a dequeue signal to the arbitrator, which is the queue number that needs to be dequeued; if the output scheduler does not currently need to schedule a queue to be dequeued, it outputs a high-impedance signal to the arbitrator.

[0045] (4) Head-drop signal from the head-drop selector: If the head-drop selector needs to schedule a queue head-drop, it outputs a head-drop signal to the arbitrator; if the head-drop selector does not need to schedule a queue head-drop, it outputs a high-impedance signal to the arbitrator.

[0046] The arbitrator uses a fixed priority: the priority of the output scheduler's dequeue signal is always higher than the head packet selector's head loss signal, ensuring that when the output scheduler needs to dequeue, it is given priority.

[0047] In addition, the arbitrator needs to resolve the conflict between the output scheduler and the head loss selector. The head loss operation of the present invention consumes the bandwidth of the data packet descriptor memory and the cell pointer memory. Therefore, when the output scheduler (OutputScheduler) and the head loss selector obtain data packets at the same time, a read conflict will occur. When encountering a conflict, the output scheduler can preempt the read operation of the head loss selector, that is, when the head loss selector is performing head loss, the current read operation should be abandoned to allow the data packet of the output scheduler to be dequeued normally. Otherwise, the present invention may not be able to guarantee line-rate forwarding, which is unacceptable.

[0048] Whenever the output scheduler needs to get a packet, the read request from the head loss selector will be blocked. This is achieved through a fixed priority arbiter. If the output scheduler schedules to a certain index in a certain cycle, any content of the head selector in this cycle will be blocked, because the priority of the head loss selector in the fixed priority arbiter will always be lower than that of the output scheduler.

[0049] Head loss actuator

[0050] The operation process of the head loss executor is basically the same as the dequeue process of the dequeue module. The difference is that the packet loss operation performed by the head loss executor does not need to read the Cell data. According to the result of the output scheduler, the data packet of this queue in the shared cache structure is dequeued and released from the shared cache structure. During the discarding process, the interaction between the head loss executor and the shared cache structure is as follows:

[0051] (5) Find the queue: find the header of the data packet descriptor list corresponding to the index number;

[0052] (6) Data packet dequeue: According to the linked list header, take out the linked list header data packet descriptor from the data packet descriptor linked list, and update the linked list header of the data packet descriptor;

[0053] (7) Get Cell pointer: According to the Cell pointer list stored in the data packet descriptor, get the Cell pointer corresponding to the data packet;

[0054] (8) Release the Cell storage space: move the Cell pointer to the free Cell pointer list;

[0055] Furthermore, consistent with the dequeue process, since a data packet may contain multiple cells, (3) and (4) need to be executed multiple times. At the same time, the data packet descriptor memory, cell pointer memory, and cell data memory may be three different memories, so access to them can be parallelized.

[0056] Demultiplexer

[0057] Receive the signal of the arbiter and the output signal of the shared cache structure, and output the signal of the shared cache structure to the head loss executor and the output module at the same time.

[0058] The present invention adds five key modules, namely, admission decision, head loss selector, arbitrator, head loss executor and demultiplexer, on the original basis, and innovatively realizes the active packet loss function of the switch cache. The system can intelligently monitor and actively clear the queue data packets that over-occupy the cache, thereby effectively avoiding the cache overload problem. Through this innovation, the switch can dynamically adjust the allocation of cache resources, which not only improves the fairness between queues, but also significantly optimizes the utilization efficiency of the cache. The present invention enhances the switch's ability to absorb burst traffic, enabling it to respond to instantaneous high-traffic requests more flexibly, while improving the performance isolation of the cache and ensuring fair resource allocation between different traffic. In addition, the system can effectively avoid cache blocking and reduce the occurrence of traffic bottlenecks, thereby significantly improving the stability and efficiency of network data transmission. In general, the present invention greatly improves the performance of the data center network, enhances the processing capability under high concurrency and complex load conditions, and effectively guarantees the efficiency and reliability of network communications. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 It is a structural diagram of the flow controller of the present invention;

[0060] Figure 2 This is the structural diagram of the head loss selector;

[0061] Figure 3 It is a shared cache structure;

[0062] Figure 4 This is a schematic diagram of the data packet dequeue pipeline;

[0063] Figure 5 (a) is a comparison chart of the average QCT of burst query flow of the present invention and the existing cache management strategy in absorbing bursts;

[0064] Figure 5 (b) is a comparison chart of the QCT99 tail quantiles of burst query flows in the present invention and the existing cache management strategy in absorbing bursts;

[0065] Figure 5 (c) is a comparison diagram of the average FCT of background flows of the present invention and the existing cache management strategy in absorbing bursts;

[0066] Figure 5 (d) is a comparison chart of the background flow FCT99 tail quantiles of the present invention and the existing cache management strategy in absorbing bursts;

[0067] Figure 6(a) is a comparison chart of the average QCT of the present invention and the existing cache management strategy in performance isolation;

[0068] Figure 6 (b) is a QCT99 tail quantile comparison chart of the present invention and the existing cache management strategy in performance isolation;

[0069] Figure 7 (a) is a comparison chart of the average QCT of the present invention and the existing cache management strategy on cache blocking.

[0070] Figure 7 (b) is a QCT99 tail quantile comparison chart of the present invention and the existing cache management strategy on cache blocking. DETAILED DESCRIPTION

[0071] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and are not intended to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms fall within the scope limited by the appended claims of the application equally.

[0072] The present invention provides a preemptive cache management system for an on-chip shared cache switching chip. The present invention is based on the improvement of the cache management system in the prior art solution. The overall architecture is as follows: Figure 1 As shown in the figure, it includes five parts: Admission Control, Head-drop Selector, Arbiter, Head Drop Executor and Demultiplexer.

[0073] The system consists of multiple modules that work together, and the modules coordinate their work through signal transmission. First, the admission decision module, as the entrance of the system, is responsible for receiving signals from the entrance queue and obtaining queue length information from the statistics module. The admission decision module determines whether to allow the data packet to enter the system based on the queue length signal. If it decides to allow admission, it will output three signals: packet descriptor signal (containing basic information of the data packet), cell pointer signal (pointing to the location of the cell data) and cell data signal (actual data content). These three signals are sent to different memories: packet descriptor memory, cell pointer memory and cell data memory for subsequent use. Next, the head loss selector module receives the queue length signal from the statistics module and passes it to the arbitrator together with the signal generated by the output scheduler. The role of the arbitrator is to comprehensively analyze the signals of the head loss selector and the output scheduler, and after making a decision, send the arbitration result to two key modules: packet descriptor memory and demultiplexer. Based on the arbitration result, the arbitrator decides which signals can be passed to subsequent modules. Then, the data in the packet descriptor memory and the cell pointer memory are sent to the demultiplexer. The task of the demultiplexer is to demultiplex the incoming packet descriptor signal and the cell pointer signal according to the signal of the arbitrator. The demultiplexer sends the processed signals to the header loss executor and the dequeue module respectively. The header loss executor is responsible for performing the header loss operation (i.e. discarding the unimportant part of the header information) according to the demultiplexing result, while the dequeue module is responsible for preparing the final data flow to the exit. Finally, the dequeue module, as the exit of the system, receives the signal from the demultiplexer and the cell data signal from the cell data memory. After the final processing, these data will be output to the exit queue to complete the processing of the data packet. The detailed workflow of each module is introduced below.

[0074] The detailed workflow of each module is introduced below.

[0075] 1. Admission decision

[0076] The admission decision module is located on the input side. When a data packet arrives at the switch, it needs to pass the admission decision before it can enter the cache. The admission decision module uses a threshold T to determine whether to allow the data packet to enter the cache. When the length of the target queue to which the data packet is to be queued exceeds the threshold T, the data packet is discarded; otherwise, the data packet is written to the cache. The threshold T is maintained by the statistical module of the switch and can change dynamically. If a dynamic threshold algorithm is used, the threshold can be obtained by the following formula:

[0077] T=α·F (3)

[0078] Wherein, F is the currently unused cache size, α is a control parameter, and the present invention recommends setting the parameter to 8.

[0079] Assuming there are N queues, the calculation method of F can be obtained as follows:

[0080]

[0081] Where B is the cache size, q i (t) is the queue length of queue i, and N is the number of queues.

[0082] In formula (3), the control parameter α affects the threshold T of the cache management strategy, and determines the cache utilization and fairness of the cache management of the present invention. In an embodiment of the present invention, the setting of the control parameter α is different from the traditional dynamic threshold strategy. The setting of α in the present invention is discussed in detail below.

[0083] According to research, the amount of free cache F retained in a steady state is given by the following formula:

[0084]

[0085] Where B is the buffer size and N is the number of congestion queues with length equal to the threshold.

[0086] Formula (5) indicates that the larger α is, the less free cache is retained by the present invention, and thus the cache utilization rate is higher. However, as α becomes larger, the improvement in cache utilization rate becomes less obvious. For example, when N=1, the free cache is retained as B / 9 (α=8) and B / 17 (α=16). The cache utilization rate is only increased by 5.2% (α is usually taken as a power of 2 to facilitate threshold calculation). Therefore, it is not necessary to set the value of α too high.

[0087] On the other hand, according to current research, when burst traffic arrives at some queues that are not congested, the present invention can fairly allocate the cache only when the following relationship is satisfied:

[0088]

[0089] Where: R is the arrival rate of burst traffic, V is the packet loss rate, M is the number of queues for burst traffic, and N is the number of queues with over-allocated buffers. Inequality (6) can be rewritten as the following inequality:

[0090]

[0091] Inequality (7) means that when the packet loss rate is high, α can be set larger. For example, for an over-allocated queue (i.e., N = 1) and a bursty queue (i.e., M = 1), inequality (7) can be rewritten as When V ≥ R / 2, α can theoretically take any large value. In summary, the setting of α needs to balance cache utilization and fairness. The higher the packet loss rate, the larger α can be set. Our experiments show that α = 8 is a better choice.

[0092] 2. Head loss selector

[0093] The head drop selector determines the queue that needs to be dropped based on the threshold T used by the admission decision module and the queue length information maintained by the statistics module. Its overall structure is as follows: Figure 2 As shown in the figure, the head loss selector consists of a bitmap, multiple comparators, and a round-robin arbiter.

[0094] The bitmap is a register composed of multiple bits, and its length is the same as the number of queues. The value of each bit of the bitmap represents whether the queue needs to be head-dropped: if the value is 1, it means that too much buffer is allocated to the queue, and the queue needs to be head-dropped; otherwise, the queue does not need to be head-dropped, and the value of the corresponding bit in the bitmap is 0. In the embodiment of the present invention, in the scenario where the number of queues is 4, the number of bits in the bitmap is equal to the queue length, which is also 4. The value of each bit of the bitmap represents whether too much buffer is allocated to this queue. If the lengths of queues 1 and 3 are greater than the threshold, and the lengths of other queues are less than or equal to the threshold, the bitmap is 0101.

[0095] The value of the bitmap is obtained by the output of multiple comparators. Each comparator compares the length of a queue with the queue leader threshold (i.e., T in the admission control module). If the queue leader exceeds the threshold, it is considered that the cache size allocated to the queue is too large and outputs 1; otherwise, it outputs 0.

[0096] The round-robin arbiter polls all bits in the bitmap that have a value of 1. Each poll outputs the index of the bit corresponding to the value 1 in the bitmap, indicating the queue that needs to be head dropped. When the head of the queue drops a packet, it continues to search for the next bit corresponding to the value 1 in the bitmap.

[0097] In an embodiment of the present invention, the specific polling method is Round Robin, i.e., polling priority arbitration, which can maintain the fairness of each queue. The input length of the polling priority arbitration is the same as the queue length. It finds the first input of 1 from high to low according to the priority of each queue and outputs it. The next time it searches, the priority of this queue will become the lowest, so that the first input of 1 will be searched from after this queue.

[0098] Let's take an example to understand polling priority arbitration. There are 4 queues, that is, the input length is 4. Req[3:0] can be used to represent the input of the polling priority arbitration. Assume that the input is Req[3:0]=0101 and the input remains unchanged. When the power is turned on, the polling arbitrator will set a default priority. We assume that the default priority is 3210, where 0 represents the highest priority and 3 represents the lowest priority. In the first cycle, the polling arbitrator will search for the position of the first 1 in the order of 0123. At this time, Req[0]=1 is found, and the output is 0001. In the second cycle, the polling arbitrator will automatically adjust the priority of each queue to 2103, so it will search for the position of the first 1 in the order of 1230, and Req[2]=1 is found, and the output is 0100.

[0099] 3. Arbitrator

[0100] The arbiter needs to select a signal output between two input signals, the two input signals are:

[0101] (1) Dequeue signal from the output scheduler: If the output scheduler needs to schedule a queue to be dequeued, it outputs a dequeue signal to the arbitrator, which is the queue number that needs to be dequeued; if the output scheduler does not currently need to schedule a queue to be dequeued, it outputs a high-impedance signal to the arbitrator.

[0102] (2) Head-drop signal from the head-drop selector: If the head-drop selector needs to schedule a queue head-drop, it outputs a head-drop signal to the arbitrator; if the head-drop selector does not need to schedule a queue head-drop, it outputs a high-impedance signal to the arbitrator.

[0103] The arbitrator uses a fixed priority: the priority of the output scheduler's dequeue signal is always higher than the head packet selector's head loss signal, ensuring that when the output scheduler needs to dequeue, it is given priority.

[0104] In addition, the arbitrator needs to resolve the conflict between the output scheduler and the head loss selector. The packet loss operation of the present invention consumes the bandwidth of the data packet descriptor memory and the cell pointer memory. Therefore, when the output scheduler (OutputScheduler) and the head loss selector obtain data packets at the same time, a read conflict will occur. When encountering a conflict, the output scheduler can preempt the read operation of the head loss selector, that is, when the head loss selector is performing head loss, it should give up discarding this data packet and let the data packet of the output scheduler be dequeued normally. Otherwise, the present invention may not be able to guarantee line-rate forwarding, which is unacceptable.

[0105] Whenever the output scheduler needs to obtain a data packet, the read request from the head loss selector will be blocked. This is achieved through a fixed priority arbiter. If the output scheduler schedules to a certain index in a certain cycle, any signal of the head loss selector in this cycle will be blocked, because the priority of the head loss selector in the fixed priority arbiter will always be lower than that of the output scheduling module.

[0106] 4. Head loss actuator

[0107] The operation process of the head loss executor is basically the same as the dequeue process of the dequeue module. The difference is that the packet loss operation performed by the head loss executor does not need to read the Cell data. According to the result of the output scheduler, the data packet in this queue in the shared cache structure is dequeued and released from the shared cache structure. During the discarding process, the interaction between the head loss executor and the shared cache structure is as follows (packet loss process):

[0108] (1) Find the queue: find the header of the data packet descriptor list corresponding to the index number;

[0109] (2) Data packet dequeue: According to the linked list header, take out the linked list header data packet descriptor from the data packet descriptor linked list, and update the linked list header of the data packet descriptor;

[0110] (3) Get Cell pointer: According to the Cell pointer list stored in the data packet descriptor, get the Cell pointer corresponding to the data packet;

[0111] (4) Release the Cell storage space: that is, move the Cell pointer to the free Cell pointer list;

[0112] Furthermore, consistent with the dequeue process, since a data packet may contain multiple cells, (3) and (4) need to be executed multiple times. Figure 3 As shown in FIG. 1 , the packet descriptor memory, the cell pointer memory and the cell data memory may be three different memories, so the access to them can be parallelized, that is, pipeline operation can be realized.

[0113] The specific implementation of pipeline operation is as follows

[0114] Figure 4 As shown (the figure describes the pipeline design when the data packet is dequeued, and the discard process here only lacks the process of reading Cell data compared to the data packet dequeuing), the following is based on

[0115] Figure 4 Introduce the operation of the pipeline in detail for each cycle:

[0116] The first cycle performs ① data packet descriptor reading operation;

[0117] In the second cycle, ② is executed to dequeue the read data packet descriptor, and ③ is executed to read the Cell pointer from the Cell pointer memory;

[0118] The third cycle continues to execute ③ to read the next Cell pointer of the data packet from the Cell pointer memory, and at the same time, ④ can be executed to release the Cell pointer read in the previous cycle, and ⑤ can be executed to read the Cell data of the Cell pointer obtained in the previous cycle, and at the same time, ① can be executed to read the next data packet descriptor.

[0119] In the fourth cycle, execution ④ releases the Cell pointer read in the previous cycle, executes ⑤, reads the Cell data of the Cell pointer obtained in the previous cycle, executes ② to dequeue the read data packet descriptor, and executes ③ to read the next Cell pointer from the Cell pointer memory.

[0120] By analogy, each subsequent cycle can complete ①, ③, and ④ or ②, ③, and ④. Therefore, a Cell can be discarded in each cycle, thus realizing the pipeline.

[0121] 5. Demultiplexer

[0122] Receive the signal of the arbiter and the output signal of the shared cache structure, and output the signal of the shared cache structure to the head loss executor and the output module at the same time.

[0123] Test Results

[0124] In order to test the effect of the present invention, we completed experiments in a real environment, tested the performance of the present invention in terms of burst absorption, performance isolation and cache blocking, and compared it with the traditional cache management strategy DT, the recently proposed non-preemptive cache management strategy ABM, and the theoretically optimal cache management strategy Pushout. The experimental results are summarized as follows:

[0125] 1) In terms of burst absorption, the present invention can reduce the average query completion time (QueryCompletion Time) of burst query flow and background by up to 55% and 42% respectively;

[0126] 2) In terms of performance isolation, the present invention can significantly reduce the average query completion time, and the present invention can significantly reduce the 99th quantile of the query completion time;

[0127] 3) In terms of cache blocking, the present invention can avoid cache blocking to the greatest extent and achieve performance similar to that of the theoretical optimal cache management strategy Pushout.

[0128] In the experiments, we use an open source traffic generator to generate two types of traffic.

[0129] 1) Burst Query Traffic: The client on each host periodically sends queries to 16 servers on other hosts (each host runs 2 servers). After receiving the query, each server generates a response to the client. The total amount of response data (called query size) varies, thus generating different burst sizes.

[0130] 2) Background Traffic: We generate background traffic based on a Poisson process. The sender and receiver are randomly selected and the traffic size follows the web-search distribution.

[0131] Figure 5 This shows the ability of the present invention to absorb bursts. In this experiment, the query traffic load is 1% and the background traffic load is 50%. The DT in the figure is a dynamic threshold strategy, and ABM is a non-preemptive cache management strategy recently proposed.

[0132] Figure 5 (a) shows that, compared with DT and ABM, the present invention can reduce the average query completion time (QueryCompletion Time) by up to 55% and 42%, respectively.

[0133] Figure 5 (b) shows that when the burst size is less than 80% of the cache size, the present invention can avoid retransmission timeout, which is 33% and 60% higher than DT and ABM respectively. Meanwhile, while absorbing burst performance better, the background flow performance under the present invention is not damaged.

[0134] Figure 5 (c) shows that the average flow completion time (FCT) of the present invention is comparable to DT and ABM.

[0135] Figure 5 (d) shows that the 99th tail quantile of flow completion time for small flows (i.e., flow size < 100kB) is comparable to ABM and up to 57% shorter than DT. This is because the present invention only discards the over-allocated buffers of background flows instead of grabbing the buffers they deserve.

[0136] Figure 6 The capability of performance isolation of the present invention is demonstrated. Figure 6 (a) and Figure 6 (b) shows the average query completion time (QCT) of the query traffic and the 99th quantile of the query completion time, respectively. Figure 6(b) shows that as the background traffic load increases, DT and ABM will cause retransmission timeouts for query traffic, thereby significantly increasing query completion time. Since background traffic and query traffic are placed in different queues, retransmission timeouts are mainly due to the inability to quickly adjust cache allocation. In contrast, the present invention can quickly adjust cache allocation by actively dropping packets that over-allocate cache queues, thereby achieving a significant reduction in the 99th quantile of query completion time.

[0137] Figure 7 It shows that the present invention can effectively alleviate the cache congestion problem. In this experiment, we set two priority queues for the host and the switch. The query flow is assigned to the high priority queue, while the background flow is assigned to the low priority queue. For the high priority queue, we set DT, ABM, and α of the present invention to 8 to allocate more cache for the high priority flow. For the low priority queue, we set α to 1. The other settings remain unchanged. We let the host receive query flows and background flows from other hosts so that the two priority queues in the same port are congested at the same time. It is hoped that the low priority background flow will not affect the high priority query flow. Figure 7 (a) and Figure 7 (b) shows the average query completion time and the 99th quantile of the query completion time, respectively. The solid line represents the query completion time without background flow, while the dashed line represents the query completion time with background flow. It can be observed that the present invention achieves similar performance to Pushout, and even if they share the same cache, the background flow does not have much impact on the performance of the query flow. However, the background flow can significantly prolong the query completion time of DT and ABM.

Claims

1. A preemptive cache management system for an on-chip shared cache switch chip, characterized in that: It consists of five parts: Admission Control, Head-drop Selector, Arbiter, Head-drop Executor and Demultiplexer. The system works in coordination through multiple modules. The admission decision serves as the entrance of the system, receives the output signal from the entrance queue, and receives the queue length signal provided by the statistics module. The admission decision is made according to the queue length signal, and the data packet descriptor signal, the cell pointer signal and the cell data signal are output to the data packet descriptor memory, the cell pointer memory and the cell data memory respectively; the head loss selector receives the queue length signal of the statistics module and generates signals with the output scheduler, and both are output to the arbitrator for decision making; The arbiter integrates the signals of the head loss selector and the scheduler, and sends the arbitration results to the data packet descriptor memory and the demultiplexer respectively; the data packet descriptor memory and the Cell pointer memory output the stored data to the demultiplexer, and the demultiplexer demultiplexes the data packet descriptor signal and the Cell pointer signal according to the output signal of the arbitrator, and sends the results to the head loss executor and the dequeue module respectively; the dequeue module serves as the exit of the system, receives the output signal of the demultiplexer and the Cell data signal of the Cell data memory, and finally outputs the processed Cell data to the exit queue.

2. The preemptive cache management system for an on-chip shared cache switching chip according to claim 1, characterized in that: The admission decision module is located on the input side. When a data packet arrives at the switch, it needs to pass the admission decision before entering the cache. The admission decision module uses a threshold T to determine whether the data packet is allowed to enter the cache. When the length of the target queue where the data packet is to be entered exceeds the threshold T, the data packet is discarded; otherwise, the data packet is written into the cache. The threshold T is maintained by the statistical module of the switch and can change dynamically, using a dynamic threshold algorithm.

3. The preemptive cache management system for an on-chip shared cache switching chip according to claim 2, characterized in that: The threshold is obtained by the following formula: T=α·F (2) Among them, F is the currently unused cache size, α is a control parameter, and the parameter is set to 8.

4. The preemptive cache management system for an on-chip shared cache switching chip according to claim 1, characterized in that: The head drop selector determines the queue that needs to be head dropped based on the threshold T used by the admission decision module and the queue length information maintained by the statistics module. The head drop selector consists of a bitmap, multiple comparators, and a round-robin arbiter.

5. The preemptive cache management system for an on-chip shared cache switching chip according to claim 4, characterized in that: The bitmap is a register composed of multiple bits, and its length is the same as the number of queues. The value of each bit in the bitmap represents whether the queue needs to be head-dropped: if the value is 1, it means that too much buffer is allocated to the queue, and the queue needs to be head-dropped; otherwise, the queue does not need to be head-dropped, and the value of the corresponding bit in the bitmap is 0; The value of the bitmap is obtained by the output of multiple comparators. Each comparator compares the length of a queue with the queue leader threshold (i.e., T in the admission control module). If the queue leader exceeds the threshold, it is considered that the cache size allocated to the queue is too large and outputs 1; otherwise, it outputs 0. The round-robin arbiter polls all bits in the bitmap whose value is 1. Each time it polls, it outputs the index of the bit corresponding to the value 1 in the bitmap, indicating the queue that needs to be head dropped. When the head of the queue drops a packet, it continues to look for the next bit in the bitmap corresponding to the value 1.

6. The preemptive cache management system for an on-chip shared cache switching chip according to claim 1, characterized in that: The arbiter needs to select a signal output between two input signals, the two input signals are: (1) Dequeue signal from the output scheduler: If the output scheduler needs to schedule a queue to be dequeued, it outputs a dequeue signal to the arbitrator, which is the queue number that needs to be dequeued; if the output scheduler does not currently need to schedule a queue to be dequeued, it outputs a high-impedance signal to the arbitrator; (2) Head-drop signal from the head-drop selector: If the head-drop selector needs to schedule a queue head-drop, it outputs a head-drop signal to the arbitrator; if the head-drop selector does not need to schedule a queue head-drop, Then output a high impedance signal to the arbitrator.

7. The preemptive cache management system for an on-chip shared cache switching chip according to claim 6, characterized in that: The arbitrator uses a fixed priority: the priority of the output scheduler's dequeue signal is always higher than the head packet selector's head loss signal, ensuring that when the output scheduler needs to dequeue, it is given priority to dequeue; The head drop operation consumes the bandwidth of the packet descriptor memory and the cell pointer memory. When the output scheduler and the head drop selector obtain data packets at the same time, a read conflict occurs. When a conflict occurs, the output scheduler can preempt the read operation of the head drop selector. That is, when the head drop selector is performing head drop, it should abandon the current read operation and let the data packet of the output scheduler be dequeued normally. Otherwise, line-rate forwarding may not be guaranteed, which is unacceptable. Whenever the output scheduler needs to obtain a data packet, the read request from the head loss selector will be blocked. This is achieved by the fixed priority arbiter. If the output scheduler schedules to a certain index in a certain cycle, any content of the head selector in this cycle will be blocked, because the priority of the head loss selector in the fixed priority arbiter will always be lower than that of the output scheduler.

8. The preemptive cache management system for an on-chip shared cache switching chip according to claim 1, characterized in that: The operation process of the head loss executor is basically the same as the dequeue process of the dequeue module. The difference is that the packet loss operation performed by the head loss executor does not need to read the Cell data. According to the result of the output scheduler, the data packet in this queue in the shared cache structure is dequeued and released from the shared cache structure.

9. The preemptive cache management system for an on-chip shared cache switching chip according to claim 8, characterized in that: During the discard process, the interaction between the head drop executor and the shared cache structure is as follows: (1) Find the queue: find the header of the data packet descriptor list corresponding to the index number; (2) Packet dequeue: According to the linked list header, take out the linked list header packet descriptor from the linked list of packet descriptors. And update the linked list header of the data packet descriptor; (3) Get Cell pointer: According to the Cell pointer list stored in the data packet descriptor, get the Cell pointer corresponding to the data packet; (4) Release the Cell storage space: that is, move the Cell pointer to the free Cell pointer list; Furthermore, consistent with the dequeue process, since a data packet may contain multiple cells, (3) and (4) need to be executed multiple times. At the same time, the data packet descriptor memory, cell pointer memory, and cell data memory may be three different memories, so access to them can be parallelized.

10. The preemptive cache management system for an on-chip shared cache switching chip according to claim 1, characterized in that: The demultiplexer receives the signal of the arbiter and the output signal of the shared cache structure, and outputs the signal of the shared cache structure to the head loss executor and the output module at the same time.