An in-chip cache-based exact matching flow table lookup method, chip and network card
Patent Information
- Application Number
- CN202611039387.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-14
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2046-07-14
AI Technical Summary
然而,现有Cache方案多采用基于地址的简单映射方式,当多个不同流表条目经哈希运算映射至同一Cache地址时,后写入条目会覆盖先写入条目,导致哈希冲突严重的场景下Cache命中率急剧下降
[0012] The present invention provides a precise matching flow table lookup method based on on-chip cache. First, it generates a group index address and a feature value through double hashing. Multiple slots are set in each storage row of the cache hash table to store different entry feature values. This ensures that even if multiple flow table entries map to the same group index address, they can coexist in different slots of the same storage row as long as their feature values are different. This significantly improves the cache's tolerance for hash collisions and the storage flexibility of hot entries, effectively increasing the cache hit rate. Second, by independently maintaining an aging count and an operation count for each slot, the aging count dynamically tracks the slot's hotness/coldness, and the operation count indicates whether the slot is in an operational state. During replacement and backfilling, only slots with zero aging count and zero operation count are selected as the replacement slots. This achieves fine-grained slot-level replacement management and provides a lightweight slot operation lock through the operation count, avoiding pipeline blockage while ensuring data consistency and guaranteeing lookup throughput.
Smart Images

Figure CN122570374B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of precise matching flow table lookup technology, and in particular to a precise matching flow table lookup method, chip, and network interface card based on on-chip cache. Background Technology
[0002] Precise flow table matching is a core component in network chips for achieving fast packet forwarding. It matches the corresponding forwarding strategy in the flow table based on the lookup key value (such as a 5-tuple or virtual network identifier) extracted from the packet header. Since flow tables typically reach millions to tens of millions of entries, the majority of flow table entries need to be stored in off-chip dynamic random access memory (DRAM) due to capacity limitations of on-chip static random access memory (SRAM). To reduce the long latency and high bandwidth pressure of off-chip access, existing solutions typically use on-chip caches to cache frequently accessed data for fast hits. However, existing cache solutions often employ simple address-based mapping. When multiple different flow table entries are mapped to the same cache address via hash operations, later entries will overwrite earlier ones, leading to a sharp drop in cache hit rate in scenarios with severe hash collisions. Furthermore, in pipelined hardware lookup architectures, if a subsequent request accesses the same cache address before off-chip backfilling is complete, there is a read-after-write risk, which can easily lead to reading outdated data. Some solutions address this problem by blocking the pipeline and waiting for backfilling to complete, but this severely sacrifices lookup throughput. In addition, existing cache replacement strategies are mostly based on simple least recently used algorithms or their approximate implementations, which have large hardware overhead and are difficult to efficiently couple with pipelined lookup timing. Summary of the Invention
[0003] To address the aforementioned technical problems, the technical solution adopted by this invention is as follows:
[0004] According to the present application, a method for precise matching flow table lookup based on on-chip cache includes the following steps:
[0005] Q100 receives search requests and retrieves search key values;
[0006] Q200, perform a first hash operation on the lookup key value to obtain the group index address, and perform a second hash operation on the lookup key value to obtain the feature value;
[0007] Q300, using the group index address as the read address, reads multiple entry feature values pre-stored in the corresponding storage row from the Cache hash table, with each entry feature value corresponding to a slot in the corresponding storage row;
[0008] Q400, compare the feature value of the search key value with the feature values of multiple read entries in parallel. If a matching entry feature value exists, read the Cache key value table entry of the slot corresponding to the matching entry feature value.
[0009] Q500 compares the lookup key value with the complete key value stored in the Cache key value table entry. If they match, it is determined that the Cache has hit. The Cache data table entry of the corresponding slot is read, the lookup result is output, and the aging count of the hit slot is updated.
[0010] Q600: If the feature value comparison shows no match, or the complete key value comparison shows no consistency, it is determined as a cache miss, generating an access request to the off-chip main storage area. The search result is obtained from the off-chip main storage area, and a replacement and backfill operation is performed based on the aging count and operation count of each slot in the storage row. The cache hash table, cache key-value table, and cache data table constitute a cache subset of the off-chip main storage area. The aging count is used to characterize the length of time the corresponding slot has been missed, and the operation count is used to characterize whether the corresponding slot is currently in an operational state. During the replacement and backfill operation, the slot with an aging count of zero and an operation count of zero is selected as the slot to be replaced.
[0011] The present invention has at least the following beneficial effects:
[0012] The present invention provides a precise matching flow table lookup method based on on-chip cache. First, it generates a group index address and a feature value through double hashing. Multiple slots are set in each storage row of the cache hash table to store different entry feature values. This ensures that even if multiple flow table entries map to the same group index address, they can coexist in different slots of the same storage row as long as their feature values are different. This significantly improves the cache's tolerance for hash collisions and the storage flexibility of hot entries, effectively increasing the cache hit rate. Second, by independently maintaining an aging count and an operation count for each slot, the aging count dynamically tracks the slot's hotness / coldness, and the operation count indicates whether the slot is in an operational state. During replacement and backfilling, only slots with zero aging count and zero operation count are selected as the replacement slots. This achieves fine-grained slot-level replacement management and provides a lightweight slot operation lock through the operation count, avoiding pipeline blockage while ensuring data consistency and guaranteeing lookup throughput. Attached Figure Description
[0013] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 A flowchart illustrating the precise matching flow table lookup method based on on-chip cache provided in this embodiment of the invention. Detailed Implementation
[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0016] It should be noted that, based on this disclosure, those skilled in the art will understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects set forth herein can be used to implement the device and / or practice the method. Furthermore, this device and / or practice the method can be implemented using other structures and / or functionalities besides one or more of the aspects set forth herein.
[0017] Example 1:
[0018] The following will refer to Figure 1 The flowchart shown illustrates a precise matching flow table lookup method based on on-chip cache, which introduces such a method.
[0019] The precise matching flow table lookup method based on on-chip cache may include the following steps:
[0020] Q100 receives a search request and retrieves the search key value.
[0021] The lookup request is generated by the packet parsing unit inside the network chip. When a packet enters the chip, the parsing unit extracts predetermined fields from the packet header and concatenates them to form a lookup key. The components and concatenation order of the lookup key are predefined by the flow table configuration file. Typical fields include source IP address, destination IP address, source port number, destination port number, protocol type, and virtual network identifier. The lookup request receiving unit latches the lookup key into the input register and simultaneously requests a free buffer pointer from the sequence control module to temporarily store the packet's context information during the lookup process.
[0022] This step utilizes a hardware-configurable multi-field concatenation mechanism, enabling the same cache lookup pipeline to adapt to the differentiated definition requirements of lookup keys in different network application scenarios. The key composition can be adjusted without modifying the hardware circuitry, thus improving the chip's flexibility and versatility.
[0023] Q200, perform a first hash operation on the lookup key value to obtain the group index address, and perform a second hash operation on the lookup key value to obtain the feature value.
[0024] Furthermore, in step Q200, a second hash operation is performed on the lookup key to obtain a feature value, including the following steps:
[0025] Q210, input the search key value into the second hash function to obtain a hash result value with a bit width of W.
[0026] In this embodiment, the second hash function is implemented in hardware using combinational logic circuits and operates in parallel with the first hash function without introducing additional clock cycles. The input to the second hash function is the lookup key value obtained in step S100, and its bit width is determined according to the network application scenario; typical values can be 128 bits, 256 bits, or 320 bits. The output bit width W of the second hash function is jointly determined by the number of slots n per row of the on-chip hash table and the bit width of each feature value segment, satisfying W = n × w, where w is the bit width of each feature value segment.
[0027] The second hash function can employ a randomized hash algorithm based on XOR-Shift. Its hardware implementation consists of a multi-stage XOR shift network: first, the input key value is split into multiple equal-width fields and XORed to produce intermediate results; then, the intermediate results are sequentially subjected to several stages of shift and XOR operations, with each shift amount being a coprime prime offset (e.g., shift by 6, 13, or 27 bits) to fully diffuse the influence of the input bits; finally, the lower W bits of the shift network are taken as the hash result value for output.
[0028] This step uses a second hash function to compress lookup keys of arbitrary width into hash results of fixed width W. W is determined by the number of slots in each row and the feature value width of each slot, allowing a single hash calculation to provide a complete data source for parallel feature value matching of all subsequent slots. The combinational logic of the XOR-Shift algorithm completes the hash calculation within a single cycle without increasing lookup latency. Simultaneously, its multi-level shift-XOR structure effectively reduces the probability of different keys producing the same hash result, providing a high-resolution basis for subsequent slot-level matching.
[0029] Q220, the hash result value is divided into n feature value segments, each feature value segment corresponding to a slot in the same storage row of the Cache hash table; where n is the maximum number of entries that can be stored in a row of the Cache hash table.
[0030] In this embodiment, the partitioning operation requires no logic gates or registers in hardware and can be completed via hardwired bus. The physical connections of the W-bit hash result output in step S210 are divided into n groups according to the principle of equal width, with each group containing w bits. The 0th group (lowest w bits) corresponds to the feature value segment of slot 0, the 1st group corresponds to the feature value segment of slot 1, and so on, with the (n-1)th group (highest w bits) corresponding to the feature value segment of slot (n-1). Each group of w-bit signals is directly connected to the comparator input port of the corresponding slot in the subsequent step S400.
[0031] When W is not divisible by n, the following method is used: The bit width w of each segment is taken as floor(W / n), i.e., rounded down. For the first (W mod n) segments, an additional 1 bit is allocated to each, making the bit width of these segments w+1. At this point, the bit widths of each feature value segment may differ by 1 bit. Each segment is still compared with the feature value of the corresponding slot for equal bit width, and the comparator bit width is configured according to the actual bit width of each segment.
[0032] This step directly divides the result of a single hash calculation into n independent feature value segments via hardwired bus, providing matching criteria for each of the n slots without additional logic overhead. Compared to the traditional approach of calculating a hash independently for each slot, this avoids the latency accumulation and hardware area expansion caused by multiple hash calculations. Compared to the approach where all slots share the same feature value, each slot uses independent feature value segments, ensuring that the same lookup key value generates different feature values in different slots. When a feature value match fails in one slot, other slots may still match successfully, thus providing multiple candidate hit positions for the same key value within the same storage row, effectively reducing the probability of lookup failures due to single feature value conflicts. Furthermore, the zero-latency characteristic of hardwired partitioning ensures that the entire feature value generation and allocation process does not occupy additional pipeline truncation time, guaranteeing a consistently high throughput for the lookup pipeline.
[0033] Q300 uses the group index address as the read address to read multiple pre-stored entry feature values in the corresponding storage row from the Cache hash table. Each entry feature value corresponds to a slot in the corresponding storage row.
[0034] Furthermore, each storage row of the Cache hash table contains n slots, and each slot stores an entry feature value and a corresponding valid flag bit; when the group index address is used as the read address, the Cache hash table reads out the entry feature values and corresponding valid flag bits of all n slots in the storage row in parallel within one read cycle.
[0035] Using the group index address generated in step Q200 as the read address, a single-cycle read operation is initiated on the Cache hash table. The Cache hash table is implemented using static random access memory, with a depth of H rows. Each row contains n slots, and each slot stores an entry feature value and its corresponding validity flag. A validity flag of 1 indicates that the entry feature value stored in that slot is valid, while 0 indicates that the slot is idle. After the group index address is sent to the address input of the Cache hash table, after one read cycle, the entry feature values and validity flags of all n slots in that storage row are read in parallel and latched into n sets of output registers. Simultaneously, using the same group index address, the aging count values and operation count values corresponding to all n slots in that storage row are read from the aging count table and operation count table, respectively, for use in subsequent steps.
[0036] This step reads the entry feature value, validity flag, aging count value and operation count value of all slots in the same storage line in parallel in a single cycle, compressing the slot polling that requires multiple serial accesses in the traditional scheme into a single parallel access, which greatly reduces the cache lookup latency; the multi-slot group associative structure provides flexible conflict tolerance space under limited storage depth, balancing storage efficiency and lookup performance.
[0037] Q400: The feature value of the search key value is compared in parallel with the feature values of multiple read entries. If a matching entry feature value exists, the Cache key value table entry of the slot corresponding to the matching entry feature value is read.
[0038] This step is implemented in hardware using an array of n independent bit-by-bit comparators, which simultaneously compares the feature values of the n slot entries read from the Cache hash table in step Q300 with the corresponding feature value fragments of the lookup key generated in step Q200.
[0039] Furthermore, step Q400 includes the following steps:
[0040] Q410, compare the feature values of the n entries read in step Q300 with the feature values of the search key values obtained in step Q200 in parallel to generate n comparison result signals.
[0041] In this embodiment, the data read in step Q300 contains n sets of entry feature values. The i-th set of entry feature values corresponds to the entry feature value stored in the i-th slot, denoted as bin_tbl[i], with a bit width of w bits. The feature value of the lookup key generated in step Q200 has been divided into n feature value segments, denoted as bin_key[i], with a bit width of w bits.
[0042] The hardware uses an array of n w-bit XOR gates. For the i-th slot, bin_tbl[i] and bin_key[i] are XORed bit by bit, generating w XOR result lines. These w XOR result lines are then input into a w-input NOR gate, and the output of the NOR gate is the i-th comparison result signal cmp[i]. When every bit of bin_tbl[i] and bin_key[i] is equal, all w XOR result lines are low, and the NOR gate outputs a high level, indicating that the feature value comparison result is consistent. When any bit is unequal, the NOR gate outputs a low level, indicating that the feature value comparison result is inconsistent.
[0043] The n comparator arrays are fully parallel in hardware and have no data dependency on each other. The combinational logic delay of all comparators from input to output is only one XOR gate delay plus one NOR gate delay, which can be completed in one clock cycle.
[0044] This step achieves fully parallel comparison of all slot feature values through n independent hardware comparator arrays. The comparison latency is constant at two gate levels, independent of the number of slots (n), thus not increasing the lookup latency when the number of slots expands, ensuring the timing convergence of the cache lookup pipeline. Each slot uses an independent feature value fragment for matching. Compared to a scheme where all slots share the same feature value, the matching results of each slot are independent, allowing the same lookup key value to match multiple different slots within the same storage line. This provides a flexible candidate hit selection space for subsequent steps, effectively improving the cache's tolerance to hash collisions.
[0045] Q420: For any comparison result signal RA, perform a logical AND judgment with the valid flag bit of the slot corresponding to RA. If the valid flag bit is valid and the feature value comparison result is consistent, then the slot corresponding to RA is determined to be a candidate hit slot.
[0046] In this embodiment, step Q300 reads the entry feature value and the valid flag bit vld[i] corresponding to each slot. The value of i ranges from 0 to n-1. The valid flag bit is a 1-bit signal. vld[i]=1 indicates that the entry feature value currently stored in the i-th slot is valid, and vld[i]=0 indicates that the slot is idle.
[0047] The hardware uses n two-input AND gates. For the i-th slot, the two inputs of the AND gate are connected to the comparison result signal cmp[i] and the valid flag bit vld[i], respectively. The output of the AND gate is the candidate hit indicator signal hit_cand[i]. hit_cand[i] = 1 if and only if cmp[i] = 1 and vld[i] = 1, indicating that the i-th slot is a candidate hit slot; otherwise, hit_cand[i] = 0.
[0048] The hardware logic in this step forcibly masks the comparison results of free slots. When vld[i]=0 for a certain slot, hit_cand[i] is always 0 regardless of the value of cmp[i]. When vld[i]=1 for a certain slot but cmp[i]=0, hit_cand[i] is also 0, indicating that although the slot is occupied, it is not the entry matched in this search.
[0049] This step incorporates the validity flag into the candidate hit judgment, requiring only one gate delay to complete the parallel validity screening of n slots. This mechanism fundamentally avoids the problem of false hits caused by idle slots coincidentally matching key-value features due to stored residual values or power-on initial values, ensuring the correctness of cache lookup results. In cache mode, slots are frequently allocated and released due to replacement operations. Accurate screening of the validity flag ensures that newly released idle slots will not produce false matches due to residual old data, maintaining the consistency of the cache state.
[0050] Q430: If there is at least one candidate hit slot, then select one from the candidate hit slots as the final hit slot according to the preset priority order.
[0051] In this embodiment, firstly, the n candidate hit indication signals hit_cand[0] to hit_cand[n-1] generated in step Q420 are input into an n-input OR gate to generate a group hit signal group_hit. group_hit=1 indicates that at least one candidate hit slot exists; group_hit=0 indicates that no candidate hit exists in any slot.
[0052] When `group_hit=1`, the priority arbitration logic is activated. The default priority order is: the smaller the slot number, the higher the priority. That is, the priorities from high to low are slot 0, slot 1, ..., slot (n-1). This priority order is chosen for the simplicity of hardware implementation, as the fixed-priority encoder does not need to maintain dynamic priority states.
[0053] The hardware implementation of priority arbitration uses a fixed priority encoder. The encoder takes an n-bit candidate hit indication signal hit_cand[n-1:0] as input and outputs the final hit slot number hit_idx, with a bit width of ceil(log2n) bits. The encoding logic is as follows: scan from the low bit (0) to the high bit (n-1) and output the position number of the first bit that is 1. The specific logical expression is: when hit_cand[0]=1, hit_idx=0 regardless of the value of other bits; when hit_cand[0]=0 and hit_cand[1]=1, hit_idx=1; and so on.
[0054] Under normal cache operating conditions, since the feature values of entries stored in each slot come from the second hash operation results of different flow table entries, the probability of multiple slots in the same storage row simultaneously matching the same lookup key is extremely low. This arbitration mechanism is mainly used to cover this extremely low probability boundary case.
[0055] This step uses a fixed-priority encoder to complete a unique arbitration of multiple candidate hits within a single cycle, with hardware overhead consisting of only a few gate circuits and no timing bottlenecks. The fixed-priority design avoids the additional state storage and dynamic update circuitry required for complex arbitration logic such as round-robin or least recently used, aligning with the cache hardware design goal of low latency and high throughput. In slot replacement operations, new entries prioritize filling lower-numbered free slots, naturally increasing the probability of lower-numbered slots being hit; therefore, the fixed-priority rule is statistically reasonable.
[0056] Q440 uses the group index address as the read address and the slot number of the final hit slot as the path selection signal to read the corresponding key-value table entry from the Cache key-value table.
[0057] In this embodiment, the Cache key-value table adopts the same storage structure as the Cache hash table: a depth of H rows, with each row containing n slots, indexed by both the group index address and the slot number. Each slot in the key-value table stores the complete key value of a flow table entry, with a bit width of K bits, where K is equal to the bit width for looking up the key value in step Q100.
[0058] The read operation consists of two sub-actions: First, the group index address generated in step Q200 is applied to the row address input of the Cache key-value table to select the corresponding storage row. Then, the final hit slot number hit_idx output in step Q430 is used as a column selection signal to select the key-value table entry for the hit_idx-th slot from the n read slot data. Column selection of the key-value table is implemented by an n-to-1 multiplexer, whose data input is connected to the output ports of the n slots in that row of the key-value table, and whose selection control is connected to hit_idx.
[0059] In the pipelined implementation, the group index address is ready in step Q200 and can be sent to the key-value table address input in advance, initiating synchronously with the cache hash table read. hit_idx is latched after arbitration in step Q430 and applied to the multiplexer as a column strobe signal. The key-value table read latency is the same as the cache hash table read latency, both being one memory access cycle.
[0060] The cache key-value table and the cache hash table use a unified slot mapping relationship. When a first-order table entry is written to the addr-th row and i-th slot of the cache hash table through a replacement and backfill operation, the complete key-value pair of that entry is simultaneously written to the addr-th row and i-th slot of the cache key-value table. This mapping relationship is guaranteed by the replacement and backfill control unit during write operations.
[0061] This step employs a strict isomorphic design between the key-value table and the hash table in terms of depth and slot mapping. Using the group index address and slot number as a common index, it allows direct location of the corresponding entry in the key-value table after a successful hash table feature value match, eliminating the need for secondary address calculation or traversal of conflict chains. The column gating latency of the n-to-1 multiplexer remains constant and does not deteriorate with increasing slot length, ensuring the deterministic timing of key-value table reads. The key-value table stores complete key values, not just compressed signatures, providing a precise basis for the full-width comparison in step Q500. This fundamentally eliminates the possibility of cache misses due to hash collisions, ensuring the absolute correctness of cache lookup results.
[0062] Q500 compares the lookup key value with the complete key value stored in the Cache key value table entry. If they match, it is determined that the Cache has hit. The Cache data table entry of the corresponding slot is read, the lookup result is output, and the aging count of the hit slot is updated.
[0063] Furthermore, step Q500 includes the following steps:
[0064] Q510: Compare the search key value with the complete key value stored in the Cache key value table entry read in step Q400. If every bit is the same, the key value is determined to be a successful match, and proceed to Q520; otherwise, it is determined to be a Cache miss, and proceed to step Q600.
[0065] In this embodiment, a K-bit XOR gate array and a K-input NOR gate are configured in the hardware. The lookup key value `key_search` obtained in step Q100 is XORed bit by bit with the cache key value entry `key_tbl` read in step Q400, generating K XOR result lines. All K XOR result lines are input to a K-input NOR gate, and the output is the key-value matching signal `key_match`. When every bit of `key_search` and `key_tbl` is equal, all XOR result lines are low, `key_match` = 1, indicating a successful key-value match and a cache hit; if any bit is unequal, `key_match` = 0, indicating a failed key-value match and a cache miss.
[0066] For example: Suppose K = 128 bits. The first 127 bits of key_search and key_tbl are equal. The 128th bit is key_search 1 and key_tbl 0. Then the XOR gate outputs 1, the NOR gate outputs 0, key_match = 0, and it is determined that there is a cache miss.
[0067] This step achieves precise matching verification between the search key and the stored key through a full-width XOR comparison. Since the feature comparison in the previous step Q400 only compares the w-bit compressed signature, different keys may produce the same feature value (i.e., hash feature value collision). This step extends the discrimination criterion to the complete K-bit key value, reducing the probability of a false hit from the feature value collision probability to zero, thus ensuring the correctness of the cache lookup result. The full-width comparison is implemented using pure combinational logic, and the latency is proportional to the logarithm of the key width. Even for a 320-bit wide key value, the latency remains within an acceptable range per cycle and does not affect the timing convergence of the cache lookup pipeline.
[0068] Q520 uses the group index address as the read address and the slot number of the finally hit slot as the path selection signal to read the data table entry of the corresponding slot from the Cache data table as the lookup result output.
[0069] After key-value verification is successful, this step reads the final lookup result from the Cache data table. The indexing method is completely consistent with the key-value table read in step Q440: the group index address generated in step Q200 is used as the row address to select the corresponding storage row, and the hit slot number determined in step Q430 is used as the column strobe signal. An n-to-1 multiplexer selects the corresponding data table entry from the n slots of that row for output. The Cache data table, Cache hash table, and Cache key-value table have the same depth and slot mapping relationship; these three tables together constitute a complete on-chip cache storage structure.
[0070] This step leverages the design advantages of isomorphic mapping across the three tables, using a unified group index address and slot number as a common index. This allows the reading of search result data to share the same address generation logic as key-value reading, eliminating the need for additional address calculations. The hardware implementation is simple and efficient, completing data return within a single memory access cycle. The separate storage architecture of the three tables allows the hash table, key-value table, and data table to be physically implemented using different memory types and bit widths, optimizing their area and power consumption characteristics respectively.
[0071] Q530, reset the aging count corresponding to the final hit slot to the initial value, and decrement the aging count of the remaining slots in the same storage row respectively.
[0072] This step involves dynamically updating the aging count after a cache hit, and it is the core mechanism for tracking hot and cold caches in the cache replacement strategy.
[0073] The aging counter table has the same depth and slot mapping as the cache hash table. Each slot maintains an aging counter value with a width of A bits, where A is typically 4 to 8 bits. The aging counter is updated by the counter update unit in the next clock cycle after the hit signal is generated.
[0074] For a hit in slot i, its aging count age_cnt[i] is reset to the initial value init_value. The initial value init_value is usually set to the maximum value that the counter's bit width can represent (e.g., when A=4, init_value=4'b1111, i.e., 15), or configured according to the system's preset initial heat level. The reset operation is implemented by writing init_value to the i-th slot in the addr-th row of the aging count table.
[0075] For all slots j (j≠i) within the same storage row except for the hit slot, their aging count age_cnt[j] is decremented by 1. The decrement operation is implemented using a subtractor: if the current age_cnt[j]>0, then age_cnt[j]-1 is written; if the current age_cnt[j]=0, then it remains at 0 and is not decremented further to prevent the aging state from being disordered due to counter flipping. The saturation process of decrement ensures that cold entries that have not been hit for a long time remain stable at a value of 0, and will not be mistakenly identified as hot entries due to continuous decrementing and flipping to the maximum value.
[0076] For example: Let addr=0x1A, n=4, and the hit slot is the 2nd slot. The current aging count values are: age_cnt[0]=5, age_cnt[1]=0, age_cnt[2]=3, age_cnt[3]=8. After the update: age_cnt[2] is reset to 15 (init_value), age_cnt[0] is decremented to 4, age_cnt[1] remains unchanged at 0, and age_cnt[3] is decremented to 7.
[0077] This step achieves fine-grained hot / cold status tracking at the slot level through a differentiated update strategy that resets the initial value of hit slots and decrements the value of other slots. Hot items, due to frequent hits, maintain a consistently high aging count; cold items, having not been hit for a long time, gradually decay their aging count to zero. Slots with zero aging counts are prioritized for replacement in the Q600's replacement and backfilling operation, thus naturally achieving a near-Least Recently Used (LRU) replacement effect. Compared to traditional LRU linked list implementations, this scheme has minimal hardware overhead for counter increment / decrement operations, requiring only adders and comparators, eliminating the need to maintain complex linked list pointers and sorting logic; furthermore, the update operation is tightly coupled with the search process, completing in the next clock cycle after a hit, eliminating the need for background periodic scanning, simplifying the hardware state machine design, and reducing power consumption.
[0078] Q600: If the feature value comparison shows no match, or the complete key value comparison shows no consistency, it is determined as a cache miss, generating an access request to the off-chip main storage area. The search result is obtained from the off-chip main storage area, and a replacement and backfill operation is performed based on the aging count and operation count of each slot in the storage row. The cache hash table, cache key-value table, and cache data table constitute a cache subset of the off-chip main storage area. The aging count is used to characterize the length of time the corresponding slot has been missed, and the operation count is used to characterize whether the corresponding slot is currently in an operational state. During the replacement and backfill operation, the slot with an aging count of zero and an operation count of zero is selected as the slot to be replaced.
[0079] Example 2:
[0080] Based on the cache pattern lookup process described in Embodiment 1, this embodiment provides a cache status entry management mechanism to further improve data consistency and off-chip storage interface bandwidth utilization efficiency in cache replacement operations. This mechanism maintains a corresponding cache status entry for each slot in the cache hash table, including three fields: a data validity flag, an off-chip validity flag, and an off-chip address field. Before the replacement and backfill operation in Embodiment 1 is executed, the cache status entry needs to be queried to determine the status of the replaced entry: if the off-chip validity flag is valid and the data has been modified, the data needs to be written back to the off-chip main storage area according to the off-chip address field; if the off-chip validity flag is invalid, the data is discarded directly without needing to be written back. It should be noted that this cache status entry management mechanism works in conjunction with the aging count and operation count in Embodiment 1—the aging count and operation count are used to select the slot to be replaced, while the cache status entry guides the write-back strategy for the replaced slot; together, they constitute a complete cache replacement management system.
[0081] Furthermore, the method also includes: maintaining a corresponding cache state table entry for each slot of the cache hash table, each cache state table entry including:
[0082] In this embodiment, the data validity flag is used to indicate whether the cache data table entry corresponding to the slot has been backfilled from the off-chip main memory area and is available.
[0083] The data validity flag, denoted as data_vld, has a bit width of 1 bit and is used to indicate whether the cache data table entry corresponding to the slot has been backfilled from the off-chip main memory area and is in an available state.
[0084] The hardware storage of data_vld is located at bit 0 of the corresponding slot in the cache status table. Its setting and clearing operations are performed by the hardware controller under specific event triggers. After the off-chip main memory returns the lookup result data and completes the writing to the cache data table entry, the hardware controller sets data_vld to 1 in the next clock cycle, indicating that the data is ready and available. When a slot is selected as the replacement slot and enters the replacement process, data_vld is cleared in step Q630, marking that the data in that slot is no longer valid, preventing subsequent lookup requests from reading incomplete or outdated data.
[0085] During the lookup process, when step Q300 reads the Cache hash table using the group index address, it simultaneously reads the Cache status table entries corresponding to all N slots within that storage row. When step Q500 determines a Cache hit, the hardware checks the data_vld value corresponding to the hit slot. If data_vld=1, it indicates that the data is ready and can be directly read from the Cache data table and returned; if data_vld=0, it indicates that although the entry has been allocated a slot, the corresponding data has not yet been backfilled from the off-chip main memory area, and the process of waiting for backfilling must be initiated.
[0086] This data validity flag decouples the slot allocation status from the data readiness status, allowing the search pipeline to correctly identify entries even in the intermediate state of "allocated but data pending backfilling," thus avoiding the problem of incorrectly returning empty or invalid data due to data not yet being ready.
[0087] The external valid flag is used to indicate whether the flow table entry corresponding to the slot has a backup copy in the external main memory area.
[0088] In this embodiment, the off-chip valid flag bit, denoted as ddr_vld, has a bit width of 1 bit and is used to indicate whether the flow table entry corresponding to the slot has a backup copy in the off-chip main storage area.
[0089] The hardware storage of `ddr_vld` is located at the first bit of the corresponding slot in the cache status table. Its setting and clearing operations depend on the source of the entry. When a flow table entry is filled back into the cache from off-chip main memory, a backup copy of the entry naturally exists in off-chip main memory, and the hardware controller sets `ddr_vld` to 1 while writing the entry data. When a flow table entry is written directly to the on-chip cache from the control plane and designated as a temporary entry, there is no corresponding copy of the entry in off-chip main memory, and the hardware controller sets `ddr_vld` to 0. When a flow table entry comes from on-chip escape memory, there is also no backup copy of the entry in off-chip main memory, and `ddr_vld` is set to 0.
[0090] The ddr_vld value plays a crucial role in the replacement operation. During the replacement write-back in step Q700, the hardware controller reads the ddr_vld value of the slot being replaced. If ddr_vld = 1, the contents of the cache data table entry for the slot being replaced must be written back to the off-chip main memory to ensure the data in the off-chip main memory is updated; if ddr_vld = 0, the contents of the slot being replaced are discarded directly, and no write-back operation is performed.
[0091] This off-chip valid flag allows the hardware to accurately distinguish between entries that need to be written back and those that don't during replacement, avoiding the coarse-grained strategy of writing back all modified entries based on the dirty bit in traditional caches. For purely on-chip temporary entries without off-chip backups, direct discarding eliminates invalid off-chip write-back operations, significantly saving off-chip storage interface bandwidth and chip power consumption. The 1-bit width makes the storage and decision logic overhead of this flag extremely small.
[0092] The off-chip address field is used to record the physical storage address of the flow table entry in the off-chip main storage area when the off-chip valid flag is valid.
[0093] In this embodiment, the off-chip address field, denoted as ddr_addr, has a width of P bits. It is used to record the physical storage address of the flow table entry in the off-chip main memory area when the off-chip valid flag is valid. The value of P is determined by the addressing range of the off-chip main memory area, and the typical value is 24 to 32 bits.
[0094] The hardware storage of ddr_addr is located at the high P bit position of the corresponding slot in the cache state table, forming a complete cache state table entry together with data_vld and ddr_vld. When ddr_vld=1, the ddr_addr field stores the valid address; when ddr_vld=0, the content of the ddr_addr field is meaningless and is ignored by the hardware logic.
[0095] The timing of writing to `ddr_addr` is synchronized with the setting of `ddr_vld`. When a flow table entry is filled back into the cache from off-chip main memory, the hardware controller writes the physical address returned by the off-chip main memory (i.e., the address of the entry in the fixed slot of the Base area or the node of the linked list area in the off-chip main memory) into the `ddr_addr` field. This address has two uses in subsequent operations: First, during replacement write-back, if `ddr_vld=1`, the hardware controller initiates a write operation to the specified location in the off-chip main memory based on the `ddr_addr` field, writing the latest data of the replaced entry back to the original address; Second, during data backfilling, if a lookup hits but `data_vld=0`, the hardware controller initiates a read operation to the specified location in the off-chip main memory based on the `ddr_addr` field to obtain the data that has not yet been backfilled.
[0096] This off-chip address field directly binds the off-chip physical address of an entry to the on-chip cache slot. This allows the cache to directly access the off-chip storage location based on the stored address during fill and write-back operations, eliminating the need for hash calculations or looking up the off-chip index table to determine the off-chip storage location. This design eliminates the latency and hardware overhead of secondary address calculations during fill / write-back in traditional schemes, minimizing the latency of off-chip access in the replacement process. Furthermore, the combined use of ddr_addr and ddr_vld allows the address field overhead of purely on-chip entries and entries without backups to be reused, resulting in no wasted storage.
[0097] Furthermore, in step Q600, before performing the replacement backfill operation, the following steps are also included:
[0098] Q601, read the cache status table entry corresponding to the replaced slot.
[0099] This step obtains the status information of the slot to be replaced before the replacement and backfill operation is executed, providing a basis for the subsequent write-back decision.
[0100] After the replacement decision sub-step S630 in step Q600 determines the slot to be replaced, the hardware controller uses the group index address used in the current lookup operation as the row address and the slot number of the slot to be replaced as the column selection signal to read the corresponding cache status table entry from the cache status table. This read operation uses the same indexing method as the cache hash table read in step Q300, utilizing the ready group index address and slot number, without requiring additional address calculation. The data returned by the read includes three fields: data validity flag data_vld, off-chip validity flag ddr_vld, and off-chip address field ddr_addr. Simultaneously, the hardware controller also reads the modification flag dirty corresponding to the slot. This flag records whether the cache data table entry has been modified after being written. If the entry has not been modified since backfilling, dirty=0; if the control plane or data plane has updated the data table entry for that entry, dirty=1.
[0101] This step utilizes the isomorphic mapping between the cache state table and the cache hash table, along with ready address signals, to complete the reading of state information within a single memory access cycle, without introducing additional access latency. By storing the modification flag, data validity flag, off-chip validity flag, and off-chip address field together in the same state table entry, all the state information required for the replacement write-back decision can be obtained in a single read, avoiding the cumulative latency from multiple accesses to different memory units.
[0102] Q602, if the external valid flag is valid and the corresponding Cache data table entry has been modified, then the Cache data table entry content of the slot to be replaced is written back to the external main storage area according to the external address field.
[0103] This step performs a write-back operation with a backup entry, synchronizing the latest data of the replaced entry in the cache to the off-chip main storage area, ensuring the consistency of data in the off-chip main storage area.
[0104] The hardware controller makes a decision on the status field read in step Q601. The decision conditions are: ddr_vld=1 and dirty=1. Both conditions must be met simultaneously to trigger the write-back operation. If ddr_vld=1 but dirty=0, it means that the entry has a backup copy off-chip and the data in the cache is consistent with the off-chip data. No write-back is needed, and the subsequent replacement and backfilling steps can proceed directly.
[0105] When the decision condition is met, the hardware controller performs the following write-back operation: First, using the group index address as the row address and the replaced slot number as the column selection signal, it reads the data table entry content of the replaced slot from the cache data table; then, using the ddr_addr field value read in step Q601 as the write address, it initiates a write request to the external main memory controller, writing the data table entry content to the specified physical address in the external main memory area. The write request format includes the write address, write data, and write enable signal. After receiving the write request, the external main memory controller writes the data to the corresponding physical memory location through the DDR interface.
[0106] After the write-back operation is completed, the hardware controller clears the dirty flag of the slot to zero, indicating that the cache and off-chip data have been synchronized.
[0107] This step implements a precise on-demand write-back strategy through a two-condition decision based on `ddr_vld` and `dirty`. For entries with off-chip backups but not modified, write-back operations are skipped, eliminating invalid off-chip write operations; write-back is only performed on entries that truly require synchronization, minimizing the write bandwidth usage of the off-chip storage interface and chip power consumption. Utilizing the pre-stored `ddr_addr` field in the cache status table entry as the write address directly avoids the latency and hardware overhead of recalculating the off-chip physical address during write-back, allowing the write-back operation to be initiated immediately after the replacement decision, minimizing the total latency of the replacement process.
[0108] Q603, if the external valid flag is invalid, the contents of the Cache data table entry of the replaced slot are directly discarded, and no write-back operation is performed.
[0109] This step handles the replacement of entries without backups and purely on-chip temporary entries, directly freeing up slot space without any off-chip write operations.
[0110] The hardware controller checks the ddr_vld field read in step Q601. If ddr_vld=0, it means that there is no backup copy of the entry in the off-chip main memory. This situation includes two typical scenarios: first, the entry is written directly to the on-chip cache by the control plane and is specified as a temporary policy entry, so there is no need to retain a copy off-chip; second, the entry originates from the on-chip escape memory area and has not yet been allocated storage space in the off-chip main memory area.
[0111] In the above scenario, the hardware controller skips the write-back process and does not initiate any write operations to the off-chip main memory area. The cache hash table entry, cache key-value table entry, and cache data table entry of the replaced slot will be directly overwritten by the data of the new entry in subsequent step Q650, and the old data will be naturally discarded. The cache status table entry corresponding to this slot will also be reset when the new entry is written, and ddr_vld will be reassigned according to the source of the new entry.
[0112] This step, by checking the external valid flag, enables the hardware to automatically identify the special state of having no backup entry and skip the write-back operation, avoiding address out-of-bounds errors or data corruption that may result from writing to invalid external addresses. For temporary policy entries dynamically injected by the control plane, this mechanism supports zero-cost discarding, ensuring that the creation and destruction of temporary entries only involve writing to and overwriting the on-chip cache, without consuming external storage interface bandwidth. This design correlates the write-back overhead of cache replacement with the actual number of entries that need to be protected, rather than with all replaced entries. In network application scenarios with a large number of temporary policies, this can significantly save external write bandwidth and chip power consumption.
[0113] The above Embodiments 1 and 2 describe the lookup process and state management scheme of the present invention in Cache mode (i.e., inner_mode=0). It should be noted that this Cache mode is not the only operating mode of the chip. The chip described in this invention supports switching to independent table mode (i.e., inner_mode=1) on the same set of hardware resources via a mode configuration signal. The following Embodiment 3 will describe the lookup process in independent table mode. This process shares the same hardware pipeline as Embodiment 1 in the front-end steps such as double hashing, feature value matching, and key value verification. The difference is that in independent table mode, the aging count is not updated when a hit occurs, and when a miss occurs, the lookup result is obtained from the off-chip independent storage area and returned directly without performing a replacement backfill operation, and there is no need to maintain the Cache state table entries described in Embodiment 2.
[0114] Example 3:
[0115] The lookup process in the independent table schema may include the following steps:
[0116] S100 receives the search request and retrieves the search key value.
[0117] In this embodiment, the lookup request is generated by the packet parsing unit inside the network chip. When a packet enters the chip, the parsing unit extracts predetermined fields from the packet header and concatenates them to form a lookup key. The constituent fields and concatenation order of the lookup key are predefined by the flow table configuration file. Typical fields include source IP address, destination IP address, source port number, destination port number, protocol type, and virtual network identifier. The lookup request receiving unit latches the lookup key into an input register and simultaneously requests a free buffer pointer from the sequence control module to temporarily store the packet's context information during the lookup process. The hardware implementation of this step can be a multiplexer working with a set of configuration registers to select the corresponding packet fields according to the currently active flow table configuration file and concatenate them for output.
[0118] This step utilizes a hardware-configurable multi-field concatenation mechanism, enabling the same lookup pipeline to adapt to the differentiated definition requirements of lookup keys in different network application scenarios. The key composition can be adjusted without modifying the hardware circuitry, thus improving the chip's flexibility and versatility.
[0119] S200, perform a first hash operation on the lookup key value to obtain the group index address, and perform a second hash operation on the lookup key value to obtain the feature value.
[0120] Furthermore, the step of performing a second hash operation on the lookup key to obtain a feature value includes the following steps:
[0121] S210, input the search key value into the second hash function to obtain a hash result value with a bit width of W.
[0122] In this embodiment, the second hash function is implemented in hardware using combinational logic circuits and operates in parallel with the first hash function without introducing additional clock cycles. The input to the second hash function is the lookup key value obtained in step S100, and its bit width is determined according to the network application scenario; typical values can be 128 bits, 256 bits, or 320 bits. The output bit width W of the second hash function is jointly determined by the number of slots n per row of the on-chip hash table and the bit width of each feature value segment, satisfying W = n × w, where w is the bit width of each feature value segment.
[0123] The second hash function can employ a randomized hash algorithm based on XOR-Shift. Its hardware implementation consists of a multi-stage XOR shift network: first, the input key value is split into multiple equal-width fields and XORed to produce intermediate results; then, the intermediate results are sequentially subjected to several stages of shift and XOR operations, with each shift amount being a coprime prime offset (e.g., shift by 6, 13, or 27 bits) to fully diffuse the influence of the input bits; finally, the lower W bits of the shift network are taken as the hash result value for output.
[0124] This step uses a second hash function to compress lookup keys of arbitrary width into hash results of fixed width W. W is determined by the number of slots in each row and the feature value width of each slot, allowing a single hash calculation to provide a complete data source for parallel feature value matching of all subsequent slots. The combinational logic of the XOR-Shift algorithm completes the hash calculation within a single cycle without increasing lookup latency. Simultaneously, its multi-level shift-XOR structure effectively reduces the probability of different keys producing the same hash result, providing a high-resolution basis for subsequent slot-level matching.
[0125] S220, the hash result value is divided into n feature value segments; each feature value segment corresponds to a slot in the same storage row of the on-chip hash table; n is the maximum number of entries that can be stored in the row of the on-chip hash table.
[0126] In this embodiment, the partitioning operation requires no logic gates or registers in hardware and can be completed via hardwired bus. The physical connections of the W-bit hash result output in step S210 are divided into n groups according to the principle of equal width, with each group containing w bits. The 0th group (lowest w bits) corresponds to the feature value segment of slot 0, the 1st group corresponds to the feature value segment of slot 1, and so on, with the (n-1)th group (highest w bits) corresponding to the feature value segment of slot (n-1). Each group of w-bit signals is directly connected to the comparator input port of the corresponding slot in the subsequent step S400.
[0127] When W is not divisible by n, the following method is used: The bit width w of each segment is taken as floor(W / n), i.e., rounded down. For the first (W mod n) segments, an additional 1 bit is allocated to each, making the bit width of these segments w+1. At this point, the bit widths of each feature value segment may differ by 1 bit. Each segment is still compared with the feature value of the corresponding slot for equal bit width, and the comparator bit width is configured according to the actual bit width of each segment.
[0128] This step directly divides the result of a single hash calculation into n independent feature value segments via hardwired bus, providing matching criteria for each of the n slots without additional logic overhead. Compared to the traditional approach of calculating a hash independently for each slot, this avoids the latency accumulation and hardware area expansion caused by multiple hash calculations. Compared to the approach where all slots share the same feature value, each slot uses independent feature value segments, ensuring that the same lookup key value generates different feature values in different slots. When a feature value match fails in one slot, other slots may still match successfully, thus providing multiple candidate hit positions for the same key value within the same storage row, effectively reducing the probability of lookup failures due to single feature value conflicts. Furthermore, the zero-latency characteristic of hardwired partitioning ensures that the entire feature value generation and allocation process does not occupy additional pipeline truncation time, guaranteeing a consistently high throughput for the lookup pipeline.
[0129] S300, using the group index address as the read address, read multiple pre-stored entry feature values from the corresponding storage row in the on-chip hash table, with each entry feature value corresponding to a slot in the corresponding storage row.
[0130] Furthermore, each storage row of the on-chip hash table contains n slots, and each slot stores an entry feature value and a corresponding valid flag bit; reading multiple entry feature values pre-stored in the corresponding storage row includes: simultaneously reading the entry feature values and valid flag bits of n slots in the corresponding storage row.
[0131] In this embodiment, the group index address generated in step S200 is used as the read address to initiate a single-cycle read operation on the on-chip hash table. The on-chip hash table is implemented using static random access memory, with a depth of H rows. Each row contains n slots, and each slot stores an entry feature value and its corresponding valid flag bit. A valid flag bit of 1 indicates that the entry feature value stored in that slot is valid, while 0 indicates that the slot is idle. After the group index address is sent to the address input of the on-chip hash table, after one read cycle, the entry feature values and valid flag bits of all n slots in the storage row are read out in parallel and latched into n sets of output registers. The on-chip hash table adopts a single-port design, and the read operation is exclusively powered by the lookup pipeline. The number of slots n is determined based on a combination of the desired storage capacity and hash collision tolerance. The larger n is, the stronger the tolerance for collisions in the same row, but the storage area and read power consumption increase accordingly. Typical n values are 4 or 8.
[0132] This step reads the entry feature values and valid flag bits of all slots in the same storage row in parallel in a single cycle, compressing the slot polling that requires multiple serial accesses in the traditional scheme into a single parallel access, which greatly reduces the lookup latency; the multi-slot group associative structure provides flexible conflict tolerance space under limited storage depth, balancing storage efficiency and lookup performance.
[0133] S400, the feature value of the search key value is compared in parallel with the feature values of multiple read entries. If a matching entry feature value exists, the key value table entry of the slot corresponding to the matching entry feature value is read.
[0134] Furthermore, step S400 includes the following steps:
[0135] S410, compare the feature values of the n entries read in step S300 with the feature values of the search key values obtained in step S200 in parallel to generate n comparison result signals.
[0136] The data read in step S300 contains n sets of entry feature values. The i-th set of entry feature values corresponds to the entry feature value stored in the i-th slot, denoted as bin_tbl[i], with a bit width of w bits. The feature value of the lookup key generated in step S200 has been divided into n feature value segments, denoted as bin_key[i], with a bit width of w bits.
[0137] The hardware uses an array of n w-bit XOR gates. For the i-th slot, bin_tbl[i] and bin_key[i] are XORed bit by bit, generating w XOR result lines. These w XOR result lines are then input into a w-input NOR gate, and the output of the NOR gate is the i-th comparison result signal cmp[i]. When every bit of bin_tbl[i] and bin_key[i] is equal, all w XOR result lines are low, and the NOR gate outputs a high level, indicating that the feature value comparison result is consistent. When any bit is unequal, the NOR gate outputs a low level, indicating that the feature value comparison result is inconsistent.
[0138] The n comparator arrays are fully parallel in hardware and have no data dependency on each other. The combinational logic delay of all comparators from input to output is only one XOR gate delay plus one NOR gate delay, which can be completed in one clock cycle.
[0139] This step achieves fully parallel comparison of feature values for all slots using n independent hardware comparator arrays. The comparison latency is constant at two gate stages, independent of the number of slots (n), thus preventing increased lookup latency as the number of slots expands and ensuring pipeline timing convergence. Each slot uses an independent feature value fragment for matching. Compared to schemes where all slots share the same feature value, the matching results for each slot are independent, allowing the same lookup key value to match multiple different slots within the same storage row, providing a flexible candidate match selection space for subsequent steps.
[0140] S420, for any comparison result signal RA, if the valid flag bit of the slot corresponding to RA is valid and the feature value comparison result is consistent, then the slot corresponding to RA is determined to be a candidate hit slot.
[0141] In step S300, while reading the entry feature value, the valid flag bit vld[i] corresponding to each slot is also read. The value of i ranges from 0 to n-1. The valid flag bit is a 1-bit signal. vld[i]=1 indicates that the entry feature value currently stored in the i-th slot is valid, and vld[i]=0 indicates that the slot is idle.
[0142] The hardware uses n two-input AND gates. For the i-th slot, the two inputs of the AND gate are connected to the comparison result signal cmp[i] and the valid flag bit vld[i], respectively. The output of the AND gate is the candidate hit indicator signal hit_cand[i]. hit_cand[i] = 1 if and only if cmp[i] = 1 and vld[i] = 1, indicating that the i-th slot is a candidate hit slot; otherwise, hit_cand[i] = 0.
[0143] The hardware logic in this step forcibly masks the comparison results of free slots. When vld[i]=0 for a certain slot, hit_cand[i] is always 0 regardless of the value of cmp[i]. When vld[i]=1 for a certain slot but cmp[i]=0, hit_cand[i] is also 0, indicating that although the slot is occupied, it is not the entry matched in this search.
[0144] This step incorporates the gatekeeper's valid flag into the candidate hit determination, completing the parallel validity screening of n slots with only one gate delay. This mechanism fundamentally avoids the problem of erroneous hits caused by coincidental matching of key-value features in idle slots due to stored residual values or power-on initial values, thus ensuring the correctness of the search results.
[0145] S430, if there is at least one candidate hit slot, then select one from the candidate hit slots as the final hit slot according to the preset priority order.
[0146] In this embodiment, firstly, the n candidate hit indication signals hit_cand[0] to hit_cand[n-1] generated in step S420 are input into an n-input OR gate to generate a group hit signal group_hit. group_hit=1 indicates that there is at least one candidate hit slot; group_hit=0 indicates that no candidate hits are found in any slot.
[0147] When group_hit=1, the priority arbitration logic is activated. The default priority order is: the smaller the slot number, the higher the priority. That is, the priority from high to low is slot 0, slot 1, ..., slot (n-1).
[0148] The hardware implementation of priority arbitration uses a fixed-priority encoder. The encoder takes an n-bit candidate hit indication signal `hit_cand[n-1:0]` as input and outputs the final hit slot number `hit_idx`, with a bit width of `ceil(log₂n)` bits. The encoding logic is as follows: scanning from the high-order bits (n-1) to the low-order bits (0), outputting the position number of the first 1 bit, with the scanning direction corresponding to low priority to high priority, so that the valid bit with the smallest number is selected.
[0149] The specific logical expression is as follows: when hit_cand[0]=1, hit_idx=0 regardless of the value of other bits; when hit_cand[0]=0 and hit_cand[1]=1, hit_idx=1; and so on. This encoder is implemented entirely by combinational logic, with a delay of several levels of gate delay.
[0150] This step uses a fixed-priority encoder to complete a unique arbitration of multiple candidate hits within a single cycle, with hardware overhead consisting of only a few gate circuits and no timing bottlenecks. The fixed-priority design avoids the additional state storage and update circuitry required for complex round-robin or least recently used arbitration logic, consistent with the design principle of no automatic hardware state updates in the independent table mode of this embodiment. Smaller slot numbers indicate higher priority; this rule is simple and clear, allowing the software to prioritize filling lower-numbered slots when configuring entries to optimize the hardware arbitration path.
[0151] S440, using the group index address as the read address and the slot number of the final hit slot as the route selection signal, read the corresponding key value table entry from the key value table.
[0152] In this embodiment, the on-chip key-value table adopts the same storage structure as the on-chip hash table: a depth of H rows, with each row containing n slots, indexed by both group index address and slot number. Each slot in the key-value table stores the complete key value of a flow table entry, with a bit width of K bits, where K is equal to the bit width of the key value search in step S100, typically 128 bits, 256 bits, or 320 bits.
[0153] The read operation consists of two sub-actions: First, the group index address generated in step S200 is applied to the row address input of the key-value table to select the corresponding storage row. Then, the final hit slot number hit_idx output in step S430 is used as a column selection signal to select the key-value table entry of the hit_idx-th slot from the n slot data read. The column selection of the key-value table is implemented by an n-to-1 multiplexer, whose data input is connected to the output ports of the n slots in that row of the key-value table, and whose selection control is connected to hit_idx.
[0154] The entire reading process is coordinated with the hash table read in step S300. In the pipelined implementation, the group index address is ready in step S200 and can be sent to the key-value table address input in advance; hit_idx is latched after arbitration in step S430 and serves as a column strobe signal. The key-value table read latency is the same as the hash table read latency, both being one memory access cycle.
[0155] The on-chip key-value table and the on-chip hash table use a unified slot mapping relationship. When a first-order table entry is written to the addr-th row and i-th slot of the on-chip hash table, the complete key-value of that entry is simultaneously written to the addr-th row and i-th slot of the on-chip key-value table. This mapping relationship is guaranteed by the control plane when configuring entries, and is not modified during hardware lookup.
[0156] This step employs a strict isomorphic design between the key-value table and the hash table in terms of depth and slot mapping, using the group index address and slot number as a common index. This allows for direct location of the corresponding entry in the key-value table after a successful match in the hash table's feature value, eliminating the need for secondary address calculation or traversal of collision chains. The column gating latency of the n-to-1 multiplexer remains constant and does not deteriorate with the increase in the number of slots, ensuring the deterministic timing of key-value table reads. The key-value table stores complete key values rather than compressed signatures, providing a precise basis for the full-width comparison in step S500 and fundamentally eliminating the possibility of false hits due to hash collisions.
[0157] S500: If the search key value matches the complete key value stored in the key-value table entry, then the data table entry of the corresponding slot is read and the search result is output; otherwise, the search result is obtained from the off-chip independent storage area; wherein, the on-chip hash table, key-value table, and data table constitute an exact matching flow table independent of the off-chip storage, the on-chip hash table, key-value table, and data table have the same depth and slot mapping relationship, and are indexed by the group index address and slot number; the on-chip storage entries are pre-configured, and no hardware automatic replacement operation is performed on the on-chip hash table, key-value table, and data table during the search process.
[0158] Furthermore, step S500 includes the following steps:
[0159] S510, compare the search key value with the complete key value stored in the key value table entry read in step S400. If every bit is consistent, the key value is determined to be successfully matched and proceed to S520; otherwise, proceed to S530.
[0160] The hardware consists of an XOR gate array with a bit width of K and a K-input NOR gate. The search key value `key_search` obtained in step S100 is XORed bit by bit with the key value entry `key_tbl` read in step S400, generating K XOR result lines. All K XOR result lines are input to a K-input NOR gate, and the output is the key-value matching signal `key_match`. When every bit of `key_search` and `key_tbl` is equal, all XOR result lines are low, `key_match` = 1, indicating a successful key-value match; if any bit is unequal, `key_match` = 0, indicating a failed key-value match.
[0161] This step achieves precise matching verification between the search key and the stored key through a full-width XOR comparison, fundamentally eliminating the possibility of false hits caused by feature value hash collisions. Since the feature value matching in the previous step S400 only compares the w-bit compressed signature, different keys may produce the same feature value. This step extends the discrimination criterion to the complete K-bit key value, reducing the probability of false hits from the probability of feature value collisions to zero, thus ensuring the absolute correctness of the search result.
[0162] S520: Using the group index address as the read address and the slot number of the finally hit slot as the path selection signal, read the data table entry of the corresponding slot from the data table and output the data table entry as the search result.
[0163] After key-value verification is successful, this step reads the final search result from the data table. Its indexing method is completely consistent with the key-value table reading in step S440: the corresponding storage row is selected using the group index address as the row address, and the corresponding data table item is selected and output from the n slots using the hit slot number as the column strobe signal. The data table, hash table, and key-value table have the same depth and slot mapping relationship; together, they constitute a complete on-chip exact match flow table.
[0164] This step leverages the design advantages of isomorphic mapping of three tables, using a unified group index address and slot number as a common index, so that reading the lookup result data does not require additional address calculations. The hardware implementation is simple and efficient, and the data return can be completed in one memory access cycle.
[0165] S530, generate an access request to the off-chip independent storage area, retrieve the flow table entry data corresponding to the lookup key value from the off-chip independent storage area as the lookup result output, and do not perform backfilling operations on the on-chip hash table, key value table and data table.
[0166] When step S510 determines that the key value matching fails, or when the feature value comparison in the previous step S400 fails to find a match, the hardware generates an off-chip access request. This request carries a lookup key value or an off-chip address generated based on the lookup key value and is sent to the off-chip memory controller. The off-chip memory controller accesses the independent flow table in the off-chip dynamic random access memory (DDR) according to the address information in the request, obtains the corresponding flow table entry data, and returns the lookup result to the message processing unit.
[0167] Throughout the process, the hardware does not perform any write operations on the on-chip hash table, key-value table, or data table. The addition, deletion, and updating of entries in the on-chip tables are entirely pre-completed by the control plane through a separate configuration interface at a time other than the lookup operation. The hardware state machine for a missed path only includes the two actions of "reading from outside the chip → returning the result," and does not include the backfill action of "reading from outside the chip → writing to inside the chip → returning the result."
[0168] This step significantly reduces hardware design complexity and state machine size by simplifying miss handling into a single off-chip read operation.
[0169] In both the Cache mode described in Embodiment 1 and the Independent Table mode described in Embodiment 3, when an on-chip lookup misses, the off-chip storage area must be accessed to obtain the lookup result. The difference lies in that: after obtaining the result from the off-chip main storage area, the Cache mode needs to backfill the data into the on-chip Cache and update the corresponding status (aging count, operation count, Cache status table entry), while the Independent Table mode returns directly after obtaining the result from the off-chip independent storage area without performing a backfill operation. Regardless of the mode, the structural design of the off-chip storage area directly affects lookup performance and storage efficiency. Embodiment 4 below provides a method for setting up an off-chip hierarchical storage structure applicable to the above two modes. This method determines the number of slots in the base area through a probability distribution model and achieves an optimized balance between storage overhead and lookup performance with a two-level structure of the base area plus the linked list area. In the Cache mode of this invention, the off-chip storage structure is a specific implementation of the "off-chip main storage area" mentioned in step Q600. The Cache backfill operation reads entries from this structure, and the replacement write-back operation writes entries to this structure. In Example 2, the physical address recorded in the off-chip address field of the Cache status table entry points to the base area fixed slot or linked list node of the off-chip storage structure.
[0170] Example 4:
[0171] R100, obtain the flow table size parameters; the flow table size parameters include: the total number of flow table entries M to be stored, and the number of storage rows N corresponding to the preset index bit width of the off-chip hash table.
[0172] This step determines the design input parameters for the off-chip hash table. The flow table size parameters include the total number of flow table entries M to be stored and the number of storage rows N corresponding to the preset index bit width of the off-chip hash table. M is determined by the target application scenario of the network chip, with typical values being tens of millions of entries, such as 67,108,864. N is determined by the index bit width of the off-chip hash table; if the index bit width is 23 bits, then the number of storage rows N=2. 23 =8,388,608 lines. Parameters M and N can be read from the configuration file or registers during chip initialization, or they can be fixed as constants in the hardware parameters during the design phase.
[0173] This step defines the design inputs of the off-chip storage structure as two quantifiable macroscopic parameters, which allows the subsequent determination of the number of slots and capacity planning to be based on precise mathematical foundations, avoiding storage waste or overflow risks caused by setting parameters based on experience.
[0174] R200, based on the total number of flow table entries M and the number of storage rows N, calculates the average load λ=M / N within each storage row of the off-chip hash table.
[0175] Based on M and N obtained in step R100, the average row load λ = M / N within each storage row of the off-chip hash table is calculated. λ represents the average number of flow table entries carried per storage row under an ideal uniform hash distribution. For example, if M = 67,108,864 and N = 8,388,608, then λ = 8, meaning there are approximately 8 entries per row on average. This value serves as the core parameter for subsequent probability calculations.
[0176] This step transforms the macroscopic flow table size and number of storage rows into the average load of a single storage row through a simple division operation, reducing the global storage planning problem to a single-row probability distribution analysis problem, which greatly simplifies the complexity of subsequent parameter selection.
[0177] R300, based on λ, uses a probability distribution model to calculate the probability P(k) of exactly k entries falling within a single storage line for different non-negative integers k.
[0178] Furthermore, the probability distribution model is a binomial distribution model or a Poisson distribution model; when the values of M and N cause the average load λ to not exceed a preset threshold, the probability value P(k) is calculated using a Poisson distribution model:
[0179] ;
[0180] Where e is a natural constant and k is a non-negative integer.
[0181] In this embodiment, under the assumption of an ideal hash function, each flow table entry is independently and uniformly mapped to one of N storage rows, and the number of entries falling into a single storage row follows a binomial distribution. When N is large and λ is moderate, the binomial distribution can be approximated as a Poisson distribution to reduce computational complexity.
[0182] Specifically, when using the Poisson distribution model, the formula for calculating the probability value P(k) is: Where e is a natural constant and k is a non-negative integer. Hardware or software configuration tools can calculate P(k) sequentially for k=0,1,2,... until the value of P(k) is small enough to be ignored. Taking λ=8 as an example, the calculated values are P(0)≈0.0335%, P(1)≈0.2684%, P(8)≈13.96%, P(12)≈4.81%, etc.
[0183] This step quantifies the random phenomenon of hash collisions into precise probability values through a probability distribution model, enabling designers to pre-assess the overflow risk under different slot lengths, providing a scientific quantitative basis for subsequent slot length selection. The Poisson distribution approximation avoids the computational complexity of large factorials involved in the binomial distribution, is easy to implement in engineering, and the approximation error is negligible when N is sufficiently large.
[0184] R400, based on the calculated probability values P(k), selects the minimum k value that makes the cumulative probability reach the preset coverage threshold Pth, as the fixed slot number η of the base area of the off-chip hash table.
[0185] Based on the probability distribution in step R300, this step determines the number of base slots η according to the preset coverage threshold Pth. The cumulative probability is defined as the sum of each P(k) from k to η, representing the probability that the number of entries in a storage row does not exceed η. The value of Pth ranges from 90% to 99%, with a typical value of 93% to 95%.
[0186] For example: λ=8, Pth=93.62%. Looking up the cumulative probability table, the cumulative probability is 93.62% when k=12, which just reaches the threshold, so η=12. At this point, approximately 93.62% of the storage rows have no more than 12 entries, and only about 6.38% of the storage rows require overflow storage in the linked list area.
[0187] This step quantifies the trade-off between storage efficiency and overflow probability by using cumulative probability and coverage threshold. The smallest η value that satisfies the threshold is selected, minimizing the number of slots per row in the base area while ensuring that the vast majority of entries can be stored in fixed slots. This significantly reduces the storage overhead of the off-chip Key / Data table. Compared to fixed large-slot schemes, this method achieves a substantial reduction in storage area with a very low overflow probability.
[0188] Furthermore, the value of Pth ranges from 90% to 99%; the cumulative probability is the sum of the probability values P(k) from k from 0 to the current value; when selecting a fixed number of slots η, the following conditions are met:
[0189] ;
[0190] Where η is the smallest non-negative integer that satisfies the above conditions.
[0191] In practice, P(k) can be accumulated starting from k=0. The k at which the first accumulated value is greater than or equal to Pth is η. If Pth=90%, η may be 10; if Pth=99%, then η=14. The higher Pth, the more fixed slots in the base region, resulting in greater storage overhead but a lower overflow probability; the lower Pth, the opposite is true. Designers can flexibly configure Pth according to the chip area and the tolerance for overflow without modifying the hardware structure.
[0192] This step provides a configurable coverage threshold parameter, Pth, which allows the same storage architecture to be adapted to different cost and performance requirements. The programmability of Pth gives system integrators the ability to fine-tune based on actual traffic characteristics and storage budgets without redesigning the hardware.
[0193] R500 divides the off-chip storage space into a base area and a linked list area. The base area contains N storage rows, each with η fixed slots, for storing the corresponding hash map flow table entries. The linked list area is used to dynamically store overflow entries that exceed the number of fixed slots in the base area.
[0194] In this embodiment, the base area contains N storage rows, each with η fixed slots. Each fixed slot stores the complete data of a flow table entry, including the key value, result data, and status flags. The base area is implemented using static random access or dynamic random access memory, directly indexed by row address and slot number, with a determined access latency.
[0195] The linked list area is used to dynamically store overflow entries that exceed the fixed number of slots in the base area. The linked list area consists of several linked list nodes, each storing one overflow entry and a pointer to its next node. The size of a node is comparable to the data size of a single slot in the base area. The linked list area and the base area can physically reside in different address segments of the same DDR memory.
[0196] Each row in the base area maintains a linked list head pointer, initially a null pointer. When a row overflows, a node is allocated from the linked list area and linked to the head of the linked list.
[0197] For example: η=12, each row in the base area has 12 fixed slots. The total capacity of the linked list area is determined based on the expected overflow E_overflow. If a row already has 12 entries filling the base area, when the 13th entry arrives, there is no space in the base area, so a node is allocated from the linked list area and inserted into the linked list.
[0198] This approach utilizes a two-tiered storage architecture consisting of a base area and a linked list area. Fixed slots ensure low latency and determinism for main path lookups, while the linked list area dynamically expands as needed to accommodate long tails of hash collisions. Only a small amount of linked list space is required to cover low-probability overflow scenarios, resulting in 30%-50% less storage space compared to a fully fixed-slot solution. The physical separation of the base area and the linked list area also allows the base area to use higher-performance memory types, optimizing access latency in common scenarios.
[0199] Furthermore, the capacity of the linked list area is determined through the following steps:
[0200] R510, based on P(k), calculate the expected number of storage rows E_overflow for more than η entries:
[0201] ;
[0202] R520, based on the expected value E_overflow, sets the linked list area to provide storage space for at least E_overflow linked list nodes, with each linked list node used to store an overflow entry and its next node pointer.
[0203] In this embodiment, the capacity of the linked list area is determined by the expected number of rows to overflow, E_overflow. Taking λ=8 and η=12 as an example, E_overflow≈535,000. The linked list area provides space for at least E_overflow nodes, with some margin considered for allocation granularity. Each node needs to store complete entry data and a pointer to the next node; the node size is slightly larger than a single slot in the base area. Total storage capacity of the linked list area = number of nodes × node size.
[0204] For example: E_overflow≈535,000, node size 256 bits (including 192 bits of data + 64 bits of pointer), the linked list area requires approximately 535K × 256 bits ≈ 16.7MB.
[0205] This step directly links the linked list area capacity to the expected number of overflow nodes, making the size of the overflow storage area scientifically controllable and avoiding waste due to excessive reservation or insufficient storage space due to insufficient space after overflow. Through probability expectation calculation, the planning of the linked list area capacity has a clear mathematical basis, and the design results are predictable and reproducible.
[0206] R600: When a first-order table entry is mapped to a target storage row via hash operation, and all η fixed slots in the base area of the target storage row are occupied, a linked list node is allocated from the linked list area to store the overflow entry, and the linked list node is linked to the overflow linked list of the target storage row; wherein, the off-chip hash table works in conjunction with the on-chip cache, and the on-chip cache caches a subset of hot entries in the off-chip hash table.
[0207] In this embodiment, overflow processing is triggered when a first-order list entry is mapped to a target storage row via hash operation, and all η base area fixed slots of that row are occupied. The hardware controller allocates a linked list node from the free node queue of the linked list area, writes the complete data of the overflow entry to the node, and inserts the node at the head (or tail) of the overflow linked list of the target storage row. The allocation and linking operations of the linked list node are completed by the linked list management state machine, and the time required depends on the access latency of the linked list area, which is typically several DDR clock cycles.
[0208] For example: The target storage row already has 12 entries filling the base area slots. When the 13th entry arrives, the base area is full. The linked list controller retrieves a node from the free linked list queue, writes the entry data, sets the node pointer to the current linked list head (which was previously empty), and updates the row's linked list head pointer to the new node.
[0209] This step utilizes a dynamic linked list expansion mechanism to provide the off-chip hash table with elastic capacity beyond the fixed slots in the base area. This allows for the smooth absorption of localized overflows caused by hash collisions, avoiding the serious drawback of fixed-slot schemes that directly discard entries upon overflow. The on-demand allocation characteristic of the linked list ensures that overflow storage overhead is proportional to the actual degree of collision. When hash distribution is uniform, there is almost no additional overhead; when distribution is uneven, it automatically adapts, resulting in high overall storage utilization.
[0210] Furthermore, in step R600, when a first-order table entry LB is replaced from the on-chip cache and needs to be written back to off-chip storage, the method further includes the following:
[0211] R610: If the LB was originally stored in a fixed slot in the base area, then write it back to the original slot address.
[0212] R620: If the LB was originally stored in the linked list area, then the next node pointer stored in the original linked list node of the LB is used to traverse to the target node and then the write-back operation is performed.
[0213] In this embodiment, when an entry in the on-chip cache needs to be written back to off-chip storage due to replacement, its position in the off-chip base area or linked list area must be accurately located. If the entry was originally stored in a fixed slot in the base area, its off-chip address was determined during the initial allocation as "base area base address + row number × row slot offset + slot number × slot size", and the original address is directly used during the write-back. If it was originally stored in a node in the linked list area, the off-chip address field of the cache status table entry stores the starting address of the linked list node, and the hardware directly initiates a write operation to that address. If the position of the linked list node changes due to linked list reorganization, the off-chip address field in the cache needs to be updated.
[0214] For example: Entry A is located in slot 5 of row 100 in the off-chip base area. The off-chip address is fixed and known, so the replacement write-back address is directly the address of this slot. Entry B is located at node 0x1000 in the linked list area. The cache state table records the address at 0x1000. During replacement write-back, new data is written to 0x1000.
[0215] This step ensures the determinism and efficiency of write-back operations by differentiating between two write-back location methods: fixed mapping of the base area and direct addressing of the linked list area. Fixed mapping of the base area eliminates the need for additional storage and lookup of the write-back address for most entries, saving address field overhead; direct addressing of the linked list area utilizes the precise address pre-stored in the cache state table, avoiding the latency of traversing the linked list and ensuring write-back throughput.
[0216] Example 5:
[0217] In the off-chip hierarchical storage structure described in Embodiment 4 above, when the fixed slots in the base area of a storage row are full and there are no available nodes in the linked list area, newly arriving flow table entries will face the problem of having nowhere to be stored. For Cache mode, this situation corresponds to the branch "no slot with a valid replacement indicator signal" in step Q600, that is, the corresponding row in the on-chip Cache has no replacement slots and the off-chip main storage area cannot accommodate the new entry; for independent table mode, this situation corresponds to the scenario where both the base area and the linked list area of the off-chip independent storage area are full. To solve this boundary case, this embodiment provides an on-chip escape storage area mechanism as a third-level elastic storage outside the off-chip base area and linked list area. This escape storage area is composed of content addressable memory (TCAM) and random access memory (SRAM) cascaded together and is shared by all flow table groups within the chip. The basic structure, triggering conditions, entry format, and dynamic management strategy of the escape storage area will be described below. This escape storage area can be used independently of any of the above embodiments, or it can be combined with the off-chip hierarchical storage structure of Embodiment 4 to form a complete three-level storage system (on-chip Cache / independent table → off-chip base area + linked list area → on-chip escape area). It should be noted that when an entry is stored in the escape storage area, its corresponding off-chip valid flag is set to invalid in Cache mode, consistent with the state management of the entry in Embodiment 2.
[0218] The on-screen escape storage mechanism may include the following steps:
[0219] The R700 has an on-chip escape storage area, which is composed of a content addressable memory and a random access memory cascaded together; the escape storage area is shared by all flow table groups within the chip.
[0220] In this embodiment, an escape storage area is provided on-chip, consisting of cascaded Content Addressable Memory (TCAM) and Random Access Memory (SRAM). The TCAM stores the flow table group identifier and complete lookup key, while the SRAM stores the corresponding action data. The escape area is shared by all flow table groups within the chip. When the fixed slots in the base area of the off-chip hash table are full, and there are no available nodes in the linked list area of the corresponding row, the overflow entry is transferred to the escape area. Writing to the escape area is automatically triggered by hardware, requiring no software intervention.
[0221] For example: If all 12 slots in the corresponding row base area of a certain flow table group are full in the off-chip DDR, and the free nodes in the linked list area are exhausted, a new entry will arrive. The hardware will write the <flow table group ID, lookup key> of this entry to the TCAM, and the action data will be written to the cascaded SRAM.
[0222] This step provides a third level of elastic storage outside the on-chip cache through an on-chip escape area, avoiding service interruptions caused by discarding flow table entries when both the off-chip base area and linked list area are full. The escape area is shared by all flow table groups, covering extreme overflow spikes of multiple flow table groups with a relatively small total capacity, resulting in high storage resource utilization. The cascaded structure of TCAM+SRAM supports single-cycle key-value matching lookups, and entries matched in the escape area still achieve low latency similar to that of the on-chip cache.
[0223] R710: When the fixed slots in the base area of the off-chip hash table are full and there are no available linked list nodes in the linked list area of the corresponding storage row, the overflow entries are stored in the escape storage area.
[0224] Each escape entry contains three fields: a flow group identifier field (the bit width is determined by the number of flow groups, such as 8 bits), a lookup key field (the bit width is equal to the full key width K), and an action data field (the bit width is equal to the data entry bit width). During a lookup, the flow group ID to which the current packet belongs and the lookup key are concatenated as the TCAM matching input. Each TCAM entry can be configured to precisely match a specified bit; upon a match, the corresponding SRAM address is output, and the action data is read and returned.
[0225] For example: Querying the key value 0xABCD... in flow table group 3, the escape TCAM contains...<Group=3,Key=0xABCD...> When the TCAM match is successful, output SRAM address 0x00F and read the action data from SRAM.
[0226] This step leverages the parallel matching feature of TCAM to achieve single-cycle parallel lookup of escape entries for all flow table groups, maintaining a constant lookup latency regardless of the number of entries stored in the escape zone. The use of flow table group IDs in the matching process isolates the key-value spaces of different flow table groups, avoiding key-value conflicts between groups and ensuring lookup accuracy.
[0227] Furthermore, the storage format of each entry in the escape storage area includes:
[0228] The flow group identifier field indicates the flow group to which this entry belongs;
[0229] The lookup key field stores the complete lookup key value for this flow table entry;
[0230] The action data field is used to store the forwarding or processing policy corresponding to this flow table entry;
[0231] During the search, the current flow group identifier and the search key value are used as the matching input for the content addressable memory. If a match is found, the action data field corresponding to the random access memory address is read and output.
[0232] In this embodiment, each escape entry contains three fields: a flow group identifier field (the bit width is determined by the number of flow groups, such as 8 bits), a lookup key field (the bit width is equal to the full key width K), and an action data field (the bit width is equal to the data entry bit width). During the lookup, the flow group ID to which the current packet belongs and the lookup key are concatenated as the TCAM matching input. Each TCAM entry can be configured to precisely match a specified bit; upon a match, the corresponding SRAM address is output, and the action data is read and returned.
[0233] For example: Querying the key value 0xABCD... in flow table group 3, the escape TCAM contains...<Group=3,Key=0xABCD...> When the TCAM match is successful, output SRAM address 0x00F and read the action data from SRAM.
[0234] This step leverages the parallel matching feature of TCAM to achieve single-cycle parallel lookup of escape entries for all flow table groups, maintaining a constant lookup latency regardless of the number of entries stored in the escape zone. The use of flow table group IDs in the matching process isolates the key-value spaces of different flow table groups, avoiding key-value conflicts between groups and ensuring lookup accuracy.
[0235] Furthermore, the total capacity of the escape storage area is determined based on the following parameters:
[0236] The number of flow table groups supported by the chip, the expected number of overflow entries for each flow table group under the preset coverage threshold Pth, and the preset capacity of the linked list area.
[0237] When the usage rate of the off-chip linked list area corresponding to any flow table group exceeds a preset warning threshold, the hardware triggers the migration of some linked list entries to the escape storage area, or sends an expansion alarm signal to the control plane.
[0238] In this embodiment, the total capacity of the escape zone is determined comprehensively based on the number of flow table groups, the expected number of overflow entries for each flow table group under the coverage threshold Pth, and the preset capacity of the linked list area. During design, the theoretical overflow amount of each flow table group is first calculated using a probability model, summed, and multiplied by the number of flow table groups supported by the chip. Then, the expected simultaneous overflow peak is multiplied by an empirical coefficient (e.g., 1.5) to obtain the recommended capacity of the escape zone. During runtime, the hardware monitors the usage rate of the linked list area for each flow table group. When the usage rate of a flow table group's linked list area exceeds a warning threshold (e.g., 80%), the hardware triggers the migration of low-frequency access entries in the linked list to the escape zone, releasing the linked list nodes; or it reports an alarm to the control plane, requesting software expansion. When the usage rate of the escape zone itself also reaches its limit, the control plane is notified to adjust the overall flow table specifications.
[0239] For example, the chip supports 32 flow table groups, with each flow table group expected to have approximately 5,000 overflow entries. The initial escape zone capacity is set to 32 × 5,000 × 1.5 = 240,000 entries. During operation, the linked list area of the 5th flow table group reaches 85% utilization. The hardware migrates the cold entries at the tail of this group's linked list to the escape zone, and the linked list utilization drops back to 70%.
[0240] This step, through multi-level linked capacity planning and dynamic migration mechanisms, enables a closed-loop elastic scheduling between the on-chip external linked list area and the on-chip escape area. The probabilistic model not only guides static capacity planning but also provides a reference benchmark for dynamic threshold setting. Hardware-triggered automatic migration and alarms enable real-time adaptive allocation of off-chip storage resources, significantly improving the system's ability to withstand hash collision spikes while maintaining low cost.
[0241] Furthermore, the method also includes the following steps:
[0242] The H900 has an on-chip escape storage area; the escape storage area is composed of a content addressable memory and a random access memory cascaded together; the escape storage area is shared by all flow table groups within the chip.
[0243] When the corresponding storage line in the on-chip cache has no available slots, and the entry cannot be accommodated in the off-chip main memory, an escape memory area is set up on-chip as a last resort. The escape memory area consists of cascaded Content Addressable Memory (TCAM) and Random Access Memory (SRAM), shared by all flow table groups within the chip. The TCAM portion stores the flow table group identifier and lookup key, while the SRAM portion stores the corresponding action data. Writes to the escape memory area are automatically triggered by hardware. When an entry is stored in the escape memory area, its corresponding off-chip valid flag bit ddr_vld is set to invalid, indicating that there is no backup copy of the entry in the off-chip main memory. Subsequently, if the entry is migrated from the escape memory area to the off-chip main memory, ddr_vld is updated to valid and the corresponding ddr_addr is updated.
[0244] For example: If a flow table entry has its corresponding line in the on-chip cache completely full, and there is no space in the off-chip base area and linked list area, the hardware writes the entry to the escape TCAM+SRAM. The status of the corresponding slot in the cache for this entry is marked as ddr_vld=0.
[0245] This step provides a third level of elastic storage beyond the cache and chip tables through the escape storage area, avoiding the serious problem of discarding entries due to both levels of storage being full. The escape area is shared by all flow table groups, covering extreme overflow spikes from multiple groups with a smaller total capacity. The unified flag `ddr_vld=0` for escape area entries maintains consistency in state management, ensuring that subsequent replacements still follow the low-overhead discard path of H800.
[0246] H910 When a first-order table entry needs to be stored, but there are no available slots in the corresponding storage row of the on-chip cache, and the entry cannot be accommodated in the off-chip main storage area, the entry is stored in the escape storage area.
[0247] H920, when a first-order table entry is stored in the escape storage area, the off-chip valid flag position of the slot corresponding to the entry is invalid, indicating that there is no backup copy of the entry in the off-chip main storage area.
[0248] Furthermore, the storage format of each entry in the escape storage area includes:
[0249] Flow group identifier field; used to indicate the flow group to which this entry belongs;
[0250] Lookup key field; used to store the complete lookup key value for this flow table entry as a matching input to content-addressable memory;
[0251] Action data field; used to store the forwarding or processing policy corresponding to this flow table entry;
[0252] During the search, the current flow group identifier and the current search key value are used as the matching input for the content addressable memory. After a match is found, the action data field corresponding to the random access memory address is read and output.
[0253] Each entry in the escape storage area has three fields: a flow group identifier field (its width is determined by the number of flow groups supported by the chip, indicating the flow group to which the entry belongs); a lookup key field (its width is equal to the full lookup key width K, serving as the TCAM matching input); and an action data field (its width is equal to the data table entry width, used to store policy data such as forwarding port and next-hop address). During a lookup, the flow group identifier of the current packet and the lookup key are concatenated as the TCAM matching input. All TCAM entries are compared in parallel, and upon a match, the corresponding SRAM address is output, and the action data is read and returned.
[0254] For example: Querying the key value 0xABCD in flow table group 5, the escape TCAM contains...<Group=5, Key=0xABCD> When the TCAM match is successful, the SRAM address 0x010 is output, and the motion data is read and output.
[0255] This step leverages the parallel matching feature of TCAM to achieve single-cycle parallel lookup of escape entries for all flow table groups, with a constant lookup latency independent of the number of entries. The use of flow table group IDs in the matching process isolates the key-value spaces of different flow table groups, ensuring lookup accuracy.
[0256] Furthermore, the method also includes:
[0257] H930 When the used capacity of the escape storage area reaches a preset migration threshold, the hardware controller migrates at least one entry with the lowest access frequency in the escape storage area to the off-chip main storage area, updates the off-chip valid flag corresponding to the migrated entry to valid, and updates the off-chip address field to the new physical storage address of the entry in the off-chip main storage area.
[0258] When the used capacity of the escape memory reaches a preset migration threshold (e.g., 70% of the total capacity), the hardware controller automatically triggers a migration operation. The hardware scans the escape memory, identifies at least one entry with the lowest access frequency, reads it from the escape memory, allocates space in the off-chip main memory (either a free slot in the base area or a new node in the linked list area), writes it, updates the off-chip valid flag bit `ddr_vld` of the corresponding cache slot to valid, and updates the off-chip address field `ddr_addr` to the newly allocated physical address. After the migration is complete, the storage space for this entry in the escape memory is released.
[0259] This step uses an automatic migration mechanism to move infrequently accessed cold entries from expensive on-chip TCAM resources to inexpensive off-chip DDR, ensuring rapid hits for high-frequency entries while avoiding cost overruns caused by escape zone capacity expansion. After migration, synchronized updates of ddr_vld and ddr_addr maintain consistency between state management and physical storage.
[0260] H940, when the used capacity of the escape storage area reaches a preset alarm threshold, the hardware controller sends an expansion alarm signal to the control plane; wherein, the alarm threshold is greater than the migration threshold.
[0261] When the used capacity of the escape storage area reaches a preset alarm threshold (e.g., 90% of the total capacity, where the alarm threshold exceeds the migration threshold), the hardware controller sends an expansion alarm signal to the control plane, carrying diagnostic information such as the current escape area utilization rate and the number of overflow entries in each flow table group. Upon receiving the alarm, the control plane can choose to: expand the escape area capacity configuration, increase the capacity of the off-chip linked list area, or adjust the flow table splitting strategy to reduce the load on a single group. The alarm signal is reported via an internal chip interrupt or a status register.
[0262] This step achieves a seamless transition from automatic hardware mitigation to software intervention through tiered thresholds (migration threshold < alarm threshold). When the migration threshold is triggered, the hardware first performs self-healing; when the alarm threshold is triggered, the software is notified to make a decision, forming a closed-loop management system that enhances the chip's fault tolerance under extreme loads and its long-term operational stability.
[0263] Furthermore, although the steps of the method in this disclosure are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or a step may be broken down into multiple steps.
[0264] Embodiments of the present invention also provide a chip, the chip comprising:
[0265] The search request receiving unit is used to receive search requests and obtain search key values.
[0266] The hash operation unit is used to perform a first hash operation on the lookup key value to obtain a group index address, and to perform a second hash operation on the lookup key value to obtain a feature value.
[0267] The cache hash table uses the group index address as the read address and outputs multiple pre-stored entry feature values in the corresponding storage row. Each entry feature value corresponds to a slot in the corresponding storage row.
[0268] The feature value comparison unit is used to compare the feature value of the search key value with the feature values of multiple entries output by the Cache hash table in parallel, and output the matching slot number when a matching entry feature value exists.
[0269] The Cache key-value table has the same depth and slot mapping relationship as the Cache hash table. It uses the group index address and the matching slot number as a common index to output the complete key-value stored in the corresponding slot.
[0270] The key-value comparison unit is used to compare the search key-value with the complete key-value output by the Cache key-value table. If they match, a hit signal is output; otherwise, a miss signal is output.
[0271] The Cache data table has the same depth and slot mapping relationship as the Cache hash table and Cache key-value table. When the hit signal is received, the group index address and the matching slot number are used as a common index to output the data table entry of the corresponding slot as the search result.
[0272] An aging count table is used to maintain a corresponding aging count value for each slot in the Cache hash table. The aging count value represents the length of time that the corresponding slot has not been hit.
[0273] An operation count table is used to maintain a corresponding operation count value for each slot in the Cache hash table. The operation count value indicates whether the corresponding slot is currently in an operation state.
[0274] The counting update unit is used to reset the aging count value corresponding to the hit slot to the initial value when the hit signal is received, and to decrement the aging count values of the other slots in the same storage row respectively.
[0275] An off-chip access control unit is used to generate an access request to the off-chip main memory area when the miss signal is received, and to obtain the search result from the off-chip main memory area.
[0276] The replacement and backfill control unit is used to select the slot to be replaced based on the aging count value and operation count value of each slot in the storage row indexed by the group index address when the miss signal is received, and after returning the search result in the off-chip main memory area, write the returned flow table entry data into the Cache key-value table entry and Cache data table entry corresponding to the slot to be replaced; wherein, the replacement and backfill control unit selects the slot with an aging count value of zero and an operation count value of zero as the slot to be replaced.
[0277] Embodiments of the present invention also provide a network interface card (NIC), which includes the chip described in the above embodiments.
[0278] While specific embodiments of the invention have been described in detail by way of examples, those skilled in the art should understand that the examples are for illustrative purposes only and are not intended to limit the scope of the invention. Those skilled in the art should also understand that various modifications can be made to the embodiments without departing from the scope and spirit of the invention.
Claims
1. A method for precise matching flow table lookup based on on-chip cache, characterized in that, The method includes the following steps: Q100 receives search requests and retrieves search key values; Q200, perform a first hash operation on the lookup key value to obtain the group index address, and perform a second hash operation on the lookup key value to obtain the feature value; Q300, using the group index address as the read address, reads multiple entry feature values pre-stored in the corresponding storage row from the Cache hash table, with each entry feature value corresponding to a slot in the corresponding storage row; Q400, compare the feature value of the search key value with the feature values of multiple read entries in parallel. If a matching entry feature value exists, read the Cache key value table entry of the slot corresponding to the matching entry feature value. Q500 compares the lookup key value with the complete key value stored in the Cache key value table entry. If they match, it is determined that the Cache has hit. The Cache data table entry of the corresponding slot is read, the lookup result is output, and the aging count of the hit slot is updated. Q600: If the feature value comparison shows no match, or the complete key value comparison shows no consistency, it is determined as a cache miss, generating an access request to the off-chip main storage area. The search result is obtained from the off-chip main storage area, and a replacement and backfill operation is performed based on the aging count and operation count of each slot in the storage row. The cache hash table, cache key-value table, and cache data table constitute a cache subset of the off-chip main storage area. The aging count is used to characterize the length of time the corresponding slot has been missed, and the operation count is used to characterize whether the corresponding slot is currently in an operational state. During the replacement and backfill operation, the slot with an aging count of zero and an operation count of zero is selected as the slot to be replaced.
2. The precise matching flow table lookup method based on on-chip cache according to claim 1, characterized in that, Step Q600 includes the following steps: Q610, using the group index address as the read address, read the aging count value and operation count value corresponding to all n slots in the storage row corresponding to the group index address from the aging count table and the operation count table respectively; Q620 compares the aging count value of each slot with zero and the operation count value of each slot with zero to generate n replaceable indication signals; wherein, the replaceable indication signal corresponding to a slot is valid if and only if the aging count value and the operation count value of a certain slot are both zero. Q630: If there is at least one slot with a valid replacement indicator signal, select one of the valid slots as the slot to be replaced according to the preset priority order, increment the operation count value of the slot to be replaced by one, and set the data validity flag in the Cache status table entry corresponding to the slot to invalid. Q640 generates a read access request to the off-chip main memory area, the read access request carrying the lookup key value or an off-chip address generated based on the lookup key value; while waiting for the off-chip main memory area to return the lookup result, the operation count value of the replaced slot is kept non-zero; Q650: When the off-chip main memory returns the search result, the complete key value in the returned flow table entry data is written into the Cache key value table entry corresponding to the replaced slot, the result data in the returned flow table entry data is written into the Cache data table entry corresponding to the replaced slot, and the aging count value corresponding to the slot is set to the initial value. Q660, decrement the operation count value of the replaced slot by one, and set the data validity flag in the Cache status table entry corresponding to the slot to valid; wherein, if there is no slot with a valid replaceable indication signal in step S630, then abandon the replacement backfill operation, do not modify the aging count value and operation count value of any slot, and only output the search result returned by the off-chip main memory area.
3. The precise matching flow table lookup method based on on-chip cache according to claim 1, characterized in that, In step Q200, a second hash operation is performed on the lookup key to obtain a feature value, including the following steps: Q210, Input the lookup key value into the second hash function to obtain a hash result value with a bit width of W; Q220, the hash result value is divided into n feature value segments, each feature value segment corresponding to a slot in the same storage row of the Cache hash table; where n is the maximum number of entries that can be stored in a row of the Cache hash table.
4. The precise matching flow table lookup method based on on-chip cache according to claim 1, characterized in that, Each storage row of the Cache hash table contains n slots, and each slot stores an entry feature value and a corresponding valid flag bit. When the group index address is used as the read address, the Cache hash table reads the entry feature values and corresponding valid flag bits of all n slots in the storage row in parallel within one read cycle.
5. The precise matching flow table lookup method based on on-chip cache according to claim 1, characterized in that, Step Q400 includes the following steps: Q410, compare the feature values of the n entries read in step Q300 with the feature values of the search key value obtained in step Q200 in parallel to generate n comparison result signals; Q420: For any comparison result signal RA, perform a logical AND judgment with the valid flag bit of the slot corresponding to RA. If the valid flag bit is valid and the feature value comparison result is consistent, then the slot corresponding to RA is determined to be a candidate hit slot. Q430: If there is at least one candidate hit slot, then select one from the candidate hit slots as the final hit slot according to the preset priority order. Q440 uses the group index address as the read address and the slot number of the final hit slot as the path selection signal to read the corresponding key-value table entry from the Cache key-value table.
6. The precise matching flow table lookup method based on on-chip cache according to claim 1, characterized in that, Step Q500 includes the following steps: Q510, compare the search key value with the complete key value stored in the Cache key value table entry read in step Q400. If every bit is the same, the key value is determined to be a successful match and proceed to Q520; otherwise, it is determined to be a Cache miss and proceed to step Q600. Q520: Using the group index address as the read address and the slot number of the finally hit slot as the path selection signal, the corresponding slot data table entry is read from the Cache data table as the lookup result output. Q530, reset the aging count corresponding to the final hit slot to the initial value, and decrement the aging count of the remaining slots in the same storage row respectively.
7. The precise matching flow table lookup method based on on-chip cache according to claim 1, characterized in that, The method further includes: maintaining a corresponding cache state table entry for each slot of the cache hash table, wherein each cache state table entry includes: The data validity flag indicates whether the cache data table entry corresponding to the slot has been backfilled from the off-chip main memory and is available; The external valid flag is used to indicate whether a backup copy exists in the external main memory area for the flow table entry corresponding to the slot. The off-chip address field is used to record the physical storage address of the flow table entry in the off-chip main storage area when the off-chip valid flag is valid.
8. The precise matching flow table lookup method based on on-chip cache according to claim 7, characterized in that, In step Q600, before performing the replacement backfill operation, the following steps are also included: Q601, Read the Cache status table entry corresponding to the replaced slot; Q602, if the external valid flag is valid and the corresponding Cache data table entry has been modified, then the Cache data table entry content of the slot to be replaced is written back to the external main storage area according to the external address field; Q603, if the external valid flag is invalid, the contents of the Cache data table entry of the replaced slot are directly discarded, and no write-back operation is performed.
9. A chip, characterized in that, The chip includes: The search request receiving unit is used to receive search requests and obtain search key values; A hash operation unit is used to perform a first hash operation on the lookup key value to obtain a group index address, and to perform a second hash operation on the lookup key value to obtain a feature value; The cache hash table uses the group index address as the read address and outputs multiple entry feature values pre-stored in the corresponding storage row. Each entry feature value corresponds to a slot in the corresponding storage row. The feature value comparison unit is used to compare the feature value of the search key value with the feature values of multiple entries output by the Cache hash table in parallel, and output the matching slot number when a matching entry feature value exists; The Cache key-value table has the same depth and slot mapping relationship as the Cache hash table. Using the group index address and the matching slot number as a common index, it outputs the complete key-value stored in the corresponding slot. The key-value comparison unit is used to compare the search key-value with the complete key-value output by the Cache key-value table. If they match, a hit signal is output; otherwise, a miss signal is output. The Cache data table has the same depth and slot mapping relationship as the Cache hash table and Cache key-value table. When the hit signal is received, the group index address and the matching slot number are used as a common index to output the data table entry of the corresponding slot as the search result. An aging count table is used to maintain a corresponding aging count value for each slot in the Cache hash table. The aging count value represents the length of time that the corresponding slot has not been hit. An operation counter table is used to maintain a corresponding operation counter value for each slot of the Cache hash table. The operation counter value indicates whether the corresponding slot is currently in an operation state. The counting update unit is used to reset the aging count value corresponding to the hit slot to the initial value when the hit signal is received, and to decrement the aging count values of the other slots in the same storage row respectively. An off-chip access control unit is used to generate an access request to the off-chip main memory area when the miss signal is received, and to obtain the search result from the off-chip main memory area; The replacement and backfill control unit is used to select the slot to be replaced based on the aging count value and operation count value of each slot in the storage row indexed by the group index address when the miss signal is received, and after returning the search result in the off-chip main memory area, write the returned flow table entry data into the Cache key-value table entry and Cache data table entry corresponding to the slot to be replaced; wherein, the replacement and backfill control unit selects the slot with an aging count value of zero and an operation count value of zero as the slot to be replaced.
10. A network interface card (NIC), characterized in that, Includes the chip described in claim 9.
Citation Information
Patent Citations
Table entry adding, deleting and searching method of hash table and hash table storage device
CN102194002A
Off-chip one-time access management method and device for satellite-borne router key value storage
CN121283943A