Network-on-chip many-core architecture chip and table item storage and search method
By adopting on-chip network multi-core architecture and distributed entry storage management methods in the data communication chip, the chip's shortcomings in entry search efficiency are solved, and high-performance entry search and message forwarding are achieved.
Patent Information
- Application Number
- CN202510507159.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-22
AI Technical Summary
Existing data communication chips have shortcomings in table entry search efficiency, especially in the process of high-frequency message forwarding, a single SRAM controller cannot meet the high-performance requirements.
The on-chip network multi-core architecture chip is adopted to realize distributed storage and parallel search of table items through distributed storage and management methods of firmware micro-core and table lookup units. Specific measures include attaching the table search unit LM in the control layer network, and managing it through the firmware micro-core FW, and dynamically adjusting the table entry storage policy using capacity-first strategy or delay-first strategy.
The table entry search performance and message forwarding capabilities of the Digital Pass chip are improved. Through distributed storage and parallel search, the table lookup delay is reduced and the overall processing performance is improved.
Smart Images

Figure CN120017584A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data communication technology, and in particular to a network-on-chip multi-core architecture chip, and a table entry storage and search method. Background Art
[0002] Data communication chips (abbreviated as data communication chips), such as those used in network devices such as routers and switches, face a persistent problem, which is how to achieve higher forwarding performance at low cost.
[0003] With the accelerated development of network technology, the requirements for the processing performance of data communication chips are getting higher and higher. Simply increasing the hardware clock frequency and scaling up to improve chip performance has been unable to keep up with the needs of network applications.
[0004] The static storage technology used by Static Random Access Memory (SRAM) does not need to be refreshed periodically like Dynamic Random Access Memory (DRAM), and has the characteristics of high stability and fast read / write speed, but SRAM has complex manufacturing process, high power consumption and low integration. In order to improve the message processing performance, data communication chips usually choose to use SRAM as built-in memory. At the same time, in order to reduce power consumption and cost, the amount of built-in SRAM will not be too large.
[0005] Table entry lookup during message forwarding is a frequently used basic function of data communication chips. Although the read / write speed of SRAM is fast, the bandwidth of SRAM is also limited. To complete the forwarding action of a message in the data communication chip, it is usually necessary to look up multiple types of tables, such as the Media Access Control (MAC) table, LPM (Longest Prefix Match) table, Interface Table (INTF), etc. Assuming that the basic L2 / L3 layer forwarding of the message is completed, each message needs to look up the table 5 times, and the chip forwarding capacity is 3.2T. For a fixed-length 256B message (with a frame gap and preamble of 276B), the table needs to be looked up 3200G / 8 / 276*5=7.25G times per second. Even if a table lookup request is processed once per clock cycle, a single SRAM controller cannot meet such performance requirements.
[0006] The comparative document "Tightly Coupled Adaptive Coprocessing System Supporting Multi-core Network Processing Architecture" with application publication number CN104503948A discloses: a multi-core processor chip integrates multiple CPU cores programmed in C language, and the cores interact with each other through shared memory, cache consistency bus, or dedicated switching structure (such as connecting multiple cores, peripherals, coprocessors, etc. through a dedicated ring network cross network, etc.). Each core can be flexibly configured to perform a certain operation of network message processing, such as single operations such as message parsing, order preservation, table lookup, flow control, etc., to implement complex business processing. Multi-cores can also realize concurrent processing and achieve high throughput data forwarding.
[0007] The above-mentioned comparative documents only disclose that multi-core processor chips can be interconnected through shared memory or dedicated ring network cross-network, and different cores can be configured to perform different network message processing operations to achieve multi-core concurrent processing. However, the way that multi-core CPUs share memory through the bus will lead to competition for memory access. For data communication chips that require high-speed table entry lookup and message forwarding, the efficiency problem of table entry lookup still cannot be solved. Summary of the invention
[0008] In order to solve the technical problems existing in the prior art, the present invention provides a network-on-chip many-core architecture chip, and a table entry storage and search method.
[0009] Based on one aspect of an embodiment of the present invention, the present invention provides a network-on-chip multi-core architecture chip, the chip includes a control layer network-on-chip, each router of the control layer network-on-chip is mounted with a firmware micro-core FW and a table lookup unit LM; The firmware microcore FW includes a management module, which is used to manage the table lookup unit LM. When receiving the table entry sent by the upper layer, the table entry is distributedly stored in the table lookup unit LM; The table lookup unit LM comprises: The storage module is used for distributed storage of table items, and when receiving the management instruction issued by the firmware micro-core, performs management operations on the table items according to the management instruction; The table lookup module is used to receive a table entry lookup request sent by the forwarding layer, perform table entry lookup and return the lookup result to the forwarding layer.
[0010] Furthermore, the management module of the firmware micro-core adopts a capacity priority strategy to distribute and store the table entry data in the table lookup unit LM mounted on the control layer on-chip network; under the capacity priority strategy, each table entry stores only one copy; or, The management module of the firmware micro-core adopts a delay priority strategy to distribute and store the table entry data in the table lookup unit LM mounted on the control layer on-chip network; under the delay priority strategy, each table entry stores multiple copies.
[0011] Further, the table lookup unit LM mounted on the control layer on-chip network is divided into a plurality of LM groups; Under the latency priority strategy, each LM group stores a copy of the table entry; Under the capacity priority strategy, only one copy of the table entry is stored in multiple LM groups.
[0012] Furthermore, the firmware microcore FW also includes a policy adjustment module, which includes: A judgment module, used for judging whether it is necessary to adjust the table item storage strategy in the LM or in the LM group according to the switching condition of the table item storage strategy; A switching module, used to execute a switching operation corresponding to a switching condition of a table entry storage policy; The switching condition of the entry storage strategy includes a first switching condition and a second switching condition; The first switching condition is a condition for switching from the latency priority strategy to the capacity priority strategy: the total number of table entries exceeds a preset first threshold; the first switching condition corresponds to a first switching operation, and the first switching operation is used to clear redundant table entry copies from multiple LMs or multiple LM groups and switch the storage strategy to the capacity priority strategy; The second switching condition, i.e., the condition for switching from the capacity priority strategy to the delay priority strategy, is that: the total number of entries in all LMs or LM groups is lower than a preset second threshold and there is no mapping position conflict for any entry in multiple LMs or LM groups; the second switching condition corresponds to a second switching operation, which is used to copy a single entry copy in a single LM or a single LM group to other LMs or LM groups and switch the storage strategy to the delay priority strategy; The first threshold and the second threshold are a watermark value lower than the total storage space of the on-chip network table entry divided by the maximum number of replicas N under the latency priority strategy.
[0013] Furthermore, the chip also includes a forwarding layer network on chip, each router in the forwarding layer network on chip is connected to a forwarding core, and the forwarding core is used to process and forward messages; The scale of the forwarding layer network-on-chip and the control layer network-on-chip is the same, and the forwarding core has a mapping relationship with the LM or LM group; When a table lookup is required, the forwarding core extracts the table entry keyword from the message, and selects the LM or LM group closest to the forwarding core to process the table entry lookup request based on the table entry storage strategy currently used by the control layer on-chip network and the hash map storage method of the table entry.
[0014] Based on another aspect of the embodiment of the present invention, the present invention further provides a table entry storage and search method, which is applied to a data communication chip including a control layer on-chip network, wherein each router in the control layer on-chip network is connected with a firmware micro-core FW and a table lookup unit LM, and the method comprises: The firmware microcore FW manages the table lookup unit LM; when receiving the table entry sent by the upper layer, the table entry is distributedly stored in the table lookup unit LM; When the table lookup unit LM receives the management instruction issued by the firmware microcore, it performs management operations on the table entries according to the management instruction; When the table lookup unit LM receives a table entry lookup request sent by the forwarding layer, it performs a table entry lookup and returns a lookup result to the forwarding layer.
[0015] Further, the firmware micro-core adopts a capacity priority strategy to distribute and store the table entries into the table lookup unit LM mounted on the control layer on-chip network; or the management module of the firmware micro-core adopts a latency priority strategy to distribute and store the table entries into the table lookup unit LM mounted on the control layer on-chip network; Under the capacity priority strategy, only one copy is stored for each table entry; Under the latency priority strategy, each entry stores multiple copies.
[0016] Further, the table lookup unit LM mounted on the control layer on-chip network is divided into a plurality of LM groups; Under the latency priority strategy, each LM group stores a copy of the table entry; Under the capacity priority strategy, only one copy of the table entry is stored in multiple LM groups.
[0017] Furthermore, the method for the firmware microcore FW to manage the table lookup unit LM also includes: Determine whether it is necessary to adjust the storage policy of the entry in the LM or in the LM group according to the switching condition of the entry storage policy; When it is determined that the switching condition of the table entry storage policy is met, a switching operation corresponding to the switching condition of the table entry storage policy is performed; The switching condition of the entry storage strategy includes a first switching condition and a second switching condition; The first switching condition is a condition for switching from the latency priority strategy to the capacity priority strategy: the total number of table entries exceeds a preset first threshold; the first switching condition corresponds to a first switching operation, and the first switching operation is used to clear redundant table entry copies from multiple LMs or multiple LM groups and switch the storage strategy to the capacity priority strategy; The second switching condition, i.e., the condition for switching from the capacity priority strategy to the delay priority strategy, is that: the total number of entries in all LMs or LM groups is lower than a preset second threshold and there is no mapping position conflict for any entry in multiple LMs or LM groups; the second switching condition corresponds to a second switching operation, which is used to copy a single entry copy in a single LM or a single LM group to other LMs or LM groups and switch the storage strategy to the delay priority strategy; The first threshold and the second threshold are a watermark value lower than the total storage space of the on-chip network table entry divided by the maximum number of replicas N under the latency priority strategy.
[0018] Based on the embodiment of the present invention, the present invention further provides a network device, which includes the on-chip network many-core architecture chip provided in the embodiment of the present invention.
[0019] The technical solution provided by the embodiment of the present invention has the following beneficial effects: Based on the architecture of the on-chip network-attached firmware microcore FW and the table lookup unit LM, the table entries are distributed and stored in multiple LMs. When the forwarding layer needs to look up the table, multiple LMs can provide table entry lookup capabilities in parallel, thereby improving the table entry lookup and message forwarding performance of the data communication chip. In addition, by adaptively and dynamically adjusting the table entry storage strategy, the chip can reduce the table lookup latency through the multi-copy mechanism when there are fewer table entries, and increase the table entry capacity by expanding the table entry storage space when there are more table entries, so that the chip can adapt to the network status more intelligently and flexibly. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments of the present invention or the description of the prior art will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings of the embodiments of the present invention.
[0021] Figure 1 A schematic diagram of a multi-core network on chip architecture provided by an embodiment of the present invention; Figure 2 A schematic diagram of location identification-based addressing routing in one embodiment of the present invention; Figure 3 It is a schematic diagram of distributed storage of table item data in LM under the capacity priority strategy in one embodiment of the present invention; Figure 4 It is a schematic diagram of distributed storage of table item data in LM under the delay priority strategy in one embodiment of the present invention; Figure 5 This is a schematic diagram of the control layer network-on-chip structure using the LM grouping mode according to an embodiment of the present invention; Figure 6 A schematic diagram of the structure of a data communication chip using a multi-layer network-on-chip many-core architecture according to an embodiment of the present invention; Figure 7 A schematic diagram of a flow chart of a table entry storage and search method in one embodiment of the present invention; Figure 8 A schematic diagram of the change state of table entries during the switching process from the delay priority strategy to the capacity priority strategy in one embodiment of the present invention; Fig. 9 A schematic diagram of the change state of table entries during the switching process from the capacity priority strategy to the delay priority strategy in one embodiment of the present invention; Fig.10 FIG. 4 is a schematic diagram of horizontally grouping the table lookup units LM in an embodiment of the present invention. DETAILED DESCRIPTION
[0022] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following exemplary embodiments do not represent all implementations consistent with this specification. Instead, they are merely examples of devices and methods consistent with some aspects of this specification as detailed in the appended claims.
[0023] The terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit this specification. The singular forms "a", "the" and "the" used in this specification and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.
[0024] It should be understood that although the terms first, second, third, etc. may be used in this specification to describe various information, this information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of this specification, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0025] In the field of data communication technology, the firmware microkernel can be understood as a lightweight, highly customized software core that runs on the hardware and is responsible for managing hardware resources, scheduling tasks, etc. By providing services such as the hardware abstraction layer, it ensures the efficient and orderly operation of various functions in the data communication equipment. From the perspective of the entire data communication system, the firmware microkernel, as firmware (Firmware), needs to consider the coordinated optimization with the hardware and upper-level software in its design and implementation. It not only needs to efficiently manage hardware resources, but also provide a simple and efficient interface for the upper-level software to achieve the optimization of the entire system in terms of performance, power consumption, cost, etc.
[0026] Network on Chip (NOC) is an on-chip communication architecture that uses data packets as the basic transmission unit, and uses on-chip network routers (NOC Routers) and communication links to achieve high-speed, parallel, and scalable communication between functional modules in a single chip by building a topology similar to that of a computer network. The on-chip network architecture is a technical means to improve performance by scaling out, which has good scalability. The NOC architecture adopts a distributed network topology, such as a two-dimensional matrix structure and a tree structure, and regards each functional module in the chip as a network node, which is connected through network links. Under the NOC architecture, different nodes can communicate in parallel. For example, in a two-dimensional matrix structure of a NOC, multiple port processing units can simultaneously transmit data to routing units in different directions, greatly improving the parallelism of data processing, thereby improving the overall forwarding performance.
[0027] In order to solve the problem of table entry search efficiency of data communication chips, the present invention proposes a network-on-chip many-core architecture chip and a table entry storage and search method implemented on the chip, so as to improve the table entry search performance and table entry management efficiency of the data communication chip in a scale-out manner.
[0028] Example 1 Figure 1 A schematic diagram of a multi-core network-on-chip architecture provided for an embodiment of the present invention. According to the functions of the firmware micro-core, the network-on-chip where the firmware micro-core / processor core that is mainly responsible for executing the control layer functions is located can be called the control layer network-on-chip, and the network-on-chip where the firmware micro-core / processor core that is mainly responsible for executing functions such as message processing and forwarding is located can be called the forwarding layer network-on-chip. Figure 1An example of a control layer on-chip network structure is provided, in which the on-chip network routers form a 2*2 matrix structure, and each router in the structure is connected with a firmware microkernel (Firmware, FW) and a lookup memory (Lookup Memory, LM). This embodiment connects multiple lookup memory units LM for storing table entry data through the on-chip network, so that the table entry data can be distributed and stored in multiple LMs, and each LM can independently perform the table entry search task. At the same time, in order to improve the management efficiency of the table entry data, a local firmware microkernel FW is added to each LM, and the table entry data in the LM is managed through the FW.
[0029] The firmware micro-core FW and the table lookup unit LM establish a network connection through the on-chip network, also forming a matrix structure. FW and LM can be connected to the on-chip network router (R) through the local port (LocalPort) respectively, and each FW can access any LM attached to the on-chip network through the on-chip network.
[0030] In order to realize the distributed routing access of FW to LM, the location identifier of the on-chip network router can be used as the address of FW and LM attached to the router, thereby solving the routing addressing problem between FW and LM. Figure 2 Take the example of a 4*4 scale on-chip network structure, with the position identifier {0, 0} of the router in the lower left corner as the origin, the position coordinate range of the router along the X-axis direction is 0~3, and the position coordinate along the Y-axis direction is also 0~3. Assuming that the FW attached to the router at position {0,1} wants to access the LM attached to the router at position {0,3}, then the FW at position {0,1} can use the position identifier {0,1} as its own on-chip network address and the position {0,3} as the on-chip network address of the destination LM. The message / data packet can be routed and forwarded to the destination LM through the deterministic routing algorithm or the adaptive routing algorithm. The on-chip network router performs routing and forwarding based on the source and destination addresses in the message / data packet.
[0031] The functions of FW and LM in this embodiment are described in detail below.
[0032] The firmware micro-core at least includes a management module, and the management module is used to manage the table lookup unit (LM).
[0033] The management module further includes an initialization module and a maintenance module; Initialization module, used to initialize LM; The maintenance module is used to maintain the table data in the LM. The maintenance includes but is not limited to the maintenance operations such as issuing, deleting, updating, moving, and counting table entries.
[0034] Table maintenance operations usually refer to operations performed on table entries by upper-layer applications or business modules as needed. For example, due to changes in network status, the upper-layer network protocol stack updates the routes and needs to refresh the routing table entries, causing the FW to perform operations such as deleting invalid table entries in LM and adding new table entries to LM.
[0035] When the FW receives the table entry (table entry entity Entry) sent from the upper layer, it will store the table entry in a distributed manner in each table lookup unit LM attached to the on-chip network. The distributed storage of table entries in the LM can be implemented by using a hash mapping storage method. For example, the table entry keyword Key is extracted from the field of the table entry, the Key is used as the input of the hash mapping function, and the hash value calculated by the hash function is used as the mapping storage location of the table entry data. Taking a 2*2 on-chip network matrix as an example, 2 bits can be extracted from a fixed position of the hash value of the Key to determine in which LM the table entry is stored, and a preset number of bits are extracted according to the capacity space size of the LM storage table entry to determine the mapping storage location of the table entry in the LM.
[0036] The table lookup unit at least includes: a storage module and a table lookup module.
[0037] The storage module is used for distributed storage of table items. When receiving the management instruction sent by the firmware micro-core, the storage module performs management operations on the table items according to the management instruction.
[0038] For example, receive the initialization instructions issued by the firmware micro-core, initialize the parameters, algorithms, table storage strategies, etc. used by the LM; receive the table maintenance instructions issued by the firmware micro-core, perform table maintenance operations, such as writing newly issued table entries, deleting, updating, moving, and counting existing table entries in the LM.
[0039] The table lookup module is used to receive a table entry lookup request sent by the forwarding layer, perform table entry lookup and return the lookup result to the forwarding layer.
[0040] In order to improve the management and forwarding efficiency of data communication chips, the layered concept can be adopted. According to the functions executed by the core, the frequency of operation and other factors, the chip can be divided into multiple layers in hardware. The control layer on-chip network is responsible for table management and control, and the forwarding layer on-chip network is responsible for message processing and forwarding. The control layer performs management and control operations less frequently, while the forwarding layer performs message processing and forwarding at a much higher frequency. Through layering, the management and search operations of table entries can be separated at the physical level to avoid resource competition between each other, which not only makes the architecture clearer, but also improves the overall performance of the chip.
[0041] The table lookup unit LM in the embodiment of the present invention is connected to both the control layer on-chip network and the forwarding layer on-chip network. The LM is managed by the control layer and provides the forwarding layer with a table entry lookup capability based on distributed storage.
[0042] Each LM in the embodiment of the present invention includes an SRAM controller, and each LM can provide independent table entry search capabilities to the forwarding layer. Assuming that the basic L2 / L3 layer forwarding of the message is completed, each message needs to look up the table 5 times, and the chip forwarding capacity is 3.2T. For a fixed-length 256B message (with a frame gap and a preamble of 276B), 3200G / 8 / 276*5=7.25G table lookups are required per second. If a 4*4 on-chip network is used, each router node is connected to an LM with an SRAM controller. A single LM only needs to provide a table entry search capability of about 460M times per second to meet the table lookup performance requirements (460M*16=7.36G), which can be easily achieved with a hardware frequency of 1G (assuming that an average of 2 clock cycles are used to complete one search). It can be seen that by distributing the processing load of table entry search to multiple SRAM controllers through the distributed storage of table entries, the table entry search performance of the data communication chip can be greatly improved.
[0043] Example 2 Figure 3 The diagram is a schematic diagram of the distributed storage of table item data in LM under the capacity priority strategy in one embodiment of the present invention. In this embodiment, a 2*2 scale on-chip network is used to hang 4 LMs, and the FW adopts a capacity priority table item storage strategy to store the table item data in the 4 LMs. Under the capacity priority strategy, each table item is stored in only one copy in all LMs hung on the on-chip network. The FW can extract 2 bits from a preset position in the hash value of the table item keyword Key as the first index to determine in which LM the table item is stored, and according to the capacity of the LM, extract a preset number of bits that match the capacity as the second index to determine the mapping storage location of the table item in the LM. The advantage of the capacity priority strategy is that it can maximize the use of the storage space of the LM, and the space utilization rate is high.
[0044] Although the distributed storage of table entry data through the capacity priority strategy can improve the overall table entry lookup throughput performance, as the scale of the on-chip network increases, the path of the forwarding core in the forwarding layer on-chip network when accessing the control layer LM will become longer, and the latency between the shortest path and the longest path will increase, which will affect the table entry lookup performance.
[0045] Example 3 Figure 4The schematic diagram of the distributed storage of table item data in LM under the latency priority strategy in one embodiment of the present invention. In this embodiment, a 2*2 scale on-chip network is adopted to hang 4 LMs, and the FW adopts a latency priority table item storage strategy to store the table item data in the 4 LMs. Under the latency priority strategy, each table item will have 1 copy in each of all LMs attached to the on-chip network, and there will be a total of 4 table item copies. Under this strategy, the FW can ignore the first index and directly store the table item mapping to each LM attached to the on-chip network according to the second index. The advantage of the latency priority strategy is that it can maximize the parallel processing performance of table item lookup. Since a copy of the table item is stored in each LM, the forwarding core of the forwarding layer can call the nearest LM to perform table item lookup, exchanging space for time, and maximizing the overall table lookup efficiency.
[0046] Under the latency priority strategy, Figure 4 The total effective table entry storage space of the example on-chip network is only 1 / 4 of the total table entry storage space (the sum of all LM storage spaces). As the scale of the on-chip network increases, the ratio of the total effective table entry storage space to the total table entry storage space will become smaller and smaller. For example, for a 4*4 on-chip network, the ratio is 1 / 16, which will result in low utilization efficiency of the storage space.
[0047] Example 4 In order to combine the advantages of the latency priority strategy and the capacity priority strategy, the FWs and LMs attached to the on-chip network are grouped. Figure 5 The schematic diagram of the control layer on-chip network structure using the LM grouping mode in one embodiment of the present invention is shown in FIG. The LMs connected to the on-chip network form a 4*4 matrix, which is divided into 4 LM groups according to the X-axis coordinate mark, and each LM group includes 4 LMs.
[0048] Under the latency priority strategy, table entries are stored in duplicate in LM groups. Under the capacity priority strategy, it is logically equivalent to merging four LM groups into a large group, merging the first index and the second index into the index of the large group, and mapping and storing table entries in the large group. Within the LM group, the table entry storage space of all LMs is uniformly and continuously addressed.
[0049] Assume that the table entry storage space of each LM is 1K, and an LM group includes 4 LMs, with a total of 4K table entry storage space, and 4 groups have a total of 16K. 14 bits can be extracted from the hash value (hash digest) of the table entry keyword Key, and 2 bits are selected as the first index to determine which LM group the table entry is stored in; the remaining 12 bits are used as the second index to determine the mapping storage location of the table entry in the LM group. The first index of 0 corresponds to the LM group with the router location identifier X coordinate of 0 (Replica-1), the first index of 1 corresponds to the LM group with the router location identifier X coordinate of 1 (Replica-2), and so on.
[0050] In the case of the capacity priority policy, each entry stores only one copy in four LM groups.
[0051] In the case of latency priority strategy, each table entry will store a copy in 4 LM groups, that is, the number of table entry copies is equal to the number of LM groups. At this time, when writing the table entry, the FW can ignore the first index and directly write the table entry data into the storage space of each LM group according to the second index.
[0052] When executing management and maintenance operations such as issuing and deleting table items, the FW firmware micro-core needs to correctly perform the corresponding management and maintenance operations according to the table item storage strategy used. For example, under the multi-copy storage strategy with latency priority, the table items need to be issued to all LM groups according to the second index; under the single-copy strategy with capacity priority, the first index and the second index need to be combined to determine the location where the table item mapping is stored in the LM group.
[0053] In one embodiment of the present invention, there are multiple types of tables used for message matching and lookup, such as MAC table, LPM table, INTF table, etc. Each firmware micro-core FW can only be responsible for the management of various table items in the local LM (LM connected to the same router as the FW). In the case of grouping LM, the FWs in the group can also be divided into different types of table items. Different firmware micro-core FWs can be configured to manage different types of table items in the LM group, and each FW manages one or more types of table items respectively.
[0054] Example 5 In order to further improve the adaptability and availability of the chip under different network conditions, the chip can intelligently and dynamically adjust the table entry storage strategy according to the number of table entries. A policy adjustment module can be added to the control layer on-chip network. The policy adjustment module dynamically adjusts the storage strategy of the table entries according to the current number of table entries.
[0055] To achieve the above objectives, the firmware microkernel FW in the control layer on-chip network also includes: A policy adjustment module, used to adjust the table item storage policy in the LM or the LM group according to the switching condition of the table item storage policy, and perform the corresponding switching operation, wherein the table item storage policy includes a delay priority policy and a capacity priority policy; The policy adjustment module further includes: A judgment module, used for judging whether it is necessary to adjust the table item storage strategy in the LM or in the LM group according to the switching condition of the table item storage strategy; A switching module, used to execute a switching operation corresponding to a switching condition of a table entry storage policy; The latency priority strategy refers to a table entry storage strategy that aims to reduce the latency of table entry lookup. This strategy requires that table entries be stored in multiple copies in multiple LMs or LM groups to improve the concurrent search efficiency of table entry lookups.
[0056] The capacity priority strategy is a table entry storage strategy that aims to increase the number of table entry storage. This strategy requires table entries to be stored in a single copy in multiple LMs or LM groups to increase the storage capacity of table entries. This storage strategy can ensure the availability of network devices as much as possible when there are many concurrent network connections, but this strategy will increase the table entry search latency to a certain extent and reduce the table entry search efficiency.
[0057] The switching condition of the entry storage policy includes a first switching condition and a second switching condition.
[0058] The first switching condition is the condition for switching from the latency priority strategy to the capacity priority strategy, which can be configured as: the total number of entries exceeds the preset first threshold. The first threshold can be configured as a watermark value lower than the total entry storage space divided by the maximum number of replicas N under the latency priority strategy. The total entry storage space is the sum of the entry storage spaces of all LMs.
[0059] For example, a 2*2 on-chip network is connected to 4 LMs, and the table entry storage space of each LM is 1K. The total table entry storage space is the sum of the table entry storage space of the 4 LMs, that is, 4K. The maximum number of copies N under the latency priority strategy is 4. Each table entry is stored in 4 LMs. The total effective storage space of the total table entry is the total table entry storage space divided by the maximum number of copies 4 under the latency priority strategy, which is equal to 1K. The first threshold can be set to a watermark value that is less than or equal to the total effective storage space of the table entry, such as 98%, or it can be set to 100%.
[0060] When the FW determines that the first switching condition is currently met, the switching module performs a first switching operation. The purpose of the first switching operation is to remove redundant copies from multiple LMs or multiple LM groups and leave only one copy.
[0061] The second switching condition, i.e., the condition for switching from the capacity priority strategy to the delay priority strategy, can be configured as: the total number of entries in all LMs or LM groups is lower than a preset second threshold and there is no mapping position conflict for any entry in multiple LMs or LM groups.
[0062] The mapping position conflict means that, in all LMs or LM groups attached to the control layer on-chip network, there are more than one valid table entry at the mapping storage position indicated by the second index in the hash value of the table entry keyword Key.
[0063] The second threshold may also be set to a watermark value that is less than or equal to the total entry storage space divided by the maximum number of replicas N under the latency priority strategy. The second threshold may be the same as or different from the first threshold.
[0064] When the FW determines that the second switching condition is currently met, the switching module performs the second switching operation. The purpose of the second switching operation is to switch the table entry storage from a single copy mode to a multiple copy mode, and copy the table entries stored in a single LM or a single LM group to other LMs or other LM groups to achieve multiple copy storage. In some scenarios that are particularly sensitive to message forwarding delays, table lookup delays can be reduced by duplicating table entries.
[0065] Example 6 Figure 6 The schematic diagram of the structure of a data communication chip using a multi-layer network-on-chip multi-core architecture in one embodiment of the present invention. In this embodiment, the data communication chip uses a multi-layer network-on-chip architecture, including a control layer network-on-chip and a forwarding layer network-on-chip. The scale of the router matrix of the control layer network-on-chip can be the same as that of the router matrix of the forwarding layer network-on-chip. The location identifiers of the routers of the control layer network-on-chip and the forwarding layer network-on-chip have a corresponding relationship and both use the same location coordinates.
[0066] Each router node in the forwarding layer on-chip network is connected to the forwarding core (FC) through a local interface (LocalPort). The forwarding core can also have programmable features like a firmware microcore, support multiple types of business features, and perform various processing and forwarding on messages.
[0067] The forwarding core in the forwarding layer network-on-chip has a mapping relationship with the firmware micro-core FW and the table lookup unit LM in the control layer network-on-chip. The forwarding core, firmware micro-core FW and table lookup unit LM in the same position can be interconnected through a local interface. The forwarding core in the forwarding layer network-on-chip can also access any firmware micro-core and table lookup unit in the control layer network-on-chip through the interconnection mechanism between the networks-on-chip.
[0068] During the process of processing and forwarding the message, when the forwarding core in the forwarding layer network-on-chip needs to perform a table lookup operation, it can extract the table entry keyword Key from the message to form a table entry lookup request, and then send the table entry lookup request to the LM or LM group storing the table entry copy in the control layer network-on-chip to perform the table entry lookup operation, thereby obtaining the table entry lookup result.
[0069] The control layer on-chip network and the forwarding layer on-chip network need to synchronize the table entry storage strategy, hash mapping method, LM group information, etc. The forwarding core in the forwarding layer on-chip network selects the LM or LM group where the nearest table entry copy is located to perform table entry lookup based on the synchronized information.
[0070] like Figure 6 For example, a mapping correspondence relationship can be established between the LM groups of the control layer network-on-chip and the forwarding core groups of the forwarding layer network-on-chip. For example, both layers of the network-on-chip are grouped according to the X-coordinate of the router, and the groups with the same X-coordinate form one group. The control layer network-on-chip divides the four LMs with X-axis position coordinates of 0, 1, 2, and 3 into four LM groups, namely Replica-1, Replica-2, Replica-3, and Replica-4. The forwarding layer also divides the four forwarding cores with X-axis position coordinates of 0, 1, 2, and 3 into four LM groups, namely Group-1, Group-2, Group-3, and Group-4. A mapping correspondence relationship is established between Group-1 and Replica-1, a mapping correspondence relationship is established between Group-2 and Replica-2, and so on. It should be noted that the present invention does not limit the direction of dividing the LM groups. The division can be along the X-axis or along the Y-axis. The division method of the forwarding core groups is consistent with the division method of the LM groups.
[0071] When the forwarding core in Group-1 needs to look up the table, if the table entry storage policy is the latency priority policy, the forwarding core in Group-1 will select the LM in Replica-1 that is closest to it physically (through how many hops of NOC routers) to process the table lookup request based on the mapping relationship. If the table entry storage policy is the capacity priority policy, assuming that the required table entry is mapped and stored in the LM in the Replica-3 group according to the hash mapping algorithm, the table lookup request will be sent to the LM in the LM group of Replica-3 through the on-chip network for table lookup.
[0072] By allocating the control layer functions and forwarding layer functions to different on-chip networks, the management and control operations for table entries are separated from the processing and forwarding operations of messages, which can reduce the mutual impact between different types of operations and thus improve the processing performance of the entire chip.
[0073] Example 7 Figure 7 The flowchart of the table entry storage and search method in one embodiment of the present invention is shown in FIG. The method is applied to a data communication chip including the control layer on-chip network as described above, and each router in the control layer on-chip network is connected with a firmware micro-core FW and a table lookup unit. The method includes: Step 701: the firmware microcore FW manages the table lookup unit LM; when receiving a table entry sent from an upper layer, the table entry is distributedly stored in the table lookup unit LM.
[0074] Step 702: When the table lookup unit LM receives the management instruction sent by the firmware microcore, it performs a management operation on the table entry according to the management instruction.
[0075] Step 703: When receiving the table entry search request sent by the forwarding layer, the LM performs a table entry search and returns the search result to the forwarding layer.
[0076] Since table entries are stored in a distributed storage manner, when multiple LMs store a portion of table entries respectively, the forwarding layer can send table entry search requests to multiple LMs to process the table entry search requests in parallel, thereby improving table search efficiency.
[0077] Furthermore, the table entry search request is sent by the forwarding core in the forwarding layer network on chip. When the message / data packet arrives at the forwarding core, the forwarding core can first perform preliminary processing on the message and extract the keyword Key used for table entry search. For example, for IP messages, the destination IP address can be extracted as the search keyword; for Ethernet messages, the destination MAC address can be extracted.
[0078] The forwarding core sends the table entry keyword and some other information (such as table entry type, etc.) to the LM as a table entry search request, triggering the LM to perform the table entry search operation. After receiving the table entry search request sent by the forwarding core, the LM quickly starts the search logic. According to the table entry type and the configured search algorithm, it quickly searches in the table entries stored in it. For example, if it is a longest prefix match search, it will use a specific data structure (such as a Trie tree structure) to efficiently search for the table entry that has the longest match with the destination IP address. After obtaining the search result, the LM returns the search result to the forwarding core of the forwarding layer.
[0079] Example 8 In this embodiment, the on-chip network adopts a 4*2 (4 rows and 2 columns, X coordinates [0,1], Y coordinates [0,1,2,3]) scale architecture. The LM matrix is divided into two LM groups vertically, that is, 4 LMs with X-axis position coordinates of 0 are one LM group corresponding to Replica-1, and 4 LMs with X-axis position coordinates of 1 are one LM group corresponding to Replica-2. The LM storage space in the LM group is uniformly addressed, and the table entries are stored in the LM group. The single copy mode and the dual copy mode are preset. The single copy mode corresponds to the capacity priority strategy, and the dual copy mode corresponds to the latency priority strategy.
[0080] When the system is initialized, the firmware microkernel in the control layer on-chip network adopts the latency-priority multi-replica strategy by default. When the FW writes an entry to the LM group, it writes the entry to both Replica-1 and Replica-2 LM groups.
[0081] In the multi-copy mode of the latency priority policy, as the number of entries written by the FW to the LM group increases, when the conditions for switching from the latency priority policy to the capacity priority policy (the first switching condition) are met, the corresponding switching operation (the first switching operation) is performed.
[0082] Figure 8 The figure is a schematic diagram of the change state of table entries during the switching process from the delay priority strategy to the capacity priority strategy in one embodiment of the present invention.
[0083] Step 801: before publishing an entry, the firmware microcore FW determines whether a condition (a first switching condition) for switching from a latency priority strategy to a capacity priority strategy is currently met. If the condition is met, a corresponding switching operation (a first switching operation) is triggered. Figure 8 Part (a) in the figure illustrates the storage status of entries in two LM groups in dual-copy mode, and entries Entry1 to Entry8 have two copies in both LM groups.
[0084] The FW may update the entry counter each time an entry is issued, deleted, or invalidated, and trigger the first switching operation when the total number of entries in the LM group exceeds a preset first threshold.
[0085] In this embodiment, the on-chip network is connected to a total of 8 LMs (4*2), two LM groups, and the maximum number of copies in the latency priority mode is 2. Each LM group includes 4 LMs. Assuming that each LM has 1K of table entry storage space, the total table entry storage space of the on-chip network is 8K. The first threshold can be set to a watermark value that is less than or equal to the total table entry storage space of the on-chip network / the maximum number of copies.
[0086] Step 802: Execute a first switching operation, by which redundant table entry copies are cleared from a plurality of LM groups and a current table entry storage strategy is switched to a capacity priority strategy; The FW that triggers the switching operation notifies all FWs in the control layer's on-chip network. Each FW starts scanning the table entries in the local LM (the LM connected to the same on-chip network router as the FW). During the scanning process, each time a table entry is scanned, the following operations are performed: Step 8021: extract the table entry keyword Key from the table entry and recalculate the hash value (hash digest) of the key; Step 8022: extract a first index from a preset position in the hash value of the Key, and determine the LM group where the entry copy is located according to the first index; The hash value of the table entry keyword Key is used to determine the mapping storage location of the table entry. In this embodiment, there are 8 LMs. Assuming that each LM can store 1K table entries (ENTRY), the total addressing space of the 8 LMs requires 13 bits, corresponding to 8K table entry storage space. Accordingly, 13 bits can be extracted from the hash value of the table entry Key at a fixed preset position, 1 bit of which is used as the first index, and the other 12 bits are used as the second index. The second index is used to determine the mapping storage location of the table entry in the LM group, corresponding to the 4K table entry storage space with continuous addressing in the LM group. The first index is used to determine in which LM group the table entry mapping is stored.
[0087] Step 8023: According to the value of the chip select bit, the table entries belonging to the LM group are retained, and the table entries not belonging to the group are cleared.
[0088] In this embodiment, the four LMs with an X-axis position coordinate of 0 form a Replica-1 group, and the LM group corresponds to a first index value of 0; the four LMs with an X-axis position of 1 form a Replica-2 group, and the LM group corresponds to a first index value of 1. For the FW in Replica-1, when a table entry is scanned, if the first index value in the hash value of the key of the table entry is 0, it means that the table entry belongs to the LM group, and the table entry is retained in the LM group; if the first index value in the hash value of the key of the table entry is 1, it means that the table entry does not belong to Replica-1, and the table entry is invalidated or deleted. The FW in the Replica-2 group also performs the same operation.
[0089] Through the above scanning process, only one copy of the table entry is retained in multiple LM groups, thereby clearing redundant copies of the table entry.
[0090] Step 8024: After all redundant table entry copies are cleared, the current table entry storage policy is switched to a capacity priority policy.
[0091] After completing the switch from the latency-first multi-copy strategy to the capacity-first single-copy strategy, it is equivalent to merging two LM groups into a larger LM group. The total table storage space of the control layer on-chip network is theoretically expanded by N times, where N is the maximum number of copies. Figure 8 Example of part (b) in .
[0092] Step 803: When the subsequent FW sends new entries, it sends the entries according to the capacity priority strategy.
[0093] Under the capacity priority strategy, the steps for mapping and storing the hash value of the table entry key are as follows: Step 8031: extract 1 bit from a fixed preset position in the hash value of the table entry keyword Key as the first index to determine the LM group to which the table entry belongs.
[0094] Step 8032: extract 12 bits from a preset position in the hash value of the table entry keyword Key as the second index, and determine the mapping storage position of the table entry in the LM group indicated by the first index.
[0095] Step 8033: Store the table entry to the corresponding storage location in the corresponding LM group based on the first index and the second index, and only one copy of each table entry will be sent down.
[0096] like Figure 8 In the example of part (c) in FIG, under the capacity priority policy, the newly issued entries Entry9, Entry10, and Entry11 are mapped and stored in the storage space vacated in the LM group, so that each group can store more entries.
[0097] Fig. 9 The following is a schematic diagram of the table entry change state during the switching process from the capacity priority strategy to the latency priority strategy in one embodiment of the present invention. Under the capacity priority strategy, when the network load decreases, many table entries may be deleted or invalidated. As the number of table entries in the LM group decreases, when the conditions for switching from the capacity priority multi-copy strategy to the latency priority multi-copy strategy are met, the firmware micro-core FW in the control layer on-chip network can trigger the switching operation. The switching process is as follows: Step 901: After deleting the table entry in the LM group or periodically detecting, the FW determines whether the condition for switching from the capacity priority strategy to the delay priority strategy (the second switching condition) is met. If the condition is met, the corresponding switching operation (the second switching operation) is triggered. Fig. 9Part (a) in the figure illustrates the storage status of entries in two LM groups in single copy mode. There is only one copy of entries Entry1 to Entry8 in both LM groups and there is no location conflict.
[0098] The position conflict means that in multiple LM groups, there are multiple valid entries at the mapping storage location indicated by the second index in the hash value of the entry keyword Key. If the multi-copy mode with the latency priority strategy is switched, it will cause a mapping position conflict. Figure 8 For example, in part (c) of the example, there are two valid entries, Entry9 and Entry3, at the third storage position in the two LM groups. In this case, if the entry is copied and the policy is switched, a conflict will occur, causing the entry to be overwritten. Of course, in the case of position conflict, you can also consider shifting the conflicting entry to a non-conflicting position to solve the position conflict problem.
[0099] The second switching condition in this step is: the total number of entries in all LM groups is lower than the preset second threshold and there is no mapping position conflict for any entry in multiple LM groups; the second threshold can be set to a watermark value that is less than or equal to the total storage space of the on-chip network entry / the maximum number of replicas under the policy priority policy. Assuming that the total storage space of Replica-1 and Replica-2 is 8K, and the maximum number of replicas in the multi-replica mode is 2, the second threshold can be set to 4K. When the total number of entries in the two LM groups is less than or equal to 4K and there is no position conflict for any entry in multiple LM groups, it can be determined that the second switching condition is currently met.
[0100] Step 902: Execute the second switching operation. By executing the second switching operation, each table entry is copied to other LM groups, so as to realize the conversion of the table entry from a single copy to multiple copies, and switch the storage strategy to a latency priority strategy.
[0101] The FW that triggers the switching operation will notify all FWs in the control layer on-chip network. Each FW will start the scanning process for the table entries in the local LM. During the scanning process, whenever a table entry is scanned, the scanned table entry will be copied to other LM groups.
[0102] like Fig. 9 In the example of part (b) in FIG, for each table entry currently existing in the local LM unit, if the current X position coordinate of the node is 0, the FW copies the table entry to the corresponding LM unit in the other LM group with the corresponding X position coordinate of 1. Similarly, if the X position coordinate of the local LM is 1, the table entry is copied to the corresponding LM in the other LM group with the corresponding X position coordinate of 0. After all FW scans are completed, each table entry has a copy stored in two LM groups.
[0103] Step 903: When the subsequent FW sends a new table entry, it sends the table entry according to the latency priority strategy.
[0104] Under the latency priority strategy, the steps for mapping and storing the hash value of the table entry key are as follows: Step 9031: extract 12 bits from a preset position in the hash value of the table entry keyword Key as the second index to determine the mapping storage position of the table entry in the LM group.
[0105] Under the latency priority strategy, the next entry does not need to extract the first index to determine the LM group, because each entry will be stored in all LM groups, and only needs to be stored according to the second index mapping.
[0106] Step 9032: Store the table entry into the corresponding hash map storage location in each LM group based on the second index.
[0107] like Fig. 9 In the example of part (c) in FIG, under the delay priority strategy, the newly issued table entries will be mapped and stored in all LM packets, so that the forwarding core of the forwarding layer can select the nearest LM packet for table entry lookup, thereby reducing the delay and improving the efficiency of table entry lookup.
[0108] The present invention does not limit the matrix scale of the on-chip network, for example, it can also be a 4*4, 8*8, 16*16 scale on-chip network. The number of LM groups can be flexibly set according to the scale of the on-chip network and the needs of actual applications, for example, it can be 2 groups, 4 groups, 8 groups, 16 groups, etc. In addition, the present invention does not limit the specific way of dividing the groups, for example, Fig.10 As shown, the LM groups can also be divided according to the Y-axis position coordinates. The grouping method of the forwarding cores in the forwarding layer network on chip should be consistent with the LM grouping method in the control layer network on chip, which will not be repeated here.
[0109] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0110] Those skilled in the art will readily appreciate other embodiments of the present specification after considering the specification and practicing the inventions claimed herein. The present specification is intended to cover any variations, uses or adaptations of the present specification that follow the general principles of the present specification and include common knowledge or customary technical means in the art that are not claimed in the present specification.
[0111] The above description is only a preferred embodiment of this specification and is not intended to limit this specification. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of this specification should be included in the scope of protection of this specification.
Claims
1. A chip with a multi-core network architecture, characterized in that: The chip includes a control layer on-chip network, and each router of the control layer on-chip network is connected with a firmware micro-core FW and a table lookup unit LM; The firmware microcore FW includes a management module, which is used to manage the table lookup unit LM. When receiving the table entry sent by the upper layer, the table entry is distributedly stored in the table lookup unit LM; The table lookup unit LM comprises: The storage module is used for distributed storage of table items, and when receiving the management instruction issued by the firmware micro-core, performs management operations on the table items according to the management instruction; The table lookup module is used to receive a table entry lookup request sent by the forwarding layer, perform table entry lookup and return the lookup result to the forwarding layer.
2. The chip according to claim 1, characterized in that: The management module of the firmware micro-core adopts a capacity priority strategy to distribute and store the table entry data in the table lookup unit LM mounted on the control layer on-chip network; under the capacity priority strategy, only one copy is stored for each table entry; or, The management module of the firmware micro-core adopts a latency priority strategy to distribute and store the table entry data in the table lookup unit LM mounted on the control layer on-chip network; under the latency priority strategy, each table entry stores multiple copies.
3. The chip according to claim 1, characterized in that: The table lookup unit LM attached to the control layer on-chip network is divided into a plurality of LM groups; Under the latency priority strategy, each LM group stores a copy of the table entry; Under the capacity priority strategy, only one copy of the table entry is stored in multiple LM groups.
4. The chip according to claim 2 or 3, characterized in that: The firmware microcore FW also includes a policy adjustment module, which includes: A judgment module, used for judging whether it is necessary to adjust the table item storage strategy in the LM or in the LM group according to the switching condition of the table item storage strategy; A switching module, used to execute a switching operation corresponding to a switching condition of a table entry storage policy; The switching condition of the entry storage strategy includes a first switching condition and a second switching condition; The first switching condition is a condition for switching from the latency priority strategy to the capacity priority strategy: the total number of table entries exceeds a preset first threshold; the first switching condition corresponds to a first switching operation, and the first switching operation is used to clear redundant table entry copies from multiple LMs or multiple LM groups and switch the storage strategy to the capacity priority strategy; The second switching condition, i.e., the condition for switching from the capacity priority strategy to the delay priority strategy, is that: the total number of entries in all LMs or LM groups is lower than a preset second threshold and there is no mapping position conflict for any entry in multiple LMs or LM groups; the second switching condition corresponds to a second switching operation, which is used to copy a single entry copy in a single LM or a single LM group to other LMs or LM groups and switch the storage strategy to the delay priority strategy; The first threshold and the second threshold are a watermark value lower than the total storage space of the on-chip network table entry divided by the maximum number of replicas N under the latency priority strategy.
5. The chip according to claim 4, characterized in that: The chip also includes a forwarding layer network on chip, each router in the forwarding layer network on chip is connected to a forwarding core, and the forwarding core is used to process and forward messages; The scale of the forwarding layer network-on-chip and the control layer network-on-chip is the same, and the forwarding core has a mapping correspondence with the LM or LM group; When a table lookup is required, the forwarding core extracts the table entry keyword from the message, and selects the LM or LM group closest to the forwarding core to process the table entry lookup request based on the table entry storage strategy currently used by the control layer on-chip network and the hash map storage method of the table entry.
6. A table entry storage and search method, characterized in that: The method is applied to a data communication chip including a control layer on-chip network, each router in the control layer on-chip network is connected with a firmware micro-core FW and a table lookup unit LM, and the method comprises: The firmware microcore FW manages the table lookup unit LM; when receiving the table entry sent by the upper layer, the table entry is distributedly stored in the table lookup unit LM; When the table lookup unit LM receives the management instruction issued by the firmware microcore, it performs management operations on the table entries according to the management instruction; When the table lookup unit LM receives a table entry lookup request sent by the forwarding layer, it performs a table entry lookup and returns a lookup result to the forwarding layer.
7. The method according to claim 6, characterized in that The firmware micro-core adopts a capacity priority strategy to distribute and store the table entries into the table lookup unit LM mounted on the control layer on-chip network; or the management module of the firmware micro-core adopts a delay priority strategy to distribute and store the table entries into the table lookup unit LM mounted on the control layer on-chip network; Under the capacity priority strategy, only one copy is stored for each table entry; Under the latency priority strategy, each entry stores multiple copies.
8. The method according to claim 7, characterized in that The table lookup unit LM attached to the control layer on-chip network is divided into a plurality of LM groups; Under the latency priority strategy, each LM group stores a copy of the table entry; Under the capacity priority strategy, only one copy of the table entry is stored in multiple LM groups.
9. The method according to claim 7 or 8, characterized in that: The method for the firmware microcore FW to manage the table lookup unit LM also includes: Determine whether it is necessary to adjust the storage policy of the entry in the LM or in the LM group according to the switching condition of the entry storage policy; When it is determined that the switching condition of the table entry storage policy is met, a switching operation corresponding to the switching condition of the table entry storage policy is performed; The switching condition of the entry storage strategy includes a first switching condition and a second switching condition; The first switching condition is a condition for switching from the latency priority strategy to the capacity priority strategy: the total number of table entries exceeds a preset first threshold; the first switching condition corresponds to a first switching operation, and the first switching operation is used to clear redundant table entry copies from multiple LMs or multiple LM groups and switch the storage strategy to the capacity priority strategy; The second switching condition, i.e., the condition for switching from the capacity priority strategy to the delay priority strategy, is that: the total number of entries in all LMs or LM groups is lower than a preset second threshold and there is no mapping position conflict for any entry in multiple LMs or LM groups; the second switching condition corresponds to a second switching operation, which is used to copy a single entry copy in a single LM or a single LM group to other LMs or LM groups and switch the storage strategy to the delay priority strategy; The first threshold and the second threshold are a watermark value lower than the total storage space of the on-chip network table entry divided by the maximum number of replicas N under the latency priority strategy.
10. A network device, characterized in that: The network device comprises the chip according to any one of claims 1-5.
Citation Information
Patent Citations
Tightly coupled self-adaptive co-processing system supporting multi-core network processing framework
CN104503948A
Distributed storage unit-based hierarchical network on chip architecture
CN102075578A
Routing table item storage method and device and routing table item search method and device
CN113992579A
Processor core selection method and device, electronic equipment and storage medium
CN119621652A
Many-core cache consistency system and method, electronic equipment, storage medium and product
CN119669109A