Network-on-chip many-core architecture and data processing method thereof
By virtualizing multiple physical firmware microcores into one logical firmware microcore and dividing task groups to reduce resource requirements, the memory resource limitation problem of on-chip network multi-core architecture when supporting multiple business features is solved, achieving lower power consumption and higher processing performance.
Patent Information
- Application Number
- CN202510426580.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-07
AI Technical Summary
The data communication chips of the on-chip network multi-core architecture have limited the ability to support different business characteristics, resulting in their disadvantages in handling complex network tasks.
By virtualizing multiple physical firmware microcores into one logical firmware microcore and dividing business features into multiple task groups, each physical firmware microcore runs instruction codes in one task group, thereby achieving support for multiple business features while reducing the need for IRAM and DRAM resources.
While supporting the same number of business characteristics, it effectively reduces the demand for physical firmware micro-checking to build-in IRAM and DRAM resources, reduces the power consumption and cost of the chip, and improves the chip's business scalability and processing performance.
Smart Images

Figure CN119938588A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data communication technology, and in particular to a network-on-chip many-core architecture and a data processing method thereof. Background Art
[0002] How to achieve higher forwarding performance at a lower cost (smaller chip area, lower power consumption) is a persistent problem faced by data communication chips used in data communication equipment such as routers and switches. With the accelerated development of network technology, the processing performance requirements for data communication chips are getting higher and higher. Simply increasing the hardware clock frequency to improve chip performance in a scale-up manner is no longer able to keep up with the requirements of network applications.
[0003] The Network on Chip (NOC) architecture, with its good scalability, has become a technical means for chip designers to improve the performance of data communication chips by scaling out when it is difficult to continuously improve hardware performance by scaling up.
[0004] The application document with application number 201510507293.7, application publication date 2015.12.30, publication number CN105207957A, and titled "A Network-on-Chip Multi-core Architecture" is the closest prior art to the present invention. The prior art discloses a network-on-chip multi-core architecture, including a network-on-chip multi-core architecture body, which includes multiple computing units, a router, and a network interface. The multiple computing units are connected through routers and network interfaces to realize parallel data processing and data interaction, one of the computing units is used as a master control core node, and the remaining computing units are used as operation core nodes. The master control core node is responsible for data exchange with the outside of the chip, and the operation core node transmits data to the master control core node, which completes the data exchange with the outside of the chip; the storage space in the multiple computing units adopts a unified addressing so that the core in each computing unit can access the storage space in any other computing unit.
[0005] The above-mentioned closest prior art has the following technical defects: In order to control the power consumption and cost of the entire chip, the computing unit / firmware micro-core in the NOC architecture only integrates a small amount of IRAM and DRAM. The limited memory resources have greatly restricted the ability of data communication chips based on the on-chip network multi-core architecture to support different business features. Summary of the invention
[0006] In view of this, the present invention provides a network-on-chip many-core architecture and a data processing method thereof, which are used to solve the technical problem that the ability of the network-on-chip many-core architecture to support business characteristics is limited.
[0007] Based on one aspect of an embodiment of the present invention, the present invention provides a network-on-chip many-core architecture, including: A logical firmware micro-core, including a plurality of physical firmware micro-cores (Firmware micro-cores, FW), wherein the plurality of physical firmware micro-cores respectively support service characteristics in different task groups (task_group), a task group includes one or more service characteristics, and different task groups include different service characteristics; A logical multi-homed task engine, including a plurality of physical task engines (TEs), wherein the physical task engines have hardware-hardened basic packet processing (PP) actions for hardware acceleration of message processing of service characteristics orchestrated by the physical firmware micro-core; each physical task engine is controlled by each physical firmware micro-core in the logical firmware micro-core; In the on-chip network, the physical firmware micro-kernel and the physical task engine are connected to the on-chip network router (NOCRouter) through a local port (LocalPort) of the on-chip network router.
[0008] Further, the total amount of instruction codes and the total amount of data random access memory required for the service characteristics in the task group does not exceed the instruction random access memory (IRAM) capacity and data random access memory (DRAM) capacity in the physical firmware microcore; The number of physical firmware microcores included in the logical firmware microcore is determined by the total amount of instruction random access memory and data random access memory required for the business characteristics that the logical firmware microcore needs to support and the instruction random access memory capacity and data random access memory capacity of the physical firmware microcore.
[0009] Further, each network-on-chip router (NOC Router) in the network-on-chip is connected to a physical firmware micro-kernel and a physical task engine through a local port; The logical firmware micro-core and the logical multi-homed task engine are respectively composed of multiple pairs of physical firmware micro-cores and physical task engines connected to the same on-chip network router; Each physical task engine in the logical multi-homed task engine is fully connected to each physical firmware microcore in the logical firmware microcore through an on-chip network.
[0010] Furthermore, the physical firmware micro-kernel in the logical firmware micro-kernel adopts a pipeline model to transmit commands to the physical task engine in the logical multi-homing task engine.
[0011] Furthermore, the on-chip network multi-core architecture provides users with an orchestration configuration interface, through which the message processing processes of different business characteristics in the task group are orchestrated, and several basic hardware-cured basic packet processing (PP) actions are orchestrated and combined to support the business characteristics in the task group; and the business characteristics supported by each physical firmware microcore in the logical firmware microcore are configured (for example, adding, deleting, modifying, etc.).
[0012] Furthermore, the physical firmware microkernel in the logical firmware microkernel is a group of adjacent firmware microkernels in the on-chip network. The principle of proximity is adopted when configuring the logical firmware microkernel, and a group of adjacent physical firmware microkernels are configured as one logical firmware microkernel.
[0013] Based on another aspect of the embodiment of the present invention, the present invention further provides a data processing method, which is applied to a chip using the on-chip network multi-core architecture provided by the present invention, and the method includes: When a service request or service message of a certain type of service characteristic is received, the service request or service message is routed to a physical firmware microcore bound to the type of service characteristic in a logical firmware microcore supporting the type of service characteristic in the on-chip network many-core architecture; The physical firmware micro-core calls one or more physical task engines in the logical multi-homing task engine based on the arranged service control logic to perform hardware acceleration on the message processing process of this type of service characteristics.
[0014] Furthermore, the physical firmware micro-kernel in the logical firmware micro-kernel adopts a pipeline model to transmit commands to the physical task engine in the logical multi-homing task engine.
[0015] Further, receiving an orchestration instruction sent by a user through an orchestration configuration interface, and orchestrating a message processing process of a service characteristic supported by the physical firmware micro-core; Receive configuration instructions sent by the user through the orchestration configuration interface, and configure the service characteristics supported by each physical firmware micro-core in the logical firmware micro-core.
[0016] Furthermore, when the on-chip network router receives a service request or service message of a certain type of service characteristic, the position identifier of the on-chip network router in the router matrix is used as the network address, and a deterministic routing algorithm or an adaptive routing algorithm is used to route the service request or service message to a physical firmware microkernel that supports this type of service characteristic.
[0017] The present invention is based on an on-chip network. By virtualizing multiple physical firmware microkernels into a logical firmware microkernel, the division of labor and cooperation of the physical firmware microkernel within the logical firmware microkernel group is realized, and the one-to-one mapping between the physical firmware microkernel and the hardware acceleration module it manages, namely, the physical task engine TE, is converted into a many-to-many mapping relationship. This effectively reduces the demand of the physical firmware microkernel FW for built-in IRAM and DRAM resources while supporting the same number of business features, greatly reduces the power consumption of the chip, and saves chip area and cost.
[0018] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present specification. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the specification and, together with the description, serve to explain the principles of the specification.
[0020] Figure 1 A schematic diagram of a multi-core network-on-chip architecture including 4*4 router nodes adopted in an embodiment of the present invention; Figure 2 This is an example diagram of routing based on the location coordinates of an on-chip network router in one embodiment of the present invention; Figure 3 A schematic diagram of the basic structure of a router in a network-on-chip multi-core architecture used in an embodiment of the present invention; Figure 4 is a schematic diagram of a logic firmware micro-core in a network-on-chip many-core architecture according to an embodiment of the present invention; Figure 5 A schematic diagram of the binding relationship between the logic firmware micro-core and the task group in one embodiment of the present invention; Figure 6 A schematic diagram of the binding relationship between the logic firmware micro-core and the task group in another embodiment of the present invention; Figure 7 A schematic flowchart of the steps of a data processing method based on a network-on-chip many-core architecture is provided in one embodiment of the present invention. DETAILED DESCRIPTION
[0021] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following exemplary embodiments do not represent all implementations consistent with this specification. Instead, they are merely examples of devices and methods consistent with some aspects of this specification as detailed in the appended claims.
[0022] The terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit this specification. The singular forms "a", "the" and "the" used in this specification and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.
[0023] It should be understood that although the terms first, second, third, etc. may be used in this specification to describe various information, this information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of this specification, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0024] For data communication chips, with the accelerated iteration and development of network technology (user-level bandwidth guarantee, fine network traffic profiling to support AI controller to intelligently optimize network parameters), it is no longer enough to just support the most basic L2 / L3 layer message forwarding actions from the hardware, but to offload more complex message processing behaviors (packetprocess) to the chip as much as possible to provide acceleration capabilities for more complex message processing. For example: large-scale high-precision (3.3ms*3) BFD (Bidirectional Forwarding Detection) / CFD (Connectivity Fault Detection) detection capabilities, millisecond-level telemetry information collection capabilities, high-performance SRV6 (Segment Routing over IPv6, segment routing based on IPv6) tunnel processing capabilities, etc. If these capabilities are implemented by the host CPU, the bottleneck is very obvious, but based on the NOC architecture, combined with the firmware microkernel technology, this bottleneck can be broken.
[0025] Figure 1 This is a schematic diagram of a multi-core network-on-chip architecture including 4*4 router nodes used in an embodiment of the present invention. Figure 1In the example of a 4*4 router matrix of the on-chip network, each on-chip network router (R) is connected to a physical FW (Firmware microkernel) and a physical TE (Task Engine) through a local interface (LocalPort). Assuming that each FW in the data communication chip based on the multi-core architecture provides 10 3.3ms*3 BFD session processing capabilities, the data communication chip as a whole can provide 160 3.3ms*3 BFD session processing capabilities. If the NOC scale is expanded to 10*10, it can provide 1,000 3.3ms*3 BFD session processing capabilities, which far exceeds the maximum processing capability that the host CPU can provide.
[0026] Figure 1 In the example of the many-core architecture of the on-chip network, each physical firmware micro-core FW has programmable capabilities. Users can customize complex control logic and combine several basic hardware-fixed PP (Packet Process) actions to support the message processing of more complex business characteristics. The physical task engine TE is used to abstract the basic packet processing PP actions from hardware fixation, such as implementing PP actions such as message header encapsulation or table lookup algorithm in RTL (Register Transfer Level). That is, FW controls the message processing process of the supported business characteristics based on the arranged control logic, and TE acts as a worker to execute basic PP actions to achieve hardware acceleration of the message processing process of specific business characteristics.
[0027] For example, assume that the TE module abstracts the basic PP actions including PP1, PP2, PP3, PP4, and PP5. Among them, the message processing flow of service feature 1 includes the packet processing actions of PP1->PP3->PP5. The control logic of the message processing process of service feature 1 can be arranged on the FW, and TE can be called in sequence to execute PP1, PP3, and PP5 to accelerate the message service processing process of service feature 1. The message processing flow of service feature 2 includes the packet processing actions of PP2->PP3->PP4. The control logic of the message processing process of service feature 2 can be arranged on the FW, and TE can be called in sequence to execute PP2, PP3, and PP4 to accelerate the message service processing process of service feature 2. Later, if it is necessary to expand and support more service features, the message processing flow on the FW can be dynamically arranged and configured to support more services.
[0028] The programmable properties of FW can support dynamic arrangement of processing flows for different business characteristics, improving the business scalability of data communication chips. TE can accelerate message processing performance by solidifying PP actions through hardware. Figure 1The example NOC-based many-core architecture enables the chip to achieve a good balance between business flexibility and processing performance.
[0029] Figure 2 This is an example diagram of routing based on the position coordinates of the on-chip network router in one embodiment of the present invention. The routers in the on-chip network constitute a router matrix, and the position identifier {X, Y} of the on-chip network router (NOC Router, referred to as Router) in the matrix can be used as the address of the router and the FW and TE attached to the router to achieve addressing and routing. Assuming that the router in the lower left corner is the origin of the coordinate axis, the router position identifiers along the X-axis direction are 0, 1, 2, and 3 respectively, and the router position identifiers along the Y-axis direction are 0, 1, 2, and 3 respectively; then the position identifier of the router at the lower left corner located at the origin of the coordinate is {0,0}, the position identifier of the router at the upper left corner is {0,3}, and the position identifier of the router at the lower right corner is {3,0}. If it is necessary to send a message to the FW / TE attached to the router at the {0,3} position, the destination address of the message can be filled in {0,3}. When the router at the {0,3} position receives the message, it can forward the message to the local attached FW / TE through the local port.
[0030] The routing algorithm adopted by the on-chip network multi-core architecture provided by the embodiment of the present invention can be selected according to the specific business scenario, and the present invention does not make specific limitations. For example, the Dimension-Order Routing algorithm and the shortest path routing algorithm in the deterministic routing algorithm can be adopted. The fully adaptive routing algorithm and the minimum adaptive routing algorithm in the adaptive routing algorithm can also be used to realize the routing of messages between routers.
[0031] like Figure 2For example, if the source router {0,1} of the on-chip network needs to send a BFD service type message to the destination router {0,3} according to the configuration of the logical firmware microkernel, the location identifier of the source / destination router can be used as the address of the source / destination router in the on-chip network, and the adaptive routing algorithm can be used to send the packet from node {0,1} to router {0,3}. When the direct connection path between router {0,1}->router {0,2} and router {0,2} to node {0,3} is unobstructed, the shortest path can be directly selected: router {0,1}->router {0,2}->router {0,3}. If the direct connection path between router {0,1}->router {0,2} is blocked, the routing path calculated based on the adaptive routing algorithm may also select the path of router {0,1}->router {1,1}->router {1,2}->router {1,3}->router {0,3}.
[0032] Figure 3 The schematic diagram of the basic structure of the router in the on-chip network multi-core architecture used in an embodiment of the present invention. The on-chip network router provides four pairs of input and output ports in the four directions of North / East / South / West, which are used to realize the interconnection between routers in the four directions in the NOC matrix. In addition, a LocalPort is also provided to connect the physical firmware microcore FW of the node and the local physical task engine TE.
[0033] The on-chip network router uses the Credit-FIFO (credit-based first-in, first-out) mechanism for flow control. The input port in each direction is configured with a FIFO queue. The destination node provides credit to the source node. The source node determines the number of messages sent to the destination node based on the size of the credit value. The source node can send data to the downstream node only when it has credit.
[0034] In order to improve data transmission efficiency, reduce transmission delay and enhance routing flexibility in the on-chip network, the message is usually divided into smaller message slices (flits) according to the physical link bandwidth, and then combined with pipeline transmission technology to achieve efficient transmission of messages between the source and destination routers. When the physical link bandwidth is wide enough or the message itself is small, the message can also be used as the basic unit for message transmission between on-chip network routers. The present invention does not make specific limitations. For the convenience of expression, the "message" in the present invention can be understood as "message" or "message slice" according to the specific environment.
[0035] The access request generated by FW or TE enters the on-chip network router through the localport local port. After several hops of routing, it leaves the on-chip network NOC through the localport local port at the on-chip network router corresponding to the target (destination) address and reaches the destination FW or TE.
[0036] In each clock cycle, the arbitration module of the on-chip network router arbitrates between the data coming from the five directions of N / E / S / W and localport, selects one to process, and sends it to the next hop according to its destination address. The input ports in the five directions are all equipped with FIFO queues to cache the messages sent by the sender. The credit management module calculates the credit value based on the cache status of the messages in the FIFO queue and feeds back the available credit value to the sender. For example, when the router takes out a message from the FIFO queue in a certain direction, the credit management module adds 1 to the credit value in the corresponding direction, and then feeds back the current credit value to the sender. The sender can continue to send according to the fed-back available credit value.
[0037] The crossbar switch (xBAR switch) in an on-chip network router is a key switching structure used to implement data exchange and routing between different input ports and output ports.
[0038] The network-on-chip multi-core architecture chip provided by the present invention integrates many cores. If each core is equipped with a large-capacity instruction random access memory IRAM and a data random access memory DRAM, the cost, power consumption, design complexity and other indicators of the chip will increase significantly. Therefore, the many cores must be sufficiently streamlined, for example, only a small amount of SRAM (Static Random Access Memory) is used to implement the instruction random access memory IRAM for storing instructions and the data random access memory DRAM for storing data. In addition, no other memory is set, and there is no complex cache system. However, less IRAM and DRAM will limit the types of business features supported by the firmware micro-core, which creates a contradiction. If you want to control cost and power consumption, you have to lose the types of business features supported.
[0039] On the basis of the on-chip network multi-core architecture provided by the present invention, in order to reduce the resource requirements of the instruction code of the service features running on the FW firmware micro-core on IRAM and DRAM, so that when supporting the same number of service features, the FW firmware micro-core requires less built-in random access memory resources, thereby reducing the power consumption and cost of the data communication chip, the present invention further proposes the following improvement scheme: Multiple physical firmware micro-cores are virtualized into a logical firmware micro-core, and each physical firmware micro-core is equivalent to a hardware thread of the logical firmware micro-core. Then, the instruction codes of multiple business features to be run on the logical firmware micro-core are divided into multiple task groups (task_group). Each physical firmware micro-core only runs the instruction codes in one task group among the multiple task groups, that is, only runs a subset of the business feature instruction code set instead of the full set, so that the codes of different task groups are distributed on different physical firmware micro-cores in the logical firmware micro-core, thereby achieving support for multiple business features while reducing the resource requirements of each firmware micro-core on IRAM and DRAM.
[0040] Figure 4 The schematic diagram of the logical firmware micro-core in the on-chip network multi-core architecture of an embodiment of the present invention. In this example, each firmware micro-core FW and task engine TE on the hardware are connected to the router matrix composed of the on-chip network router (noc router), and any FW in the logical firmware micro-core can access any TE belonging to the logical firmware micro-core through the on-chip network. The user plans the logical firmware micro-core for the numerous firmware micro-cores mounted on the on-chip network based on the number of IRAM and DRAM resources of the firmware micro-core FW and the instruction code amount and data space requirements of the business characteristics to be supported, and divides the code running on the firmware micro-core into multiple task groups (task_group), so that the instruction code amount of the business characteristics contained in each task group (corresponding to the IRAM requirement) and the required data random access memory (corresponding to the DRAM requirement) do not exceed the resource amount of the instruction random access memory IRAM and data random access memory DRAM built into the firmware micro-core.
[0041] Assume that the many-core architecture of the example needs to support packet processing of four service characteristics, namely, bidirectional forwarding detection BFD / connectivity error detection CFD packet generation (BFD / CFD packet generation), deep packet inspection (DPI), telemetry, and network traffic statistics and analysis technology (Netstream). The IRAM and DRAM of each firmware microcore are 8KB respectively. The user plans these types of services into two service groups, namely Task_group1 and Task_group2, according to the number of instruction codes of these types of service characteristics and the amount of data random access memory space required, as shown in the example in Table 1.
[0042] Table 1: The IRAM and DRAM required for the BFD / CFD service feature are 2KB and 1.5KB respectively, the IRAM and DRAM required for the DPI service feature are 6KB and 6.5KB respectively, the IRAM and DRAM required for the Telemetry service feature are 3KB and 4KB respectively, the IRAM and DRAM required for the Netstream service feature are 5KB and 4KB respectively, and the total amount of IRAM and DRAM required for the four services is 16KB respectively. Assuming that the amount of IRAM and DRAM of each firmware microcore in the on-chip network multi-core architecture is 8KB, without using the technical solution provided by the present invention, the on-chip network multi-core architecture cannot support the message processing of so many types of service features.
[0043] Based on the technical solution provided by the present invention, in this embodiment, the four business characteristics are divided into two business groups according to the amount of instruction code and the amount of data required for each business characteristic, namely Task_group1 and Task_group2. Task_group1 includes BFD / CFD business characteristics and DPI business characteristics, and Task_group2 includes Telemetry business characteristics and Netstream business characteristics. Through such division, the amount of instruction code and the amount of data required for the business characteristics in each business group do not exceed the number of IRAM and DRAM built into each firmware microcore FW.
[0044] On the basis of completing the above planning, Figure 4 In the example of FIG. 1 , two firmware micro-cores FW_A and FW_B directly connected to the on-chip network router node {0,3} and the router node {1,3} are bound through the chip configuration interface and virtualized into a logical firmware micro-core. Task_group1 is bound to FW_A, and Task_group2 is bound to FW_B. That is, the business features in the two task groups are distributed to two different firmware micro-cores in the logical firmware micro-core. Each firmware micro-core in the logical firmware micro-core only needs to support the message processing flow of the business features in the bound task group. In this way, a logical firmware micro-core can simultaneously support the above four business features.
[0045] Figure 5 Schematic diagram of the binding relationship between the logical firmware microkernel and the task group in one embodiment of the present invention.
[0046] FW_A and TE_A are firmware micro-kernels and task engines directly connected to the same on-chip network router {0,3}; FW_B and TE_B are firmware micro-kernels FW and task engines TE directly connected to the same on-chip network router {1,3}. After FW_A and FW_B are virtualized into a logical firmware micro-kernel, FW_A and FW_B are equivalent to being fully connected to TE_A and TE_B respectively through the on-chip network, and FW_A and FW_B can respectively call TE_A and / or TE_B to perform the basic packet processing PP actions required by the service characteristics in their respective bound task groups. For example, FW_A can dispatch the PP actions required for the message processing process of the service characteristics BFD / CFD and / or DPI in Task_group1 to TE_A and / or TE_B for execution; FW_B can also dispatch the PP actions required for the message processing process of the service characteristics Telemetry and / or Netstream in Task_group2 to TE_A and / or TE_B for execution.
[0047] Multiple physical task engines directly connected to the physical firmware microkernel in the logical firmware microkernel on the same on-chip network router constitute a logical multi-homing task engine. For example, TE_A and TE_B are the logical multi-homing task engines Logic TE1{TE_A, TE_B} of the logical firmware microkernel Logic FW1{FW_A, FW_B}. Any physical firmware microkernel FW in the logical firmware microkernel can manage and call one or more physical task engines TE in the logical multi-homing task engine.
[0048] There is a binding relationship between the logic firmware micro-core (Logic FW1) and the logic multi-homing task engine (Logic TE1) in the present invention, and any firmware micro-core in the logic firmware micro-core can manage and call one or more task engines in the logic multi-homing task engine.
[0049] Assuming that FW_A in the logic firmware micro-core Logic FW1 needs to process the packets of 20 BFD session packets, FW_A can assign the PP actions required for these packet processing PP actions to one or more task engines of the logic multi-homing task engine Logic TE1 for execution based on the preset load balancing algorithm. For example, the packet processing actions required for the first 10 BFD session packets are sent to TE_A in Logic TE1 for processing through the source address {0,3} and the destination address {0,3}, and the packet processing actions required for the next 10 BFD session packets are sent to TE_B in Logic TE1 for processing through the source address {0,3} and the destination address {1,3}. The on-chip network router noc router can correctly send the packet to the TE attached to the router corresponding to the destination address according to the destination address of the data packet.
[0050] It can be seen from the above embodiments that although the total instruction code and data space requirements for the service features run by the FW are both 16KB, by managing two TEs with one FW, each FW only needs 8KB of code and 8KB of data space, that is, the required random memory resources of each FW are halved to meet the demand, thereby achieving the effect of reducing the power consumption and cost of the firmware micro-core in the on-chip network many-core architecture while supporting the same number of service features.
[0051] Figure 6 Schematic diagram of the binding relationship between the logical firmware microkernel and the task group in another embodiment of the present invention.
[0052] In this embodiment, the firmware micro-cores FW_A, FW_B, FW_C and FW_D respectively mounted on the four on-chip network routers are virtualized into a large logical firmware micro-core Virtual FW{FW_A, FW_B, FW_C, FW_D}. The four firmware micro-cores are fully connected to the four task engines TE_A, TE_B, TE_C and TE_D respectively connected to the same on-chip network router node with the four firmware micro-cores through the on-chip network. The four task engines constitute the logical multi-home task engine Virtual TE{TE_A, TE_B, TE_C, TE_D}. Any of the four firmware micro-cores in the logical firmware micro-core can manage and call one or more task engines TE in the logical multi-home task engine Virtual TE. This is equivalent to that each firmware micro-core in the logical firmware micro-core Virtual FW can manage and call the four task engines TE in the logical multi-home task engine Virtual TE.
[0053] Assuming that the four firmware micro-cores have only 8KB IRAM and 8KB DRAM, and each firmware micro-core can support two different business features, it is equivalent to increasing the number of business features that the chip can support from 2 to 8 without increasing the IRAM and DRAM resources of each firmware micro-core, thus achieving the effect of 32KB IRAM and 32KB DRAM.
[0054] Furthermore, when any firmware micro-core FW in a logical firmware micro-core manages multiple task engines TE through an on-chip network, compared to the case where each firmware micro-core FW only manages a local TE, the cmd (command) issued by the physical firmware micro-core FW reaches other target TEs in the logical multi-home task engine except the local TE, which will inevitably increase the latency, for example, the clock delay may increase by several clock cycles. In one embodiment of the present invention, in order to reduce latency, when the physical firmware micro-core FW in the logical firmware micro-core sends a (cmd) command to the physical task engine in the logical multi-home task engine, a pipeline model is used to transmit the cmd, thereby trying to avoid performance degradation caused by latency.
[0055] When the physical firmware microkernel uses a pipeline model to send cmds to TEs, assuming that FW_A needs to send commands to TE_A and TE_B to perform PP actions, after FW_A sends cmds to TE_A, it can continue to process the cmds sent to TE_B without waiting for the response returned by TE_A to FW_A, and then return to process the response returned by TE_A, and then process the response returned by TE_B. That is, by parallelizing the FW's processing of the next target TE and the transmission of the cmd to the previous target TE on the NOC, the impact of the latency caused by transmitting cmds on the NOC on performance when one FW controls multiple TEs can be minimized.
[0056] Based on the above embodiments, it can be seen that the present invention is based on an on-chip network. By virtualizing multiple physical firmware microkernels into a logical firmware microkernel, the division of labor and cooperation of the physical firmware microkernel within the logical firmware microkernel group is realized, and the one-to-one mapping between the physical firmware microkernel and the hardware acceleration module it manages, namely the physical task engine TE, is converted into a many-to-many mapping relationship, thereby effectively reducing the demand of the physical firmware microkernel FW for built-in IRAM and DRAM resources while supporting the same number of business features, thereby greatly reducing the power consumption of the chip, and saving chip area and cost.
[0057] Furthermore, the logical firmware microkernel and task group Task_group can be designed as a dynamic configuration mode. Users can dynamically configure the physical firmware microkernel included in the logical firmware microkernel, the task groups supported by each firmware microkernel, and the message processing process of the business characteristics supported by the physical firmware microkernel through the chip's orchestration configuration interface, so that users can flexibly customize and expand the chip's functions according to actual application scenarios.
[0058] Furthermore, in order to reduce the delay of transmitting cmd between FW and non-local TE through the on-chip network, the principle of proximity can be adopted when configuring the logical firmware micro-core, and a group of adjacent physical firmware micro-cores can be configured as a logical firmware micro-core.
[0059] Based on the on-chip network many-core architecture provided by the embodiment of the present invention, the present invention further proposes a data processing method, which is applied to a chip adopting the on-chip network many-core architecture provided by the present invention.
[0060] Figure 7 A schematic flow chart of the steps of a data processing method based on a multi-core architecture of an on-chip network provided in one embodiment of the present invention, in which it is assumed that two adjacent physical firmware microkernels 1 and 2 are bound and virtualized into a logical firmware microkernel in the on-chip network, a physical task engine 1 co-located with the physical firmware microkernel 1 and a physical task engine 2 co-located with the physical firmware microkernel 2 are bound and virtualized into a logical multi-homed task engine, and the physical firmware microkernel 1 in the logical firmware microkernel is bound to task group 1, and the physical firmware microkernel 2 is bound to task group 2. Task group 1 includes two service characteristics, namely service characteristic 1 and service characteristic 2. When any network-on-chip router in the network-on-chip receives a service message of service feature 1, the service request or service message of service feature 1 is routed to the network-on-chip router directly connected to the physical firmware microcore 1 based on the preset routing algorithm according to the logical firmware microcore in the chip configuration and the binding relationship between the service feature and the physical firmware microcore, and then forwarded to the physical firmware microcore 1 through the local port by the router. The physical firmware microcore 1 can call one or more physical task engines in the logical multi-homing task engine to accelerate the service processing process according to the configuration of the logical firmware microcore. Based on the example of the above application environment, the steps of the method include: Step 701. When a service request or service message of a certain type of service characteristic is received, the service request or service message is routed to a physical firmware micro-core bound to the type of service characteristic in a logical firmware micro-core supporting the type of service characteristic in the on-chip network many-core architecture; For example, when a service request or service message of service characteristic 1 is received, the service request or service message of the service characteristic is routed to the physical firmware microcore 1 bound to the service characteristic 1 in the logical firmware microcore supporting the service characteristic 1 in the on-chip network multi-core architecture.
[0061] Step 702: The physical firmware micro-core calls one or more physical task engines in the logical multi-homed task engine based on the orchestrated service control logic to perform hardware acceleration on the message processing process of the service characteristics.
[0062] For example, after receiving the service message of service characteristic 1, the physical firmware micro-core 1 calls one or more physical task engines in the logical multi-homing task engine based on the orchestrated service control logic to perform hardware acceleration on the message processing process of service characteristic 1.
[0063] Furthermore, the physical firmware micro-kernel in the logical firmware micro-kernel adopts a pipeline model to transmit commands to the physical task engine in the logical multi-homed task engine.
[0064] Furthermore, the on-chip network multi-core architecture provides an orchestration configuration interface. When the chip receives an orchestration instruction sent by the user through the orchestration configuration interface, the message processing process of the service characteristics supported by the physical firmware micro-core is orchestrated according to the orchestration instruction. When the chip receives a configuration instruction sent by the user through the orchestration configuration interface, the service characteristics supported by each physical firmware micro-core in the logical firmware micro-core are configured according to the configuration instruction.
[0065] Furthermore, when the on-chip network router receives a service request or service message of a certain type of service characteristic, the position identifier of the on-chip network router in the router matrix is used as the network address, and a deterministic routing algorithm or an adaptive routing algorithm is used to route the service request or service message to a physical firmware microkernel that supports this type of service characteristic.
[0066] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0067] Those skilled in the art will readily appreciate other embodiments of the specification after considering the specification and practicing the inventions claimed herein. The specification is intended to cover any variations, uses, or adaptations of the specification that follow the general principles of the specification and include common knowledge or customary techniques in the art that are not claimed in the specification. The specification and examples are to be considered exemplary only, and the true scope and spirit of the specification are indicated by the claims.
[0068] It should be understood that the present description is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present description is limited only by the appended claims.
[0069] The above description is only a preferred embodiment of this specification and is not intended to limit this specification. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of this specification should be included in the scope of protection of this specification.
Claims
1. A network-on-chip many-core architecture, characterized in that: include: A logical firmware micro-core, including a plurality of physical firmware micro-cores, wherein the plurality of physical firmware micro-cores respectively support service characteristics in different task groups, and a task group includes one or more service characteristics; A logical multi-homed task engine, including a plurality of physical task engines, wherein the physical task engines have hardware-hardened basic packet processing actions for hardware acceleration of message processing of service characteristics orchestrated by the physical firmware micro-core; each physical task engine is controlled by each physical firmware micro-core in the logical firmware micro-core; In the network on chip, the physical firmware micro-kernel and the physical task engine are connected to the network on chip router through a local port of the network on chip router.
2. The on-chip network many-core architecture according to claim 1, characterized in that: The total amount of instruction codes and the total amount of data random access memory required for the service characteristics in the task group shall not exceed the instruction random access memory capacity and the data random access memory capacity in the physical firmware micro-core; The number of physical firmware microcores included in the logical firmware microcore is determined by the total amount of instruction random access memory and data random access memory required for the business characteristics that the logical firmware microcore needs to support and the instruction random access memory capacity and data random access memory capacity of the physical firmware microcore.
3. The on-chip network many-core architecture according to claim 1, characterized in that: Each network-on-chip router in the network-on-chip is connected to a physical firmware micro-core and a physical task engine through a local port; The logical firmware micro-core and the logical multi-homed task engine are respectively composed of multiple pairs of physical firmware micro-cores and physical task engines connected to the same on-chip network router; Each physical task engine in the logical multi-homed task engine is fully connected to each physical firmware microcore in the logical firmware microcore through an on-chip network.
4. The on-chip network many-core architecture according to claim 3, characterized in that: The physical firmware micro-core in the logical firmware micro-core adopts a pipeline model to transmit commands to the physical task engine in the logical multi-homing task engine.
5. The on-chip network many-core architecture according to claim 1, characterized in that: The on-chip network multi-core architecture provides users with an orchestration configuration interface, through which the message processing processes of different business characteristics in the task group are orchestrated, and several basic hardware-cured basic packet processing actions are orchestrated and combined to support the business characteristics in the task group; and the business characteristics supported by each physical firmware microcore in the logical firmware microcore are configured.
6. The on-chip network many-core architecture according to claim 5, characterized in that: The physical firmware micro-core in the logical firmware micro-core is a group of adjacent firmware micro-cores in the on-chip network. The principle of proximity is adopted when configuring the logical firmware micro-core, and a group of adjacent physical firmware micro-cores are configured as one logical firmware micro-core.
7. A data processing method, characterized in that: The method is applied to a chip using an on-chip network many-core architecture according to any one of claims 1 to 6, and the method comprises: When a service request or service message of a certain type of service characteristic is received, the service request or service message is routed to a physical firmware microcore bound to the type of service characteristic in a logical firmware microcore supporting the type of service characteristic in the on-chip network many-core architecture; The physical firmware micro-core calls one or more physical task engines in the logical multi-homing task engine based on the arranged service control logic to perform hardware acceleration on the message processing process of this type of service characteristics.
8. The data processing method according to claim 7, characterized in that: The physical firmware micro-core in the logical firmware micro-core adopts a pipeline model to transmit commands to the physical task engine in the logical multi-homing task engine.
9. The data processing method according to claim 7, characterized in that: Receiving an orchestration instruction sent by a user through an orchestration configuration interface, and orchestrating a message processing process of a service characteristic supported by the physical firmware micro-core; Receive configuration instructions sent by the user through the orchestration configuration interface, and configure the service characteristics supported by each physical firmware micro-core in the logical firmware micro-core.
10. The data processing method according to claim 7, characterized in that: When the on-chip network router receives a service request or service message of a certain type of service feature, the position identifier of the on-chip network router in the router matrix is used as the network address, and a deterministic routing algorithm or an adaptive routing algorithm is used to route the service request or service message to the physical firmware microkernel that supports this type of service feature.
Citation Information
Patent Citations
Performance acceleration method of heterogeneous multi-core computing platform on chip
CN102360313A
On-chip network multi-core framework
CN105207957A
Big data cluster performance optimization method and device on physical core super multi-thread server
CN111176847A
Distributed accelerator
US20230236889A1