Processor Reconfigurable Programmable Switching System and Programmable Data Plane Chip

By designing a reconfigurable processor in a programmable switching system, it can be dynamically reconfigured into a pipeline stage or an RTC processor, solving the problems of programmable switching chip architecture in the prior art, insufficient throughput and high latency, and achieving more efficient network processing and performance improvement.

CN119071255BActive Publication Date: 2025-06-24TSINGHUA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410929824.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-11
Publication Date
2025-06-24
Estimated Expiration
2044-07-11

AI Technical Summary

Technical Problem

The architecture of existing programmable switch chips is relatively limited in programming, insufficient throughput and high latency in network communication, and cannot effectively process data packets containing a large number of computing requirements, long processing processes and dependencies.

Method used

A programmable switching system with reconfigurable processors is designed. Through the reconfigurable processor, it can be reconfigured into a pipeline stage or an RTC processor, which realizes the transmission of data packet header vectors and the transmission of other data and control signals, and supports parallel processing of multiple reconfigurable processors and dynamic switching of data paths.

Benefits of technology

It greatly improves the programmability of the switch, reduces the latency of network communication, improves the performance of network services, and reduces the demand for servers and intermediate box equipment of the data center network, and reduces the construction and maintenance costs of the data center network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119071255B_ABST
    Figure CN119071255B_ABST
Patent Text Reader

Abstract

This application relates to the field of programmable data plane technology, and particularly to a programmable switching system with reconfigurable processors and a programmable data plane chip. The structure includes: a first side path for transmitting the packet header vector (PHV); a second side path for transmitting other data and control signals other than the PHV; a reconfigurable processor sequence including multiple reconfigurable processors, which is connected to the first side path and the second side path. Each reconfigurable processor is allowed to be reconfigured into a pipeline stage or run to complete the RTC processor. There is a data path for transmitting the PHV between each adjacent pair of reconfigurable processors. Different connection operations are performed according to the reconfigurable processors at both ends of the data path. Thus, the problems in the related art such as the limited architecture programmability of the programmable switching chip, insufficient throughput, and high latency in network communication are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of programmable data planes, and in particular to a programmable switching system with reconfigurable processors and a programmable data plane chip. Background Art

[0002] Software-defined networks include data planes and control planes. The data plane is responsible for protocol parsing and reverse parsing, and packet processing, while the control plane is responsible for issuing parsing rules and flow table matching rules. From customized protocol forwarding to supporting in-network computing applications, high-performance programmable switching chips have further explored the potential of data plane devices. The architecture of programmable switching chips can be divided into three types: pipeline, RTC (Run-To-Complete), and the fusion architecture of the two.

[0003] Programmable switching chips using pipeline architecture often have relatively high throughput and relatively low processing latency, but they have the following disadvantages:

[0004] 1. Each pipeline stage can perform a limited number of computational steps;

[0005] 2. Due to the unidirectional nature of the pipeline, there is no efficient mechanism to pass data to the early stages;

[0006] In response to the above shortcomings, if the packet processing process meets any of the following conditions: it contains a large number of computing requirements, the processing process is long and cannot be parallelized due to dependencies, or the state data needs to be written back to the early stage, then the pipeline architecture needs to recirculate the data packet. Recirculation will cause the latency of the recirculated data packet to increase significantly, the system throughput to decrease, and the recirculated data packet to occupy the pipeline processing capacity multiple times, thereby reducing the background traffic processing capacity other than the recirculated traffic and generating potential disorder, thereby destroying the strong state consistency of stateful applications.

[0007] If recirculation is not used, there are generally three solutions:

[0008] 1. Submit such tasks to the control plane for processing;

[0009] 2. Forward the data packets that need to perform such tasks to the dedicated network function server for processing;

[0010] 3. Connect an FPGA (Field-Programmable Gate Array) or other dedicated hardware processing to the switch chip;

[0011] All of the above three solutions require a significant increase in the cost of the solution and still cannot solve the problems of a significant increase in the delay of the recycled data packets and the potential for out-of-order packets, which may disrupt the strong state consistency of stateful applications.

[0012] The processing of a data packet by the RTC processor can continue until the entire processing flow is completed. It can execute long and complex processing flows without the need to forward it to other hardware modules. However, it also faces problems such as insufficient throughput, high latency, and difficulty in ensuring state consistency. Summary of the Invention

[0013] This application provides a processor-reconfigurable programmable switching system and a programmable data plane chip to solve the problems in the related art, such as the limited architecture programmability of the programmable switching chip, insufficient throughput, and high latency in network communication.

[0014] The first aspect of the embodiments of this application provides a processor-reconfigurable programmable switching system, including: a first side path for transmitting PHV (Packet Header Vector); a second side path for transmitting other data and control signals except PHV; a reconfigurable processor sequence, which includes multiple reconfigurable processors. The first side path and the second side path connect multiple reconfigurable processors. Among them, each reconfigurable processor of the multiple reconfigurable processors is allowed to be reconfigured into a pipeline stage or run to complete the RTC processor. There is a data path for transmitting PHV between each adjacent pair of reconfigurable processors. When both ends of the data path are reconfigurable processors reconfigured into pipeline stages, the data path is opened to connect the two pipeline stages into a complete pipeline. When both ends of the data path are not reconfigurable processors reconfigured into pipeline stages, the data path is closed. The reconfigurable processor reconfigured into the RTC processor is connected through the first side path.

[0015] Optionally, the reconfigurable processor includes a processor that is reconfigurable for all of the following items, or a processor that is reconfigurable for some of the following items and the other items are not included in the processor, or a processor that is reconfigurable for some of the following items and the other items are not reconfigurable: a register array that is used to store data packet header fields, metadata, and intermediate signals during the pipeline stage, and serves as the register bank specified by the instruction set in the RTC processor; a data memory that is used to store lookup tables and status tables during the pipeline stage, where the lookup tables are used for lookups when data packets pass through, and the status tables are used to read, modify, and write back status when data packets pass through. In the RTC processor, the data memory is used to store data packet header fields, metadata, lookup tables, status tables, and intermediate signals. If an operating system is installed in the RTC processor, the data memory is also used to store data structures related to operating system maintenance; an instruction memory that is used to store instructions during both the pipeline stage and the RTC processor; an ALU (Arithmetic and Logic Unit) that is used to perform arithmetic and logical calculations during both the pipeline stage and the RTC processor; a SALU (Stateful Arithmetic and Logic Unit) that is used to read status from the status table in the data memory during the pipeline stage, perform arithmetic and logical calculations, and then write the calculation result back to the status table. In the RTC processor, it is used to read data from the data memory, perform arithmetic and logical calculations, and then write the calculation result back to the data memory; a matching module that is used to perform lookups and matches in the lookup tables of the data memory based on a given keyword during both the pipeline stage and the RTC processor, and return whether there is a matching table entry and the corresponding table entry content.

[0016] Optionally, when the reconfigurable processor is reconfigured into pipeline stages, it includes: after the PHV of a data packet is received by the current processor, it is written into the register array of the current processor; according to the keyword selection rule specified by the user program, a part of the fields are selected to form a keyword and sent to the matching module; the matching module performs lookup and matching in the lookup table of the data memory according to the given keyword. When there is a matching table entry, the operation ID (Identification) is determined according to the table entry content; when there is no matching table entry, the default operation ID specified by the user program is used; instructions are read from the instruction memory according to the operation ID, the instructions are parsed using a decoder, and the parsed instructions are dispatched to the ALU and SALU. Among them, the ALU reads the fields participating in the calculation from the register array or instruction parameters, and writes back the result to the register array after calculation; the SALU reads the fields participating in the calculation from the register array or instruction parameters, and reads the status participating in the calculation from the status table in the data memory, and writes back the result to the register array and / or the status table in the data memory; in the processing flow, the PHV of a data packet is passed in the registers corresponding to different pipeline beats in the register array, and after the processing is completed, the corresponding PHV is sent out of the current processor.

[0017] Optionally, when the reconfigurable processor is reconfigured into an RTC processor, it includes: after the PHV of a data packet is received by the current processor, it is written into the data memory of the current processor, and the current processor is notified to start the processing flow. The current processor reads the instructions to be executed from the instruction register according to the program counter until the processing termination instruction is read; each read instruction is parsed by a dedicated decoder and dispatched to the ALU or SALU. Among them, the ALU reads the fields participating in the calculation from the register array or instruction parameters, and writes back the result to the register array after calculation; the SALU reads the fields participating in the calculation from the register array or instruction parameters, and reads the data participating in the calculation from the data memory, and writes back the result to the register array and / or the data memory; it is allowed to connect or not connect the matching module to the RTC processor. If the matching module is connected to the RTC processor, the current processor supports matching instructions, extracts the data required by the instructions from the register array and submits it to the matching module, and the matching module writes the matching result back to the register array.

[0018] Optionally, the reconfigurable processor is allowed to be reconfigured into multiple RTC cores, and the number of RTC cores is determined according to the parallelism of the pipeline stage.

[0019] Optionally, it further includes: a parser, a payload buffer, an inverse parser, and a hardware load. Among them, the parser and the inverse parser are connected to the first side path, the payload buffer and the hardware load are connected to the second side path, and the parser is also connected to the payload buffer for parsing the data packet to obtain the PHV and the payload, and sending the payload to the payload buffer.

[0020] Optionally, the first side path includes a plurality of first path nodes. In the first side path, after the PHV goes through the complete pipeline processing flow, it is sent by the reconfigurable processor that is the last one to be reconfigured into a pipeline stage in the pipeline to the corresponding first path node. Among them, for the PHV that does not need to be sent to the RTC processor for processing, it is transmitted along the first path, and the first path starts from the first path node corresponding to the reconfigurable processor that is the last one to be reconfigured into a pipeline stage in the pipeline and is directly transmitted along the first side path to the first path node corresponding to the inverse parser; for the PHV that needs to be sent to the RTC processor for further processing, it is transmitted along the second path, and the second path starts from the first path node corresponding to the reconfigurable processor that is the last one to be reconfigured into a pipeline stage in the pipeline, is transmitted along the first side path to the first path node corresponding to the RTC processor where the RTC core it is assigned to process is located, downloaded to the RTC core it is assigned to process for processing, and after the processing is completed, the cycle when the corresponding first path node is idle is selected and uploaded to the first path node corresponding to the RTC processor where the RTC core processing this PHV is located.

[0021] Optionally, the first path and the second path have a common source point and a common sink point. Among them, the source point is the first path node corresponding to the reconfigurable processor that is the last one to be reconfigured into a pipeline stage, and the sink point is the first path node corresponding to the inverse parser. There is buffer processing congestion at the sink point. When the buffer meets the condition of impending overflow, the buffer sends a target signal to the reconfigurable processor reconfigured into an RTC processor to stop uploading new PHVs to the second path based on the target signal.

[0022] Optionally, the second side path includes at least one second path node. The second side path supports at least one type of data transfer, including: sending a request from the reconfigurable processor reconfigured into an RTC processor, reading the payload of the data packet processed by the reconfigurable processor reconfigured into an RTC processor in the payload buffer, and transferring the payload back to the reconfigurable processor reconfigured into an RTC processor, where the payload will be stored by the reconfigurable processor of the reconfigured RTC processor in the data memory for subsequent processing; sending a request from the reconfigurable processor reconfigured into an RTC processor, reading the data in the data memory or register array on other reconfigurable processors, and transferring the read data back to the data memory of the reconfigurable processor reconfigured into an RTC processor that sent the read request for processing; sending a request from the reconfigurable processor reconfigured into an RTC processor to write data to the data memory or register array on other reconfigurable processors; sending a request from the reconfigurable processor reconfigured into an RTC processor to send data to the hardware load, and the hardware load sends the processing result back to the reconfigurable processor reconfigured into an RTC processor.

[0023] The second aspect embodiment of the present application provides a programmable data plane chip, including the processor-reconfigurable programmable switching system of any one of the above embodiments.

[0024] Therefore, the present application has the following beneficial effects:

[0025] Through the design of reconfiguring the reconfigurable processor into a pipeline stage or an RTC processor in the embodiments of the present application, the programmability of the switch is greatly improved with only a small amount of chip area consumed. As a result, the network applications originally deployed on servers and middleboxes can be offloaded to the programmable switch, reducing the latency of network communication, improving the performance of network services, and reducing the requirements for server and middlebox devices in the data center network, thereby reducing the construction and maintenance costs of the data center network. Thus, the problems in the related art such as the limited architecture programmability, insufficient throughput, and high latency of network communication of programmable switching chips are solved.

[0026] The additional aspects and advantages of the present application will be partially given in the following description, partially become apparent from the following description, or be understood through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The above and / or additional aspects and advantages of the present application will become apparent and easy to understand from the following description of the embodiments in conjunction with the drawings, where:

[0028] Figure 1 It is a schematic diagram of the Banzai architecture technology in the related art;

[0029] Figure 2 It is a schematic diagram of the Trio architecture technology in the related art;

[0030] Figure 3 It is a schematic diagram of the Sirius architecture technology in the related art;

[0031] Figure 4 It is a block diagram of a processor-reconfigurable programmable switching system provided according to an embodiment of the present application;

[0032] Figure 5 It is an example diagram of the composition structure and data path of a reconfigurable processor provided according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0033] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are intended to explain the present application and should not be construed as limiting the present application.

[0034] The processor-reconfigurable programmable switching system and programmable data plane chip according to the embodiments of the present application will be described below with reference to the accompanying drawings. In view of the problems mentioned in the above background art, the present application provides a processor-reconfigurable programmable switching system. By designing the reconfigurable processor into a pipeline stage or an RTC processor, the programmability of the switch is greatly improved with only a small amount of chip area consumed. Network applications originally deployed on servers and middleboxes can be offloaded to the programmable switch, reducing network communication latency, improving the performance of network services, and reducing the requirements for server and middlebox devices in the data center network, thereby reducing the construction and maintenance costs of the data center network. Thus, the problems in the related art, such as the limited architecture programmability, insufficient throughput, and high network communication latency of programmable switching chips, are solved.

[0035] Before introducing the specific content of the present application, the related technologies involved in the present application are described in detail, including the following aspects:

[0036] 1. Related Technology 1

[0037] As Figure 1 shown, Banzai is the technical prototype of the programmable switch Tofino launched by Intel. When a data packet enters, Banzai parses the header fields of the data packet through a programmable parser and processes the header through two pipelines at the entrance and the exit. These two pipelines with the same structure constitute the programmable data packet processing logic of Banzai. Each pipeline consists of a series of stages, namely MAU (Match Action Unit), and MAU is the basic programmable unit of this prototype. The user programs the MAU in two aspects through the P4 program: what flow table the data packet needs to query at this stage and what processing logic should be executed after the query result is obtained. In Banzai, the processing logic of each stage is performed using a set of parallel atomic operations, and each atomic operation can execute a "read - add / subtract - write back" function.

[0038] Programmable switching chips with a pipeline architecture often have high throughput and low latency. However, the pipeline encounters difficulties in processing the following network tasks: (1) A large number of computational requirements. Since the processing logic in each MAU uses a set of parallel atomic operations, that is, only one cycle in each MAU actually performs calculations. Therefore, if the packet processing flow contains a large number of calculations, there is a delay of an entire MAU between two calculations, seriously affecting the packet processing delay. When the amount of computation reaches a certain level, the number of calculations provided by the pipeline length is not sufficient to accommodate the complete calculation, so it is necessary to re-loop (i.e., resend the packet leaving the pipeline to the head of the pipeline for reprocessing), resulting in a significant attenuation of throughput. (2) Complex stateful functions. Due to the simplicity of atomic operations, only simple stateful functions (such as counters) can be implemented via atomic operations. For more complex functions, the logic for determining the next state may involve a series of operations and table accesses and cannot be achieved through atomic operations. To avoid bottlenecks in the pipeline, this technical solution cannot support these stateful functions by simply expanding the atomic operation circuit and enhancing the atomic operation function. Due to the unidirectional nature of the pipeline, there is no effective mechanism to write back state data to the early stage. The only way is still to re-loop, resulting in a significant attenuation of throughput. (3) Long dependency chains between statements. Since the processing logic in each MAU uses a set of parallel atomic operations, if a statement in the packet processing flow needs to use the calculation result of another statement (i.e., depends on this statement), these two statements cannot be executed in parallel and must be arranged in different MAUs. When the length of the dependency chain between statements reaches N, N different MAUs are required to execute this processing flow, even though a large amount of storage resources and computational resources in these N MAUs are wasted. This causes serious resource waste, and for a more complex packet processing flow (such as a network service chain), the usual pipeline length (i.e., the number of MAUs) is not sufficient to layout the complete process, so the packet needs to be re-looped, resulting in a significant attenuation of throughput.

[0039] 2. Related Technology 2

[0040] Trio is the technical prototype of the MX series programmable switches launched by Juniper. The unit for processing packets in Trio is the PFE (Packet Forwarding Engine). One PFE can be integrated on a single chip. To expand the processing capacity of the switch, multiple PFEs need to be connected through a switching network. The structure of the PFE is as Figure 2As shown, it contains hundreds of PPEs (Packet Processing Engines). After the data packet enters the PFE, its payload is stored in the buffer, and the header is dispatched to the target PPE through a dispatcher via a fully connected network. Each PPE is a processor of an RTC and can access the shared memory system (including SRAM (Static Random Access Memory) cache and off-chip DRAM (Dynamic Random Access Memory)), hardware acceleration modules such as hash and filters, and the buffer storing the payload through the same fully connected network. After the data packet header is processed in the PPE, it passes through the reordering engine to ensure that the sending order is the same as the arrival order of the data packets, and then is assembled with the payload in the buffer and sent out.

[0041] Since the PPE is an RTC processor and the processing of a data packet can continue until the entire processing flow is completed, Trio can execute longer and more complex processing flows without being limited by the pipeline length. However, it also faces the following challenges: (1) Insufficient throughput. Due to chip size and power limitations, the network throughput that a single PFE chip can support is quite limited, about 1-2 orders of magnitude lower than that of programmable switching chips with a pipeline architecture. (2) Higher latency. When the PPE accesses the memory (such as looking up the flow table and status table), it needs to access the shared storage system through the fully connected network. The latency of a single memory access is about 1-2 orders of magnitude lower than that of programmable switching chips with a pipeline architecture that uses MAU exclusive memory. For the status table where read-write conflicts may occur, synchronization mechanisms such as locking will further increase the memory access latency. (3) It is difficult to ensure state consistency. Since the throughput of each PPE is very low, when processing a large-scale stateful flow, the same flow may be dispatched to different PPEs, resulting in the data packets of the flow being processed in a different order from the arrival order of the data packets. When the application requires strong state consistency or relatively strict bounded stale consistency, the state consistency will be damaged.

[0042] 3. Related Technology Three

[0043] Sirius is the technical prototype of the Pensando DPU (Data Processing Unit) launched by AMD, and its architecture is as Figure 3As shown, it includes 2 general-purpose pipelines and 3 dedicated pipelines for DMA (Direct Memory Access). Each pipeline in Sirius contains a TE (Table Engine) and several MPUs (Match Process Units). The function of the MPU is similar to that of the MAU in Banzai, but it does not include a memory and the corresponding memory access module. When the MPU needs to access memory, it accesses its shared storage system (including SRAM cache and off-chip DRAM) through the TE via a NOC (Network on Chip) interconnection network. This NOC interconnection network also connects to the PCIe port leading to the server, hardware accelerators such as compression, encryption, and CRC (Cyclic Redundancy Check), and a group of embedded ARM cores. When a data packet cannot complete the processing flow through the pipeline, it is sent through the NOC interconnection network to the ARM core working in the form of RTC for processing.

[0044] By integrating the two architectures of the pipeline and RTC, Sirius realizes the separation of tasks with simple processing flows and tasks with complex processing flows. Simple tasks can be processed through the pipeline in a high-performance form, while complex tasks can be sent to the more programmable RTC core without being restricted by the pipeline length. However, the existing integrated architectures still have the following problems: (1) It is not applicable in the switch scenario. The existing integrated architectures are all designed for scenarios such as smart network cards, DPU, or IPU (Infrastructure Processing Unit). For example, the designed bandwidth of AMD Pensando DSC2-200 is only 400 Gbps, and this throughput cannot be used in the switch scenario. (2) Insufficient scalability. The integrated architecture always connects between the pipeline and the RTC core through a fully associative bus or shared memory. These structures will be in O(N) when the number of access modules N increases. 2) The complexity of () rapidly expands, resulting in unacceptable chip area and power consumption. (3) Resource idleness phenomenon. Due to the rapid development of software-defined networks, network switch equipment manufacturers cannot accurately predict the future network services carried by the network equipment they produce, and there are significant differences in the demand for pipelines and RTC cores for different network services. Services with stricter performance requirements such as throughput and latency may require a longer pipeline length, while services with complex processing flows may require a larger number of RTC cores. In order to meet the user needs in different scenarios, network switch equipment manufacturers with a fusion architecture need to provide the maximum length of the pipeline and the maximum number of RTC cores; however, when the user service only requires one type of computing resource, the other computing resource is idle, thus reducing the effective utilization rate of the chip area.

[0045] Specifically, Figure 4 It is a block diagram of a processor-reconfigurable programmable switching system provided by an embodiment of the present application.

[0046] As Figure 4 shown, the processor-reconfigurable programmable switching system includes: a first side path, a second side path, and a sequence of reconfigurable processors.

[0047] Among them, the first side path is used to transmit the packet header vector PHV. PHV is a signal composed of the fields that need to be processed in the packet header and other metadata fields generated during the packet processing process, arranged in a certain order, and its typical bit width value is 4096bit; the second side path is used to transmit other data and control signals other than PHV; the sequence of reconfigurable processors includes a plurality of reconfigurable processors. The first side path and the second side path connect the plurality of reconfigurable processors. Among them, each reconfigurable processor of the plurality of reconfigurable processors allows to be reconfigured into a pipeline stage or an RTC processor. RTC means that a packet runs through all the processing processes on one processor until completion without being sent to other processors, corresponding to the process where each stage in the pipeline structure only processes one link. There is a data path for transmitting PHV between each adjacent reconfigurable processor. When both ends of the data path are reconfigurable processors reconfigured into pipeline stages, the data path is opened to connect the two pipeline stages into a complete pipeline; when both ends of the data path are not reconfigurable processors reconfigured into pipeline stages, the data path is closed, and the reconfigurable processor reconfigured into an RTC processor is connected through the first side path.

[0048] In the embodiment of the present application, the reconfigurable processor is a special processor structure, which can be reconfigured into a stage of the pipeline (for example, a MAU) or an RTC processor according to the content of the configuration register issued by the compiler. The structure of the reconfigurable processor is asFigure 5 As shown, it includes a series of hardware components and the connections between the hardware components. When it is reconstructed into a pipeline or RTC architecture, specific components and the connections between components are made effective, thereby realizing the reuse of hardware resources.

[0049] Furthermore, the embodiment of the present application adopts a structure such as Figure 4 to connect each reconfigurable processor in series. It mainly includes a set of linearly arranged reconfigurable processor sequences, a first side path for transmitting the PHV, and a second side path for transmitting other data and control signals.

[0050] Among them, after the data packet passes through the parser, the parsed PHV enters the first reconfigurable processor in the reconfigurable processor sequence (if all reconfigurable processors are reconstructed into RTC processors, it enters the first node of the first side path), and the remaining payloads are stored in the buffer SRAM. There is a data path for transmitting the PHV between the reconfigurable processors. When both ends of the data path are reconfigurable processors reconstructed into pipeline stages, the data path is opened to connect the two pipeline stages into a complete pipeline; when both ends of the path are not reconfigurable processors reconstructed into pipeline stages, the data path is closed and does not transmit the PHV. Therefore, in the entire reconfigurable processor sequence, the first several (which can be 0) reconfigurable processors are reconstructed into pipeline stages and are connected into a complete pipeline through the data paths between the reconfigurable processors. The subsequent several (which can be 0) reconfigurable processors are reconstructed into RTC processors, and each RTC processor is relatively independent and is only connected through the first side path.

[0051] It should be noted that the parallelism of the reconfigurable processor in the embodiment of the present application when it is reconstructed into a pipeline stage determines the number of RTC cores when it is reconstructed into an RTC processor. When the reconfigurable processor uses a larger chip area, it has a greater parallelism when it is in the pipeline stage (such as being able to search more tables simultaneously), and at the same time it has more cores when it is an RTC processor. This enables the embodiment of the present application to have good scalability. When the manufacturer is willing to pay more chip area and power consumption, (under the limitation of hardware wiring capabilities), the performance can be linearly improved. In addition, the side path includes, but is not limited to, ring interconnects, one-way or two-way linear interconnects, and other interconnects with similar principles. The ring interconnect is as shown in Figure 4 and the side path in the embodiment of the present application takes the ring interconnect shown in Figure 4 as an example.

[0052] In one embodiment of the present application, a reconfigurable processor includes a processor that is reconfigurable for all of the following items, or a processor in which some of the following items are reconfigurable and other items are not included in the processor, or a processor in which some of the following items are reconfigurable and other items are not reconfigurable, and further includes a register array, a data memory, an instruction memory, an ALU, an SALU, and a matching module. Each reconfigurable processor is allowed to be reconfigured into a pipeline stage or an RTC processor. The components that can be reused between the pipeline stage and the RTC processor in the embodiments of the present application are as follows:

[0053] 1) Register array. In the pipeline stage, the register array is used to store data packet header fields, metadata, and intermediate signals used in the processing flow. To avoid pipeline stalls, this data is required to be stored in the register array rather than in the RAM. In the RTC processor, the register array is used to store general and special register sets.

[0054] 2) Data memory. In the pipeline stage, the data memory is used to store lookup tables for packets to look up when passing through, and perform corresponding operations according to the lookup results; the data memory is also used to store status tables for packets to read, modify, and write back the status when passing through. In the RTC processor, the data memory is used to store data packet header fields, metadata, lookup tables, and status tables, as well as intermediate signals used in the processing flow. If an operating system is installed in the RTC processor, the data memory is also used to store data structures such as stacks maintained by the operating system.

[0055] 3) Instruction memory. It is used to store instructions in both the pipeline stage and the RTC processor, but the instruction formats are different.

[0056] 4) ALU. It is used to perform arithmetic and logical calculations in both the pipeline stage and the RTC processor.

[0057] 5) SALU. In the pipeline stage, it is used to read the status from the status table in the data memory, perform arithmetic and logical calculations, and write the calculation results back to the status table. In the RTC processor, it is used to read data from the data memory (not limited to the status table), perform arithmetic and logical calculations, and write the calculation results back to the data memory.

[0058] 6) Matching module. It is used to perform lookup and matching in the lookup table of the data memory according to a given keyword in both the pipeline stage and the RTC processor, and return whether there is a matching table entry and the corresponding table entry content.

[0059] It should be noted that in order to implement the reconfigurable processor, the embodiments of the present application reuse the above-mentioned multiple hardware components. Any one of these hardware components is reused to implement the reconfigurable technology, rather than being limited to reusing each reusable component mentioned in the embodiments of the present application.

[0060] In one embodiment of the present application, when the reconfigurable processor is reconfigured into pipeline stages, after the parsed packet header and its carried metadata and other fields (i.e., PHV) are received by the current processor, they are written into the register array of the current processor. According to the keyword selection rule specified by the user program, a part of the fields are selected to form keywords and sent to the matching module. The matching module performs lookup and matching in the lookup table of the data memory according to the given keywords. When there is a matching entry, the operation ID to be executed is determined according to the entry content; when there is no matching entry, the default operation ID specified by the user program is used. According to the operation ID, the corresponding instruction is read from the instruction memory, parsed by a dedicated decoder, and dispatched to the ALU and SALU. The ALU reads the fields participating in the calculation from the register array or instruction parameters, performs the calculation, and writes the result back to the register array. The SALU reads the fields participating in the calculation from the register array or instruction parameters, and reads the status participating in the calculation from the status table in the data memory, performs the calculation, and writes the result back to the register array and / or the status table in the data memory.

[0061] In the above entire processing flow, the PHV of a packet is passed in the registers corresponding to different pipeline beats in the register array to achieve an uninterrupted pipeline function. After the processing is completed, the corresponding PHV is sent out of the current processor.

[0062] In one embodiment of the present application, when the reconfigurable processor is reconfigured into an RTC processor, after the PHV is received by the current processor, it is written into the data memory of the current processor, and the current processor is notified to start executing the processing flow. The RTC processor continuously reads the instructions to be executed from the instruction register according to the PC (Program Counter), until a processing termination instruction is read.

[0063] Each read instruction is parsed by a dedicated decoder and dispatched to the ALU or SALU. The ALU reads the fields participating in the calculation from the register array or instruction parameters, performs the calculation, and writes the result back to the register array. The SALU reads the fields participating in the calculation from the register array or instruction parameters, and reads the data participating in the calculation from the data memory, performs the calculation, and writes the result back to the register array and / or the data memory. The branch jump instruction in the instruction changes the content of the PC.

[0064] The RTC processor can write the extracted key codes into the register array and submit them to the matching module through specific instructions. The matching module works as a coprocessor and writes the result back to the register array. After reading the processing termination instruction, the corresponding PHV is sent out of the current processor.

[0065] It should be noted that a reconfigurable processor can be reconfigured into multiple RTC cores, and the number of convertible RTC cores is determined by the parallelism of the pipeline stage. In the pipeline stage, instructions may be given in the form of VLIW (Very Long Instruction Word), and are parsed and dispatched to multiple ALUs or SALUs for parallel operation. At the same time, the pipeline stage provides multiple memory access channels and multiple functional accelerators (such as hash, usually integrated inside the matching module). Therefore, multiple groups of parallel functional components can be assigned to different cores in the RTC processor, so as to realize the reconfiguration of one pipeline stage into multiple RTC cores.

[0066] The RTC processor does not participate in the process of receiving and sending PHVs, and can process other PHVs while receiving and sending PHVs. Each RTC core contains a PHV buffer, which is used to temporarily store the received PHVs that have not been processed in time, and the PHVs that have been processed but have not been sent in time.

[0067] It should be noted that the above content is the part that takes effect when the reconfigurable processor is reconfigured into the pipeline stage and when the reconfigurable processor is reconfigured into the RTC processor, rather than all parts of the reconfigurable processor. All parts of the reconfigurable processor also include:

[0068] 1. An instruction decoder dedicated to the pipeline stage and an instruction decoder dedicated to the RTC processor;

[0069] 2. Configuration registers dedicated to the pipeline stage, configuration registers dedicated to the RTC processor, and other registers (including the program counter);

[0070] 3. A circuit that connects all the above modules in a specific way;

[0071] 4. Configuration registers for the above circuit. According to the status of the configuration register, a part of the circuit is connected while another part is disconnected. The circuit connected to the switching fabric can enable the reconfigurable processor to be connected to the entire switching fabric as a pipeline stage or an RTC processor.

[0072] Furthermore, the processor reconfigurable programmable switching system also includes: a parser, a payload buffer, an inverse parser, and a hardware load.

[0073] Specifically, the parser, the inverse parser, and each reconfigurable processor are connected to the first side path and hold a corresponding external path node. The first side path includes a plurality of first path nodes. After the PHV undergoes a complete pipeline processing flow, it is sent by the last reconfigurable processor in the pipeline that is reconfigured into a pipeline stage to its corresponding first path node. This node determines whether the PHV needs to be sent to the RTC processor for further processing and to which RTC processor it needs to be sent for further processing based on a dedicated field in the PHV.

[0074] For a PHV that does not need to be sent to the RTC processor for further processing, it starts from the first path node corresponding to the last reconfigurable processor in the pipeline that is reconfigured into a pipeline stage and is directly transmitted along the first side path to the first path node corresponding to the inverse parser, which is the first path.

[0075] For a PHV that needs to be sent to the RTC processor for further processing, it starts from the first path node corresponding to the last reconfigurable processor in the pipeline that is reconfigured into a pipeline stage, is transmitted along the first side path to the corresponding RTC processor, downloaded to the RTC core assigned to it for processing, and after the processing is completed, it selects a cycle when the corresponding first path node is idle, is sent to the first path node and continues to be transmitted forward to the first path node corresponding to the inverse parser, which is the second path.

[0076] Furthermore, the first path and the second path in the embodiments of the present application have a common source point and a common sink point. The source point is the first path node corresponding to the last reconfigurable processor reconfigured into a pipeline stage, and the sink point is the first path node corresponding to the inverse parser. Other nodes are not shared, so as to ensure that packets passing through the first path that do not require complex processing can always be processed with low latency. When the RTC processor reconfigured from a reconfigurable processor contains more than 1 core, a scheduling module is included at the first path node corresponding to each reconfigurable processor to schedule the PHVs sent to the first path node by each RTC core on the reconfigurable processor in a Round-Robin manner.

[0077] In addition, the first path and the second path of the embodiments of the present application have a common sink point, that is, the first path node corresponding to the inverse parser. There is a buffer at the sink point to handle temporary congestion. When this buffer is about to overflow, it sends a signal to the reconfigurable processor reconfigured as an RTC processor, requesting it not to upload new PHVs to the second path. This mechanism enables the data memories for storing PHVs in all RTC processors to be used as buffers for handling sink point congestion. For each data packet entering this buffer, it means that the first path is idle in one cycle, so a data packet on the second path can be scheduled, thus ensuring that the number of data packets on the second path will not increase when this mechanism is triggered, and therefore ensuring that the congestion at the sink point can be resolved without packet loss.

[0078] In an embodiment of the present application, each reconfigurable processor, payload buffer, and other hardware loads (such as shared memory and accelerators, etc.) are connected to the second side path and hold a corresponding second path node. The second side path includes at least one second path node. The second side path is used to control the transmission of signals and data and can support at least one of the following data transmissions:

[0079] 1) Send a request from the reconfigurable processor reconfigured as an RTC processor to read the payload of the data packet processed by the reconfigurable processor reconfigured as an RTC processor in the payload buffer, and transfer the payload back to the reconfigurable processor reconfigured as an RTC processor for implementing DPI (Deep Packet Inspection), where the payload will be stored in the data memory by the reconfigurable processor of the reconfigured RTC processor for subsequent processing.

[0080] 2) Send a request from the reconfigurable processor reconfigured as an RTC processor to read the data in the data memory or register array on other reconfigurable processors, and transfer the read data back to the data memory of the reconfigurable processor reconfigured as an RTC processor that sent the read request for processing.

[0081] 3) Send a request from the reconfigurable processor reconfigured as an RTC processor to write data to the data memory or register array on other reconfigurable processors.

[0082] 4) Send a request from the reconfigurable processor reconfigured as an RTC processor to send data to the hardware load, and the hardware load sends the processing result back to the reconfigurable processor reconfigured as an RTC processor.

[0083] Next, the embodiments of the present application will be illustrated by taking in-network computing applications that the current chip cannot support but the present application can support, including the following application scenarios:

[0084] Application Scenario 1 (In-network Aggregation):

[0085] As the large model technology stands out in various artificial intelligence application scenarios, distributed neural networks are widely deployed in data center networks. The computing nodes in the network need to conduct a large amount of communication to synchronize training parameters, and an important operation among them is to aggregate the gradients calculated by each computing node. The traditional approach is to use a server for aggregation, which will impose a relatively large communication burden on the server network interface, thereby reducing the convergence speed of the entire neural network training. ATP proposed a solution for in-network aggregation on programmable switches, but due to the limitation of the pipeline length of current chips, the in-network aggregation ability of this solution on programmable switches is limited. At the same time, current chips cannot achieve effective functional isolation when used for in-network aggregation. When processing in-network aggregation traffic, it will significantly reduce the ability of the programmable switch to process non-in-network aggregation traffic, resulting in limited local available bandwidth for other uses and reducing the overall service quality of the cloud data center network. Using Trio can break through the pipeline length limitation, but still cannot achieve functional isolation, and in-network aggregation traffic will still significantly reduce the overall performance of the switch. The present invention can break through the limitation of the pipeline length on the aggregation ability by sending in-network aggregation traffic to the RTC core for processing and effectively isolate it from other traffic processed on the pipeline.

[0086] Application Scenario 2 (In-network Caching):

[0087] In-network caching technology is very effective for accelerating the caching of query servers. By caching a portion of the server query results on the switch, if a newly arrived query packet hits the cache, the query result in the cache is directly returned without the need to forward it to the server for querying; if the cache is not hit, when the server returns the query result and passes through the switch, the corresponding query result is written into the cache to replace the outdated entry. NetCache uses a controller to handle the cache replacement function, resulting in a performance bottleneck. P4LRU implements an LRU cache replacement function on the pipeline, but due to the limitations of the pipeline programmability of existing chips, each group of LRU can only support 3 entries, making the hit rate of this method quite different from that of the ideal LRU when the data locality is not particularly good. At the same time, under the existing technology, the memory resources on the pipeline cannot be fully utilized to store in-network cache entries, preventing the in-network cache from fully leveraging the memory potential of the switch. The present invention uses the pipeline as the first-level cache and forwards it to the RTC core to search for the second-level cache when the first-level cache is not hit. With the help of the RTC core's ability to maintain complex data structures, a better cache replacement algorithm can be implemented for the second-level cache on each core; at the same time, the RTC core can use all the available memory on the processor as the cache, enabling more in-network cache entries to be stored on the switch under the same scale of resources; ultimately, significantly improving the hit rate of the in-network cache and further reducing the network load and computing load on the server side.

[0088] Application scenario 3 (network service integration):

[0089] In data center networks and edge networks, to enhance network security, performance, etc., specialized middleboxes and accelerators that support various network functions are usually deployed. For example, a gateway switch requires DDoS detection and mitigation (such as SYN flood detection). Other basic network functions also include server load balancing and network address translation. In addition, telemetry and measurement aimed at improving network visibility (such as heavy hit detection and in-band network telemetry) are becoming indispensable for improving efficiency and ensuring service quality. Deploying multiple single-purpose middleboxes is not only costly but also makes network configuration and management complex. Integrating certain network functions into programmable data plane devices helps reduce costs. On existing pipelined chips, due to the complexity of network functions, packets may need to be looped multiple times, resulting in a huge loss of throughput and latency. On existing chips with a multi-core RTC architecture, the processing performance of the chip will also be significantly reduced due to complex network functions. The present invention combines the pipeline and the multi-core RTC processor organically by partitioning network services between them. Basic and common functions applicable to all traffic are processed by the pipeline, while complex functions applicable only to selected packets are processed by the RTC core. Thanks to the reconfigurable processor design of the present invention, even without prior knowledge of application and traffic distribution, the partitioning of network services between the pipeline and the multi-core RTC processor can be changed through reconfiguration, so that the overall processing performance reaches the optimal.

[0090] According to the programmable switching system with a reconfigurable processor proposed in the embodiments of the present application, through the design of reconfiguring the reconfigurable processor into a pipeline stage or an RTC processor, the programmability of the switch is greatly improved with only a small amount of chip area consumed. Network applications originally deployed on servers and middleboxes can be offloaded to the programmable switch, reducing network communication latency, improving the performance of network services, and reducing the demand for server and middlebox devices in the data center network, thus reducing the construction and maintenance costs of the data center network. Thereby, the problems in the related art such as the limited architecture programmability of programmable switching chips, insufficient throughput, and high network communication latency are solved.

[0091] The embodiments of the present application also provide a programmable data plane chip, including the programmable switching system with a reconfigurable processor in the above embodiments.

[0092] In the description of this specification, the descriptions with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic expressions of the above terms are not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0093] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of these features. In the description of this application, the meaning of "N" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.

[0094] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of code including one or N executable instructions for implementing a customized logic function or process, and the scope of the preferred embodiments of this application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in the reverse order according to the involved functions, rather than in the order shown or discussed, which should be understood by those skilled in the technical field to which the embodiments of this application belong.

[0095] It should be understood that each part of this application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, the steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one of the following well-known technologies in the art or a combination of them can be used: discrete logic circuits having logic gate circuits for implementing logic functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays, field-programmable gate arrays, etc.

[0096] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the methods for implementing the above embodiments can be completed by instructing relevant hardware through a program, and the above program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0097] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. A processor-reconfigurable programmable switching system, characterized in that: include: A first side channel for transmitting a data packet header vector PHV; A second side channel is used to transmit other data and control signals other than the PHV; A reconfigurable processor sequence, the reconfigurable processor sequence comprising a plurality of reconfigurable processors, the first side path and the second side path connecting the plurality of reconfigurable processors, wherein each of the plurality of reconfigurable processors is allowed to be reconfigured into a pipeline stage or run to complete an RTC processor, a data path for transmitting the PHV exists between each adjacent reconfigurable processor, when both ends of the data path are reconfigurable processors reconfigured into pipeline stages, the data path is opened to connect two pipeline stages into a complete pipeline, and when both ends of the data path are not reconfigurable processors reconfigured into pipeline stages, the data path is closed, and the reconfigurable processors reconfigured into RTC processors are connected via the first side path.

2. The processor reconfigurable programmable switching system according to claim 1, characterized in that: The reconfigurable processor includes a processor in which all of the following items are reconfigurable, or a processor in which some of the following items are reconfigurable and the other items are not included in the processor, or a processor in which some of the following items are reconfigurable and the other items are not reconfigurable: A register array, in the pipeline stage, the register array is used to store data packet header fields, metadata and intermediate signals, and in the RTC processor, the register array serves as a register group specified by the instruction system; Data storage, in the pipeline stage, the data storage is used to store lookup tables and state tables, the lookup tables are used for lookups when data packets pass through, and the state tables are used for reading, modifying, and writing back states when data packets pass through. In the RTC processor, the data storage is used to store data packet header fields, metadata, lookup tables and state tables, and intermediate signals. If the RTC processor is equipped with an operating system, the data storage is also used to store data structures related to operating system maintenance; Instruction memory, used to store instructions in both the pipeline stage and the RTC processor; Arithmetic logic unit ALU, used to perform arithmetic logic calculations in both the pipeline stage and the RTC processor; The stateful arithmetic logic unit SALU is used to read the state from the state table of the data memory in the pipeline stage, perform arithmetic logic calculations, and write the calculation results back to the state table; In the RTC processor, it is used to read data from the data memory, perform arithmetic and logical calculations, and write the calculation results back to the data memory; The matching module is used in both the pipeline stage and the RTC processor to search and match in the lookup table of the data memory according to the given keyword, and return whether there is a matching table entry and the corresponding table entry content.

3. The processor reconfigurable programmable switching system according to claim 2, characterized in that: When the reconfigurable processor is reconfigured into a pipeline stage, it includes: After the PHV of the data packet is received by the current processor, it is written into the register array of the current processor; according to the keyword selection rule specified by the user program, a part of the fields is selected to form the keyword and transmitted to the matching module; The matching module searches and matches in the lookup table of the data storage according to the given keywords. When there is a matching table entry, the execution operation ID is determined according to the table entry content; when there is no matching table entry, the default operation ID specified by the user program is used; Read instructions from the instruction memory according to the operation ID, parse the instructions using the decoder, and dispatch the parsed instructions to the ALU and SALU, wherein the ALU reads the fields involved in the calculation from the register array or the instruction parameter, and writes them back to the register array after calculation; the SALU reads the fields involved in the calculation from the register array or the instruction parameter, and reads the states involved in the calculation from the state table in the data memory, and writes them back to the register array and / or the state table in the data memory after calculation; In the processing flow, the PHV of a data packet is transferred between registers corresponding to different pipeline beats in the register array, and after the processing is completed, the corresponding PHV is sent out of the current processor.

4. The processor reconfigurable programmable switching system according to claim 2, characterized in that: When the reconfigurable processor is reconfigured into an RTC processor, it includes: After the PHV of the data packet is received by the current processor, it is written into the data memory of the current processor and notifies the current processor to start executing the processing flow. The current processor reads the instructions to be executed from the instruction register according to the program counter until the processing termination instruction is read; Each instruction read is parsed by a dedicated decoder and dispatched to the ALU or SALU, wherein the ALU reads the fields involved in the calculation from the register array or instruction parameters, and writes them back to the register array after calculation; the SALU reads the fields involved in the calculation from the register array or instruction parameters, and reads the data involved in the calculation from the data memory, and writes them back to the register array and / or data memory after calculation; The matching module is allowed to be connected to or not connected to the RTC processor. If the matching module is connected to the RTC processor, the current processor supports matching instructions, extracts data required by the instructions from the register array and submits it to the matching module, and the matching module writes the matching results back to the register array.

5. The processor-reconfigurable programmable switching system according to any one of claims 1 to 4, characterized in that: The reconfigurable processor allows being reconfigured into multiple RTC cores, and the number of RTC cores is determined according to the parallelism of the pipeline stage.

6. The processor reconfigurable programmable switching system according to claim 1, characterized in that: Also includes: A parser, a payload buffer, a reverse parser and a hardware payload, wherein the parser and the reverse parser are connected to the first side path, the payload buffer and the hardware payload are connected to the second side path, and the parser is also connected to the payload buffer for parsing data packets to obtain PHVs and payloads, and sending the payloads to the payload buffer.

7. The processor reconfigurable programmable switching system according to claim 6, characterized in that: The first side path includes a plurality of first-way nodes, in which the PHV is sent to the corresponding first-way node by the last reconfigurable processor in the pipeline that is reconfigured as a pipeline stage after the PHV passes through the processing flow of the complete pipeline, wherein: For PHVs that do not need to be sent to the RTC processor for processing, they are transmitted along the first path, where the first path starts from the first way node corresponding to the last reconfigurable processor in the pipeline that is reconfigured as a pipeline stage, and is directly transmitted to the first way node corresponding to the inverse resolver along the first side path; For the PHV that needs to be sent to the RTC processor for further processing, it is transmitted along the second path, which starts from the first node corresponding to the last reconfigurable processor in the pipeline that is reconstructed into a pipeline stage, and is transmitted along the first side path to the first node corresponding to the RTC processor where the RTC core assigned to it for processing is located, and downloaded to the RTC core assigned to it for processing. After the processing is completed, the idle cycle of the corresponding first node is selected and uploaded to the first node corresponding to the RTC processor where the RTC core that processes the PHV is located.

8. The processor reconfigurable programmable switching system according to claim 7, characterized in that: The first path and the second path have a common source point and a sink point, wherein the source point is a first-way node corresponding to the last reconfigurable processor reconstructed as a pipeline stage, the sink point is a first-way node corresponding to the inverse parser, and a buffer is processed at the sink point. When the buffer meets an overflow condition, the buffer sends a target signal to the reconfigurable processor reconstructed as an RTC processor, so as not to upload a new PHV to the second path based on the target signal.

9. The processor reconfigurable programmable switching system according to claim 6, characterized in that: The second side path includes at least one second path node, and the second side path supports at least one data transmission, including: Sending a request from a reconfigurable processor reconfigured as an RTC processor, reading a payload of a data packet processed by the reconfigurable processor reconfigured as the RTC processor in a payload buffer, and transferring the payload back to the reconfigurable processor reconfigured as the RTC processor, wherein the payload will be stored in a data memory by the reconfigurable processor of the reconfigured RTC processor for subsequent processing; Sending a request from the reconfigurable processor reconfigured as the RTC processor to read data in a data memory or a register array on other reconfigurable processors, and transferring the read data back to the data memory processed by the reconfigurable processor reconfigured as the RTC processor; Sending a request from the reconfigurable processor reconfigured as the RTC processor to write data to the data memory or register array on the other reconfigurable processors; A request is sent from the reconfigurable processor reconfigured as the RTC processor, data is sent to a hardware load, and the hardware load sends a processing result back to the reconfigurable processor reconfigured as the RTC processor.

10. A programmable data plane chip, characterized in that: A programmable switching system with a processor that can be reconfigured comprising the processor as described in any one of claims 1-9.

Citation Information

Patent Citations

  • Reconfigurable processor chip driven by data stream and reconfigurable processor cluster

    CN116303225A

  • Reconfigurable packet protocol parser equipment of network switching chip with hundred gigabit rate

    CN117880395A