Network Interface Device
The network interface device addresses inefficiencies in data packet processing by employing a configurable hardware module with multiple processing units for parallel operations, enhancing flexibility and efficiency in performing functions like filtering and routing.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- XILINX INC
- Filing Date
- 2019-11-05
- Publication Date
- 2026-04-22
AI Technical Summary
Existing network interface devices lack flexibility and efficiency in processing data packets, particularly in performing functions such as filtering, tunneling, encapsulation, and routing, due to limitations in processing unit configuration and synchronization.
A network interface device with a configurable hardware module comprising multiple processing units, each performing a predetermined operation in a single step, allowing for parallel processing and flexible interconnection to form data processing pipelines for various functions like filtering, tunneling, and routing.
Enhances the processing efficiency and flexibility of data packets by enabling parallel operations and dynamic reconfiguration of processing pipelines, optimizing performance for different network functions.
Smart Images

Figure 0007849969000001 
Figure 0007849969000002 
Figure 0007849969000003
Abstract
Description
Technical Field
[0001] Field This application relates to a network interface device for performing functions related to data packets.
Background Art
[0002] Background Network interface devices are known and are typically used to provide an interface between a computing device and a network. A network interface device can be configured to process data received from the network and / or data placed on the network.
Summary of the Invention
Means for Solving the Problems
[0003] Summary According to one aspect, there is provided a network interface device for interfacing a host device with a network, the network interface device comprising: a first interface configured to receive a plurality of data packets; and a configurable hardware module comprising a plurality of processing units, each processing unit being associated with a predetermined type of operation executable in a single step, wherein at least some of the plurality of processing units are associated with different predetermined types of operations, and the hardware module is configurable to interconnect at least some of the plurality of processing units to provide a first data processing pipeline for processing one or more of the plurality of data packets and performing a first function with respect to the one or more of the plurality of data packets.
[0004] In some embodiments, the first function includes a filtering function. In some embodiments, the function includes at least one of tunneling, encapsulation, and routing functions. In some embodiments, the first function includes an extended Berkley packet filtering function.
[0005] In some embodiments, the first function includes a distributed denial-of-service scrubbing operation. In some embodiments, the first function includes firewall operation.
[0006] In some embodiments, the first interface is configured to receive a first data packet from the network.
[0007] In some embodiments, the first interface is configured to receive a first data packet from a host device.
[0008] In some embodiments, at least two of several of the processing units are configured to perform at least one predetermined operation related to them in parallel.
[0009] In some embodiments, at least two of several of the processing units are configured to perform their associated predetermined types of operations in accordance with a common clock signal of the hardware module.
[0010] In some embodiments, each of at least two of a plurality of processing units is configured to perform its associated predetermined type of operation within a predetermined time length defined by a clock signal.
[0011] In some embodiments, at least two of several processing units are configured to access a first data packet within a predetermined time period and, in response to the end of the predetermined time period, transfer the result of each of the at least one of the above operations to the next processing unit.
[0012] In some embodiments, the results include at least one or more of the following: values from one or more of a plurality of data packets, updates to the map state, and metadata.
[0013] In some embodiments, each of the processing units includes an application-specific integrated circuit configured to perform at least one operation associated with the respective processing unit.
[0014] In some embodiments, each processing unit includes a field-programmable gate array. In some embodiments, each processing unit includes any other type of soft logic.
[0015] In some embodiments, at least one of a plurality of processing units comprises a digital circuit and a memory for storing a state related to the processing performed by the digital circuit, and the digital circuit is configured to communicate with the memory to perform a predetermined type of operation associated with each processing unit.
[0016] In some embodiments, the network interface device includes a memory accessible to two or more of a plurality of processing units, the memory is configured to store state associated with a first data packet, and during the execution of a first function by the hardware module, two or more of the plurality of processing units are configured to access and modify the state.
[0017] In some embodiments, a first processing unit among at least some of the multiple processing units is configured to stall while a second processing unit among the multiple processing units accesses a state value.
[0018] In some embodiments, one or more of the processing units can be individually configured to perform operations specific to their respective pipelines based on their associated predetermined types of operations.
[0019] In some embodiments, the hardware module is configured to interconnect at least some of the plurality of processing units, to receive instructions and provide a data processing pipeline for processing one or more of the plurality of data packets in response to the instructions; to cause one or more of the plurality of processing units to perform a predetermined type of associated operation with respect to one or more data packets; to add one or more of the plurality of processing units to the data processing pipeline; and to remove one or more of the plurality of processing units from the data processing pipeline.
[0020] In some embodiments, a predetermined operation includes at least one of loading at least one value of a first data packet from memory, storing at least one value of a data packet in memory, and performing a lookup in a lookup table to determine an action to be performed with respect to the data packet.
[0021] In some embodiments, a hardware module is configured to receive instructions, and the hardware module can be configured to interconnect at least some of the plurality of processing units to provide a data processing pipeline for processing one or more of the plurality of data packets in response to the instructions, and the instructions include data packets transmitted through a third processing pipeline.
[0022] In some embodiments, one or more of the processing units can be configured to perform a selected operation from among predetermined types of operations associated with one or more of the data packets in response to the instruction.
[0023] In some embodiments, the plurality of components includes a second component of the plurality of components configured to provide a first function in a circuit different from the hardware module, and the network interface device comprises at least one controller configured such that data packets passing through a processing pipeline are processed by the first component of the plurality of components and one of the second component of the plurality of components.
[0024] In some embodiments, the network interface device includes at least one controller configured to issue a command to a hardware module to initiate the execution of a first function for a data packet, the command being configured to cause a first component of a plurality of components to be inserted into the processing pipeline.
[0025] In some embodiments, the network interface device comprises at least one controller configured to issue an instruction to cause a hardware module to initiate execution of a first function on a data packet, the instruction being transmitted through a processing pipeline and including a control message configured to cause a first component among a plurality of components to be activated.
[0026] In some embodiments, for one or more of at least some of a plurality of processing units, at least one associated operation includes at least one of loading at least one value of a first data packet from a memory of the network interface device, storing at least one value of the first data packet in the memory of the network interface device, and performing a lookup in a lookup table to determine an action to be taken with respect to the first data packet.
[0027] In some embodiments, one or more of at least some of the plurality of processing units are configured to pass at least one result of at least one associated predetermined operation to a next processing unit in a first processing pipeline, and the next processing unit is configured to execute a next predetermined operation in response to the at least one result.
[0028] In some embodiments, each of different predetermined types of operations is defined by a different template.
[0029] In some embodiments, the type of a predetermined operation includes at least one of accessing a data packet, accessing a lookup table stored in a memory of a hardware module, performing a logical operation on data loaded from the data packet, and performing a logical operation on data loaded from the lookup table.
[0030] In some embodiments, the hardware module includes routing hardware, and the hardware module can be configured to interconnect at least some of the processing units in order to provide the first data processing pipeline by configuring the routing hardware to route data packets between the processing units in a specific order defined by the first data processing pipeline.
[0031] In some embodiments, the hardware module can be configured to interconnect at least some of the plurality of processing units to provide a second data processing pipeline for processing one or more of the plurality of data packets to perform a second function different from the first function.
[0032] In some embodiments, the hardware module can be configured to interconnect at least some of the plurality of processing units to provide a first data processing pipeline, and then to interconnect at least some of the plurality of processing units to provide a second data processing pipeline.
[0033] In some embodiments, the network interface device includes additional circuitry, separate from the hardware module, configured to perform a first function for one or more of the aforementioned data packets.
[0034] In some embodiments, the further circuitry includes a field-programmable gate array and at least one of a plurality of central processing units.
[0035] In some embodiments, the network interface device comprises at least one controller, and further circuitry is configured to perform a first function on data packets during a compilation process to enable the first function to be performed in a hardware module, and at least one controller is configured to control the hardware module to begin performing the first function on data packets in response to the completion of the compilation process.
[0036] In some embodiments, the further circuitry includes a plurality of central processing units. In some embodiments, at least one controller is configured to control further circuitry to stop executing the first function on data packets in response to the above-mentioned decision that the compilation process for executing the first function in the hardware module is complete.
[0037] In some embodiments, the network interface device comprises at least one controller, the hardware module is configured to perform a first function on data packets during a compilation process to enable the first function to be performed in further circuitry, and the at least one controller is configured to determine that the compilation process to enable the first function to be performed in further circuitry is complete and, in response to the determination, to control the further circuitry to begin performing the first function on data packets.
[0038] In some embodiments, the further circuitry includes a field-programmable gate array.
[0039] In some embodiments, at least one controller is configured to control a hardware module to stop performing the first function on data packets in response to the above-mentioned decision that a compilation process for performing the first function in further circuitry has been completed.
[0040] In some embodiments, the network interface device comprises at least one controller configured to perform a compilation process to enable the first function to be executed in a hardware module.
[0041] In some embodiments, the compilation process includes providing instructions for providing a control plane interface within a hardware module that responds to control messages.
[0042] In another embodiment, a data processing system is provided comprising a network interface device according to the first embodiment and a host device, the data processing system comprising at least one controller configured to perform a compilation process to enable the first function to be performed in a hardware module.
[0043] In some embodiments, at least one controller is provided by one or more of the network interface device and host devices.
[0044] In some embodiments, the compilation process is performed in response to a decision by at least one controller that a computer program representing a first function is to be safely executed in kernel mode on the host device.
[0045] In some embodiments, at least one controller is configured to perform the compilation process by assigning each of at least some of a plurality of processing units to perform at least one operation from a plurality of operations represented by a sequence of computer code instructions in a particular order of a first data processing pipeline, wherein the plurality of operations provide a first function to one or more of a plurality of data packets.
[0046] In some embodiments, at least one controller is configured to send a first instruction to further circuitry of a network interface device to perform a first function on data packets before the compilation process is complete, and to send a second instruction to the hardware module to begin performing the first function on data packets after the compilation process is complete.
[0047] In another embodiment, a method for implementation in a network interface device is provided, the method comprising the steps of receiving a plurality of data packets on a first interface, and configuring a hardware module to interconnect at least some of a plurality of processing units of the hardware module to provide a first data processing pipeline for processing one or more of the plurality of data packets and performing a first function on one or more of the plurality of data packets, wherein each processing unit is associated with a predetermined type of operation that can be performed in a single step, and at least some of the plurality of processing units are associated with different predetermined types of operations.
[0048] In another embodiment, a non-temporary computer-readable medium is provided which includes program instructions for causing a network interface device to carry out the method, the method comprising the steps of receiving a plurality of data packets at a first interface, and configuring a hardware module to interconnect at least some of a plurality of processing units of a hardware module to provide a first data processing pipeline for processing one or more of the plurality of data packets and performing a first function on one or more of the plurality of data packets, each processing unit being associated with a predetermined type of operation that can be performed in a single step, and at least some of the plurality of processing units being associated with different predetermined types of operations.
[0049] In another embodiment, a processing unit is provided, which is connected to a first further processing unit configured to perform at least one predetermined operation on a first data packet received by a network interface device, and to perform at least one further first predetermined operation on the first data packet, and is connected to a second further processing unit configured to perform at least one further second predetermined operation on the first data packet, and is configured to receive the result of the first further at least one predetermined operation from the first further processing unit, perform at least one predetermined operation in accordance with the result of the first further at least one predetermined operation, and transmit the result of at least one predetermined operation to the second further processing unit for processing in the second further at least one predetermined operation.
[0050] In some embodiments, the processing unit is configured to receive a clock signal for timing at least one predetermined operation, and the processing unit is configured to perform at least one predetermined operation in at least one cycle of the clock signal.
[0051] In some embodiments, the processing unit is configured to perform at least one predetermined operation in a single cycle of the clock signal.
[0052] In some embodiments, at least one predetermined operation, a first further at least one predetermined operation, and a second further at least one predetermined operation form part of a function performed on a first data packet received by the network interface device.
[0053] In some embodiments, the first data packet is received from a host device, and the network interface device is configured to interface the host device to the network.
[0054] In some embodiments, a first data packet is received from the network, and the network interface device is configured to interface the host device to the network.
[0055] In some embodiments, the function is filtering. In some embodiments, the filtering function is an extended Berkley packet filtering function.
[0056] In some embodiments, the processing unit includes an application-specific integrated circuit configured to perform at least one predetermined operation.
[0057] In some embodiments, the processing unit includes a digital circuit configured to perform at least one predetermined operation, and a memory for storing a state related to the at least one predetermined operation to be performed.
[0058] In some embodiments, the processing unit is configured to access a memory accessible to a first further processing unit and a second further processing unit, the memory is configured to store the state associated with the first data packet, and at least one predetermined operation includes modifying the state stored in the memory.
[0059] In some embodiments, the processing unit is configured to read the value of the state from memory during a first clock cycle and provide the value to a second further processing unit for modification by a second further processing unit, and the processing unit is configured to stall during a second clock cycle after the first clock cycle.
[0060] In some embodiments, at least one predetermined operation includes at least one of loading a first data packet from the memory of a network interface device, storing the first data packet in the memory of the network interface device, and performing a lookup in a lookup table to determine an action to be performed with respect to the first data packet.
[0061] In another embodiment, a method is provided which is performed in a processing unit, the method comprising: performing at least one predetermined operation with respect to a first data packet received in a network interface device; connecting to a first further processing unit configured to perform a first further at least one predetermined operation with respect to the first data packet; connecting to a second further processing unit configured to perform a second further at least one predetermined operation with respect to the first data packet; receiving the result of the first further at least one predetermined operation from the first further processing unit; performing at least one predetermined operation in accordance with the result of the first further at least one predetermined operation; and transmitting the result of the at least one predetermined operation to the second further processing unit for processing in the second further at least one predetermined operation.
[0062] In another embodiment, a computer-readable non-temporary storage device is provided that, when executed by a processing unit, stores instructions causing the processing unit to perform a method, the method comprising: performing at least one predetermined operation with respect to a first data packet received in a network interface device; connecting to a first further processing unit configured to perform a first further at least one predetermined operation with respect to the first data packet; connecting to a second further processing unit configured to perform a second further at least one predetermined operation with respect to the first data packet; receiving the result of the first further at least one predetermined operation from the first further processing unit; performing at least one predetermined operation in response to the result of the first further at least one predetermined operation; and transmitting the result of the at least one predetermined operation to the second further processing unit for processing in the second further at least one predetermined operation.
[0063] In another embodiment, a network interface device is provided for interface a host device to a network, the network interface device comprising at least one controller, a first interface configured to receive data packets, a first circuit configured to perform a first function on data packets received on the first interface, and a second circuit, wherein the first circuit is configured to perform a first function on data packets received on the first interface during a compilation process to enable the first function to be performed on the second circuit, and at least one controller is configured to determine that the compilation process to enable the first function to be performed on the second circuit is complete, and in response to the determination, to control the second circuit to begin performing the first function on data packets received on the first interface.
[0064] In some embodiments, at least one controller is configured to control the first circuit to stop performing the first function on data packets received at the first interface in response to the above-mentioned decision that the compilation process for performing the first function at the second circuit has been completed.
[0065] In some embodiments, at least one controller is configured to control the first circuit to start executing the first function on data packets of a first data flow received at the first interface, and to stop executing the first function on data packets of a first data flow, in response to the above-mentioned decision that a compilation process for executing the first function in the second circuit has been completed.
[0066] In some embodiments, the first circuit comprises at least one central processing unit, each of which is configured to perform a first function for at least one data packet received at a first interface.
[0067] In some embodiments, the second circuit comprises a field-programmable gate array configured to initiate the execution of a first function for data packets received at the first interface.
[0068] In some embodiments, the second circuit comprises a hardware module having a plurality of processing units, each processing unit associated with at least one predetermined operation, the first interface is configured to receive a first data packet, and the hardware module is configured to cause at least some of the plurality of processing units to perform their associated at least one predetermined operation in a specific order so as to perform the first function on the first data packet, after a compilation process for the first function to be executed in the second circuit.
[0069] In some embodiments, the first circuit comprises a hardware module having a plurality of processing units, each processing unit associated with at least one predetermined operation, the first interface configured to receive first data packets, and the hardware module configured to cause at least some of the plurality of processing units to perform their associated at least one predetermined operation in a specific order to perform the first function on the first data packets during a compilation process to enable the first function to be performed in the second circuit.
[0070] In some embodiments, at least one controller is configured to perform a compilation process for compiling a first function to be performed by a second circuit.
[0071] In some embodiments, at least one controller is configured to instruct a first circuit to perform a first function on data packets received at a first interface before the compilation process is complete.
[0072] In some embodiments, a compilation process for compiling a first function to be performed by a second circuit is performed by a host device, and at least one controller is configured to determine that the compilation process is complete in response to receiving an instruction from the host device that the compilation process is complete.
[0073] In some embodiments, the system includes a processing pipeline for processing data packets received at a first interface, the processing pipeline comprising a plurality of components, each configured to perform one of a plurality of functions on data packets received at the first interface, the first of the plurality of components configured to provide a first function when provided by a first circuit, and the second of the plurality of components configured to provide a first function when provided by at least one second processing unit.
[0074] In some embodiments, at least one controller is configured to control a second circuit to initiate the execution of a first function on data packets received at a first interface by inserting a second component of a plurality of components into a processing pipeline.
[0075] In some embodiments, at least one controller is configured to control the first circuit to stop performing the first function on data packets received at the first interface by removing the first component of a plurality of components from the processing pipeline in response to the above-mentioned decision that the compilation process for performing the first function at the second circuit has been completed.
[0076] In some embodiments, at least one controller is configured to control a second circuit to initiate the execution of a first function on data packets received at a first interface by sending a control message through a processing pipeline to activate a second component of a plurality of components.
[0077] In some embodiments, at least one controller is configured to control the first circuit to stop performing the first function on data packets received at the first interface by sending a control message through a processing pipeline to deactivate the second component of a plurality of components in response to the above-mentioned decision that a compilation process for performing the first function in the second circuit has been completed.
[0078] In some embodiments, a first component among the plurality of components is configured to provide a first function to a first data flow of data packets passing through a processing pipeline, and a second component among the plurality of components is configured to provide a first function to data packets of a second data flow passing through a processing pipeline.
[0079] In some embodiments, the first function includes filtering data packets.
[0080] In some embodiments, the first interface is configured to receive data packets from the network.
[0081] In some embodiments, the first interface is configured to receive data packets from a host device.
[0082] In some embodiments, the compilation time for the first function of the second circuit is longer than the compilation time for the first function of the first circuit.
[0083] In another embodiment, a method is provided, which includes the steps of receiving a data packet on a first interface of a network interface device, and performing a first function on the data packet received on the first interface in a first circuit of the network interface device, wherein the first circuit is configured to perform the first function on the data packet received on the first interface during a compilation process for which the first function is performed in a second circuit, and the method includes the steps of determining that the compilation process for which the first function is performed in the second circuit is complete, and, in response to the determination, controlling the second circuit of the network interface device to begin performing the first function on the data packet received on the first interface.
[0084] In another embodiment, a non-temporary computer-readable medium is provided which includes program instructions for causing a data processing system to carry out a method, the method comprising the steps of receiving a data packet on a first interface of a network interface device, and performing a first function on the data packet received on the first interface in a first circuit of the network interface device, the first circuit being configured to perform the first function on the data packet received on the first interface during a compilation process for which the first function is performed in a second circuit, the method comprising the steps of determining that the compilation process for which the first function is performed in the second circuit is complete, and, in response to the determination, controlling the second circuit of the network interface device to begin performing the first function on the data packet received on the first interface.
[0085] In another embodiment, a non-temporary computer-readable medium is provided, the medium including a program instruction for causing a data processing system to perform a compilation process for compiling a first function to be performed by a second circuit of a network interface device; before the completion of the compilation process, sending a first instruction to the first circuit of the network interface device to perform a first function with respect to data packets received at the first interface of the network interface device; and after the completion of the compilation process, sending a second instruction to the second circuit to initiate the performance of a first function with respect to data packets received at the first interface.
[0086] In some embodiments, the non-temporary computer-readable medium includes program instructions causing a data processing system to perform a further compilation process to compile a first function to be executed by a first circuit, wherein the time required for the compilation process is longer than the time required for the further compilation process.
[0087] In some embodiments, the data processing system includes a host device, and the network interface device is configured to interface the host device with a network.
[0088] In some embodiments, the data configuration system includes a network interface device, which is configured to interface a host device with the network.
[0089] In some embodiments, the data processing system comprises a host device and a network interface device, the network interface device being configured to interface the host device with a network.
[0090] In some embodiments, the first function includes filtering data packets received from the network on the first interface.
[0091] In some embodiments, the non-temporary computer-readable medium includes a configuration program instruction that causes a data processing system to send a third instruction to a first circuit that causes the first circuit to stop performing a function on data packets received at the first interface, after the compilation process is complete.
[0092] In some embodiments, the non-temporary computer-readable medium includes a program instruction causing a data processing system to transmit an instruction to a second circuit to perform a first function on data packets of a first data flow, and an instruction to the first circuit to stop performing the first function on data packets of a first data flow.
[0093] In some embodiments, the first circuit comprises at least one central processing unit, each of which is configured to perform a first function on at least one data packet received at the first interface before the completion of the second compilation process.
[0094] In some embodiments, the second circuit comprises a field-programmable gate array configured to initiate the execution of a first function for data packets received at the first interface.
[0095] In some embodiments, the second circuit comprises a hardware module having a plurality of processing units, each processing unit associated with at least one predetermined operation, data packets received at the first interface include a first data packet, and the hardware module is configured to perform a first function on the first data packet by at least some of the plurality of processing units, each performing its respective at least one operation on the first data packet after the completion of a second compilation process.
[0096] In some embodiments, the first circuit comprises a hardware module having a plurality of processing units configured to provide a first function to a data packet, each processing unit associated with at least one predetermined operation, the data packet received at the first interface includes a first data packet, and the hardware module is configured to perform the first function to the first data packet by at least some of the plurality of processing units, each performing its respective at least one operation to the first data packet, before the completion of a second compilation process.
[0097] In some embodiments, the compilation process includes assigning each of a plurality of processing units of a second circuit to perform, in a specific order, at least one operation associated with one of a plurality of processing stages in a sequence of computer code instructions.
[0098] In some embodiments, a first function provided by a first circuit is provided as a component of a processing pipeline for processing data packets received at a first interface, and a first function provided by a second circuit is provided as a component of a processing pipeline.
[0099] In some embodiments, the first instruction includes an instruction configured such that a first component of a plurality of components is inserted into the processing pipeline.
[0100] In some embodiments, the second instruction includes an instruction configured such that the second component of a plurality of components is inserted into the processing pipeline.
[0101] In some embodiments, the non-transient computer-readable medium includes a configuration program instruction causing a data processing system to send a third instruction to a first circuit after the completion of the compilation process, the third instruction being configured such that a first component of a plurality of components is removed from the processing pipeline.
[0102] In some embodiments, the first instruction includes a control message sent through the processing pipeline to activate a second component of a plurality of components.
[0103] In some embodiments, the second instruction includes a control message sent through the processing pipeline to activate a second component of a plurality of components.
[0104] In some embodiments, the non-transient computer-readable medium includes a program instruction causing a data processing system to send a third instruction to a first circuit after the completion of the compilation process, the third instruction including a control message via a processing pipeline for deactivating the first component of a plurality of components.
[0105] In another embodiment, a data processing system is provided comprising at least one processor and at least one memory containing computer program code, wherein the at least one memory and the computer program code are configured to cause the data processing system to perform a compilation process for compiling functions to be executed by a second circuit of a network interface device using at least one processor; before the compilation process is completed, instructing a first circuit of the network interface device to perform functions on data packets received at the first interface of the network interface device; and after the second compilation process is completed, instructing at least one second processing unit to begin performing functions on data packets received at the first interface.
[0106] In another embodiment, a method is provided for implementation in a data processing system, the method comprising: performing a compilation process for compiling a function to be performed by a second circuit of a network interface device; before the completion of the compilation process, sending a first instruction to a first circuit of the network interface device to perform a function with respect to data packets received on a first interface of the network interface device; and after the completion of the compilation process, sending a second instruction to a second circuit to initiate the performance of a function with respect to data packets received on the first interface.
[0107] In another embodiment, a data processing system is provided with a non-temporary computer-readable medium containing program instructions for assigning each of a plurality of processing units to perform in a particular order at least one operation associated with one of a plurality of processing stages in a sequence of computer code instructions, wherein the plurality of processing stages provide a first function to a first data packet received at a first interface of a network interface device, each of the plurality of processing units is configured to perform one of a plurality of types of processing, at least some of the plurality of processing units are configured to perform different types of processing, and for each of the plurality of processing units, the assignment is performed in accordance with the determination that the processing unit is configured to perform a type of processing suitable for performing each of its at least one operations.
[0108] In some embodiments, each type of processing is defined by one of several templates.
[0109] In some embodiments, the type of processing includes at least one of accessing data packets received by a network interface device, accessing lookup tables stored in the memory of a hardware module, performing logical operations on data loaded from data packets, and performing logical operations on data loaded from lookup tables.
[0110] In some embodiments, two or more of at least some of a plurality of processing units are configured to perform at least one of their associated operations in accordance with a common clock signal of the hardware module.
[0111] In some embodiments, the assignment includes assigning each of at least two of a plurality of processing units to perform at least one of its associated operations within a predetermined time length defined by a clock signal.
[0112] In some embodiments, the assignment includes assigning at least two of several processing units to access a first data packet within a predetermined time period.
[0113] In some embodiments, the assignment includes assigning each of two or more of at least some of a plurality of processing units to transfer the results of at least one of their operations to the next processing unit in response to the end of a period of a predetermined time length.
[0114] In some embodiments, the non-transient computer-readable medium includes program instructions causing a data processing system to perform the task of allocating at least some of a plurality of stages to occupy a single clock cycle.
[0115] In some embodiments, the non-temporary computer-readable medium includes program instructions causing a data processing system to assign two or more of a plurality of processing units to perform at least one of their assigned operations, which are executed in parallel.
[0116] In some embodiments, the network interface device includes a hardware module comprising multiple processing units.
[0117] In some embodiments, the non-temporary computer-readable medium includes computer program instructions causing a data processing system to perform a compilation process including assignment; before the compilation process is complete, to send a first instruction to the circuitry of a network interface device to perform a first function on data packets received on a first interface; and after the compilation process is complete, to send a second instruction to a plurality of processing units to begin performing a first function on data packets received on the first interface.
[0118] In some embodiments, a non-temporary computer-readable medium is included, and for one or more of at least some of a plurality of processing units, at least one assigned operation includes loading at least one value of a first data packet from the memory of a network interface device, storing at least one value of a first data packet in the memory of the network interface device, and performing a lookup in a lookup table to determine an action to be taken with respect to the first data packet.
[0119] In some embodiments, the non-transient computer-readable medium includes computer program instructions causing a data processing system to issue instructions that configure the routing hardware of a network interface device to route a first data packet between multiple processing units in a specific order in order to perform a first function for the first data packet.
[0120] In some embodiments, the first function provided by multiple processing units is provided as a component of a processing pipeline for processing data packets received at the first interface.
[0121] In some embodiments, the non-temporary computer-readable medium includes computer program instructions that cause a data processing system to issue instructions causing components to be inserted into a processing pipeline, thereby causing a plurality of processing units to begin performing a first function on data packets received at a first interface.
[0122] In some embodiments, the non-transient computer-readable medium includes computer program instructions that cause a data processing system to issue instructions causing components to operate within a processing pipeline, thereby causing a plurality of processing units to initiate the execution of a first function on data packets received at a first interface.
[0123] In some embodiments, the data processing system includes a host device, and the network interface device is configured to interface the host device with a network.
[0124] In some embodiments, the data processing system includes a network interface device.
[0125] In some embodiments, the data processing system comprises a network interface device and a host device, wherein the network interface device is configured to interface the host device with a network.
[0126] In another embodiment, a data processing system is provided comprising at least one processor and at least one memory containing computer program code, wherein the at least one memory and the computer program code are configured to cause the data processing system to use at least one processor to assign each of a plurality of processing units to perform in a particular order at least one operation associated with one of a plurality of processing stages in a sequence of computer code instructions, wherein the plurality of processing stages provide a first function for a first data packet received at a first interface of a network interface device, each of the plurality of processing units is configured to perform one of a plurality of types of processing, at least some of the plurality of processing units are configured to perform different types of processing, and for each of the plurality of processing units, the assignment is performed in accordance with the determination that the processing unit is configured to perform a type of processing suitable for performing its respective at least one operation.
[0127] In another embodiment, a method is provided which includes the step of assigning each of a plurality of processing units to perform in a particular order at least one operation associated with one of a plurality of processing stages in a sequence of computer code instructions, wherein the plurality of processing stages provide a first function to a first data packet received on a first interface of a network interface device, each of the plurality of processing units is configured to perform one of a plurality of types of processing, at least some of the plurality of processing units are configured to perform different types of processing, and for each of the plurality of processing units, the assignment is performed in accordance with the determination that the processing unit is configured to perform a type of processing suitable for performing each of the at least one operations.
[0128] The processing units of the hardware module are described as performing those types of operations in a single step. However, those skilled in the art will recognize that this feature is merely a desirable feature and is not essential or indispensable to the functionality of the present invention.
[0129] According to one embodiment, a method is provided which includes the steps of receiving a bitfile description and a program in a compiler, wherein the bitfile description includes a description of routing for a part of a circuit, and compiling the program using the bitfile description to output a bitfile for the program.
[0130] The method may include the step of using the bit file to configure at least a portion of the above-mentioned part of the above-mentioned circuit to perform a function associated with the above-mentioned program.
[0131] The bitfile description may include information about routing between multiple processing units in the aforementioned part of the circuit.
[0132] The bitfile description may include routing information indicating, for at least one of the above-mentioned processing units, at least one of which of the one or more other processing units it can output data to, and which of the one or more other processing units it can receive data from.
[0133] A bitfile description may include routing information that indicates one or more routes between two or more processing units.
[0134] A bitfile description may include information indicating only the roots available to the compiler when compiling a program to provide a bitfile for the program.
[0135] The bitfile may include information for each processing unit indicating at least one of the following: that input should be provided from one or more of the one or more other processing units in the bitfile description of that processing unit; or that output should be provided to one or more of the one or more other processing units in the bitfile description of that processing unit.
[0136] A portion of the circuit may include at least a portion of a configurable hardware module comprising a plurality of processing units, each processing unit associated with a predetermined type of operation that can be performed in a single step, at least some of the plurality of processing units associated with different predetermined types of operation, the bitfile description includes information regarding routing between at least some of the plurality of processing units, and the method may include the step of using the bitfile to cause the hardware to interconnect at least some of the plurality of processing units in order to provide a first data processing pipeline for processing one or more of the plurality of data packets and performing a first function with respect to one or more of the plurality of data packets.
[0137] The bitfile description may be at least a part of the FPGA. The bitfile description may be part of a dynamically programmable FPGA.
[0138] The program may include one of the following: an eBPF program or a P4 program.
[0139] The compiler and FPGA may be located within the network interface device.
[0140] In another embodiment, a device is provided comprising at least one processor and at least one memory containing computer code for one or more programs, wherein the at least one memory and computer code are configured to cause the device, using at least one processor, to receive at least a bitfile description and a program, wherein the bitfile description includes a description of the routing of a portion of a circuit, and to compile the program using the bitfile description to output a bitfile for the program.
[0141] At least one memory and computer code can be configured, using at least one processor, to cause the device to configure at least a portion of the circuit to perform a function associated with the program using the bit file.
[0142] The bitfile description may include information about routing between multiple processing units in the aforementioned part of the circuit.
[0143] The bitfile description may include routing information indicating, for at least one of the above-mentioned processing units, at least one of which of the one or more other processing units it can output data to, and which of the one or more other processing units it can receive data from.
[0144] A bitfile description may include routing information that indicates one or more routes between two or more processing units.
[0145] A bitfile description may include information indicating only the roots available to the compiler when compiling a program to provide a bitfile for the program.
[0146] The bitfile may include information for each processing unit indicating at least one of the following: that input should be provided from one or more of the one or more other processing units in the bitfile description of that processing unit; or that output should be provided to one or more of the one or more other processing units in the bitfile description of that processing unit.
[0147] A portion of the circuit may include at least a portion of a configurable hardware module comprising a plurality of processing units, each processing unit associated with a predetermined type of operation that can be performed in a single step, at least some of the plurality of processing units associated with different predetermined types of operation, the bitfile description includes information relating to routing between at least some of the plurality of processing units, and at least one memory and computer code are configured to cause the device, using at least one processor, to perform the step of interconnecting at least some of the plurality of processing units, using the bitfile, to provide a first data processing pipeline for processing one or more of the plurality of data packets and performing a first function with respect to one or more of the plurality of data packets.
[0148] The bitfile description may be at least a part of the FPGA. The bitfile description may be part of a dynamically programmable FPGA.
[0149] The program may include one of the following: an eBPF program or a P4 program.
[0150] In another embodiment, a network interface device is provided, the network interface device comprising a first interface configured to receive a plurality of data packets, and a configurable hardware module comprising a plurality of processing units, each processing unit being associated with a predetermined type of operation that can be performed in a single step, and a compiler configured to receive a bitfile description containing a routing description of at least a portion of the configurable hardware module and a program, and to compile the program using the bitfile description to output a bitfile for the program, wherein the hardware module is configurable using the bitfile to perform a first function associated with the program.
[0151] A network interface device may also be used to interface a host device to a network.
[0152] At least some of the above-mentioned processing units may be associated with different predetermined types of operations.
[0153] The hardware module may be configured to interconnect at least some of the plurality of processing units to process one or more of the plurality of data packets and to provide a first data processing pipeline for performing a first function on one or more of the plurality of data packets.
[0154] In some embodiments, the first function includes a filtering function. In some embodiments, the function includes at least one of tunneling, encapsulation, and routing functions. In some embodiments, the first function includes an extended Berkley packet filtering function.
[0155] In some embodiments, the first function includes a distributed denial-of-service scrubbing operation. In some embodiments, the first function includes firewall operation.
[0156] In some embodiments, the first interface is configured to receive a first data packet from the network.
[0157] In some embodiments, the first interface is configured to receive a first data packet from a host device.
[0158] In some embodiments, at least two of several of the processing units are configured to perform at least one predetermined operation related to them in parallel.
[0159] In some embodiments, at least two of several of the processing units are configured to perform their associated predetermined types of operations in accordance with a common clock signal of the hardware module.
[0160] In some embodiments, each of at least two of a plurality of processing units is configured to perform its associated predetermined type of operation within a predetermined time length defined by a clock signal.
[0161] In some embodiments, at least two of several processing units are configured to access a first data packet within a predetermined time period and, in response to the end of the predetermined time period, transfer the result of each of the at least one of the above operations to the next processing unit.
[0162] In some embodiments, the results include at least one or more of the following: values from one or more of a plurality of data packets, updates to the map state, and metadata.
[0163] In some embodiments, each of the processing units includes an application-specific integrated circuit configured to perform at least one operation associated with the respective processing unit.
[0164] In some embodiments, each processing unit includes a field-programmable gate array. In some embodiments, each processing unit includes any other type of soft logic.
[0165] In some embodiments, at least one of a plurality of processing units comprises a digital circuit and a memory for storing a state related to the processing performed by the digital circuit, and the digital circuit is configured to communicate with the memory to perform a predetermined type of operation associated with each processing unit.
[0166] In some embodiments, the network interface device includes a memory accessible to two or more of a plurality of processing units, the memory is configured to store state associated with a first data packet, and during the execution of a first function by the hardware module, two or more of the plurality of processing units are configured to access and modify the state.
[0167] In some embodiments, a first processing unit among at least some of the multiple processing units is configured to stall while a second processing unit among the multiple processing units accesses a state value.
[0168] In some embodiments, one or more of the processing units can be individually configured to perform operations specific to their respective pipelines based on their associated predetermined types of operations.
[0169] In some embodiments, the hardware module is configured to interconnect at least some of the plurality of processing units, to receive instructions and provide a data processing pipeline for processing one or more of the plurality of data packets in response to the instructions; to cause one or more of the plurality of processing units to perform a predetermined type of associated operation with respect to one or more data packets; to add one or more of the plurality of processing units to the data processing pipeline; and to remove one or more of the plurality of processing units from the data processing pipeline.
[0170] In some embodiments, a predetermined operation includes at least one of loading at least one value of a first data packet from memory, storing at least one value of a data packet in memory, and performing a lookup in a lookup table to determine an action to be performed with respect to the data packet.
[0171] In some embodiments, a hardware module is configured to receive instructions, and the hardware module can be configured to interconnect at least some of the plurality of processing units to provide a data processing pipeline for processing one or more of the plurality of data packets in response to the instructions, and the instructions include data packets transmitted through a third processing pipeline.
[0172] In some embodiments, one or more of the processing units can be configured to perform a selected operation from among predetermined types of operations associated with one or more of the data packets in response to the instruction.
[0173] In some embodiments, the plurality of components includes a second component of the plurality of components configured to provide a first function in a circuit different from the hardware module, and the network interface device comprises at least one controller configured such that data packets passing through a processing pipeline are processed by the first component of the plurality of components and one of the second component of the plurality of components.
[0174] In some embodiments, the network interface device includes at least one controller configured to issue a command to a hardware module to initiate the execution of a first function for a data packet, the command being configured to cause a first component of a plurality of components to be inserted into the processing pipeline.
[0175] In some embodiments, the network interface device comprises at least one controller configured to issue a command to a hardware module to initiate the execution of a first function for a data packet, the command including a control message which is transmitted through a processing pipeline and is configured to activate a first component of a plurality of components.
[0176] In some embodiments, for one or more of at least some of a plurality of processing units, at least one relevant operation includes at least one of loading at least one value of a first data packet from the memory of a network interface device, storing at least one value of a first data packet in the memory of the network interface device, and performing a lookup in a lookup table to determine an action to be taken with respect to the first data packet.
[0177] In some embodiments, one or more of at least some of a plurality of processing units are configured to pass at least one result of its associated at least one predetermined operation to the next processing unit in a first processing pipeline, and the next processing unit is configured to perform the next predetermined operation in response to at least one result.
[0178] In some embodiments, each of the different predetermined types of operation is defined by a different template.
[0179] In some embodiments, a predetermined type of operation includes at least one of accessing a data packet, accessing a lookup table stored in the memory of a hardware module, performing a logical operation on data loaded from a data packet, and performing a logical operation on data loaded from a lookup table.
[0180] In some embodiments, the hardware module includes routing hardware, and the hardware module can be configured to interconnect at least some of the processing units in order to provide the first data processing pipeline by configuring the routing hardware to route data packets between the processing units in a specific order defined by the first data processing pipeline.
[0181] In some embodiments, the hardware module can be configured to interconnect at least some of the plurality of processing units to provide a second data processing pipeline for processing one or more of the plurality of data packets to perform a second function different from the first function.
[0182] In some embodiments, the hardware module can be configured to interconnect at least some of the plurality of processing units to provide a first data processing pipeline, and then to interconnect at least some of the plurality of processing units to provide a second data processing pipeline.
[0183] In some embodiments, the network interface device includes additional circuitry, separate from the hardware module, configured to perform a first function for one or more of the aforementioned data packets.
[0184] In some embodiments, the further circuitry includes a field-programmable gate array and at least one of a plurality of central processing units.
[0185] In some embodiments, the network interface device comprises at least one controller, and further circuitry is configured to perform a first function on data packets during a compilation process to enable the first function to be performed in a hardware module, and at least one controller is configured to control the hardware module to begin performing the first function on data packets in response to the completion of the compilation process.
[0186] In some embodiments, the further circuitry includes a plurality of central processing units. In some embodiments, at least one controller is configured to control further circuitry to stop executing the first function on data packets in response to the above-mentioned decision that the compilation process for executing the first function in the hardware module is complete.
[0187] In some embodiments, the network interface device comprises at least one controller, the hardware module is configured to perform a first function on data packets during a compilation process to enable the first function to be performed in further circuitry, and the at least one controller is configured to determine that the compilation process to enable the first function to be performed in further circuitry is complete and, in response to the determination, to control the further circuitry to begin performing the first function on data packets.
[0188] In some embodiments, the further circuitry includes a field-programmable gate array.
[0189] In some embodiments, at least one controller is configured to control a hardware module to stop performing the first function on data packets in response to the above-mentioned decision that a compilation process for performing the first function in further circuitry has been completed.
[0190] In some embodiments, the network interface device comprises at least one controller configured to perform a compilation process to enable the first function to be executed in a hardware module.
[0191] In some embodiments, the compilation process includes providing instructions for providing a control plane interface within a hardware module that responds to control messages.
[0192] In another embodiment, a computer implementation method is provided, the method comprising the step of determining routing information for at least some of a configurable hardware module having a plurality of processing units, each processing unit associated with a predetermined type of operation that can be performed in a single step, at least some of the plurality of processing units associated with different predetermined types of operations, and the routing information provides information about available routes between at least the plurality of processing units.
[0193] The configurable hardware module may include substantially static and substantially dynamic parts, and the step of determining above includes determining routing information for the substantially dynamic parts.
[0194] Determining the routing information for the substantially dynamic portion may include determining the routing for the substantially dynamic portion used by one or more of the processing units for the substantially static portion.
[0195] The determination may include analyzing the bitfile descriptions of at least some of the configurable hardware modules in order to determine the routing information described above.
[0196] In another embodiment, a non-temporary computer-readable medium is provided, the medium includes program instructions for determining routing information for at least some of a configurable hardware module comprising a plurality of processing units, each processing unit being associated with a predetermined type of operation that can be performed in a single step, at least some of the plurality of processing units being associated with different predetermined types of operation, and the routing information provides information about available routes between at least the plurality of processing units.
[0197] Computer programs may also be provided that include program code means adapted to perform the method(s). The computer programs may be stored and / or otherwise embodied by a carrier medium.
[0198] Many different embodiments have been described above. It should be understood that further embodiments may be provided by any combination of two or more of the embodiments described above.
[0199] Various other aspects and further embodiments are also described in the following detailed description and appended claims.
[0200] Brief explanation of the drawing Herein, several embodiments will be described as mere examples with reference to the attached drawings. [Brief explanation of the drawing]
[0201] [Figure 1] This is a schematic diagram of a data processing system connected to a network. [Figure 2] This is a schematic diagram of a data processing system that includes a filtering application configured to operate in user mode on a host computing device. [Figure 3] This is a schematic diagram of a data processing system with filtering operations configured to operate in kernel mode on a host computing device. [Figure 4] This is a schematic diagram of a network interface device equipped with multiple CPUs for performing functions related to data packets. [Figure 5] This is a schematic diagram of a network interface device equipped with a field-programmable gate array that runs an application to perform functions related to data packets. [Figure 6] This is a schematic diagram of a network interface device that includes hardware modules for performing functions related to data packets. [Figure 7] This is a schematic diagram of a network interface device comprising a field-programmable gate array and at least one processing unit for performing functions on data packets. [Figure 8] This figure shows a method implemented within a network interface device according to several embodiments. [Figure 9] This figure shows a method implemented within a network interface device according to several embodiments. [Figure 10] This figure shows an example of processing data packets using a series of programs. [Figure 11] This figure shows an example of processing data packets using multiple processing units. [Figure 12] This figure shows an example of processing data packets using multiple processing units. [Figure 13] This figure shows an example of a processing stage pipeline for processing data packets. [Figure 14] This figure shows an example of a slice architecture with multiple pluggable components. [Figure 15] This figure shows an illustrative representation of the configuration and sequence of processing units. [Figure 16] This diagram illustrates an example of how to compile a function. [Figure 17] This figure shows an example of a stateful processing unit. [Figure 18] This figure shows an example of a stateless processing unit. [Figure 19] This figure shows a method according to several embodiments. [Figure 20a] This diagram shows the routing between slices in an FPGA. [Figure 20b] This diagram shows the routing between slices in an FPGA. [Figure 21] This diagram schematically shows the partitions on the FPGA. [Modes for carrying out the invention]
[0202] Detailed explanation The following description is provided to enable those skilled in the art to create and use the present invention, and is provided in the context of a particular application. Various modifications to the disclosed embodiments will be readily apparent to those skilled in the art.
[0203] The general principles defined herein can be applied to other embodiments and uses without departing from the spirit and scope of the invention. Therefore, the invention is not intended to be limited to the embodiments shown, but rather should be given the broadest scope consistent with the principles and features disclosed herein.
[0204] When data is transferred between two data processing systems via a data channel, such as a network, each data processing system has a suitable network interface that enables communication over the channel. Often, the network is based on Ethernet® technology. Data processing systems communicating over a network have a network interface that can support the physical and logical requirements of the network protocol. The physical hardware component of a network interface is called a network interface device or network interface card (NIC).
[0205] Most computer systems include an operating system (OS) for user-level applications to communicate with the network. Part of the operating system, known as the kernel, includes a protocol stack for translating commands and data between applications and device drivers specific to network interface devices. Device drivers can directly control network interface devices. By providing these functions to the operating system kernel, the complexity and differences of network interface devices can be hidden from user-level applications. Network hardware and other system resources (such as memory) can be securely shared by many applications, and the system can be protected from flawed or malicious applications.
[0206] A typical data processing system 100 for performing transmission over a network is shown in Figure 1. The data processing system 100 comprises a host computing device 101 coupled to a network interface device 102 configured to interface the host to a network 103. The host computing device 101 includes an operating system 104 that supports one or more user-level applications 105. The host computing device 101 may also include a network protocol stack (not shown). For example, the protocol stack may be components of the application, a library to which the application is linked, or provided by the operating system. In some embodiments, two or more protocol stacks may be provided.
[0207] The network protocol stack may be a Transmission Control Protocol (TCP) stack. An application 105 can send and receive TCP / IP messages by opening a socket and reading and writing data to the socket, and the operating system 104 ensures that the messages are forwarded over the network. For example, an application can invoke a system call (syscall) to send data to the network 103 via a socket, and then via the operating system 104. This interface for sending messages may be known as a message passing interface.
[0208] Instead of implementing the stack on the host 101, some systems offload the protocol stack to a network interface device 102. For example, if the stack is a TCP stack, the network interface device 102 may have a TCP Offload Engine (TOE) to perform TCP protocol processing. Performing protocol processing within the network interface device 102 rather than the host computing device 101 reduces the demand on the processor(s) of the host system 101. Data transmitted over the network may be sent by application 105 via a TOE-enabled virtual interface driver, partially or completely passing the kernel TCP / IP stack. Thus, data transmitted along this high-speed path only needs to be formatted to meet the requirements of the TOE driver.
[0209] The host computing device 101 may comprise one or more processors and one or more memories. In some embodiments, the host computing device 101 and the network interface device 102 can communicate via a bus, such as the Peripheral Interconnection Express (PCIe bus).
[0210] During the operation of the data processing system, data to be transmitted over the network can be transferred from the host computing device 101 to the network interface device 102 for transmission. In one example, data packets may be transferred directly from the host to the network interface device by the host processor. The host can provide data to one or more buffers 106 located on the network interface device 102. The network interface device 102 can then prepare the data packets and transmit them over the network 103.
[0211] Alternatively, the data may be written to a buffer 107 in the host system 101. The data can then be retrieved from the buffer 107 by a network interface device and transmitted over the network 103.
[0212] In both of these cases, the data is temporarily stored in one or more buffers before being transmitted over the network. The data transmitted over the network may be returned to the host (in a lookback).
[0213] When data packets are sent and received over network 103, there are many processing tasks that can be expressed as operations on the data packets, either on the data packets being sent over the network or on the data packets being received over the network. For example, a filtering process can be performed on received data packets to protect the host system 101 from distributed denial-of-service (DDOS) filtering. Such a filtering process can be performed by simple pack checking or an enhanced Berkley packet filter (eBPF). As another example, encapsulation and forwarding may be performed on data packets being sent over network 103. These processes consume many CPU cycles and can be burdensome for traditional OS architectures.
[0214] Refer to Figure 2, which illustrates one method by which filtering or other packet processing operations may be performed in the host system 220. The processes performed by the host system 220 are shown as running in either user space or kernel space. A receive path exists in kernel space for delivering data packets received from the network by the network interface device 210 to the terminal application 250. This receive path comprises a driver 235, a protocol stack 240, and a socket 245. The filtering operation 230 is performed in user space. Incoming packets provided to the host system 220 by the network interface device 210 bypass the kernel (where protocol processing takes place) and are provided directly to the filtering operation 230.
[0215] Filtering operation 230 is provided with a virtual interface (which may be an Ether Fabric Virtual Interface (EFVI), a Data Plane Development Kit (DPDK), or any other suitable interface) for exchanging data packets with other elements within the host system 220. Filtering operation 230 can perform DDOS scrubbing and / or other forms of filtering. The DDOS scrubbing process can be performed on all packets that are readily recognized as DDOS candidates, e.g., sample packets, packet copies, and packets that have not yet been classified. Packets that are not delivered to filtering operation 230 can be passed directly to driver 235 from the network interface. Operation 230 can provide an Extended Berkeley Packet Filter (eBPF) for performing filtering. If an incoming packet passes the filtering provided by operation 230, operation 230 is configured to reinject the packet into the receive path in the kernel for processing the incoming packet. Specifically, the packet is delivered to driver 235 or stack 240. The packet is then protocol-processed by the protocol stack 240. The packet is then passed to socket 245 associated with the terminating application 250. The terminal application 250 issues a recv() call to retrieve the data packet from the buffer of the associated socket.
[0216] However, this approach has several problems. First, filtering operation 230 runs on the host CPU. In order for filtering 230 to work, the host CPU must process data packets at the rate they are received from the network. If data is sent and received from the network at a high rate, this can constitute a significant loss of processing resources for the host CPU. High data traffic to filtering operation 230 can lead to heavy consumption of other limited resources such as I / O bandwidth and internal memory / cache bandwidth.
[0217] To perform the reinjection of data packets into the kernel, the filtering operation 230 must be provided with a privileged API for performing the reinjection. The reinjection process is complex and may require careful consideration of packet ordering. To perform the reinjection, operation 230 may often require a dedicated CPU core.
[0218] The steps of providing data to the operation and reinjecting it require copying the data to memory and then copying it back from memory. This copying represents a resource burden on the system.
[0219] Similar problems can arise when providing other types of behavior besides filtering for data transmitted and received over a network.
[0220] Some operations (such as DPDK-type operations) may require the processing of packets to be forwarded back onto the network.
[0221] Refer to Figure 3, which illustrates another approach. Similar elements are referred to by similar reference symbols. In this example, an additional layer known as the Express Data Path (XDP) 310 is inserted into the transmit and receive paths in the kernel. The extension to XDP 310 enables insertion into the transmit path. The XDP helper enables packets to be transmitted (as a result of receive operations). XDP 310 is inserted at the operating system driver level, enabling programs to run at this level to perform operations on data packets received from the network before they are protocol-processed by stack 240. XDP 310 also enables programs to run at this level to perform operations on data packets transmitted over the network. Thus, eBPF programs and other programs can operate in both the transmit and receive paths.
[0222] As illustrated in Figure 3, a filtering operation 320 may be injected into XDP from user space to form a program 330, which is part of XDP 310. Operation 320 is injected using the XDP control plane, which will run on the data receiving path, to provide a program 330 that performs a filtering operation (e.g., a DDoS scrub) on packets on the receiving path. Such a program 330 may be an eBPF program.
[0223] Program 330 is shown as being injected into the kernel between driver 235 and protocol stack 240. However, in other examples, program 330 may be injected at other points in the receive path within the kernel. Program 330 may also be part of a separate control path that receives data packets. Program 330 may be provided by an application by providing an extension to the application programming interface (API) of socket 245 for that application.
[0224] The program 330 can, additionally or alternatively, perform one or more operations on data being transmitted over the transmission path. The XDP 310 then invokes the transmit function of the driver 235 to transmit data over the network via the network interface device 210. In this case, the program 330 can provide load balancing or routing operations for data packets to be transmitted over the network. The program 330 can also provide segment reencapsulation and forwarding operations for data packets to be transmitted over the network.
[0225] Program 330 can be used for firewalling, virtual switching, or other operations that do not require protocol termination or application processing.
[0226] One of the advantages of using XDP 310 in this way is that program 330 can directly access the memory buffers that are handled by the driver without intermediate copies.
[0227] In order to insert program 330 into the kernel for operation, it is necessary to ensure that program 330 is safe. If an unsafe program is inserted into the kernel, this introduces certain risks such as infinite loops that can crash the kernel, buffer overflows, uninitialized variables, compiler errors, and performance problems caused by large programs.
[0228] In order to ensure that program 330 is safe before being inserted into XDP 310, a verifier can be operated on the host system 220 to verify the safety of program 330. The verifier can be configured to ensure that no loops exist. Backward jump operations may be permitted as long as they do not create loops. The verifier can be configured to ensure that program 330 has no more than a predetermined number of instructions (e.g., 4000). The verifier can perform a check of register usage validity by traversing the data path of program 330. If there are too many possible paths, program 330 is rejected as unsafe to run in kernel mode. For example, if there are more than 1000 branches, program 330 may be rejected.
[0229] XDP is one example of how a secure program 330 can be installed in the kernel, and it will be understood by those skilled in the art that there are other ways to achieve this.
[0230] The method described above for Figure 3 may be as efficient as the method described above for Figure 2, for example, if the operation can be expressed in a safe (or sandboxed) language necessary for the kernel to execute the code. The eBPF language can run efficiently on x86 processors, and JIT (just-in-time) compilation techniques enable the compilation of eBPF programs into native machine code. The language is designed to be safe, for example, state is restricted to mapping only structures that are shared data structures (such as hash tables). Permitted loops are restricted, and instead, one eBPF program can tail-call another eBPF program. The state space is constrained.
[0231] However, in some implementations, this method can result in significant resource losses on the host system 220 (e.g., I / O bandwidth and internal memory / cache bandwidth, host CPU). The operations on data packets are still performed by the host CPU, which must perform such operations at the rate in which the data is being sent / received.
[0232] Another suggestion is to perform the above operations on the network interface device rather than the host system. Doing so would free up CPU cycles used by the host CPU when performing the operations, in addition to the consumed I / O bandwidth, memory, and cache bandwidth. Shifting the execution of processing operations from the host to the network interface device hardware may present several challenges.
[0233] One proposal for performing processing in network hardware is to provide a network processing unit (NPU) with multiple CPUs specialized for packet processing and / or operational tasks to the network interface device.
[0234] See Figure 4, which shows an example of a network interface device 400 having an array 410 of central processing units (CPUs), such as CPU 420. The CPUs are configured to perform functions such as filtering data packets sent to and received from the network. Each CPU in the array 410 may be an NPU. Although not shown in Figure 4, the CPUs may additionally or alternatively be configured to perform operations such as load balancing on data packets received from a host for transmission over the network. These CPUs are specialized for such packet processing / manipulation operations. The CPUs execute instruction sets optimized for such packet processing / manipulation operations.
[0235] The network interface device 400 is shared among the CPU arrays 410 and further includes memory (not shown) accessible to the arrays.
[0236] The network interface device 400 includes a network medium access control (MAC) layer 430 for interface the network interface device 400 with the network. The MAC layer 430 is configured to receive data packets over the network and to transmit data packets over the network.
[0237] The operation on packets received by the network interface device 400 is parallelized across the CPUs. As shown in the figure, when a data flow is received at the MAC layer 430, it is passed to the spreading function 440, which is configured to extract data packets from the flow and distribute them across multiple CPUs in the NPU 410 for the CPUs to process these data packets, for example, by filtering. The spreading function 440 can parse the received data packets to identify the data flow to which they belong. For each packet, the spreading function 440 generates an instruction for the location of each packet within the data flow to which the packet belongs. The instruction may be, for example, a tag. The spreading function 440 adds the respective instruction to the relevant metadata of each packet. The relevant metadata of each data packet can be appended to the data packet. The relevant metadata can be passed to the spreading function 440 as sideband control information. The instructions are added according to the flow to which the data packets belong so that the order of data packets in any particular flow can be reconstructed.
[0238] After being programmed by multiple CPUs 410, the data packets are then passed to a reordering function 450, which reorders the data flow packets into the correct order before passing them to the host interface layer 460. The reordering function 450 can reorder the data packets in a flow by comparing the instructions (e.g., tags) within the data packets in the flow to reconstruct the order of the data packets. The reordered data packets then traverse the host interface 460 and are delivered to the host system 220.
[0239] Figure 4 shows an array of CPUs 410 that operates only on data packets received from the network, but similar principles (including spreading and reordering) can be applied to data packets received from hosts for transmission over the network, and the array of CPUs 410 performs functions (e.g., load balancing) on these data packets received from hosts.
[0240] The program executed by the CPU may be a compiled or transcoded version of the program executed on the host CPU in the example described above with respect to Figure 3. In other words, the instruction set executed on the host CPU to perform the operation is translated for execution on each CPU array of the dedicated CPUs within the network interface 400.
[0241] To achieve parallelism across CPUs, multiple instances of a program are compiled and executed in parallel on multiple CPUs. Each instance of the program can be responsible for processing a different set of data packets received by a network interface device. However, each individual data packet is processed by a single CPU when providing the program's functionality to that data packet. The overall effect of parallel program execution may be the same as running a single program (e.g., program 330) on the host CPU.
[0242] One of the dedicated CPUs can process data packets at a rate of approximately 50 million packets per second. This operating speed may be slower than that of the host CPU. Therefore, parallelization can be used to achieve the same performance as achieved by running an equivalent program on the host CPU. To perform parallelization, data packets are spread across the CPUs and then reordered after processing by the CPUs. The requirement to process data packets in each flow sequentially, along with the reordering step 450, can introduce a bottleneck, increase memory resource overhead, and limit the available throughput of the device. This requirement and the reordering step 450 can increase device jitter because processing throughput may fluctuate depending on the content of the network traffic and the degree to which parallelism can be applied.
[0243] One advantage of using such a dedicated CPU could be shorter compilation times. For example, it might be possible to compile a filtering application so that it runs on such a CPU in less than a second.
[0244] If this method scales to higher link speeds, there may be issues with the use of the CPU array. The host network interface may need to reach terabits per second in the near future. If such a CPU array 410 is scaled up to these higher speeds, the required power consumption could become a problem.
[0245] Another suggestion is to include a field-programmable gate array (FPGA) in the network interface device and use the FPGA to perform actions on data packets received from the network.
[0246] Refer to Figure 5, which shows an example of the use of FPGA 510 in network interface device 500, having FPGA application 515 for performing operations on data packets received by network interface device 500. Elements similar to those in Figure 4 are referred to by the same reference numerals.
[0247] Figure 5 shows an FPGA application 515 that operates only on data packets received from a network, but such an FPGA application 515 may be used to perform functions (e.g., load balancing and / or firewall functions) on these data packets received from a host for transmission over the network or for return to another network interface on the host or system.
[0248] FPGA application 515 can be provided by compiling a program written in a common system-level language such as C, C++, or Scala to run on FPGA 510.
[0249] The FPGA 510 may have network interface functionality and FPGA functionality. The FPGA functionality can provide FPGA applications 515 that can be programmed into FPGA 510 according to the needs of the network interface device user. FPGA applications 515 can, for example, provide message filtering on the receiving path from network 230 to the host. FPGA applications 515 can provide a firewall.
[0250] FPGA 510 can be programmed to provide FPGA application 515. Some of the network interface device functionality may be implemented as “hard” logic within FPGA 510. For example, the hard logic may be gates of an application-specific integrated circuit (ASIC). FPGA application 515 may be implemented as “soft” logic. The soft logic may be provided by programming FPGA LUTs (lookup tables). The hard logic may be clocked at a higher rate than the soft logic.
[0251] The network interface device 500 includes a host interface 505 configured to send and receive data with a host. The network interface device 520 includes a network medium access control (MAC) interface 520 configured to send and receive data with the network.
[0252] When a data packet is received from the network via the MAC interface 520, the data packet is passed to the FPGA application 515, which is configured to perform functions such as filtering on the data packet. The data packet (if it passes any filtering) is then passed to the host interface 505, and from there to the host. Alternatively, the FPGA application 515 can decide to drop or retransmit the data packet.
[0253] One problem with this approach of using FPGAs to perform functions on data packets is the relatively long compilation time required. FPGAs consist of many logic elements (e.g., logic cells) that individually represent primitive logic operations such as AND, OR, and NOT. These logic elements are arranged and configured in a matrix with programmable interconnections. To provide functionality, these logic cells may need to work together to enforce circuit definitions and synchronous clock timing constraints. Arranging each logic cell and routing between them can be an algorithmically challenging task. Compiling on FPGAs with lower utilization levels may take less than 10 minutes. However, as FPGA devices become more widely used by various applications, the challenges of location and routing can increase, leading to longer compilation times for a given function on the FPGA. Therefore, adding additional logic to an FPGA where a significant portion of its routing resources are already consumed can take several hours to compile.
[0254] One approach is to design the hardware using specific processing primitives, such as parsing primitives, matching primitives, and action primitives. These can be used to construct a processing pipeline in which every packet undergoes each of three processes. First, the packet is parsed to construct a metadata representation of the protocol header. Second, the packet is flexibly matched against rules held in a table. Finally, when a match is found, the packet is acted upon according to the entry from the table selected in the matching operation.
[0255] The P4 programming language (or a similar language) can be used to implement functionality using analysis / matching / action models. The P4 programming language is target-independent, meaning that programs written in P4 can be compiled and run on different types of hardware such as CPUs, FPGAs, ASICs, and NPUs. Each different type of target has its own compiler that maps the P4 source code to the appropriate target switch model.
[0256] P4 can be used to provide a programming model that allows high-level programs to express the packet processing operations of a packet processing pipeline. This approach works well for operations that naturally express themselves in a declarative style. In the P4 language, the programmer expresses the parsing, matching, and action stages as operations performed on an incoming data packet. These operations are grouped together with dedicated hardware for efficient execution. However, this declarative style may not be suitable for expressing programs of an imperative nature, such as eBPF programs.
[0257] Network interface devices may require a sequence of eBPF programs to be executed sequentially. In this case, a chain of eBPF programs calling each other is generated. Each program can modify its state, and the output will appear as if the entire chain of programs had been executed sequentially. It can be difficult for the compiler to collect all the parsing, matching, and action steps. However, even if a chain of eBPF programs is already installed, it may be necessary to install, remove, or modify the chain, which can present further challenges.
[0258] To provide an example of such a program requiring repeated execution, see Figure 10, which shows an example sequence of programs e1, e2, and e3 configured to process data packets. Each program may be, for example, an eBPF program. Each program is configured to parse an incoming data packet, perform a lookup in table 1010 to determine an action in a matching entry in table 1010, and then perform an action on the data packet. The action may include modifying the packet. Each eBPF program may also perform an action depending on the local and shared state. Data packet P0 is first processed by eBPF program e1 before being passed to the next program in the pipeline, e2, and modified. The output of the sequence of programs is the output of the final program in the pipeline, i.e., e3.
[0259] Combining the effects of n such programs into a single P4 program can be complex for a compiler. Furthermore, certain programming models (such as XDP) may require the dynamic insertion and removal of programs at any point in the program sequence in response to changing circumstances.
[0260] According to some embodiments of this application, a network interface device comprising a plurality of processing units is provided. Each processing unit is configured in hardware to perform at least one predetermined operation. Each processing unit includes memory for storing its own local state. Each processing unit includes digital circuitry for modifying this state. The digital circuitry may be an application-specific integrated circuit. Each processing unit is configured to execute a program containing configurable parameters to perform each of the plurality of operations. Each processing unit may be an atom. An atom is defined by specific programming and routing of a predefined template. This defines its specific operational behavior and logical location in the flow provided by the plurality of connected processing units. Where the term “atom” is used herein, it may be understood to refer to a data processing unit configured to perform its operation in a single step; that is, an atom performs its operation as an atom operation.
[0261] An atom can be viewed as a set of hardware structures that can be configured to repeatedly perform one of a series of computations that take one or more inputs and produce one or more outputs.
[0262] Atoms are provided by hardware. Atoms may be configured by a compiler. Atoms may be configured to perform computations.
[0263] During compilation, at least some of the processing units are configured to perform operations such that functions are performed with respect to data packets received by the network interface device by at least some of the processing units. Each of at least some of the processing units is configured to perform at least one predetermined operation with respect to the data packets. In other words, the operations configured to be performed by the connected processing units are performed on the received data packets. The operations are performed sequentially by at least some of the processing units. Collectively, the execution of each of the operations provides a function with respect to the received packets, such as filtering.
[0264] By arranging each atom to perform at least one predetermined operation to execute a function, compilation time can be reduced compared to the FPGA application example described above with respect to Figure 5. Furthermore, by executing functions using processing units specifically dedicated to performing particular operations within the hardware, the speed at which functions can be executed can be improved with respect to using a CPU running software within a network interface device to execute the function of each data packet, as described above with respect to Figure 4.
[0265] Refer to Figure 6, which shows an example of a network interface device 600 according to an embodiment of the present application. The network interface device includes a hardware module 610 configured to perform processing of data packets received on the interface of the network interface device 600. Figure 6 shows a hardware module 610 that performs a function for data packets on the receiving path (e.g., filtering), but the hardware module 610 may also be used to perform a function for data packets on the transmitting path received from a host (e.g., load balancing or firewall).
[0266] The network interface device 600 includes a host interface 620 for sending and receiving data packets with a host, and a network MAC interface 630 for sending and receiving data packets with the network.
[0267] The network interface device 600 comprises a hardware module 610 having a plurality of processing units 640a, 640b, 640c, and 640d. Each processing unit may be an atom processing unit. The term atom is used herein to refer to a processing unit. Each processing unit is configured in hardware to perform at least one operation. Each processing unit comprises a digital circuit 645 configured to perform at least one operation. The digital circuit 645 may be an application-specific integrated circuit. Each processing unit further comprises a memory 650 for storing state information. The digital circuit 645 updates the state information when performing each of the plurality of operations. In addition to local memory, each processing unit can access a shared memory 660 which can also store state information accessible to each of the plurality of processing units.
[0268] The state information in shared memory 660 and / or the state information in the processing unit's memory 650 may include at least one of the following: metadata passed between processing units, temporary variables, data packet contents, and the contents of one or more shared maps.
[0269] Multiple processing units together can provide functions to be performed on data packets received by the network interface device 600. The compiler outputs instructions to configure the hardware module 610 to perform functions on incoming data packets by configuring at least some of the multiple processing units to perform at least one predetermined operation of their respective nature with respect to each incoming data packet. This can be achieved by linking (i.e., connecting) at least some of the processing units 640a, 640b, 640c, and 640d together, so that each of the connected processing units performs at least one operation of its respective nature with respect to each incoming data packet. Each processing unit performs at least one operation of its respective nature in a specific order in order to perform its function. The order may be such that two or more processing units execute in parallel, i.e., simultaneously. For example, one processing unit may read from a data packet during a period (defined by a periodic signal of the hardware module 610 (e.g., a clock signal)) in which a second processing unit also reads from a different location within the same data packet.
[0270] In some embodiments, the data packet passes through each stage represented by a processing unit in sequence. In this case, each processing unit completes its processing before passing the data packet to the next processing unit to perform its processing.
[0271] In the example shown in Figure 6, processing units 640a, 640b, and 640d are connected to each other at compile time, and as a result, each of them performs at least one operation, such as performing a function, filtering, with respect to the received data packets. Processing units 640a, 640b, and 640d form a pipeline for processing data packets. Data packets can travel along this pipeline in multiple stages, each having an equal duration. Durations can be defined according to a periodic signal or beat. Durations may also be defined by a clock signal. Several periods of the clock can define one duration for each stage of the pipeline. Data packets travel along one stage in the pipeline at the end of each occurrence of a repeating duration. Durations may be fixed intervals. Alternatively, each duration of a stage in the pipeline may take a variable amount of time. A signal indicating the next stage in the pipeline may be generated when the previous processing stage has finished its operation, and this may take a variable amount of time. A stall may be introduced at any stage of the pipeline by delaying the signal for some predetermined amount of time.
[0272] Each of the processing units 640a, 640b, and 640d may be configured to access shared memory 660 as part of at least one of their operations. Each of the processing units 640a, 640b, and 640d may be configured to pass metadata to each other as part of at least one of their operations. Each of the processing units 640a, 640b, and 640d may be configured to access data packets received from the network as part of at least one of their operations.
[0273] In this example, processing unit 640c is not used to perform processing of incoming data packets to provide functionality and is therefore omitted from the pipeline.
[0274] Data packets received at the network MAC layer 630 may be passed to the hardware module 610 for processing. Although not shown in Figure 6, the processing performed by the hardware module 610 may be part of a larger processing pipeline that provides additional functionality for data packets beyond the functionality provided by the hardware module 610. This is shown with reference to Figure 14 and will be explained in more detail below.
[0275] The first processing unit 640a is configured to perform at least one first operation on a data packet. This at least one first operation may include at least one of reading from the data packet, reading and writing to shared state in memory 660, and / or performing a lookup in a table to determine an action. The first processing unit 640a is then configured to produce a result from its at least one operation. The result may be in the form of metadata. The result may include modifications to the data packet. The result may include modifications to the shared state in memory 660. The second processing unit 640b is configured to perform at least one operation on the first data packet depending on the result of the operation performed by the first processing unit 640a. The second processing unit 640b produces a result from its at least one operation and passes the result to the third processing unit 640d, which is configured to perform that at least one operation on the first data packet. The first processing unit 640a, the second processing unit 640b, and the third processing unit 640d are all configured to provide functionality related to the data packet. Next, the data packet may be passed to the host interface 620, and then from the host interface to the host system.
[0276] Therefore, it can be seen that the connected processing unit forms a pipeline for processing data packets received at the network interface device. This pipeline can provide the processing of eBPF programs. The pipeline can provide the processing of multiple eBPF programs. The pipeline can provide the processing of multiple modules executed in sequence.
[0277] The interconnection of the processing units within the hardware module 610 may be performed by programming the routing function of the pre-synthesized interconnection fabric of the hardware module 610. This interconnection fabric provides connections between various processing units of the hardware module 610. The interconnection fabric is programmed according to the topology supported by the fabric. Possible exemplary topologies are described below with reference to FIG. 15.
[0278] The hardware module 610 supports at least one bus interface. The at least one bus interface receives data packets at the hardware module 610 (e.g., from a host or network). The at least one bus interface outputs data packets from the hardware module 610 (e.g., to a host or network). The at least one bus interface receives control messages at the hardware module 610. The control messages may be for configuring the hardware module 610.
[0279] The example shown in FIG. 6 has the advantage of reduced compilation time compared to the FPGA application 515 shown in FIG. 5. The hardware module 610 of FIG. 6, for example, may require less than 10 seconds to compile the filtering function. The example shown in FIG. 6 has the advantage of improved processing speed compared to the example of the CPU array shown in FIG. 4.
[0280] An application can be compiled for execution on such a hardware module 610 by mapping a general-purpose program (or multiple programs) to a pre-composed data path. The compiler constructs the data path by linking any number of processing stage instances, each instance being constructed from one of the pre-composed processing stage atoms.
[0281] Each atom is constructed from a circuit. Each circuit can be defined using RTL (Register Transfer Language) or a high-level language. Each circuit is synthesized using a compiler or toolchain. An atom may be synthesized into hard logic and therefore available as a hard (ASIC) resource within a hardware module of a network interface device. An atom may be synthesized into soft logic. Atoms in soft logic can be constrained to assign and maintain location and routing information for the synthesized logic on a physical device. An atom can be designed with configurable parameters that specify the atom's behavior. Each parameter may be a variable that specifies at least one operation to be performed by a processing unit during a clock cycle of the processing pipeline, or even a sequence of operations (microprogram). The logic implementing the atom may be clocked synchronously or asynchronously.
[0282] The atom's processing pipeline itself may be configured to operate according to a periodic signal. In this case, each data packet and metadata moves through one stage along the pipeline in response to each occurrence of the signal. The processing pipeline can operate asynchronously. In this case, a higher level of back pressure in the pipeline causes each downstream stage to begin processing only when data from the upstream stage is presented to it.
[0283] When compiling a function executed by multiple such atoms, a sequence of computer code instructions is separated into multiple operations, each operation being mapped to a single atom. Each operation can represent a single line of disassembled instructions within the computer code instruction. Each operation is assigned to one of the atoms to be executed by one of the atoms. There may be one atom for each representation within the computer code instruction. Each atom is associated with one type of operation and is selected to perform at least one operation within the computer code instruction based on the type of operation it is associated with. For example, an atom may be pre-configured to perform a load operation from a data packet. Thus, such an atom is assigned to execute an instruction in the computer code representing a load operation from a data packet.
[0284] Within a computer code instruction, one atom can be selected per line. Therefore, when implementing a function within a hardware module containing such atoms, there may be 100 such atoms, each performing its own operation to execute the function for its data packet.
[0285] Each atom can be constructed according to one of a set of processing stage templates that determine the type of its associated operation(s). The compilation process is configured to generate instructions to control each atom to perform at least one specific operation based on its associated type. For example, if an atom is pre-configured to perform a packet access operation, the compilation process can assign that atom an operation to load specific information (e.g., the packet's source ID) from the packet header. The compilation process is configured to send instructions to a hardware module, and the atom is configured to perform the operation assigned to it by the compilation process.
[0286] The processing stage templates that specify the behavior of an atom include logical stage templates (e.g., providing behavior for registers, scratchpad memory, and stacks, as well as for branching), packet access state templates (e.g., providing packet data loading and / or packet data storage), and map access stage templates (e.g., map lookup algorithm, map table size).
[0287] The packet access stage may include at least one of the following: reading a byte sequence from a data packet, replacing one byte sequence in a data packet with a different byte sequence, inserting a byte into a data packet, and deleting a byte in a data packet.
[0288] Map access stages can be used to access different types of maps, including direct index arrays and associative arrays (e.g., lookup tables). A map access stage may include at least one of the following: reading a value from a location, writing a value to a location, or replacing a value at a location in the map with a different value. A map access stage may include a comparison operation in which a value is read from a location in the map and compared with a different value. If the value read from that location is less than the different value, a first action may be performed (e.g., do nothing, swap the value at that location with the different value, or add the values together). Otherwise, a second action may be performed (e.g., do nothing, swap or add the values). In either case, the value read from that location may be provided to the next processing stage.
[0289] Each map access stage may be performed in a stateful processing unit. Refer to Figure 17 for an example of circuitry 1700 that may be included in an atom configured to perform the processing of a map access stage. Circuitry 1700 may include a hash function 1710 configured to hash input values used as input to a lookup table. Circuitry 1700 includes a memory 1720 configured to store states associated with the atom's operation. Circuitry 1700 includes an arithmetic logic unit 1730 configured to perform operations.
[0290] A logical stage can perform calculations on values provided by preceding stages. The processing units configured to perform the logical stages may be stateless processing units. Each stateless processing unit can perform simple operations. Each processing unit may, for example, perform 8-bit operations.
[0291] Each logic stage may be implemented in a stateless processing unit. Refer to Figure 18, which shows an example of a circuit 1800 that may be included in an atom configured to perform processing of the logic stages. The circuit 1800 comprises an array of arithmetic logic units (ALUs) and multiplexers. The ALUs and multiplexers are arranged in layers, and the output of one layer of processing by the ALUs is used by the multiplexers to provide input to the next layer of the ALUs.
[0292] The stage pipeline implemented in the hardware module may include a first packet access stage (pkt0), followed by a first logical stage (logic0), followed by a first map access stage (map0), followed by a second logical stage (logic1), followed by a second packet access stage (pkt1), and so on. Therefore, it can take the following form: pkt0->logic0->map0->logic1->pkt1.
[0293] In some cases, stage pkt0 extracts the necessary information from the packet and passes this information to stage logic0. Stage logic0 determines whether the packet is a valid IP packet. In some cases, logic0 forms a map request and sends it to map0, which performs the map operation. Stage map0 may perform an update to the lookup table. Next, stage logic1 collects the results from the map operation and decides whether to drop the packet as a result.
[0294] In some cases, the map request is disabled to cover situations where a map operation should not be performed on this packet. If a map operation is not performed, logic0 tells logic1 whether the packet should be dropped, depending on whether the packet is a valid IP packet. In some examples, the lookup table contains 256 entries, each an 8-bit value.
[0295] The example described here involves only five stages. However, as mentioned above, many more can be used. Furthermore, not all operations need to be performed sequentially; several operations on the same data packet may be performed simultaneously by different processing units.
[0296] The hardware module 610 shown in Figure 6 represents a single pipeline of atoms for performing functions with respect to data packets. However, the hardware module 610 may have multiple pipelines for processing data packets. Each of the multiple pipelines can perform a different function with respect to data packets. The hardware module 610 can be configured to interconnect a first set of atoms of the hardware module 610 to form a first data processing pipeline. The hardware module 610 can also be configured to interconnect a second set of atoms of the hardware module 610 to form a second data processing pipeline.
[0297] A series of steps starting from a sequence of computer code may be executed to compile the functions implemented in a hardware module having a plurality of processing units. A compiler that can be executed on a processor on a host device or a network interface device can access the disassembled sequence of the computer code.
[0298] First, the compiler is configured to split a sequence of computer code instructions into separate stages. Each stage can include operations according to one of the processing stage templates described above. For example, one stage can provide reading from a data packet. One stage can provide an update of map data. Another stage can make a path drop decision. The compiler assigns each of the plurality of operations represented by the code to one of the plurality of stages.
[0299] Second, the compiler is configured to assign each of the processing stages determined from the code to be executed by different processing units. This means that each of at least one operation of each processing stage is executed by a different processing stage. Then, using the output of the compiler, the processing units can be made to execute the operations of each stage in a specific order to perform the function.
[0300] The output of the compiler includes the generated instructions used to cause the processing units of the hardware module to execute the operations associated with each processing stage.
[0301] The output of the compiler may also be used to generate the logic within the hardware module that responds to control messages for configuring the hardware module 610. Such control messages are described in more detail below with respect to FIG. 14.
[0302] The compilation process for compiling functionality to run on the network interface device 600 can be executed in response to the determination that the process providing the functionality is safe to run in the host device's kernel. The determination of program safety can be performed by an appropriate verifier, as described above with respect to Figure 3. Once the process is determined to be safe for execution in the kernel, it can be compiled for execution on the network interface device.
[0303] Refer to Figure 15, which shows representations of at least some of a plurality of processing units that each perform at least one operation to perform a function with respect to a data packet. Such representations can be generated by a compiler and used to configure a hardware module to perform a function. The representations show the order in which the operations may be performed and how some of the processing units perform their operations in parallel.
[0304] Representation 1500 is in the form of a table having rows and columns. Some of the entries in the table represent atoms, e.g., atom 1510a, configured to perform their respective operations. The row to which a processing unit belongs indicates the timing of the operation performed by that processing unit for a particular data packet. Each row can correspond to a single period represented by one or more cycles of a clock signal. Processing units belonging to the same row perform their operations in parallel.
[0305] Input to the logical stage is provided in row 0, and the computation flow proceeds to the following rows. By default, an atom receives the results of processing by an atom in the same column but in a previous row. For example, atom 1510b receives the results of processing by atom 1510a and performs its own processing according to these results.
[0306] When using local routing resources, atoms can also access output from previous atoms in rows where the column number differs by only 2 or less. For example, atom 1510d can receive the results from processing performed by atom 1510c.
[0307] When using global routing resources, atoms can also access output from atoms in the previous two rows and any columns. This can be done using global routing resources. For example, atom 1510f can receive the results from processing performed by atom 1510e.
[0308] These constraints on routing between atoms are given as examples, and other constraints may apply. Applying stricter constraints can make routing information between atoms easier. Applying less restrictive constraints can make scheduling easier. If the number of atoms of a given type (e.g., map, logic, or packet access) is exhausted, or if routing between atoms is not possible, the compilation of the functionality into the hardware module will fail.
[0309] Specific constraints are determined by the topology supported by the interconnect fabric, which is supported by the hardware modules. The interconnect fabric is programmed to cause the atoms of the hardware modules to perform their operations in a specific order and provide data to each other within the constraints. Figure 15 shows an example of how an interconnect fabric can be programmed in this way.
[0310] (As shown in Figure 5) A place-and-route algorithm is used during the synthesis of FPGA application 515 onto the FPGA. However, in this case, the solution space is constrained, and therefore the algorithm has a short bounded execution time.
[0311] A trade-off exists between processing speed or efficiency and compilation time. According to embodiments of the present application, it may be desirable to first compile and run a program on at least one processing unit (which may be a CPU or atom as described above with respect to Figure 6) in order to provide functionality with respect to received data packets. The at least one processing unit can then activate and execute functionality with respect to received data packets during a first period. While the network interface device is operating, a second at least one processing unit (which may be an FPGA application or template-type processing unit as described above with respect to Figure 6) can be configured to execute functionality with respect to data packets. Functionality can then be transferred from the first at least one processing unit to the second at least one processing unit so that the second at least one processing unit can then execute functionality with respect to data packets received by the network interface device. Thus, the slower compilation time of the second at least one processing unit does not prevent the network interface device from executing functionality with respect to data packets before functionality is compiled for the second at least one processing unit, as the first at least one processing unit can be compiled faster and can be used to execute functionality with respect to data packets while functionality is being compiled for the second at least one processing unit. Since the second processing unit typically has a faster processing time, migrating to the second processing unit during compilation enables faster processing of data packets received by the network interface device.
[0312] According to embodiments of the present application, the compilation process can be configured to operate on at least one processor of a data processing system, the at least one processor being configured to send instructions for at least one first processing unit and at least one second processing unit to perform at least one function with respect to a data packet when appropriate. The at least one processor may include a host CPU. The at least one processor may include a control processor on a network interface device. The at least one processor may include a combination of one or more processors on a host system and one or more processors on a network interface device.
[0313] Therefore, at least one processor is configured to perform a first compilation process for compiling functions to be executed by at least one first processing unit of the network interface device. At least one processor is also configured to perform a second compilation process for compiling functions to be executed by at least one second processing unit of the network interface device. Before the completion of the second compilation process, at least one processing unit instructs at least one first processing unit to perform functions with respect to data packets received from the network. Then, after the completion of the second compilation process, at least one processing unit instructs at least one second processing unit to begin performing functions with respect to data packets received from the network.
[0314] By performing these steps, the network interface device can perform functions using at least one first processing unit (which may have a shorter compilation time but slower and / or less efficient processing) while waiting for the second compilation process to complete. Once the second compilation process is complete, the network interface device can perform functions using at least one second processing unit (which may have a longer compilation time but faster and / or more efficient processing) in addition to, or instead of, the first processing unit.
[0315] Refer to Figure 7, which shows an exemplary network interface device 700 according to an embodiment of the present application. Reference elements similar to those shown in the previous figures are indicated by the same reference numerals.
[0316] The network interface device comprises at least one first processing unit 710. The at least one first processing unit 710 may include a hardware module 610, as shown in Figure 6, which comprises multiple processing units. The at least one first processing unit 710 may include one or more CPUs, as shown in Figure 4.
[0317] The function is compiled to run on the first at least one processing unit 710 so that the function is executed by the first at least one processing unit 710 with respect to data packets received from the network during a first period. The first at least one processing unit 710 is instructed by the first at least one processor to execute the function with respect to data packets received from the network before the completion of the second compilation process of the second at least one processing unit.
[0318] The network interface device comprises a second processing unit 720. The second processing unit 720 may comprise an FPGA having an FPGA application (as shown in Figure 5), or it may comprise a hardware module 610 as shown in Figure 6, which comprises multiple processing units.
[0319] During the first period, the second compilation process is performed to compile a function to be executed on at least one second processing unit. That is, the network interface device is configured to compile the FPGA application 515 on the fly.
[0320] After the first period (i.e., after the completion of the second compilation process), at least one second processing unit 720 is configured to begin executing functions related to data packets received from the network.
[0321] After the first period, the first at least one processing unit 710 may stop performing functions related to data packets received from the network. In some embodiments, the first at least one processing unit 710 may partially stop performing functions related to data packets. For example, if the first at least one processing unit includes multiple CPUs, after the first period, one or more CPUs may stop performing processing related to data packets received from the network, while the remaining CPUs of the multiple CPUs continue to perform processing.
[0322] The first processing unit 710 can be configured to perform functions with respect to data packets of the first data flow. Once the second compilation process is complete, the second processing unit 720 can begin performing functions with respect to data packets of the first data flow. Once the second compilation process is complete, the first processing unit can stop performing functions with respect to data packets of the first data flow.
[0323] Different combinations are possible for the first and second processing units. For example, in some embodiments, the first and second processing units 710 include a plurality of CPUs (as shown in Figure 4), while the second and second processing units 720 include a hardware module having a plurality of processing units (as shown in Figure 6). In some embodiments, the first and second processing units 720 include an FPGA (as shown in Figure 5). In some embodiments, the first and second processing units 720 include a hardware module having a plurality of processing units (as shown in Figure 6), while the second and second processing units 720 include an FPGA (as shown in Figure 5).
[0324] Refer to Figure 11, which illustrates how multiple connected processing units 640a, 640b, and 640d can perform at least one of their respective operations on data packets. Each processing unit is configured to perform at least one of its respective operations on incoming data packets.
[0325] At least one operation of each processing unit can represent a logical stage within a function (e.g., a function of an eBPF program). At least one operation of each processing unit can be represented by an instruction executed by the processing unit. The instruction can determine the behavior of an atom.
[0326] Figure 11 shows how a packet (P0) progresses through the processing stages performed by each processing unit.
[0327] Each processing unit performs processing on packets in a specific order specified by the compiler. The order may be such that some processing units are configured to perform their processing in parallel. This processing may include accessing at least a portion of the packets held in memory. Additionally or alternatively, this processing may include performing lookups to a lookup table to determine what action should be taken on the packets. Additionally or alternatively, this processing may include modifying state 1110.
[0328] The processing units exchange metadata M0, M1, M2, and M3 with each other. The first processing unit 640a is configured to perform at least one predetermined operation and generate metadata M1 accordingly. The first processing unit 640a is configured to pass metadata M1 to the second processing unit 640b.
[0329] At least some of the processing units perform at least one operation depending on at least one of the following: the contents of the data packet, its own stored state, the global shared state, and metadata associated with the data packet (e.g., M0, M1, M2, M3). Some of the processing units may be stateless.
[0330] Each processing unit can perform its associated type of operation for a data packet (P0) within at least one clock cycle. In some embodiments, each processing unit can perform its associated type of operation within a single clock cycle. Each processing unit may be clocked individually to perform its operations. This clocking can be added to the clocking of the processing pipeline of the processing units.
[0331] A more detailed examination of the operation of the second processing unit 640b reveals that it is configured to connect to a first processing unit 640a, which is configured to perform at least one predetermined operation on a first data packet. The second processing unit 640b is configured to receive the result of the at least one predetermined operation from a further processing unit. The second processing unit 640b is configured to perform at least one second predetermined operation in response to the result of the at least one predetermined operation. The second processing unit 640b is configured to connect to a third processing unit 640d, which is configured to perform at least one third predetermined operation on a first data packet. The second processing unit 640b is configured to transmit the result of the at least one predetermined operation to the third processing unit 640d for processing in the at least one predetermined operation.
[0332] The processing units can similarly operate in an order that provides functionality for each of the multiple data packets.
[0333] The embodiments of this application are such that, if the functionality allows, multiple packets can be pipelined simultaneously.
[0334] Refer to Figure 12, which illustrates the pipelined processing of data packets. As shown, different packets may be processed simultaneously by different processing units. The first processing unit 640a performs at least one operation on the third data packet (P2) at the first time point (t0). The second processing unit 640b performs at least one operation on the second data packet (P1) at the first time point (t0). The third processing unit 640d performs at least one operation on the first data packet (P0) at the first time point (t0).
[0335] After each processing unit has performed at least its respective operation, each packet moves along one stage in the sequence. For example, at a subsequent second time point (t1), the first processing unit 640a has performed at least its respective operation at the first time point (t0) for the fourth data packet (P3). The second processing unit 640b has performed at least its respective operation at the first time point (t0) for the third data packet (P2). The third processing unit 640d has performed at least its respective operation at the first time point (t0) for the first data packet (P1).
[0336] It should be understood that in some embodiments, multiple packets may exist in a given stage.
[0337] In some embodiments, packets can move from one stage to the next without necessarily having to go through a lock step.
[0338] As long as there are no pipeline hazards, such pipelines operating at a fixed clock can have a constant bandwidth. This can reduce jitter in the system.
[0339] To avoid risks when executing instructions (such as conflicts when accessing shared state), each processing unit may be configured to execute no-action (i.e., the processing unit stalls) instructions as needed.
[0340] In some embodiments, operations (such as simple arithmetic, increments, addition / subtraction of constant values, shifts, and addition / subtraction of values from data packets or metadata) require one clock cycle to be performed by the processing unit. This may mean that a shared state value needed by one processing unit has not yet been updated by another processing unit. Thus, the old values of the shared state 1110 may be read by the processing unit that needs them. Therefore, there is a potential risk when reading and writing values to the shared state. On the other hand, operations on intermediate values can be passed as metadata without any risk.
[0341] An example of a risk to read from and write to the shared state 1110 that can be avoided can be given in the context of an increment operation. Such an increment operation may be an operation that increments a packet counter in the shared state 1110. In one embodiment of the increment operation, during a first time slot of the pipeline, a second processing unit 640b is configured to read the value of the counter from the shared state 1110 and provide the output of this read operation (e.g., as metadata M2) to a third processing unit 640d. The third processing unit 640d is configured to receive the value of the counter from the second processing unit 640b. During the second time slot, the third processing unit 640d increments this value and writes the newly incremented value to the shared state 1110.
[0342] Problems can arise when performing such an increment operation, namely, if the second processing unit 640b attempts to access the counter stored in the shared state 1110 during the second time slot, the second processing unit 640b may read the previous value of the counter before the counter value in the shared state 1110 is updated by the third processing unit 640d.
[0343] Therefore, to address this problem, the second processing unit 640b may be stalled during the second time slot (through execution by the second processing unit 640b of a no-action instruction or a pipeline bubble). The stall can be understood as a delay in the execution of the next instruction. This delay can be achieved by executing a "no-action" instruction instead of the next instruction. The second processing unit 640b then reads the counter value from the shared state 1110 during the subsequent third time slot. During the third time slot, the counter in the shared state 1110 is updated, so it is guaranteed that the second processing unit 640b will read the updated value.
[0344] In some embodiments, each atom is configured to read from a state, update its state, and write the updated state within a single pipeline time slot. In this case, stalling the processing unit described above is not necessary. However, stalling the processing unit can reduce the cost of the required memory interface.
[0345] In some embodiments, to avoid danger, processing units in a pipeline may wait until other processing units in the pipeline have finished their processing before performing their own operations.
[0346] As described above, the compiler constructs a data path by linking any number of processing stage instances, each instance being constructed from one of a predetermined number (three in the given example) of pre-composed processing stage templates. These processing stage templates include logical stage templates (e.g., providing operations on registers, scratchpad memory, and metadata), packet access state templates (e.g., providing packet data loading and / or packet data storage), and map access stage templates (e.g., map lookup algorithm, map table size).
[0347] Each processing stage instance may be implemented by a single processing unit; that is, each processing stage includes at least one operation performed by the processing unit.
[0348] Figure 13 shows an example of how processing stages can be connected to each other within pipeline 1300 to process incoming data packets. As shown in Figure 13, the first data packet is received and stored in FIFO 1305. In the first logical stage 1310, one or more call arguments are received. The call arguments may include a program selector that identifies a function to be performed on the incoming data packet. The call arguments may include an indication of the packet length of the incoming data packet. The first logical stage 1310 is configured to process the call arguments and provide output to the first packet access stage 1315.
[0349] The first packet access stage 1315 loads data from the first packet at the network tap 1320. The first packet access stage 1315 can also write data to the first packet according to the output of the first logical stage 1310. The first packet access stage 1315 can write data to the beginning of the first data packet. The first packet access stage 1315 can overwrite data within the data packet.
[0350] The loaded data and any other metadata and / or arguments are then provided to the second logical stage 1325, which processes the first data packet and provides the output arguments to the first map access stage 1330. The first map access stage 1330 uses the output from the second logical stage 1325 to perform a lookup into a lookup table and determine the action to be performed on the first data packet. The output is then passed to the third logical stage 1335, which processes this output and passes the result to the second packet access stage 1340.
[0351] The second packet access stage 1340 can read data from the first data packet and / or write data to the first data packet, depending on the output of the third logical stage 1335. The result of the second packet access stage 1340 is then passed to a fourth logical stage 1345, which is configured to perform processing on the incoming input.
[0352] The pipeline may include multiple packet access stages, logical stages, and map access stages. The final logical stage 1350 is configured to output a return argument. The return argument may include a pointer that identifies the start of a data packet. The return argument may include instructions for an action to be performed on the data packet. The action instructions may indicate whether the packet should be dropped or not. The action instructions may indicate whether the packet should be forwarded to the host system or not. The network interface device may include at least one processing unit configured to drop each data packet in response to an instruction that the packet should be dropped.
[0353] The pipeline 1300 may further include one or more bypass FIFOs 1355a, 1355b, 1355c. The bypass FIFOs may be used to pass processing data, such as data from a first data packet, around the map access stage and / or packet access stage. In some embodiments, the map access stage and / or packet access stage do not require data from the first data packet to perform at least one of their respective operations. The map access stage and / or packet access stage may perform at least one of their respective operations depending on the input arguments.
[0354] Refer to Figure 8, which shows a method 800 performed by network interface devices 600, 700 according to embodiments of this application.
[0355] In the S810, a functional hardware module of a network interface device is configured to perform a function. The hardware module comprises a plurality of processing units, each configured to perform a certain operation in hardware with respect to a data packet. The S810 includes configuring at least some of the plurality of processing units to perform their respective predetermined types of operations in a specific order to provide functionality for each incoming data packet. Configuring the hardware module in this way includes connecting at least some of the plurality of processing units so that incoming data packets are processed by each of at least some of the plurality of operations of the plurality of processing units. Connecting can be achieved by configuring the routing hardware of the hardware module to route data packets and associated metadata between processing units.
[0356] In S820, the first data packet is received from the network on the first interface of the network interface device.
[0357] In S830, the first data packet is processed by each of at least several processing units connected during the compilation process of S810. Each of at least several processing units performs a type of operation that it has pre-configured to perform for at least one data packet. Thus, a function is performed for the first data packet.
[0358] In S840, the first processed data packet is forwarded toward its destination. This may include sending the data packet to a host. This may include sending the data packet across a network.
[0359] Refer to Figure 9, which shows a method 900 that can be performed in the network interface device 700 according to an embodiment of this application.
[0360] In S910, at least one first processing unit (i.e., a first circuit) of the network interface device is configured to receive and process data packets received over the network. This processing includes performing functions with respect to the data packets. The processing is performed during a first period.
[0361] In S920, a second compilation process is performed during the first period to compile a function for execution on at least one second processing unit (i.e., a second circuit).
[0362] In S930, it is determined whether the second compilation process is complete. If it is not complete, the process returns to S910 and S920, and at least one of the first processing units continues to process data packets received from the network, while the second compilation process continues.
[0363] In S940, in response to the decision that the second compilation is complete, at least one of the first processing units stops executing functions related to the received data packets. In some embodiments, at least one of the first processing units may stop executing functions only with respect to specific data flows. Then, at least one of the second processing units may instead execute functions with respect to those specific data flows (S950).
[0364] In S950, once the second compilation process is complete, at least one second processing unit is configured to begin executing functions related to data packets received from the network.
[0365] Refer to Figure 16 illustrating Method 1600 according to an embodiment of this application. Method 1600 can be performed on a network interface device or a host device.
[0366] In S1610, a compilation process is performed to compile a function so that it is executed by at least one first processing unit.
[0367] In S1620, a compilation process is performed to compile a function to be executed by at least one second processing unit. This process includes assigning each of several processing units of the at least one second processing unit to perform at least one operation associated with one of several stages for processing data packets to provide a first function. Each of the several processing units is configured to perform a certain type of processing, and the assignment is made in accordance with the determination that each processing unit is configured to perform a type of processing suitable for performing its respective at least one operation. In other words, the processing units are selected according to their templates.
[0368] In 1630, prior to the completion of the compilation process in S1620, an instruction is sent to at least one first processing unit to perform a function. This instruction may be sent before the compilation process in S1620 begins.
[0369] In S1640, after the compilation process in S1620 is completed, an instruction is sent to the second circuit to cause it to perform a function related to the data packet. This instruction may include the compiled instruction generated in S1620.
[0370] The functionality according to embodiments of this application may be provided as a pluggable component of a processing slice within a network interface. Refer to Figure 14, which shows an example of how slice 1425 may be used in the network interface device 600. Slice 1425 may be referred to as a processing pipeline.
[0371] The network interface device 600 includes a transmit queue 1405 for receiving and storing data packets from a host that are processed by slice 1425 and then transmitted over the network. The network interface device 600 includes a receive queue 1410 for storing data packets received from network 1410 that are processed by slice 1425 and then delivered to a host. The network interface device 600 includes a receive queue 1415 for storing data packets received from the network that are processed by slice 1425 and intended for delivery to a host. The network interface device 600 includes a transmit queue for storing data packets received from a host that are processed by slice 1425 and intended for delivery to the network.
[0372] Slice 1425 of the network interface device 600 comprises multiple processing functions for processing data packets on the receive and transmit paths. Slice 1425 may include a protocol stack configured to perform protocol processing of data packets on the receive and transmit paths. In some embodiments, multiple slices may exist within the network interface device 600. At least one of the multiple slices may be configured to process received data packets received from the network. At least one of the multiple slices may be configured to process transmitted data packets for transmission over the network. The slices may be implemented by hardware processing devices such as at least one FPGA and / or at least one ASIC.
[0373] Accelerator components 1430a, 1430b, 1430c, and 1430d can be inserted into different stages within the slice as shown in the figure. Each accelerator component provides functionality related to data packets traversing the slice. Accelerator components can be inserted or removed on the fly, i.e., while the network interface device is operating. Therefore, accelerator components are pluggable components. The accelerator components are logical regions and are allocated to slice 1425. Each of them supports a streaming packet interface that allows packets traversing the slice to be streamed into and out of the component.
[0374] For example, one type of accelerator component may be configured to provide encryption of data packets on the receiving or transmitting path. Another type of accelerator component may be configured to provide decryption of data packets on the receiving or transmitting path.
[0375] The functions described above, provided by performing operations carried out by multiple connected processing units (as described above with reference to Figure 6), can be provided by accelerator components. Similarly, functions provided by arrays of network processing CPUs (as described above with reference to Figure 4) and / or FPGA applications (as described above with reference to Figure 5) may be provided by accelerator components.
[0376] As described, during the operation of the network interface device, processing performed by at least one first processing unit (such as multiple connected processing units) may be migrated from at least one second processing unit. To carry out this migration, components of slice 1425 that are for processing by at least one first processing unit can be replaced with components that are for processing by at least one second processing unit.
[0377] The network interface device may include a control processor configured to insert and remove components from slice 1425. During the first period described above, components from the execution of functions by the first at least one processing unit may be present in slice 1425. After the first period, the control processor may be configured to remove pluggable components that provide functionality by the first at least one processing unit from slice 1425 and to insert pluggable components that provide functionality by the second at least one processing unit into slice 1425.
[0378] In addition to, or instead of, inserting and removing components from slices, the control processor can load programs into components and issue control plane commands to control the flow of frames to those components. In this case, components may be operated without being inserted into or removed from the pipeline, or they may not be operated at all.
[0379] In some embodiments, control plane or configuration information is carried over the data path without requiring a separate control bus. In some embodiments, requests to update the configuration of data path components are encoded as messages carried over the same bus as network packets. Thus, the data path can carry two types of packets: network packets and control packets.
[0380] Control packets are formed by the control processor and injected into slice 1425 using the same mechanism used to send or receive data packets using slice 1425. This same mechanism may be a transmit queue or a receive queue. Control packets can be distinguished from network packets in any suitable way. In some embodiments, different types of packets may be distinguished by one or more bits in a metadata word.
[0381] In some embodiments, a control packet includes a routing field in a metadata word that determines the path the control packet takes through slice 1425. A control packet can carry a sequence of control commands. Each control command may target one or more components of slice 1425. Each data path component is identified by a component ID field. Each control command encodes a request for each identified component. The request may involve making a change to the configuration of that component. The request can control whether the component is activated, i.e., whether the component performs its function with respect to data packets traversing the slice.
[0382] Therefore, in some embodiments, the control processor of the network interface device 600 is configured to send a message to one of the slice components to initiate the execution of a function with respect to data packets received in the network interface device. This message is a control plane message sent through a pluggable component that causes an atomic switchover of the frame to the component to perform the function. This component then performs the function for all received data packets traversing the slice until it is switched out. The control processor is configured to send a message to another of the slice components to cause this component to stop the execution of a function with respect to data packets received in the network interface device 600.
[0383] Sockets can be located at various points in the ingress and egress data paths to switch components in and out of data slice 1425. The control processor can inspect additional logic for in and out of slice 1425. This additional logic can take the form of a FIFO (First-In, First-Out) placed between components.
[0384] The control processor can send control plane messages to the configured components of slice 1425 through slice 1425. The configuration can determine the functions performed by the components of slice 1425. For example, a control message sent through slice 1425 can configure a hardware module to perform a function with respect to data packets. Such a control message can cause atoms of the hardware module to be interconnected in the pipeline of the hardware module to provide a specific function. Such a control message can cause individual atoms of the hardware module to select an operation to be performed by an individually selected atom. Since each atom is pre-configured to perform a certain type of operation, the selection of an operation for each atom depends on the type of operation that each atom is pre-configured to perform.
[0385] Next, several further embodiments will be described with reference to Figures 19 to 21. In these embodiments, a packet processing program or feedforward pipeline is operated in the FPGA. A method for implementing the packet processing program or feedforward pipeline in a subunit of the FPGA will be described. The packet processing program or feedforward pipeline may be an eBPF program, a P4 program, or any other suitable program.
[0386] This FPGA may be provided in a network interface device. In some embodiments, the packet processing program is deployed or run only after the network interface device has been installed on its host.
[0387] A packet processing program or feedforward pipeline can implement a loop-free logical flow.
[0388] In some embodiments, the program may be written in a non-privileged domain, such as the user level, or a lower privileged domain. The program may run in a privileged domain, such as the kernel, or a higher privileged domain. The hardware running the program may require that there be no arbitrary loops.
[0389] The following embodiments refer to example eBPF programs. However, it should be understood that other embodiments may be used with any other suitable programs.
[0390] It should be understood that one or more of the following embodiments can be used in combination with one or more of the embodiments described above.
[0391] Some embodiments may be provided in the context of FPGAs, ASICs, or any other suitable hardware devices. Some embodiments utilize subunits such as FPGAs or ASICs. The following examples are described with reference to FPGAs. It should be understood that similar processes may be performed by ASICs or any other suitable hardware devices.
[0392] Subunits may also be atoms. Several examples of atoms have been mentioned above. It should be understood that any of the aforementioned examples of atoms may be used as subunits, either alternatively or additionally. Alternatively or additionally, these subunits may be called “slices” or configurable logic blocks.
[0393] Each of these subunits may be configured to execute a single instruction or multiple related instructions. In the latter case, the related instructions may provide a single output (which may be defined by one or more bits).
[0394] A subunit can be thought of as a computing unit. Subunits may be arranged in a pipeline in which packets are processed sequentially. In some embodiments, subunits can be dynamically assigned to execute each instruction (or multiple instructions) in a program.
[0395] In some embodiments, a subunit may be all or part of a unit used to define, for example, a block of an FPGA. In some FPGAs, a block of the FPGA is called a slice. In some embodiments, a subunit or atom is equivalent to a slice.
[0396] By mapping each atom or subunit to a corresponding block or slice of the FPGA, improved resource utilization can be achieved compared to methods that map RTL atoms to FPGA resources. As a result of such a latter method, RTL atoms may require a relatively large number of individual blocks or slices of the FPGA.
[0397] In some embodiments, compilation may be at the atom level. This may have the advantage of pipelined processing. Packets can be processed sequentially. The compilation process can be executed relatively quickly.
[0398] In some embodiments, arithmetic operations may require one slice per byte. Logical operations may require half a slice per byte. Shift operations may require a set of slices depending on the width of the shift operation. Comparison operations may require one slice per byte. Selection operations may require half a slice per byte.
[0399] As part of the compilation process, placement and routing are performed. Placement is the assignment of specific physical subunits to execute a particular instruction or set of instructions. Routing ensures that one or more outputs of a particular subunit are routed to the correct destination, which may be, for example, another one or more subunits.
[0400] Deployment and routing can utilize a process in which operations are assigned to specific subunits starting from one end of the pipeline. In some embodiments, the most important operations may be placed before less important operations. In some embodiments, routing may be assigned simultaneously with the deployment of specific operations. In some embodiments, routes may be selected from a limited set of pre-calculated routes, which will be discussed in more detail later.
[0401] In some embodiments, if a route cannot be assigned, the operation is postponed for later.
[0402] In some embodiments, the pre-calculated route may be a byte-width route. However, this is merely an example, and in other embodiments, different route widths may be defined. In some embodiments, multiple routes of different sizes may be provided.
[0403] In some embodiments, routing may be limited to routing between nearby subunits.
[0404] In some embodiments, the subunits may be physically arranged on the FPGA in a regular structure.
[0405] In some embodiments, rules can be established regarding how subunits can communicate in order to facilitate routing. For example, a subunit may only provide output to adjacent, above, or below it.
[0406] Alternatively or additionally, restrictions can be placed on how far apart the following subunits must be for routing purposes. For example, a subunit can only output data to adjacent subunits or subunits within a specified distance (e.g., without two or more intervening subunits).
[0407] Refer to Figure 19, which illustrates the methods of several embodiments. In some embodiments, an FPGA may have one or more "static" regions and one or more "dynamic" regions. The static regions provide a standard configuration, while the dynamic regions can provide functionality according to the requirements of the end user. The static regions may be defined, for example, before the end user receives the network interface device, for example, before the network interface device is deployed to the host. For example, the static regions may be configured to cause the network interface device to provide a specific function. The static regions provide pre-calculated routes between atoms. There may be routing between one or more static regions that passes through one or more dynamic regions, as will be described in more detail later. The dynamic regions may be configured by the end user according to their requirements when the network interface device is deployed to the host. The dynamic regions may be configured to perform different functions for the end user over time.
[0408] In step S1, a first compilation process is performed to provide a first bitfile called the main bitfile 50 and a tool checkpoint 52. This is, in some embodiments, a bitfile of at least a portion of the static area. Once downloaded to the FPGA, the bitfile causes the FPGA to function so that the bitfile is specified in the program compiled from it. In some embodiments, the program used in the first compilation process may be any one or more programs, or it may be a test program specifically designed to assist in routing decisions within a portion of the FPGA. In some embodiments, a set of simple programs may be used alternatively or additionally.
[0409] The program may have a reconfigurable partition that can be modified or used by the compiler. The program may be modified to make the compiler's job easier by moving nets from the reconfigurable partition.
[0410] Step S1 may be performed in a design tool. For example, the Vivado tool may be used with the XilinXFPGA. A checkpoint file may be provided by the design tool. The checkpoint file represents a snapshot of the design at the time the bit file was generated. The checkpoint file may contain one or more of the following: a composite netlist, design constraints, placement information, and routing information.
[0411] In step S2, the bitfile is analyzed, taking into account the checkpoint file, in order to provide a bitfile description 54. The analysis may be for one or more of the following purposes: to discover resources, to generate roots, to check timings, to generate one or more subbyte files, and to generate a bitfile description.
[0412] The analysis may be configured to extract routing information from a bitfile. The analysis may be configured to determine which wire or route the signal propagated along.
[0413] The analysis phase can be performed, at least partially, within the synthesis or design tool. In some embodiments, Vivado's scripting tool can be used. The scripting tool may be TCL (Tool Command Language). TCL can be used to add or modify Vivado's functionality. Vivado's functionality can be invoked and controlled by TCL scripts.
[0414] The bitfile description 54 defines how a given part of the FPGA can be used. For example, the bitfile description indicates which atoms can be routed to which other atoms, and one or more routes that can be routed between those atoms. For example, for each atom, the bitfile description indicates where inputs to that atom may originate, and where outputs from that atom can be routed, along with one or more routes for data outputs. The bitfile description is independent of any program.
[0415] A bitfile description may include one or more of the following: root information, instructions on which root pairs conflict, and a description of how to generate the bitfile from the required configuration of the atom.
[0416] A bitfile description can provide a set of available routes between a set of atoms, before any particular instruction is executed by a given atom.
[0417] A bitfile description may be for a part of an FPGA. A bitfile description may be for the dynamic part of an FPGA. A bitfile description includes which routes are available and / or which routes are unavailable. For example, a bitfile may indicate which routes are available to the dynamic part of an FPGA, taking into account any routing across the dynamic part of the FPGA required by the static part(s) of the FPGA.
[0418] It should be understood that in some embodiments, the bitfile description can be obtained by any suitable method. For example, the bitfile description may be provided by the FPGA or ASIC provider.
[0419] In some embodiments, the bitfile description may be provided by the design tool. In this embodiment, the analysis step may be omitted. The design tool can output a bitfile description. The bitfile description may be for the static part of the FPGA, including any necessary routing across the dynamic part of the FPGA.
[0420] It should be recognized that any other suitable technique may be used to generate the bitfile description. In the example above, the tool used to design the FPGA is used to provide the analysis used to generate the bitfile.
[0421] It should be understood that different tools may be used in other embodiments. In some embodiments, the tools may be specific to a product or set of products. For example, an FPGA provider may provide associated tools for managing its FPGAs.
[0422] In other embodiments, a general-purpose scripting tool can be used. In some embodiments, different tools or techniques can be used to determine the sub-bitfiles. For example, the main bitfile can be analyzed to determine which features correspond to which. This may require the generation of multiple sub-bitfiles.
[0423] Step S3 should be understood to occur when the network interface device is installed on the host and runs on the physical FPGA device. Steps S1 and S2 can be performed as part of the design synthesis process to generate a bitfile image that implements the network interface device. In some embodiments, steps S1 and / or S2 are used to characterize the operation of the FPGA. Once the FPGA is characterized, the bitfile description is stored in the memory of all physical network interface devices that will operate in a given predefined manner.
[0424] In step S3, compilation is performed using the bitfile description and the eBPF program. The output of the compilation is a partial bitfile of the eBPF program. The compilation adds roots to the partial bitfile and the programming that will be executed by the individual slices.
[0425] It should be understood that the bitfile description may be provided within the system being deployed. The bitfile description may be stored in memory. The bitfile description may be stored in an FPGA, a network interface device, or a host device. In some embodiments, the bitfile description is stored in flash memory connected to an FPGA on a network interface device, etc. The flash memory may also include the main bitfile.
[0426] The eBPF program may be stored together with the bitfile description or separately. The eBPF program may be stored on an FPGA, a network interface device, or a host. In the case of eBPF, the program can be transferred from a user-mode program running on the host to the kernel. The kernel then transfers the program to a device driver, which transfers the program to a compiler running on the host or network interface device. In some embodiments, the eBPF program may be stored on the network interface device so that it can run before the host OS boots.
[0427] The compiler may be located on a network interface device, an FPGA, or any suitable location on the host. As just one example, the compiler may run on the CPU on the network interface device.
[0428] Next, the compiler flow will be described. The compiler frontend receives the eBPF program. The eBPF program may be written in any suitable language. For example, the eBPF program may be written in a C language. The compiler is configured in the frontend to convert the program into an intermediate representation IR. In some embodiments, the IR may be LLVM-IR or any other suitable IR.
[0429] In some embodiments, pointer analysis can be performed to create packet / map access primitives.
[0430] It should be understood that in some embodiments, IR optimization may be performed by the compiler. This may be optional in some embodiments.
[0431] The compiler's high-level synthesis backend is configured to divide the program pipeline into stages, generate packet access taps, and emit C code. In some embodiments, the HLS portion of a design tool and / or the design tool being used can be invoked to synthesize the output of the HLS phase.
[0432] The FPGA atom's compiler backend divides the pipeline into stages and generates packet access taps. IF conversions may be performed to convert control dependencies to data dependencies. The design is deployed and routed. A partial bit file of the eBPF program is output.
[0433] When routing conflicts exist, routing problems like the one shown in Figure 20a can occur. For example, slice A can communicate with slice C, and slice B can communicate with slice D. In the configuration of Figure 20a, the common routing unit 60 is assigned to communication between slice A and slice C, and between slice B and slice D. In some embodiments, this routing conflict can be avoided. In this regard, refer to Figure 20b. As can be seen from the figure, a separate route 64 is provided between slice A and slice C, compared with route 62 between slice B and slice D.
[0434] In some embodiments, a bitfile description may contain multiple different routes to at least several pairs of subunits. The compilation process checks for routing conflicts, as shown in Figure 20a. In the event of a routing conflict, the compiler may resolve or avoid such a conflict by selecting one of the appropriate alternative routes.
[0435] Figure 21 shows a partition 66 in the FPGA for executing an eBPF program. The partition interfaces with the static part of the FPGA, for example, through a series of input flip-flops 68 and a series of output flip-flops. In some embodiments, as described above, there may be routing 70 throughout the design.
[0436] The compiler may need to handle routing across the FPGA regions configured by the compiler. The compiler needs to generate sub-bitfiles that fit within the reconfigurable partitions in the main bitfile. When the main bitfile is generated using the reconfigurable partitions, the design tool can avoid using the logical resources within the reconfigurable partitions, and as a result, those resources can be used by the sub-bitfiles. However, the design tool may not always be able to avoid using routing resources within the reconfigurable partitions.
[0437] As a result, the analysis tool needs to avoid using routing resources used by the design tool within the main bitfile. The analysis tool may need to ensure that its list of available routes within the bitfile description does not include any resources used by the main bitfile. Available routes can be defined in terms of route templates that can be used in numerous locations within the FPGA because the FPGA is highly regular. The routing resources used by the main bitfile break this regularity, which means that the analysis tool needs to avoid using those templates in locations that conflict with the main bitfile. The analysis tool may need to generate new route templates that can be used in those locations, and / or prevent certain route templates from being used in certain locations.
[0438] Here are some examples of the features that the compiler provides when translating several exemplary eBPF program fragments into instructions that are executed by atoms.
[0439] Some embodiments may use any suitable synthesis tool to generate the bitfile description. For example, some embodiments may use the Bluespec tool based on a mode that uses atomic transactions for hardware.
[0440] In the first example, the eBPF program fragment has the following two instructions: Instruction 1:r1+=r2 Instruction 2: r1+=r3 The first instruction adds the number in register 1 (r1) to the number in register 2 (r2) and places the result in r1. The second instruction adds r1 to r3 and places the result in r1. Both instructions in this example use 64-bit registers, but only the least significant 32 bits are used. The upper 32 bits of the result are filled with zeros.
[0441] The compiler translates these into instructions that are executed by atoms. A 32-bit add instruction requires 32 pairs of lookup tables (LUTs), a 32-bit carry chain, and 32 flip-flops.
[0442] Each pair in a lookup table adds 2 bits to produce a 2-bit result. A carry chain is a structure that allows bits to be carried from one digit column to the next during addition, and allows bits to be borrowed from the next column during subtraction.
[0443] The 32 flip-flops are memory elements that receive a value in one clock cycle and retrieve that value in the next. They can be used to limit the amount of work performed per clock cycle and to simplify timing analysis.
[0444] In some embodiments, the FPGA may include several slices. In some exemplary slices, the carry chain propagates from the bottom (CIN) of the slice to the top (COUT) of the slice, and then connects to the CIN input of the next slice up.
[0445] In an example where each slice has a 4-bit carry chain, eight slices are used to perform a 32-bit addition. In this embodiment, an atom can be thought of as being provided by a pair of slices, because in some embodiments it may be convenient for the atom to operate on an 8-bit value.
[0446] In an example where each slice has an 8-bit carry chain, four slices are used to perform a 32-bit addition. In this embodiment, atoms can be thought of as being provided by slices.
[0447] This is merely an example, and as mentioned above, it should be understood that atoms can be defined in any appropriate way.
[0448] In this example, a case where the FPGA has a slice that supports an 8-bit carry chain is used in compiling the first exemplary eBPF program fragment.
[0449] There are three 32-bit wide input values and one 32-bit wide output value. There may be other preceding instructions that generated these three input values. Below, we assume several arbitrary locations within a slice (atom).
[0450] The following numbering rules are used. Slices (atoms) are arranged in a regular row and column sequence. XnYm indicates the position of the atom in the sequence, where Xn is the column and Ym is the row. X6Y0 indicates that the slice is in column 6 and row 0. It should be understood that in other embodiments, any other suitable numbering scheme may be used.
[0451] Assume that the initial values were generated simultaneously in the following locations. r1: Slices X6Y0, X6Y1, X6Y2 and X6Y3 r2: Slices X6Y4, X6Y5, X6Y6 and X6Y7 r3: Slices X6Y8, X6Y9, X6Y10 and X6Y11 The result of the first instruction must be computed by four adjacent slices in the same column so that the carry chain is correctly connected. The compiler may choose to compute the result in slices X7Y0, X7Y1, X7Y2, and X7Y3. For this to work, the inputs must be connected. There is a connection from X6Y0 to X7Y0, another connection from X6Y1 to X7Y1, one connection from X6Y2 to X7Y2, and one connection from X6Y3 to X7Y3. Corresponding connections from X6Y4-X6Y7 to X7Y0-X7Y3 are also required.
[0452] These are full-byte connections, meaning each of the eight input bits is connected to the corresponding output bit. For example, the output from slice X6Y0 flip-flip-0 is connected to input 0 of slice X7Y0 LUT 0.
[0453] The output from slice X6Y0 flip flip 1 is connected to input 0 of slice X7Y0 LUT 1.
[0454] The same applies thereafter. The output from slice X6Y0 flip flip 7 is connected to input 0 of slice X7Y0 LUT 7.
[0455] During the first clock cycle, the r1 and r2 values from slices X6Y0-X6Y7 are transferred to the inputs of slices X7Y0-X7Y3, processed by the LUT and carry chain, and the results are stored in the flip-flip (X7Y0-X7Y3) of those slices, ready for use in the next cycle.
[0456] Moving on to instruction 2. The compiler needs to choose where to compute the result of instruction 2. It could choose slice X7Y4~X7Y7. Here again, there will be a full byte connection from the result of instruction 1 (X7Y0~X7Y3) to the input of instruction 2 (X7Y4~X7Y7).
[0457] The value of r3 is also needed. If r1, r2, and r3 are generated in cycle 0, then r1 + r2 is generated in cycle 1. The value of r3 needs to be delayed by one clock cycle so that it is generated in cycle 1. The compiler may choose to generate r3 in cycle 1 using slices X7Y8~X7Y11. Next, a connection is needed from the original slice (X6Y8~X6Y11) that generated r3 in cycle 0 to a new slice (X7Y8~X7Y11) that generates the same value in cycle 1. Once that is done, a connection is now needed from those new slices to the slice for instruction 2. Thus, the output from slice X7Y8 is connected to the input of slice X7Y4, and so on.
[0458] At this time, the FPGA bitfile includes the following functions: - Full byte connection from X6Y0 to X7Y0 input 0 (initial r1 byte 0) - Full byte connection from X6Y1 to X7Y1 input 0 (initial r1 byte 1) - Full byte connection from X6Y2 to X7Y2 input 0 (initial r1 byte 2) - Full byte connection from X6Y3 to X7Y3 input 0 (initial r1 byte 3) - Full byte connection from X6Y4 to X7Y0 input 1 (initial r2 byte 0) - Full byte connection from X6Y5 to X7Y1 input 1 (initial r2 byte 1) - Full byte connection from X6Y6 to X7Y2 input 1 (initial r2 bytes 2) - Full byte connection from X6Y7 to X7Y3 input 1 (initial r2 bytes 3) - Full byte connection from X6Y8 to X7Y8 input 0 (initial r3 bytes 0) - Full byte connection from X6Y9 to X7Y9 input 0 (initial r3 byte 1) - Full byte connection from X6Y10 to X7Y10 input 0 (initial r3 bytes 2) -Full byte connection from X6Y11 to X7Y11 input 0 (initial r3 bytes 3) - Slice X7Y0 (instruction byte 0) configured to add input 0 to input 1. - Slice X7Y1 (instruction 1 byte 1) configured to add input 0 to input 1 - Slice X7Y2 (instruction 1 byte 2) configured to add input 0 to input 1. - Slice X7Y3 (instruction 1 byte 3) configured to add input 0 to input 1 - A slice X7Y8 (r3 delayed byte 0) configured to copy input 0 to output. - A slice X7Y9 (r3 delay byte 1) configured to copy input 0 to output. - Slice X7Y10 (r3 delay byte 2) configured to copy input 0 to output - Slice X7Y11 (r3 delay byte 3) configured to copy input 0 to output - Full byte connection from X7Y0 to X7Y4 input 0 (instruction byte 0) - Full byte connection from X7Y1 to X7Y5 input 0 (1 instruction byte 1) - Full byte connection from X7Y2 to X7Y6 input 0 (1 instruction byte 2) - Full byte connection from X7Y3 to X7Y7 input 0 (1 byte 3 instruction) - Full byte connection from X7Y8 to X7Y4 input 1 (r3 delay byte 0) - Full byte connection from X7Y9 to X7Y5 input 1 (r3 delay byte 1) - Full byte connection from X7Y10 to X7Y6 input 1 (r3 delay, byte 2) - Full byte connection from X7Y11 to X7Y7 input 1 (r3 delay, byte 3) - Slice X7Y4 (instruction 2 bytes 0) configured to add input 0 to input 1. -Slice X7Y5 (Instruction 2 Byte 1) configured to add Input 0 to Input 1 -Slice X7Y6 (Instruction 2 Byte 2) configured to add Input 0 to Input 1 -Slice X7Y7 (Instruction 2 Byte 3) configured to add Input 0 to Input 1 The compiler does not need to generate the upper 32 bits of the result of Instruction 2 because it is known that they are 0. Noting that fact, 0 can be used whenever they are used.
[0459] Next, a second example of compiling an eBPF fragment will be described. Instruction 1: r1 &= 0xff Instruction 2: r2 &= 0xff Instruction 3: If r1 < r2, proceed to L1 Instruction 4: r1 = r2 Label L1.
[0460] The first instruction performs a bitwise AND of r1 and the constant 0xff and places the result in r1. A given bit in the result is set to 1 if the corresponding bit was originally set to 1 in r1 and the corresponding bit is set to 1 in the constant; otherwise, it is set to 0. The constant 0xff has bits 0 to 7 set and bits 8 to 63 cleared, so as a result, bits 0 to 7 of r1 are unchanged, but bits 8 to 63 are set to 0. This simplifies things for the compiler because it understands that bits 8 to 63 are 0 and does not need to generate them. The second instruction does the same for r2.
[0461] Instruction 3 checks whether r1 is less than r2 and, if so, jumps to label L1. This skips Instruction 4. Instruction 4 simply copies the value from r2 to r1. This instruction sequence finds the minimum of r1 byte 0 and r2 byte 0 and places the result in r1 byte 0.
[0462] Compilers can use a technique known as "if conversion" to transform conditional jumps into select instructions.
[0463] Instruction 1:r1&=0xff Instruction 2: r2&=0xff Instruction 5:c1=(r1 <r2) Instruction 6:r1=c1?r1:r2 Instruction 5 compares r1 with r2 and sets c1 to 1 if r1 is less than r2, otherwise sets c1 to 0. Instruction 6 is a selection instruction that copies r1 to r1 if c1 is set (this has no effect), and copies r2 to r1 otherwise. If c1 is equal to 1, instruction 3 has skipped instruction 4, which means r1 retains its value from instruction 1. In this case, the selection instruction also retains r1 unchanged. If c1 is equal to 0, instruction 3 has not skipped instruction 4, so r2 is copied to r1 by instruction 4. Here again, the selection instruction copies r2 to r1, so the new sequence has the same effect as the old sequence.
[0464] Instruction 6 is not a valid eBPF instruction. However, it is represented in LLVM-IR while the compiler is working on it. Instruction 6 is a valid instruction in LLVM-IR.
[0465] These instructions need to be assigned to atoms here. Assume that input r1 is available in slices X0Y0 to X0Y7 and r2 is available in slices X0Y8 to X0Y15. Instructions 1 and 2 cause the compiler to note that the upper 7 bytes of r1 and r2 should be set to 0.
[0466] Next, the compiler may choose to compute the result of instruction 5 in slice X1Y0. This requires a full byte connection from the output of slice X0Y0 to input 0 of slice X1Y0, and a full byte connection from the output of slice X0Y8 to input 1 of slice X1Y0. The way to compare the two values is to subtract one from the other and check whether the calculation overflows by attempting to borrow from the next bit up. The result of this comparison is then stored in flip-flop 7 of slice X1Y1.
[0467] Similar to the first example, r1 and r2 need to be delayed by one cycle to present their values in a timely manner for instruction 6. The compiler may use slices X1Y1 and X1Y2 for r1 and r2, respectively.
[0468] The select instruction requires three inputs: c1, r1, and r2. Note that r1 and r2 are 1 byte wide, but c1 is only 1 bit wide. Assume that the compiler computes the result of the select instruction slice X2Y0. The select is performed bit by bit, with each LUT in slice X2Y0 handling 1 bit.
[0469] If c1 is set, the resulting bit 0 is bit r1 0. Otherwise, the r2 bit is 0.
[0470] If c1 is set, the resulting bit 1 is r1 bit 1. Otherwise, bit r2 is 1.
[0471] ...and the same applies to the following. If c1 is set, the resulting bit 7 is r1 bit 7. Otherwise, the r2 bit is 7.
[0472] Each LUT may need to access the corresponding bits from r1 and r2, but all LUTs need to access c1. This means that c1 needs to be duplicated across the bits of input 0 of the slice. Thus the connections for the input of instruction 6 are as follows:
[0473] Duplicate bit 7 of the output of slice X1Y0 to input 0 of slice X2Y0. Full byte connection from the output of slice X1Y1 to input 1 of slice X2Y0.
[0474] Full byte connection from the output of slice X1Y2 to input 2 of slice X2Y0. Another issue that needs to be addressed concerns shift instructions. Consider the following example.
[0475] A 16-bit shift 5 bits to the left is: Set output bit 0 to 0, Set output bit 1 to 0, Set output bit 2 to 0, Set output bit 3 to 0, Set output bit 4 to 0, Copy input bit 0 to output bit 5, Copy input bit 1 to output bit 6, ... This requires copying input bit 10 to output bit 15.
[0476] Note that the inputs and outputs here are those of the connection. The input to the connection is from the output of the first slice. The output of the connection then proceeds to the input of the second slice.
[0477] It may not always be possible to make this type of connection within a slice, but it may be possible through interconnections between slices. A compiler can assume that a 16-bit input value is produced by two adjacent slices in the same column, because the compiler can verify that the value is produced there.
[0478] As an example, let's assume the input is generated by slices X0Y4 and X0Y5, and the output is directed to slices X1Y4 and X1Y5. In that case, the following connections are required.
[0479] Since we know that the slice X1Y4 bit 0 is 0, it is not necessary. Since we know that the bit 1 in slice X1Y4 is 0, it is not necessary. Since we know that the slice X1Y4 bit 2 is 0, it is not necessary. Since we know that the slice X1Y4 bit 3 is 0, it is not necessary. Since we know that bit 4 in slice X1Y4 is 0, it is not necessary. Slice X1Y4 bit 5 is derived from slice X0Y4 bit 0. Slice X1Y4 bit 6 is derived from slice X0Y4 bit 1. Slice X1Y4 bit 7 is derived from slice X0Y4 bit 2. Slice X1Y5 bit 0 comes from slice X0Y4 bit 3. Slice X1Y5 bit 1 comes from slice X0Y4 bit 4. Slice X1Y5 bit 2 is derived from slice X0Y4 bit 5. Slice X1Y5 bit 3 is derived from slice X0Y4 bit 6. Slice X1Y5 bit 4 is derived from slice X0Y4 bit 7. Slice X1Y5 bit 5 is derived from slice X0Y5 bit 0. Slice X1Y5 bit 6 is derived from slice X0Y5 bit 1. Slice X1Y5 bit 7 is derived from slice X0Y5 bit 2. The eight connections to the input of slice X1Y5 can be thought of as shifted connections or shifted roots. The same structure can be used for slice X1Y4, but it has inputs from X1Y3 and X1Y4. This is because bits 5-7 are matched, and the slice can ignore bits 0-4, so it doesn't matter which input is presented there.
[0480] It may be necessary to be able to shift by any amount from 1 to 7 bits. A connection that shifts by 0 bits or 8 bits is exactly the same as a full-byte connection, in which case each bit is connected to the corresponding bit in another slice.
[0481] The shift of a variable quantity can be performed in two or three stages, depending on the range of the value being shifted. The stages are as follows:
[0482] Stage 1: Shift only 0, 1, 2, or 3. Stage 2: Shift only 0, 4, 8, or 12.
[0483] Stage 3: Shift by 0, 16, 32, or 48 (32-bit or 64-bit only).
[0484] As another example, if we have an arithmetic right shift of a variable number of bytes, the value to be shifted is generated by slice X3Y2, and the amount of the shift is generated by X3Y3.
[0485] An arithmetic right shift requires an "arithmetic right shift" type connection. This type of connection takes the outputs of one slice and connects them to the inputs of another slice, but in the process shifts them to the right by a certain amount and duplicates the sign bit if necessary.
[0486] For example, the "arithmetic right 3 shift" connection has the following: Output bit 0 comes from input bit 3. Output bit 1 comes from input bit 4. Output bit 2 comes from input bit 5. Output bit 3 comes from input bit 6. Output bit 4 comes from input bit 7. Output bit 5 is derived from input bit 7 (sign bit). Output bit 6 is derived from input bit 7 (sign bit). Output bit 7 is derived from input bit 7 (sign bit). Stage 1 can be calculated in slice X4Y2, in which case the following connections are required.
[0487] Full byte from slice X3Y2 to slice X4Y2 input 0 Arithmetic right shift from slice X3Y2 to slice X4Y2 input 1 Arithmetic right shift from slice X3Y2 to slice X4Y2 input 2 Arithmetic right shift 3 from slice X3Y2 to slice X4Y2 input 3 Duplicate slice X3Y3 bit 0 to slice X4Y2 input 4. Duplicate slice X3Y3 bit 1 to slice X4Y2 input 5. Next, slice X4Y2 is configured to select one of the first four inputs based on inputs 4 and 5, as follows:
[0488] Input 4 is 0 and Input 5 is 0: Select Input 0 Input 4 is 1 and Input 5 is 0: Select Input 1 Input 4 is 0 and Input 5 is 1: Select Input 2 Input 4 is 1 and Input 5 is 1: Select Input 3 By copying the shift amount from slice X3Y3 to slice X4Y3, a delayed version can be provided.
[0489] Stage 2 can be calculated at slice X5Y2, in which case the following connections are required.
[0490] Full byte from slice X4Y2 to slice X5Y2 input 0 Arithmetic right shift 4 from slice X4Y2 to slice X5Y2 input 1 Duplicate slice X4Y3 bit 2 to slice X5Y2 input 2. Next, slice X5Y2 is configured to select either input 0 or input 1 based on input 2, as follows:
[0491] Input 2 is 0: Selects Input 0 Input 2 is 1: Selects Input 1 The output of slice X5Y2 is the result of a variable arithmetic right shift operation.
[0492] The bitfile of a given atom may be as follows: Atom identification information A list of other atoms that a given atom can accept as an input and the routes available for that input.
[0493] A list of other atoms that a given atom can provide an output and routes available to that output. Because FPGAs have a regular structure, it's important to understand that there may be a common template that can be used for multiple atoms, with modifications to individual atoms as needed.
[0494] As an example, the bitfile description for slice X7Y1 can specify the following possible inputs and outputs:
[0495] Input from X6Y1 via Root A or Root B Input from X6Y5 via Root C or Root D Input from X7Y0 via Root E or Root F Output to X8Y1 via Root G or Root H Output to X7Y2 via Route I or Route J Output to X7Y5 via Root K or Root L.
[0496] The compiler uses this bitfile description to provide partial bitfiles for the input and output of slice X7Y1 in the first eBPF example described below.
[0497] Input from X6Y1 via Root A Input from X6Y5 via root C Output to X7Y5 via Root K or Root L.
[0498] As an example, the bitfile description of slice XnYm can specify the following possible inputs and outputs:
[0499] Input from Xn-1Ym via Route A or Route B Input from Xn-1Ym+4 via Root C or Root D Input from XnYm-1 via Root E or Root F Output to Xn+1Ym via √G or √H Output to XnYm+1 via Root I or Root J Output to XnYm+4 via the root of K or the root of L.
[0500] As mentioned earlier, this bitfile description can be modified to remove one or more routes that are unavailable for the compiler to use. This may be because the route is being used by another atom or for routing across partitions.
[0501] It should be understood that a compiler can be implemented by a computer program containing computer-executable instructions that can be executed by one or more computer processors. A compiler can run on hardware such as at least one processor working in conjunction with one or more memory devices.
[0502] While the above describes exemplary embodiments, it should be noted that there are several variations and modifications that can be made to the disclosed solutions without departing from the scope of the present invention.
[0503] Accordingly, the embodiments may vary within the scope of the appended claims. Generally, some embodiments can be implemented in hardware, dedicated circuitry, software, logic, or any combination thereof. For example, some embodiments may be implemented in hardware, while others may be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device, but the embodiments are not limited to these.
[0504] The embodiments can be stored in memory and implemented by computer software that can be executed by at least one data processor of the involved entities, by hardware, or by a combination of software and hardware.
[0505] Software can be stored on physical media such as memory chips or memory blocks implemented within a processor, magnetic media such as hard disks or floppy disks, and optical media such as DVDs and their data variants, and CDs.
[0506] The memory may be of any type suitable for the local technical environment and can be implemented using any appropriate data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory.
[0507] The data processor may be of any type suitable for the local technical environment and may include, in non-limiting examples, one or more of the following: general-purpose computers, dedicated computers, microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), gate-level circuits, and processors based on multicore processor architectures.
[0508] In conjunction with the accompanying drawings and claims, various modifications and adaptations may become apparent to those skilled in the art, considering the foregoing description. However, all such and similar modifications of this teaching still fall within the scope defined in the accompanying claims.
Claims
1. A network interface device for interface a host device to a network, A first interface, the first interface being configured to receive multiple data packets, A configurable hardware module comprising multiple processing units, each processing unit comprising a configurable hardware module associated with a predetermined type of operation that can be performed in a single step, At least some of the plurality of processing units are associated with different predetermined types of operations, The configurable hardware module, when executed by the configurable hardware module, is configured to receive instructions to interconnect at least some of the plurality of processing units in order to provide a first data processing pipeline for processing one or more of the plurality of data packets, and the operation of at least some of the processing units, each of a predetermined type of operation, in the first data processing pipeline results in performing a first function in relation to one or more of the plurality of data packets. With respect to the first data packet among the one or more of the plurality of data packets, at least two or more of the plurality of processing units in the first data processing pipeline are each configured to perform their respective predetermined types of operations in parallel with respect to the first data packet. Multiple data packets among the aforementioned multiple data packets are processed simultaneously by different processing units in the first data processing pipeline. The system includes a shared memory accessible to two or more of the at least some of the plurality of processing units, the shared memory being configured to store a state associated with the first data packet, and during the execution of the first function by the configurable hardware module, two or more of the plurality of processing units being configured to access and modify the state associated with the first data packet. A network interface device in which, of the two or more of the at least some of the plurality of processing units, the first processing unit is configured to stall during access by the second processing unit, of the two or more of the at least some of the plurality of processing units, to the state value associated with the first data packet.
2. Of the plurality of processing units, at least two of the above-mentioned two or more are Each of the predetermined types of operations is performed within a predetermined time length defined by the clock signal. The network interface device according to claim 1, configured to transfer the results of each of the predetermined types of operations to the next processing unit in response to the completion of the predetermined time period.
3. The network interface device according to claim 1 or 2, wherein each of the plurality of processing units includes an application-specific integrated circuit configured to perform the respective predetermined type of operation associated with each of the processing units.
4. The network interface device according to any one of claims 1 to 3, wherein at least one of the plurality of processing units comprises a digital circuit and a memory for storing a state related to processing performed by the digital circuit, and the digital circuit is configured to communicate with the memory to perform the predetermined type of operation associated with each of the processing units.
5. The network interface device according to any one of claims 1 to 4, wherein one or more of the plurality of processing units can be individually configured to perform operations specific to each pipeline based on each predetermined type of operation.
6. The configurable hardware module receives an instruction and responds to the instruction, To cause one or more of the aforementioned processing units to perform the respective predetermined types of operations with respect to one or more data packets, Adding one or more of the aforementioned processing units to the data processing pipeline, A network interface device according to any one of claims 1 to 5, configured to perform at least one of removing one or more of the plurality of processing units from a data processing pipeline.
7. The aforementioned predetermined type of operation is, Loading at least one value of the first data packet from memory, Store at least one value of the data packet in memory, and A network interface device according to any one of claims 1 to 6, comprising at least one of performing a lookup in a lookup table to determine an action to be performed with respect to a data packet.
8. The network interface device according to any one of claims 1 to 7, wherein one or more of the plurality of processing units is configured to pass the result of at least one of the predetermined types of operations to the next processing unit in a first data processing pipeline, and the next processing unit is configured to perform the next predetermined type of operation in response to the at least one result.
9. The network interface device according to any one of claims 1 to 8, wherein each of the different predetermined types of operation is defined by a different template.
10. The aforementioned predetermined type of operation is, Accessing data packets, Accessing the lookup table stored in the memory of the configurable hardware module, Performing logical operations on data loaded from data packets, and A network interface device according to any one of claims 1 to 9, comprising at least one of performing a logical operation on data loaded from the lookup table.
11. The network interface device according to any one of claims 1 to 10, wherein the configurable hardware module comprises routing hardware, and the configurable hardware module is configurable to interconnect at least some of the plurality of processing units in order to provide the first data processing pipeline by configuring the routing hardware to route data packets between the plurality of processing units in a specific order defined by the first data processing pipeline.
12. The network interface device according to any one of claims 1 to 11, wherein the configurable hardware module can be configured to interconnect at least some of the plurality of processing units to provide a second data processing pipeline for processing one or more of the plurality of data packets to perform a second function different from the first function.
13. The network interface device according to any one of claims 1 to 12, wherein the configurable hardware module can be configured to interconnect at least some of the plurality of processing units to provide a first data processing pipeline, and then to interconnect at least some of the plurality of processing units to provide a second data processing pipeline.
14. The network interface device according to any one of claims 1 to 13, further comprising a separate circuit configured to perform the first function for one or more of the plurality of data packets, apart from the configurable hardware module.
15. The aforementioned further circuit is Field-programmable gate array, and The network interface device according to claim 14, comprising at least one of a plurality of central processing units.
16. The network interface device according to claim 14 or 15, further comprising at least one controller, the further circuitry configured to perform the first function on data packets during a compilation process to enable the first function to be performed in the configurable hardware module, and the at least one controller configured to control the configurable hardware module to begin performing the first function on data packets in response to the completion of the compilation process.
17. The network interface device according to claim 16, wherein the at least one controller is configured to control the further circuitry to stop performing the first function on data packets in response to the completion of the compilation process.
18. The network interface device according to any one of claims 14 or 15, wherein the network interface device comprises at least one controller, the configurable hardware module is configured to perform the first function on data packets during a compilation process for which the first function is performed in the further circuit, and the at least one controller is configured to determine that the compilation process for which the first function is performed in the further circuit is complete, and to control the further circuit to start performing the first function on data packets in response to the completion of the compilation process.
19. The network interface device according to claim 18, wherein the at least one controller is configured to control the configurable hardware module to stop performing the first function on data packets in response to the completion of the compilation process.
20. A network interface device according to any one of claims 1 to 19, comprising at least one controller configured to perform a compilation process to enable the first function to be performed in the configurable hardware module.
21. A data processing system comprising a network interface device according to any one of claims 1 to 20 and a host device, wherein the data processing system comprises at least one controller configured to perform a compilation process to enable the first function to be performed in the configurable hardware module.
22. The aforementioned at least one controller is The aforementioned network interface device, and The data processing system according to claim 21, provided by one or more of the host devices.
23. The data processing system according to claim 21 or 22, wherein the compilation process is performed in response to a decision by the at least one controller that a computer program representing the first function is safely executed in kernel mode on the host device.
24. The aforementioned at least one controller is Before the completion of the compilation process, a first instruction is sent to further circuitry of the network interface device to perform the first function on data packets, the further circuitry being configured to perform the first function on one or more of the plurality of data packets, separately from the configurable hardware module. The data processing system according to any one of claims 21 to 23, wherein after the completion of the compilation process, it is configured to send a second instruction to the configurable hardware module to initiate the execution of the first function on a data packet.
25. A method for implementation in a network interface device, The first interface includes the step of receiving multiple data packets, The steps include configuring the configurable hardware module of the network interface device such that, upon receiving an instruction, at least some of the processing units of the configurable hardware module interconnect to provide a first data processing pipeline for processing one or more of the multiple data packets and performing a first function on one or more of the multiple data packets, Each processing unit is associated with a predetermined type of operation that can be performed in a single step. At least some of the plurality of processing units are associated with different predetermined types of operations, With respect to the first data packet among the one or more of the plurality of data packets, at least two or more of the plurality of processing units are configured to perform their respective predetermined types of operations in parallel with respect to the first data packet. Multiple data packets among the aforementioned multiple data packets are processed simultaneously by different processing units in the first data processing pipeline. During the execution of the first function, two or more of the at least some of the plurality of processing units are configured to access and modify the state associated with the first data packet in the shared memory of the network interface device, and the shared memory is accessible to two or more of the at least some of the plurality of processing units in the first data processing pipeline. A method in which, during the execution of the first function, the first processing unit, of at least some of the plurality of processing units, is configured to stall during the access of the state value associated with the first data packet by the second processing unit, of at least some of the plurality of processing units, in the first data processing pipeline.
26. A non-temporary computer-readable medium containing program instructions for causing a network interface device to carry out a method, wherein the method is The first interface includes the step of receiving multiple data packets, The steps include configuring the configurable hardware module of the network interface device such that, upon receiving an instruction, at least some of the processing units of the configurable hardware module interconnect to provide a first data processing pipeline for processing one or more of the multiple data packets and performing a first function on one or more of the multiple data packets, Each processing unit is associated with a predetermined type of operation that can be performed in a single step. At least some of the plurality of processing units are associated with different predetermined types of operations, With respect to the first data packet among the one or more of the plurality of data packets, at least two or more of the plurality of processing units are configured to perform their respective predetermined types of operations in parallel with respect to the first data packet. Multiple data packets among the aforementioned multiple data packets are processed simultaneously by different processing units in the first data processing pipeline. During the execution of the first function, at least two of the plurality of processing units are configured to access and modify the state associated with the first data packet in the shared memory of the network interface device, and the shared memory is accessible to at least two of the plurality of processing units in the first data processing pipeline. During the execution of the first function, a non-transient computer-readable medium is configured such that, among at least some of the plurality of processing units, two or more of the first processing units are configured to stall during access by a second processing unit among at least some of the plurality of processing units in the first data processing pipeline to the state value associated with the first data packet.
Citation Information
Patent Citations
Adjustable cycle pipeline system and method
US7861067B1
Streaming interconnect architecture
US9940284B1