Network interface device
By introducing configurable hardware modules and controllers into the network interface device, a parallel processing pipeline of multiple processing units is realized, which solves the problem of inefficiency of existing equipment when processing multiple operations, and improves the efficiency and flexibility of data packet processing.
Patent Information
- Application Number
- CN201980087757.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-04-25
- Filing Date
- 2019-11-05
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2039-12-23
AI Technical Summary
When processing data packets, existing network interface devices are difficult to efficiently process multiple predetermined operations in parallel, resulting in insufficient resource utilization and inefficient processing efficiency.
It adopts a configurable hardware module, including multiple processing units, each unit is associated with a specific operation type, and through the interconnection of the hardware module and the coordination of the controller, a parallel processing pipeline of multiple data packets is realized, supporting functions such as filtering, encapsulation, routing and firewall operations.
It improves the efficiency and flexibility of data packet processing, can dynamically adjust the processing pipeline according to needs, meet different functional needs, and improves the performance of network interface devices.
Smart Images

Figure CN113272793B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to a network interface device for performing functions on data packets. Background Art
[0002] Network interface devices are known and are commonly used to provide an interface between a computing device and a network. A network interface device may be configured to process data received from a network and / or process data to be placed on a network. Summary of the Invention
[0003] According to one aspect, a network interface device for connecting a host to a network is provided, the network interface device comprising: a first interface configured to receive a plurality of data packets; and a configurable hardware module comprising a plurality of processing units, each processing unit being associated with a predetermined operation type that can be performed in a single step, wherein at least some of the plurality of processing units are associated with different predetermined operation types, wherein the hardware module is configurable to interconnect at least some of the plurality of processing units to provide a first data processing pipeline for processing one or more of the plurality of data packets so as to perform a first function on the one or more of the plurality of data packets.
[0004] In some embodiments, the first function comprises a filtering function. In some embodiments, the function comprises at least one of a tunnel, an encapsulation, and a routing function. In some embodiments, the first function comprises an extended Berkeley packet filter function.
[0005] In some embodiments, the first function comprises a distributed denial of service cleanup operation.
[0006] In some embodiments, the first function includes firewall operations.
[0007] In some embodiments, the first interface is configured to receive a first data packet from the network.
[0008] In some embodiments, the first interface is configured to receive a first data packet from the host device.
[0009] In some embodiments, two or more of at least some of the plurality of processing units are configured to perform their associated at least one predetermined operation in parallel.
[0010] In some embodiments, two or more of at least some of the plurality of processing units are configured to perform their associated predetermined types of operations according to a common clock signal of the hardware module.
[0011] In some embodiments, each of two or more of at least some of the plurality of processing units is configured to perform its associated predetermined type of operation for a predetermined length of time specified by the clock signal.
[0012] In some embodiments, two or more of at least some of the plurality of processing units are configured to: access the first data packet within a time period of a predetermined length of time; and transmit a result of the respective at least one operation to a next processing unit in response to an end of the predetermined length of time.
[0013] In some embodiments, the result includes at least one or more of: at least one value from one or more of the plurality of data packets; an update to a mapping state; and metadata.
[0014] In some embodiments, each of the plurality of processing units includes an application specific integrated circuit configured to perform at least one operation associated with the respective processing unit.
[0015] In some embodiments, each processing unit comprises a field programmable gate array. In some embodiments, each processing unit comprises any other type of soft logic.
[0016] In some embodiments, at least one of the plurality of processing units includes a digital circuit and a memory for storing a state associated with processing performed by the digital circuit, wherein the digital circuit is configured to perform a predetermined type of operation associated with the respective processing unit in communication with the memory.
[0017] In some embodiments, the network interface device includes two or more processing unit memories accessible to the plurality of processing units, wherein the memories are configured to store a state associated with the first data packet, wherein during performance of the first function by the hardware module, the two or more processing units of the plurality of processing units are configured to access and modify the state.
[0018] In some embodiments, a first processing unit of at least some of the plurality of processing units is configured to cease execution while a second one of the plurality of processing units accesses the value of the state.
[0019] In some embodiments, one or more processing units of the plurality of processing units are individually configured to perform specific operations for respective pipelines based on their associated predetermined operation types.
[0020] In some embodiments, the hardware module is configured to receive instructions and, in response to the instructions, perform at least one of the following: interconnect at least some of the plurality of processing units to provide a data processing pipeline for processing one or more of the plurality of data packets; cause one or more of the plurality of processing units to perform their associated predetermined operation types related to the one or more data packets; add one or more of the plurality of processing units to the data processing pipeline; and remove one or more of the plurality of processing units from the data processing pipeline.
[0021] In some embodiments, the predetermined operation includes at least one of: loading at least one value of the first data packet from a memory; storing at least one value of the data packet in a memory; and performing a lookup in a lookup table to determine an action to be performed on the data packet.
[0022] In some embodiments, the hardware module is configured to receive an instruction, wherein the hardware module is configurable to interconnect at least some of the plurality of processing units in response to the instruction to provide a data processing pipeline for processing one or more packets of the plurality of packets, wherein the instruction includes a packet sent through the third processing pipeline.
[0023] In some embodiments, one or more of at least some of the plurality of processing units may be configured to, in response to the instruction, perform a selected operation of an associated predetermined operation type on one or more of the plurality of data packets.
[0024] In some embodiments, the plurality of components includes a second one of the plurality of components configured to provide the first functionality in a circuit distinct from a hardware module, wherein the network interface device includes at least one controller configured to pass packets through a processing pipeline for processing by one of the first one of the plurality of components and the second one of the plurality of components.
[0025] In some embodiments, a network interface device includes at least one controller configured to issue instructions to cause a hardware module to begin performing a first function on a data packet, wherein the instructions are configured to cause a first component of a plurality of components to be inserted into the processing pipeline.
[0026] In some embodiments, a network interface device includes at least one controller configured to issue instructions to cause a hardware module to begin performing a first function with respect to a data packet, wherein the instructions include a control message sent through a processing pipeline and configured to cause a first component of a plurality of components to start.
[0027] In some embodiments, for one or more processing units among at least some of the multiple processing units, the associated at least one operation includes at least one of the following: loading at least one value of the first data packet from a memory of a network interface device; storing at least one value of the first data packet in a memory of the network interface device; and performing a query in a lookup table to determine an action to be performed on the first data packet.
[0028] In some embodiments, one or more processing units among at least some of the multiple processing units are configured to pass at least one result of at least one associated predetermined operation to the next processing unit in the first processing pipeline, and the next processing unit is configured to perform a next predetermined operation based on the at least one result.
[0029] In some embodiments, each of the different predetermined operation types is defined by a different template.
[0030] In some embodiments, the predetermined operation type includes at least one of the following operations: accessing a data packet; accessing a lookup table stored in a memory of the hardware module; performing a logical operation on data loaded from the data packet; and performing a logical operation on data loaded from the lookup table.
[0031] In some embodiments, the hardware module includes routing hardware, wherein the hardware module is configured to: route data packets between the plurality of processing units in a specific order specified by the first data processing pipeline by configuring the routing hardware, thereby interconnecting at least some of the plurality of processing units to provide the first data processing pipeline.
[0032] In some embodiments, the hardware module may be configured to interconnect at least some of the plurality of processing units to provide a second data processing pipeline for processing one or more of the plurality of data packets to perform a second function different from the first function.
[0033] In some embodiments, the hardware module may be configured to interconnect at least some of the plurality of processing units to provide a second data processing pipeline after interconnecting at least some of the plurality of processing units to provide a first data processing pipeline.
[0034] In some embodiments, the network interface device includes additional circuitry separate from the hardware module and configured to perform a first function on one or more packets of the plurality of packets.
[0035] In some embodiments, the further circuitry includes at least one of: a field programmable gate array; and a plurality of central processing units.
[0036] In some embodiments, the network interface device includes at least one controller, wherein the further circuitry is configured to perform execution of the first function on the data packet during a compilation process for the first function to be executed in the hardware module, wherein the at least one controller is configured to control the hardware module to begin executing the first function on the data packet in response to completion of the compilation process.
[0037] In some embodiments, the further circuitry comprises a plurality of central processing units.
[0038] In some embodiments, the at least one controller is configured to, in response to determining that a compilation process for a first function to be executed in the hardware module has been completed, control the further circuit to stop executing the first function on the data packet.
[0039] In some embodiments, the network interface device includes at least one controller, wherein the hardware module is configured to perform the first function on a data packet during a compile process for the first function to be performed in the hardware circuit, wherein the at least one controller is configured to determine that the compile process for the first function to be performed in the additional circuit is completed, and in response to the determination, control the additional circuit to start performing the first function with respect to the data packet.
[0040] In some embodiments, the further circuitry comprises a field programmable gate array.
[0041] In some embodiments, the at least one controller is configured to, in response to determining that a compilation process for the first function to be executed in the further circuit is completed, control the hardware module to stop executing the first function on the data packet.
[0042] In some embodiments, the network interface device includes at least one controller configured to perform a compilation process to provide the first function to be executed in the hardware module.
[0043] In some embodiments, the compiling process includes providing instructions to provide a control plane interface in the hardware module that is responsive to the control message.
[0044] According to another aspect, there is provided a data processing system comprising a network interface device and a host device according to the first aspect, and wherein the data processing system comprises at least one controller configured to perform a compilation process to provide a first function to be executed in a hardware module.
[0045] In some embodiments, the at least one controller is provided by one or more of: a network interface device and a host device.
[0046] In some embodiments, the compiling process is performed in response to a determination by the at least one controller that the computer program representing the first function is safe for execution in kernel mode of the host device.
[0047] In some embodiments, the at least one controller is configured to perform compilation processing by specifying each of at least some of the plurality of processing units to execute at least one operation represented by a series of computer code instructions in a specific order of the first data processing pipeline, wherein the plurality of operations provide the first function for the one or more data packets of the plurality of data packets.
[0048] In some embodiments, the at least one controller is configured to: send a first instruction to cause another circuit of the network interface device to perform the first function on the data packet before the compilation process is completed; and send a second instruction to cause the hardware module to start performing the first function on the data packet after the compilation process is completed.
[0049] According to another aspect, a method for implementation in a network interface device is provided, the method comprising: receiving a plurality of data packets at a first interface; and configuring a hardware module to interconnect at least some of a plurality of processing units of the hardware module to provide a first data processing pipeline for processing one or more of the plurality of data packets to perform a first function on the one or more of the plurality of data packets, wherein each processing unit is associated with a predetermined operation type that can be performed in a single step, and wherein at least some of the plurality of processing units are associated with different predetermined operation types.
[0050] According to another aspect, a non-transitory computer-readable medium is provided, the non-transitory computer-readable medium including program instructions for causing a network interface device to perform a method, the method comprising: receiving a plurality of data packets at a first interface; and configuring a hardware module to interconnect at least some of a plurality of processing units of the hardware module to provide a first data processing pipeline for processing one or more of the plurality of data packets to perform a first function on the one or more of the plurality of data packets, wherein each processing unit is associated with a predetermined operation type that can be performed in a single step, and wherein at least some of the plurality of processing units are associated with different predetermined operation types.
[0051] According to another aspect, a processing unit is provided, which is configured to: perform at least one predetermined operation on a first data packet received at a network interface device; be connected to a first additional processing unit, which is configured to perform at least one additional first predetermined operation on the first data packet; be connected to a second additional processing unit, which is configured to perform at least one second additional predetermined operation on the first data packet; receive a result of the first additional at least one predetermined operation from the first additional processing unit; perform the at least one predetermined operation based on the result of the first additional at least one predetermined operation; and send the result of the at least one predetermined operation to the second additional processing unit for processing in the at least one second additional predetermined operation.
[0052] In some embodiments, the processing unit is configured to receive a clock signal for timing at least one predetermined operation, wherein the processing unit is configured to perform the at least one predetermined operation in at least one cycle of the clock signal.
[0053] In some embodiments, the processing unit is configured to perform at least one predetermined operation in a single cycle of a clock signal.
[0054] In some embodiments, the at least one predetermined operation, the first further at least one predetermined operation, and the second further at least one predetermined operation form part of a function performed with respect to a first data packet received at the network interface device.
[0055] In some embodiments, a first data packet is received from a host device, wherein the network interface device is configured to interface the host device to a network.
[0056] In some embodiments, a first data packet is received from a network, wherein the network interface device is configured to interface the host device to the network.
[0057] In some embodiments, the function is a filtering function.
[0058] In some embodiments, the filtering functionality is an extended Berkeley packet filtering functionality.
[0059] In some embodiments, the processing unit includes an application specific integrated circuit configured to perform at least one predetermined operation.
[0060] In some embodiments, the processing unit includes: a digital circuit configured to perform at least one predetermined operation; and a memory storing a state related to the at least one predetermined operation performed.
[0061] In some embodiments, the processing unit is configured to access a memory accessible by the first additional processing unit and the second additional processing unit, wherein the memory is configured to store a state associated with the first data packet, wherein the at least one predetermined operation includes modifying the state stored in the memory.
[0062] In some embodiments, the processing unit is configured to read the value of the state from the memory during a first clock cycle and provide the value to a second further processing unit for modification by the second further processing unit, wherein the processing unit is configured to stop during a second clock cycle following the first clock cycle.
[0063] In some embodiments, the at least one predetermined operation includes at least one of: loading the first data packet from a memory of the network interface device; storing the first data packet in the memory of the network interface device; and performing a lookup in a lookup table to determine an action to be performed on the first data packet.
[0064] According to another aspect, a method implemented in a processing unit is provided, the method comprising: performing at least one predetermined operation on a first data packet received at a network interface device; connecting to a first additional processing unit, the first additional processing unit being configured to perform the first additional at least one predetermined operation on the first data packet; connecting to a second additional processing unit, the second additional processing unit being configured to perform the second additional at least one predetermined operation on the first data packet; receiving a result of the first additional at least one predetermined operation from the first additional processing unit; performing the at least one predetermined operation based on the result of the first additional at least one predetermined operation; and sending the result of the at least one predetermined operation to the second additional processing unit for processing in the second additional at least one predetermined operation.
[0065] According to another aspect, a computer-readable non-volatile storage device is provided, which stores instructions that, when executed by a processing unit, cause the processing unit to perform a method comprising: performing at least one predetermined operation with respect to a first data packet received at a network interface device; connecting to a first additional processing unit, the first additional processing unit being configured to perform the first additional at least one predetermined operation with respect to the first data packet; connecting to a second additional processing unit, the second additional processing unit being configured to perform the second additional at least one predetermined operation with respect to the first data packet; receiving a result of the first additional at least one predetermined operation from the first additional processing unit; performing the at least one predetermined operation based on the result of the first additional at least one predetermined operation; and sending the result of the at least one predetermined operation to the second additional processing unit for processing in the second additional at least one predetermined operation.
[0066] According to another aspect, a network interface device for interfacing a host device with a network is provided, the network interface device comprising: at least one controller; a first interface configured to receive data packets; a first circuit configured to perform a first function on data packets received at the first interface; and a second circuit, wherein the first circuit is configured to perform the first function with respect to data packets received at the first interface during a compilation process for the first function to be performed in the second circuit, wherein the at least one controller is configured to determine that the compilation process for the first function to be performed in the second circuit has been completed, and in response to the determination, control the second circuit to start performing the first function on data packets received at the first interface.
[0067] In some embodiments, the at least one controller is configured to control the first circuit to stop executing the first function on the data packet received at the first interface in response to the determination that the compilation process of the first function to be executed in the second circuit has been completed.
[0068] In some embodiments, the at least one controller is configured to, in response to the determination that the compilation process for the first function to be performed in the second circuit has been completed: start performing the first function on data packets of the first data stream received at the first interface; and control the first circuit to stop performing the first function on data packets of the first data stream.
[0069] In some embodiments, the first circuit includes at least one central processing unit, wherein each of the at least one central processing unit is configured to perform a first function on at least one data packet received at the first interface.
[0070] In some embodiments, the second circuit includes a field programmable gate array configured to initiate execution of the first function for a data packet received at the first interface.
[0071] In some embodiments, the second circuit includes a hardware module, which includes multiple processing units, each processing unit is associated with at least one predetermined operation, wherein the first interface is configured to receive a first data packet, wherein the hardware module is configured to: after compiling a first function to be executed in the second circuit, cause at least some of the multiple processing units to perform their associated at least one predetermined operation in a specific order, thereby executing the first function for the first data packet.
[0072] In some embodiments, the first circuit includes a hardware module, the hardware module including a plurality of processing units, each processing unit being associated with at least one predetermined operation, wherein the first interface is configured to receive a first data packet, wherein the hardware module is configured to: during a compilation process for a first function to be executed in the second circuit, cause at least some of the plurality of processing units to perform their associated at least one predetermined operation in a specific order, thereby executing the first function for the first data packet.
[0073] In some embodiments, the at least one controller is configured to perform a compilation process for compiling the first function to be performed by the second circuit.
[0074] In some embodiments, the at least one controller is configured to, before completing the compilation process, instruct the first circuit to perform a first function on a data packet received at the first interface.
[0075] In some embodiments, a compilation process for compiling the first function to be executed by the second circuit is performed by a host device, wherein the at least one controller is configured to determine that the compilation process has been completed in response to receiving an indication of completion of the compilation process from the host device.
[0076] In some embodiments, the invention includes: a processing pipeline for processing a data packet received at a first interface, wherein the processing pipeline includes multiple components, each component is configured to perform one of multiple functions for the data packet received at the first interface, wherein a first component of the multiple components is configured to provide the first function when provided by a first circuit, and wherein a second component of the multiple components is configured to provide the first function when provided by a second at least one processing unit.
[0077] In some embodiments, the at least one controller is configured to control the second circuit to begin performing the first function with respect to the data packet received at the first interface by inserting a second component of the plurality of components into the processing pipeline.
[0078] In some embodiments, the at least one controller is configured to, in response to the determination that a compilation process for the first function to be executed in the second circuit is complete, control the first circuit to stop executing the first function for the data packet received at the first interface by removing a first component of the plurality of components from a processing pipeline.
[0079] In some embodiments, the at least one controller is configured to control the second circuit to begin performing the first function on the data packet received at the first interface by sending a control message through the processing pipeline to start a second component of the plurality of components.
[0080] In some embodiments, the at least one controller is configured to, in response to the determination that a compilation process for the first function to be executed in the second circuit is complete, control the first circuit to deactivate a second component of the plurality of components by sending a control message through the processing pipeline to stop executing the first function for the data packet received at the first interface.
[0081] In some embodiments, a first component of the plurality of components is configured to provide a first function with respect to packets of a first data stream passing through the processing pipeline, wherein a second component of the plurality of components is configured to provide the first function with respect to packets of a second data stream passing through the processing pipeline.
[0082] In some embodiments, the first function includes filtering data packets.
[0083] In some embodiments, the first interface is configured to receive a data packet from the network.
[0084] In some embodiments, the first interface is configured to receive a data packet from the host device.
[0085] In some embodiments, a compile time of the first function of the second circuit is greater than a compile time of the first function of the first circuit.
[0086] According to another aspect, a method is provided, comprising: receiving a data packet at a first interface of a network interface device; performing a first function on the data packet received at the first interface in a first circuit of the network interface device; wherein the first circuit is configured to perform the first function on the data packet received at the first interface during a compilation process for the first function to be performed in a second circuit, the method comprising: determining that the compilation process for the first function to be performed in the second circuit is completed; and in response to the determination, controlling the second circuit of the network interface device to start performing the first function on the data packet received at the first interface.
[0087] According to another aspect, a non-transitory computer-readable medium including program instructions for causing a data processing system to perform a method, the method comprising: receiving a data packet at a first interface of a network interface device; and receiving the data packet at the first interface of the network interface device; performing a first function on the data packet received at the first interface in first circuitry of the network interface device, wherein the first circuitry is configured to perform the first function on the data packet received at the first interface during a compilation process. The first function is to be performed in second circuitry, the method comprising: determining that the compilation process for the first function to be performed in the second circuitry is complete; and in response to the determination, controlling the second circuitry of the network interface device to begin performing the first function on the data packet received at the first interface.
[0088] According to another aspect, a non-transitory computer-readable medium is provided, comprising program instructions for causing a data processing system to perform the following operations: executing a compilation process to compile a first function to be performed by a second circuit of a network interface device; before completing the compilation process, sending a first instruction to cause the first circuit of the network interface device to perform the first function on a data packet received at a first interface of the network interface device; and sending a second instruction to cause the second circuit to begin performing the first function on the data packet received at the first interface after completing the compilation process.
[0089] In some embodiments, the non-transitory computer-readable medium includes program instructions for causing a data processing system to perform a further compilation process to compile a first function to be executed by a first circuit, wherein the compilation process takes longer than the further compilation process.
[0090] In some embodiments, a data processing system includes a host device, wherein the network interface device is configured to interface the host device with a network.
[0091] In some embodiments, the data composition system includes a network interface device, wherein the network interface device is configured to interface the host device with a network.
[0092] In some embodiments, a data processing system includes a host device and a network interface device, wherein the network interface device is configured to interface the host device with a network.
[0093] In some embodiments, the first function includes filtering packets received from the network at the first interface.
[0094] In some embodiments, the non-transitory computer-readable medium includes program instructions for causing the data processing system to execute the following operations: sending a third instruction to cause the first circuit to stop executing the function for the data packet received at the first interface after the compilation process is completed.
[0095] In some embodiments, a non-transitory computer-readable medium includes program instructions for causing a data processing system to: send an instruction to cause the second circuit to perform a first function on a packet of a first data stream; and send an instruction to cause the first circuit to stop performing the first function on a packet of the first data stream.
[0096] In some embodiments, the first circuit includes at least one central processing unit, wherein each of the at least one central processing unit is configured to perform a first function on at least one data packet received at the first interface before completing the second compilation process.
[0097] In some embodiments, the second circuit includes a field programmable gate array configured to initiate execution of the first function for a data packet received at the first interface.
[0098] In some embodiments, the second circuit includes a hardware module, the hardware module including a plurality of processing units, each processing unit being associated with at least one predetermined operation, wherein the data packet received at the first interface includes a first data packet, wherein the hardware module is configured to, after completing the second compilation process, have each processing unit perform a first function on the first data packet, and at least some of the plurality of processing units perform their respective at least one operation on the first data packet.
[0099] In some embodiments, the first circuit includes a hardware module comprising a plurality of processing units configured to provide a first function for a data packet, each processing unit being associated with at least one predetermined operation, wherein the data packet received at the first interface includes a first data packet, wherein the hardware module is configured to perform the first function for the first data packet by each of at least some of the plurality of processing units that perform its respective at least one operation for the first data packet before the second compilation process is completed.
[0100] In some embodiments, the compilation process includes assigning each of the plurality of processing units of the second circuit to perform at least one operation associated with one of the plurality of processing stages in the series of computer code instructions in a particular order.
[0101] In some embodiments, the first functionality provided by the first circuitry is provided as a component of a processing pipeline for processing packets received at the first interface, wherein the first functionality provided by the second circuitry is provided as a component of the processing pipeline.
[0102] In some embodiments, the first instructions include instructions configured to cause a first component of the plurality of components to be inserted into the processing pipeline.
[0103] In some embodiments, the second instructions include instructions configured to cause a second component of the plurality of components to be inserted into the processing pipeline.
[0104] In some embodiments, a non-transitory computer-readable medium includes program instructions for causing a data processing system to perform the following operations: sending a third instruction to cause the first circuit to stop performing the first function for a data packet received at the first interface after a compilation process is completed, wherein the third instruction includes an instruction configured to cause a first component of the plurality of components to be removed from a processing pipeline.
[0105] In some embodiments, the first instruction comprises a control message sent through the processing pipeline to activate a second component of the plurality of components.
[0106] In some embodiments, the second instruction comprises a control message sent through the processing pipeline to activate a second component of the plurality of components.
[0107] In some embodiments, a non-transitory computer-readable medium includes program instructions for causing a data processing system to perform the following operations: sending a third instruction to cause the first circuit to stop performing functions on data packets received at the first interface after a compilation process is completed, wherein the third instruction includes a control message that passes through a processing pipeline to disable a first component of a plurality of components.
[0108] According to another aspect, a data processing system is provided, comprising at least one processor and at least one memory comprising computer program code, wherein the at least one memory and the computer program code are configured, together with the at least one processor, to cause the data processing system to: execute a compilation process to compile a function to be executed by a second circuit of a network interface device; before completing the compilation process, instruct the first circuit of the network interface device to execute the function with respect to a data packet received at a first interface of the network interface device; and after completing the second compilation process, instruct the second at least one processing unit to start executing the function on the data packet received at the first interface.
[0109] According to another aspect, a method for implementation in a data processing system is provided, the method comprising: performing a compilation process to compile a function to be performed by a second circuit of a network interface device; before completing the compilation process, sending a first instruction to cause the first circuit of the network interface device to perform the function on a data packet received at a first interface of the network interface device; and sending a second instruction to cause the second circuit to begin performing the function on the data packet received at the first interface after the compilation process is completed.
[0110] According to another aspect, a non-transitory computer-readable medium is provided, comprising program instructions for causing a data processing system to assign each of a plurality of processing units to perform, in a particular order, at least one operation associated with one of a plurality of processing stages in a series of computer code instructions, wherein the plurality of processing stages provide a first function for a first data packet received at a first interface of a network interface device, wherein each of the plurality of processing units is configured to perform one of a plurality of processing types, wherein at least some of the plurality of processing units are configured to perform different types of processing, wherein for each of the plurality of processing units, the assigning is performed based on a determination regarding the processing type that the processing unit is configured to perform that is suitable for performing the respective at least one operation.
[0111] In some embodiments, each treatment type is defined by one of a plurality of templates.
[0112] In some embodiments, the processing type includes at least one of: accessing a data packet received at the network interface device; accessing a lookup table stored in a memory of the hardware module; performing a logical operation on data loaded from the data packet; and performing a logical operation on data loaded from the lookup table.
[0113] In some embodiments, two or more of at least some of the plurality of processing units are configured to perform their associated at least one operation according to a common clock signal of the hardware module.
[0114] In some embodiments, the allocating includes allocating each of two or more of at least some of the plurality of processing units to perform its associated at least one operation within a predetermined length of time defined by a clock signal.
[0115] In some embodiments, the allocating includes allocating two or more of at least some of the plurality of processing units to access the first data packet for a time period of a predetermined length of time.
[0116] In some embodiments, the allocating includes allocating each of two or more of at least some of the plurality of processing units to transfer a result of the corresponding at least one operation to a next processing unit in response to an end of a time period of a predetermined length of time.
[0117] In some embodiments, the non-transitory computer-readable medium includes program instructions for causing a data processing system to: allocate at least some of the plurality of stages to occupy a single clock cycle.
[0118] In some embodiments, a non-transitory computer-readable medium includes program instructions for causing a data processing system to allocate two or more of a plurality of processing units to perform their allocated at least one operation to be performed in parallel.
[0119] In some embodiments, the network interface device comprises a hardware module including a plurality of processing units.
[0120] In some embodiments, a non-transitory computer-readable medium includes computer program instructions for causing a data processing system to: execute a compilation process including distribution; before the compilation process is completed, send a first instruction to cause circuitry of a network interface device to perform a first function on a data packet received at a first interface; and after the compilation process is completed, send a second instruction to cause a plurality of processing units to begin performing the first function on the data packet received at the first interface.
[0121] In some embodiments, the non-transitory computer-readable medium includes: for one or more of at least some of the plurality of processing units, at least one operation of the assignment includes at least one of: loading at least one value of the first data packet from a memory of a network interface device; storing at least one value of the first data packet in a memory of the network interface device; and performing a lookup table query to determine an action to be performed for the first data packet.
[0122] In some embodiments, a non-transitory computer-readable medium includes computer program instructions for causing a data processing system to issue instructions to configure routing hardware of a network interface device to route a first data packet between a plurality of processing units in a specific order to perform a first function on the first data packet.
[0123] In some embodiments, the first functionality provided by the plurality of processing units is provided as a component of a processing pipeline for processing packets received at the first interface.
[0124] In some embodiments, a non-transitory computer-readable medium includes computer program instructions for causing a plurality of processing units to begin performing a first function on a data packet received at a first interface by causing a data processing system to issue an instruction causing a component to be inserted into a processing pipeline.
[0125] In some embodiments, a non-transitory computer-readable medium includes computer program instructions for causing a plurality of processing units to begin performing a first function on a data packet received at a first interface by causing a data processing system to issue an instruction causing a component to be activated in a processing pipeline.
[0126] In some embodiments, a data processing system includes a host device, wherein the network interface device is configured to interface the host device with a network.
[0127] In some embodiments, a data processing system includes a network interface device.
[0128] In some embodiments, a data processing system includes: a network interface device; and a host device, wherein the network interface device is configured to connect the host device to a network interface.
[0129] According to another aspect, a data processing system is provided, comprising at least one processor and at least one memory including computer program code, wherein the at least one memory and the computer program code are configured, together with the at least one processor, to cause the data processing system to allocate each of a plurality of processing units to perform, in a particular order, at least one operation associated with one of a plurality of processing stages in a series of computer code instructions, wherein the plurality of processing stages provide a first function with respect to a first data packet received at a first interface of a network interface device, wherein each of the plurality of processing units is configured to perform one of a plurality of processing types, wherein at least some of the plurality of processing units are configured to perform different processing types, wherein for each of the plurality of processing units, the allocation is performed based on a determination that the processing unit is configured to perform a processing type that is suitable for performing the corresponding at least one operation.
[0130] According to another aspect, a method is provided, comprising: assigning each of a plurality of processing units to perform at least one operation associated with one of a plurality of processing stages in a series of computer code instructions in a particular order, wherein the plurality of processing stages provide a first function for a first data packet received at a first interface of a network interface device, wherein each of the plurality of processing units is configured to perform one of a plurality of processing types, wherein at least some of the plurality of processing units are configured to perform different types of processing, wherein for each of the plurality of processing units, the assigning is performed based on a determination that the processing unit is configured to perform a type of processing suitable for performing the corresponding at least one operation.
[0131] The processing units of the hardware modules have been described as performing their types of operations in a single step. However, those skilled in the art will recognize that this feature is merely a preferred feature and is not necessary or essential to the functionality of the present invention.
[0132] According to one aspect, a method is provided that includes receiving a bitfile description and a program at a compiler, the bitfile description including a description of routing of a portion of a circuit; and compiling the program using the bitfile description to output a bitfile for the program.
[0133] The method may include using the bitfile to configure at least a portion of the portion of the circuitry to perform a function associated with the program.
[0134] The bitfile description may include information regarding routing between the plurality of processing units of the portion of the circuit.
[0135] The bitfile description may include routing information for at least one of the plurality of processing units, the routing information indicating at least one of: data to which one or more other processing units may be output, and data from which one or more other processing units may be received.
[0136] The bit file description may include routing information indicating one or more routes between two or more corresponding processing units.
[0137] The bitfile description may include information indicating only routes that may be used by the compiler when compiling a program to provide a bitfile for the program.
[0138] The bit file may include information indicating at least one of the following for each processing unit: input is provided from one or more of the one or more other processing units in the bit file description for the corresponding processing unit; output is provided to one or more of the one or more other processing units in the bit file description.
[0139] A portion of the circuitry may include at least a portion of a configurable hardware module, the configurable hardware module including a plurality of processing units, each processing unit being associated with a predetermined operation type that can be performed in a single step, at least some of the plurality of processing units being associated with different predetermined operation types, the bit file description including information regarding routing between at least some of the plurality of processing units, wherein the method may include using the bit file to hardware interconnect at least some of the plurality of processing units to provide a first data processing pipeline for processing one or more of the plurality of data packets to perform a first function on the one or more of the plurality of data packets.
[0140] The bit file description may be at least a portion of an FPGA.
[0141] The bit file description may be part of a dynamically programmable FPGA.
[0142] The program may include one of an eBPF program and a P4 program.
[0143] The compiler and the FPGA may be provided in the network interface device.
[0144] According to another aspect, a device is provided, comprising at least one processor and at least one memory, the at least one memory comprising computer code for one or more programs, the at least one memory and the computer code being configured to, with the at least one processor, cause the device to: receive a bit file description and a program, the bit file description comprising a description of routing of a portion of a circuit; and compile the program using the bit file description to output a bit file for the program.
[0145] The at least one memory and the computer code may be configured with the at least one processor to cause the apparatus to configure at least a portion of the portion of the circuitry using the bit file to perform functions associated with the program.
[0146] The bitfile description may include information regarding routing between the plurality of processing units of the portion of the circuit.
[0147] The bit file description may include: for at least one of the plurality of processing units, routing information indicating at least one of the following: data of one or more other processing units may be output thereto; and data of one or more other processing units may be received therefrom.
[0148] The bitfile description may include routing information indicating one or more routes between two or more corresponding processing units.
[0149] A bitfile description may include information indicating only routes that are available to a compiler when compiling a program to provide a bitfile for the program.
[0150] The bit file may include information for indicating the corresponding processing unit, at least one of which: provides input to the one or more other processing units in the bit file description of the corresponding processing unit; and provides output to one or more of the one or more other processing units in the bit file description of the corresponding processing unit.
[0151] The portion of the circuitry may include at least a portion of a configurable hardware module, the configurable hardware module including a plurality of processing units, each processing unit associated with a predetermined operation type that can be performed in a single step, at least some of the plurality of processing units being associated with different predetermined operation types, the bit file description including information about routing between at least some of the plurality of processing units, wherein the at least one memory and the computer code are configured, with the at least one processor, to cause the apparatus to use the bit file to hardware interconnect at least some of the plurality of processing units to provide a first data processing pipeline for processing one or more of the plurality of data packets to perform a first function on the one or more of the plurality of data packets.
[0152] The bit file description may be at least a portion of an FPGA.
[0153] The bit file description may be part of a dynamically programmable FPGA.
[0154] The program may include one of an eBPF program and a P4 program.
[0155] According to another aspect, a network interface device is provided, comprising: a first interface configured to receive a plurality of data packets; a configurable hardware module comprising a plurality of processing units, each processing unit being associated with a predetermined type of operation that can be performed in a single step; and a compiler configured to receive a bit file description and a program, the bit file description comprising a description of routing for at least a portion of the configurable hardware module, and to compile the program using the bit file description to output a bit file for the program, wherein the hardware module can be configured using the bit file to perform a first function associated with the program.
[0156] A network interface device may be used to connect a host device to a network interface.
[0157] At least some of the plurality of processing units may be associated with different predetermined operation types.
[0158] The hardware module may be configured to interconnect at least some of the plurality of processing units to provide a first data processing pipeline for processing one or more of the plurality of data packets to perform a first function on the one or more of the plurality of data packets.
[0159] In some embodiments, the first function includes a filtering function. In some embodiments, the function includes at least one of a tunnel, an encapsulation, and a routing function. In some embodiments, the first function includes an extended Berkeley packet filtering function.
[0160] In some embodiments, the first function comprises a distributed denial of service cleanup operation.
[0161] In some embodiments, the first function includes firewall operations.
[0162] In some embodiments, the first interface is configured to receive a first data packet from a network.
[0163] In some embodiments, the first interface is configured to receive a first data packet from a host device.
[0164] In some embodiments, two or more of at least some of the plurality of processing units are configured to perform their associated at least one predetermined operation in parallel.
[0165] In some embodiments, two or more of at least some of the plurality of processing units are configured to perform their associated predetermined operation types according to a common clock signal of the hardware module.
[0166] In some embodiments, each of two or more of the at least some of the plurality of processing units is configured to perform its associated predetermined type of operation within a predetermined length of time defined by a clock signal.
[0167] In some embodiments, two or more of the at least some of the multiple processing units are configured to: access the first data packet within a time period of a predetermined time length; and transmit the result of the corresponding at least one operation to the next processing unit in response to the end of the predetermined time length.
[0168] In some embodiments, the result includes at least one or more of: at least one value from one or more of the plurality of data packets; an updated value of the mapping state; and metadata.
[0169] In some embodiments, each of the plurality of processing units includes an application specific integrated circuit configured to perform at least one operation associated with the corresponding processing unit.
[0170] In some embodiments, each processing unit comprises a field programmable gate array. In some embodiments, each processing unit comprises any other type of soft logic.
[0171] In some embodiments, at least one of the plurality of processing units includes digital circuitry and memory storing state related to processing performed by the digital circuitry, wherein the digital circuitry is configured to perform predetermined operation types associated with the respective processing units in communication with the memory.
[0172] In some embodiments, the network interface device includes a memory accessible by two or more of the plurality of processing units, wherein the memory is configured to store a state associated with the first data packet, wherein during performance of the first function by the hardware module, the two or more of the plurality of processing units are configured to access and modify the state.
[0173] In some embodiments, a first one of the at least some of the plurality of processing units is configured to stall during access of the state value by a second one of the plurality of processing units.
[0174] In some embodiments, one or more of the plurality of processing units are independently configurable to perform operations specific to respective pipelines based on their associated predetermined operation types.
[0175] In some embodiments, the hardware module is configured to receive instructions and, in response to the instructions, perform at least one of the following: interconnect at least some of the plurality of processing units to provide a data processing pipeline for processing one or more data packets; cause one or more of the plurality of processing units to perform their associated predetermined operation types on the one or more data packets; add one or more of the plurality of processing units to the data processing pipeline; and remove one or more of the plurality of processing units from the data processing pipeline.
[0176] In some embodiments, the predetermined operation includes at least one of: loading at least one value of the first data packet from memory; storing at least one value of the data packet in memory; and performing a lookup in a lookup table to determine an action to be performed on the data packet.
[0177] In some embodiments, the hardware module is configured to receive an instruction, wherein the hardware module is configurable to interconnect at least some of the plurality of processing units in response to the instruction to provide a data processing pipeline for processing one or more of the plurality of data packets, wherein the instruction includes sending a data packet through the third processing pipeline.
[0178] In some embodiments, one or more of at least some of the plurality of processing units may be configured to, in response to the instruction, perform a selected operation of the associated predetermined operation type on one or more of the plurality of data packets.
[0179] In some embodiments, the plurality of components includes a second one of the plurality of components configured to provide the first function in a circuit distinct from a hardware module, wherein the network interface device includes at least one controller configured to cause packets transmitted through the processing pipeline to be processed by one of the first one of the plurality of components and the second one of the plurality of components.
[0180] In some embodiments, a network interface device includes at least one controller configured to issue instructions to cause a hardware module to begin performing a first function on a data packet, wherein the instructions are configured to cause a first component of a plurality of components to be inserted into a processing pipeline.
[0181] In some embodiments, a network interface device includes at least one controller configured to issue instructions to cause a hardware module to begin performing a first function on a data packet, wherein the instructions include a control message sent through a processing pipeline and configured to cause a first component of a plurality of components to activate.
[0182] In some embodiments, for one or more of at least some of the plurality of processing units, the associated at least one operation includes at least one of: loading at least one value of the first data packet from a memory of the network interface device; storing at least one value of the first data packet in the memory of the network interface device; and performing a lookup table query to determine an action to be performed for the first data packet.
[0183] In some embodiments, one or more of at least some of the multiple processing units are configured to pass at least one result of at least one predetermined operation associated therewith to a next processing unit in the first processing pipeline, and the next processing unit is configured to perform a next predetermined operation based on the at least one result.
[0184] In some embodiments, each of the different predetermined operation types is defined by a different template.
[0185] In some embodiments, the predetermined operation type includes at least one of: accessing a data packet; accessing a lookup table stored in a memory of the hardware module; performing a logical operation on data loaded from the data packet; and performing a logical operation on data loaded from the lookup table.
[0186] In some embodiments, the hardware module includes routing hardware, wherein the hardware module is configurable to interconnect at least some of the plurality of processing units to provide a first data processing pipeline by configuring the routing hardware to route data packets between the plurality of processing units in a specific order defined by the first data processing pipeline.
[0187] In some embodiments, the hardware module may be configured to interconnect at least some of the plurality of processing units to provide a second data processing pipeline for processing one or more of the plurality of data packets to perform a second function different from the first function.
[0188] In some embodiments, the hardware module may be configured to interconnect at least some of the plurality of processing units to provide a second data processing pipeline after interconnecting at least some of the plurality of processing units to provide a first data processing pipeline.
[0189] In some embodiments, the network interface device includes additional circuitry separate from the hardware module and configured to perform the first function on one or more of the plurality of data packets.
[0190] In some embodiments, the additional circuitry includes at least one of: a field programmable gate array; and a plurality of central processing units.
[0191] In some embodiments, the network interface device includes at least one controller, wherein the further circuitry is configured to perform the first function with respect to the data packet during a compilation process for the first function to be performed in the hardware module, wherein the at least one controller is configured to control the hardware module to begin performing the first function with respect to the data packet in response to completion of the compilation process.
[0192] In some embodiments, the further circuitry comprises a plurality of central processing units.
[0193] In some embodiments, the at least one controller is configured to control the further circuit to stop executing the first function on the data packet in response to the determination that the compilation process of the first function to be executed in the hardware module is completed.
[0194] In some embodiments, the network interface device includes at least one controller, wherein the hardware module is configured to perform the first function on the data packet during a compilation process for the first function to be performed in the additional circuit, wherein the at least one controller is configured to determine that the compilation process of the first function to be performed in the additional circuit has been completed, and in response to the determination, control the additional circuit to start performing the first function on the data packet.
[0195] In some embodiments, the further circuitry comprises a field programmable gate array.
[0196] In some embodiments, the at least one controller is configured to, in response to the determination that a compilation process for the first function to be executed in the further circuit has been completed, control the hardware module to stop executing the first function on the data packet.
[0197] In some embodiments, the network interface device includes at least one controller configured to perform a compilation process to provide a first function to be executed in a hardware module.
[0198] In some embodiments, the compilation process includes providing instructions to provide a control plane interface in the hardware module that is responsive to the control messages.
[0199] According to another aspect, a computer-implemented method is provided, the method comprising determining routing information for at least a portion of a configurable hardware module comprising a plurality of processing units, each processing unit being associated with a predetermined type of operation executable in a single step, at least some of the plurality of processing units being associated with different predetermined types of operations, the routing information providing information regarding available routes between at least the plurality of processing units.
[0200] The configurable hardware module may include a substantially static portion and a substantially dynamic portion, and the determining may include determining routing information for the substantially dynamic portion.
[0201] The routing information for the substantially dynamic portion may comprise determining routes in the substantially dynamic portion used by one or more of the processing units in the substantially static portion.
[0202] The determining may include analyzing a bitfile description of at least a portion of the configurable hardware module to determine the routing information.
[0203] According to another aspect, a non-transitory computer-readable medium including program instructions is provided, the program instructions being configured to determine routing information for at least a portion of a configurable hardware module including a plurality of processing units, each processing unit being associated with a predetermined type of operation executable in a single step, at least some of the plurality of processing units being associated with different predetermined types of operations, the routing information providing information regarding available routes between at least the plurality of processing units.
[0204] A computer program comprising program code means adapted to perform the method may also be provided.The computer program may be stored and / or embodied in other ways by means of a carrier medium.
[0205] In the above, many different embodiments have been described. It should be understood that other embodiments can be provided by combining any two or more of the above embodiments.
[0206] Various other aspects and further embodiments are described in the following detailed description and in the appended claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0207] Some embodiments will now be described, by way of example only, with reference to the accompanying drawings, in which:
[0208] Figure 1 A schematic diagram showing a data processing system connected to a network;
[0209] Figure 2 A schematic diagram of a data processing system including a filter operation application configured to run in user mode on a host computing device is shown;
[0210] Figure 3 A schematic diagram of a data processing system including a filtering operation configured to run in kernel mode on a host computing device is shown;
[0211] Figure 4 A schematic diagram of a network interface device including multiple CPUs that perform functions on data packets is shown;
[0212] Figure 5 shows a schematic diagram of a network interface device including a field programmable gate array running an application for performing functions on data packets;
[0213] Figure 6 A schematic diagram of a network interface device including hardware modules for performing functions on data packets is shown;
[0214] Figure 7 shows a schematic diagram of a network interface device comprising a field programmable gate array and at least one processing unit for performing functions on data packets;
[0215] Figure 8 A method implemented in a network interface device according to some embodiments is shown;
[0216] Figure 9 A method implemented in a network interface device according to some embodiments is shown;
[0217] Figure 10 An example of processing a data packet through a series of procedures is shown;
[0218] Figure 11 An example of processing a data packet by multiple processing units is shown;
[0219] Figure 12 An example of processing a data packet by multiple processing units is shown;
[0220] Figure 13 An example of a pipeline of processing stages for processing a data packet is shown;
[0221] Figure 14An example of a slicing architecture with multiple pluggable components is shown;
[0222] Figure 15 An example representation showing the arrangement and order of processing of a plurality of processing units is shown;
[0223] Figure 16 An example method of compiling functionality is shown.
[0224] Figure 17 An example of a stateful processing unit is shown.
[0225] Figure 18 An example of a stateless processing unit is shown.
[0226] Figure 19 The methods of some embodiments are shown;
[0227] Figure 20a and 20b shows routing between slices in the FPGA; and
[0228] Figure 21 The partitioning on the FGPA is schematically shown. Specific implementation plan
[0229] The following description is provided to enable any person skilled in the art to make and use the invention, and is provided in the context of a particular application. Various modifications to the disclosed embodiments will be apparent to those skilled in the art.
[0230] Without departing from the spirit and scope of the present invention, the general principles defined herein may be applied to other embodiments and applications. Therefore, the present invention is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features disclosed herein.
[0231] When data is transferred between two data processing systems over a data channel, such as a network, each data processing system must have a suitable network interface to allow communication across the channel. Networks are typically based on Ethernet technology. Data processing systems communicating over a network are equipped with network interfaces that support the physical and logical requirements of the network protocol. The physical hardware component of a network interface is called a network interface device or network interface card (NIC).
[0232] Most computer systems include an operating system (OS), through which user-level applications communicate with the network. Part of the operating system, called the kernel, includes a protocol stack for translating commands and data between applications and device drivers specific to network interface devices. Device drivers can directly control network interface devices. Providing these functions within the operating system kernel hides the complexity and differences of network interface devices from user-level applications. This allows network hardware and other system resources, such as memory, to be securely shared by many applications, protecting the system from erroneous or malicious applications.
[0233] exist Figure 1 1 shows a typical data processing system 100 for performing transmissions over a network. Data processing system 100 includes a host computing device 101 coupled to a network interface device 102, which is configured to interface the host to a network 103. Host computing device 101 includes an operating system 104 that supports one or more user-level applications 105. Host computing device 101 may also include a network protocol stack (not shown). For example, the protocol stack may be a component of the application, a library linked with the application, or a library provided by the operating system. In some embodiments, more than one protocol stack may be provided.
[0234] The network protocol stack may be a Transmission Control Protocol (TCP) stack. Application 105 may send and receive TCP / IP messages by opening a socket and reading and writing data to and from the socket, and operating system 104 transmits the message across the network. For example, an application may invoke a system call (syscall) to transmit data through a socket, which is then transmitted to network 103 by operating system 104. This interface for transmitting messages may be referred to as a message passing interface.
[0235] Instead of implementing the stack in the host computer 101, some systems offload the protocol stack to the network interface device 102. For example, if the stack is a TCP stack, the network interface device 102 may include a TCP offload engine (TOE) for performing TCP protocol processing. By performing protocol processing in the network interface device 102 rather than in the host computing device 101, the demands on the host system 101 processor can be reduced. Data to be transmitted over the network can be sent by the application 105 through a TOE-enabled virtual interface driver, partially or completely bypassing the kernel TCP / IP stack. Therefore, data sent along this fast path only needs to be formatted to meet the requirements of the TOE driver.
[0236] Host computing device 101 may include one or more processors and one or more memories. In some embodiments, host computing device 101 and network interface device 102 may communicate via a bus, such as a Peripheral Component Interconnect Express (PCIe) bus.
[0237] During operation of the data processing system, data to be transmitted over a network may be transferred from the host computing device 101 to the network interface device 102 for transmission. In one example, data packets may be transmitted directly from the host to the network interface device by the host processor. The host may provide data to one or more buffers 106 located on the network interface device 102. The network interface device 102 may then prepare the data packets and send them over the network 103.
[0238] Alternatively, the data may be written to a buffer 107 in the host system 101. The network interface device may then retrieve the data from the buffer 107 and transmit it over the network 103.
[0239] In both cases, data is temporarily stored in one or more buffers before being transmitted over the network. Data sent over the network can be returned to the host (backtracking).
[0240] When sending and receiving packets over the network 103, there are many processing tasks that can be represented as operations on packets to be sent over the network or on packets received over the network. For example, filtering processing can be performed on received packets to protect the host system 101 from distributed denial of service (DDOS) filtering. Such filtering processes can be performed by simple packet inspection or extended Berkeley Packet Filter (eBPF). As another example, encapsulation and forwarding can be performed on packets to be transmitted over the network 103. These processes can consume many CPU cycles and are heavy for conventional OS architectures.
[0241] refer to Figure 2 , which illustrates one way in which filtering operations or other packet processing operations can be implemented in the host system 220. The processes performed by the host system 220 are shown as being executed in either user space or kernel space. In kernel space, there is a receive path for delivering packets received from the network at the network interface device 210 to the terminal application 250. This receive path includes a driver 235, a protocol stack 240, and a socket 245. The filtering operation 230 is implemented in user space. Incoming packets provided to the host system 220 by the network interface device 210 bypass the kernel (where protocol processing occurs) and are provided directly to the filtering operation 230.
[0242] Filtering operation 230 is provided with a virtual interface (which may be an Ethernet Fabric Virtual Interface (EFVI), a Data Plane Development Kit (DPDK), or any other suitable interface) for exchanging packets with other components in host system 220. Filtering operation 230 may perform DDOS scrubbing and / or other forms of filtering. The DDOS scrubbing process may be performed on all packets that are easily identified as DDOS candidates—for example, sample packets, packet copies, and packets that have not yet been classified. Packets not passed to filtering operation 230 may be passed directly from the network interface to driver 235. Operation 230 may provide an extended Berkeley Packet Filter (eBPF) for performing filtering. If a received packet passes the filtering provided by operation 230, operation 230 is configured to reinject the packet into the kernel's receive path for processing. Specifically, the packet is provided to driver 235 or stack 240. Protocol stack 240 then performs protocol processing on the packet. The packet is then passed to socket 245 associated with terminating application 250. The terminating application 250 issues a recv() call to retrieve the data packet from the associated socket's buffer.
[0243] However, this approach presents several issues. First, filtering operation 230 runs on the host CPU. To run filtering 230, the host CPU must process data packets at the rate at which they are received from the network. This can consume significant processing resources on the host CPU if the rate at which data is sent and received from the network is high. The high data flow rate to filtering operation 230 can lead to significant consumption of other limited resources, such as I / O bandwidth and internal memory / cache bandwidth.
[0244] In order to reinject packets into the kernel, it is necessary to provide a privileged API for performing reinjection to the filter operation 230. The reinjection process can be cumbersome and requires attention to the ordering of packets. In order to perform the reinjection, the operation 230 may require a dedicated CPU core in many cases.
[0245] The steps of providing data for operations and re-injecting it require copying the data into and out of memory. This copying is a resource burden on the system.
[0246] Similar issues may arise when providing other types of operations instead of filtering the data to be sent / received over the network.
[0247] Certain operations, such as DPDK-type operations, may require that processed packets be forwarded back to the network.
[0248] refer to Figure 3, which shows another approach. Like elements are labeled with the same reference numerals. In this example, an additional layer called the Express Data Path (XDP) 310 is inserted into the send and receive paths in the kernel. Extensions to XDP 310 allow insertion into the transmit path. XDP helpers allow packets to be sent (as a result of receive operations). XDP 310 is inserted at the driver level of the operating system and allows programs to execute at this level to perform operations on packets received from the network before they are processed by the protocol stack 240. XDP 310 also allows programs to execute at this level to perform operations on packets to be sent over the network. Thus, eBPF programs and other programs can run in the send and receive paths.
[0249] like Figure 3 As shown, a filtering operation 320 can be inserted into XDP from user space to form a program 330 as part of XDP 310. Operation 320 is inserted using the XDP control plane to be executed on the data receive path to provide a program 330 that performs filtering operations (e.g., DDOS cleaning) on packets on the receive path. Such a program 330 can be an eBPF program.
[0250] Program 330 is shown inserted into the kernel between driver 235 and protocol stack 240. However, in other examples, program 330 can be inserted elsewhere in the receive path within the kernel. Program 330 can be part of a separate control path for receiving packets. Program 330 can be provided by an application program through an extension of the application programming interface (API) that provides socket 245 to the application program.
[0251] Program 330 may additionally or alternatively perform one or more operations on data sent via the transmission path. XDP 310 then invokes the send function of driver 235 to send the data over the network via network interface device 210. In this case, program 330 may provide load balancing or routing operations for data packets to be sent over the network. Program 330 may also provide segment repackaging and forwarding operations for data packets to be sent over the network.
[0252] Program 330 may be used for firewall and virtual switching or other operations that do not require protocol termination or application processing.
[0253] One advantage of using XDP 310 in this manner is that program 330 can directly access the memory buffers handled by the driver without an intermediate copy.
[0254] In order to insert a program 330 to be run into the kernel in this manner, it is necessary to ensure that the program 330 is safe. If an unsafe program is inserted into the kernel, certain risks are introduced, such as infinite loops that may cause kernel crashes, buffer overflows, uninitialized variables, compiler errors, and performance issues caused by large programs.
[0255] In order to ensure that the program 330 is safe before being inserted into the XDP 310 in this manner, a verification program can be run on the host system 220 to verify the safety of the program 330. The verification program can be configured to ensure that there are no loops. If no loop is caused, a backward jump operation can be allowed. The verification program can be configured to ensure that the program 330 has no more than a predetermined number of instructions (e.g., 4000). The verification program can perform a check on the validity of register usage by traversing the data path of the program 330. If there are too many possible paths, the program 330 will be rejected because it is not safe to run in kernel mode. For example, if there are more than 1000 branches, the program 330 can be rejected.
[0256] Those skilled in the art will appreciate that XDP is one example of a security program 330 that may be installed in the kernel, and that there are other ways to accomplish this.
[0257] If, for example, the operations can be expressed in a safe (or sandboxed) language required to execute code in the kernel, then the above explanation about Figure 3 The method discussed above can be compared with Figure 2 The methods discussed are just as efficient. The eBPF language can be executed efficiently on x86 processors, and JIT (Just in Time) compilation technology allows eBPF programs to be compiled into native machine code. The language is designed to be safe; for example, state is limited to structures that only map to shared data structures (such as hash tables). Bounded loops are allowed, rather than allowing one eBPF program to make tail calls to another. The state space is restricted.
[0258] However, in some embodiments, using this approach can significantly deplete the resources of host system 220 (e.g., I / O bandwidth and internal memory / cache bandwidth, host CPU). Operations on the data packets are still performed by the host CPU, which is required to perform such operations at the rate at which the data is sent / received.
[0259] Another suggestion is to perform the above operations in the network interface device rather than in the host system. In addition to the I / O bandwidth, memory, and cache bandwidth consumed, doing so may free up CPU cycles used by the host CPU when performing the operations. Moving the execution of processing operations from the host to the network interface device hardware may introduce some challenges.
[0260] One proposal to implement processing in network hardware is to provide a network processing unit (NPU) in the network interface device that includes multiple CPUs dedicated to packet processing and / or manipulation operations.
[0261] refer to Figure 4 , which shows an example of a network interface device 400 that includes an array 410 of central processing units (CPUs), such as CPU 420. The CPUs are configured to perform functions such as filtering packets sent and received from a network. Each CPU in the CPU array 410 may be an NPU. Although in Figure 4 Although not shown, the CPUs may additionally or alternatively be configured to perform operations such as load balancing packets received from hosts for transmission over the network. These CPUs are dedicated to such packet processing / manipulation operations. The CPUs execute an instruction set optimized for such packet processing / manipulation operations.
[0262] The network interface device 400 additionally includes memory (not shown) that is shared between and accessible by the array of CPUs 410 .
[0263] The network interface device 400 includes a network medium access control (MAC) layer 430 for interfacing the network interface device 400 with a network. The MAC layer 430 is configured to receive data packets from the network and to transmit data packets on the network.
[0264] Operations on packets received at the network interface device 400 are parallelized on the CPU. As shown, when a data stream is received at the MAC layer 430, the data stream is passed to the expansion function 440, which is configured to extract packets from the data stream and distribute them across multiple CPUs in the NPU 410, allowing the CPUs to perform processing, such as filtering, on the packets. The expansion function 440 can parse the received packets to identify the data streams to which they belong. For each packet, the expansion function 440 generates an indication of its position within the data stream to which it belongs. This indication can be, for example, a tag. The expansion function 440 adds the corresponding indication to the metadata associated with each packet. The metadata associated with each packet can be appended to the packet. The associated metadata can be passed to the expansion function 440 as sideband control information. Adding an indication based on the stream to which the packet belongs allows the sequence of packets for any particular stream to be reconstructed.
[0265] After being programmed by the plurality of CPUs 410, the data packets are then passed to the reordering function 450, which reorders the data packets of the data stream into their correct order before passing them to the host interface layer 460. The reordering function 450 can reorder the data packets of the data stream by comparing indications (e.g., tags) within the data packets of the data stream to reconstruct the order of the data packets. The reordered data packets are then passed through the host interface 460 and transmitted to the host system 220.
[0266] although Figure 4 The CPU array 410 is illustrated as operating only on packets received from the network, but similar principles (including expansion and reordering) can be performed on packets received from a host for transmission over the network, with the CPU array 410 performing functions (e.g., load balancing) on these packets received from the host.
[0267] The program executed by the CPU may be the one described above with respect to Figure 3 In the example described, a compiled or transcoded version of the program is executed on the host CPU. In other words, the instruction set to be executed on the host CPU for operation is converted to be executed on each CPU of the dedicated CPU array in the network interface 400.
[0268] To achieve parallelization on the CPU, multiple instances of a program are compiled and executed in parallel on multiple CPUs. Each instance of the program can be responsible for processing a different set of data packets received at the network interface device. However, when providing program functionality with respect to that data packet, each individual data packet is processed by a single CPU. The overall effect of executing the parallel programs can be the same as the execution of a single program (e.g., program 330) on the host CPU.
[0269] One of the dedicated CPUs can process 50 million packets per second. This operating speed may be slower than the host CPU's. Therefore, parallelization can be used to achieve the same performance as an equivalent program executed on the host CPU. To implement parallelization, packets are distributed across the CPUs and then reordered after being processed by the CPUs. The requirement to process packets for each flow sequentially, along with reordering step 450, can create bottlenecks, increase memory resource overhead, and potentially limit the device's available throughput. This requirement and reordering step 450 can increase jitter on the device, as processing throughput can fluctuate depending on the content of the network traffic and the degree of applicable parallelism.
[0270] One advantage of using such a specialized CPU may be the short compilation time. For example, it may be possible to compile a filtering application to run on such a CPU in less than 1 second.
[0271] When this approach is scaled to higher link speeds, using a CPU array may become problematic. In the near future, host network interfaces may be required to reach terabit per second speeds. When scaling such a CPU array 410 to these higher speeds, the amount of power required may become problematic.
[0272] Another proposal is to include a field programmable gate array (FPGA) in the network interface device and use the FPGA to perform operations on data packets received from the network.
[0273] refer to Figure 5 , which shows an example of using an FPGA 510 with an FPGA application 515 in a network interface device 500 for performing operations on data packets received at the network interface device 500. Figure 4 The same elements in FIG. 1 are designated by the same reference numerals.
[0274] although Figure 5 An FPGA application 515 is shown operating only on packets received from a network, but such an FPGA application 515 can be used to perform functions (e.g., load balancing and / or firewall functions) on these packets received from a host for transmission over the network or returned to the host or another network interface on the system.
[0275] FPGA application 515 may be provided by compiling a program written in a general-purpose system-level language, such as C or C++ or Scala, to run on FPGA 510 .
[0276] The FPGA 510 may have both network interface functionality and FPGA functionality. The FPGA functionality may provide an FPGA application 515, which may be programmed into the FPGA 510 based on the needs of the network interface device user. The FPGA application 515 may, for example, provide filtering of messages on the receive path from the network 230 to the host. The FPGA application 515 may also provide a firewall.
[0277] FPGA 510 can be programmed to provide FPGA applications 515. Some network interface device functions can be implemented as "hard" logic within FPGA 510. For example, hard logic can be application specific integrated circuit (ASIC) gates. FPGA applications 515 can be implemented as "soft" logic. Soft logic can be provided by programming FPGA LUTs (look-up tables). Hard logic can be clocked at a higher rate than soft logic.
[0278] The network interface device 500 includes a host interface 505 configured to send and receive data with a host. The network interface device 520 includes a network medium access control (MAC) interface 520 configured to send and receive data with a network.
[0279] When a packet is received from the network at MAC interface 520, it is passed to FPGA application 515, which is configured to perform functions such as filtering on the packet. The packet (if it passes any filtering) is then passed to host interface 505, from where it is delivered to the host. Alternatively, packet FPGA application 515 can decide to discard or resend the packet.
[0280] One issue with using FPGAs to perform functions on packets is the relatively long compile times required. FPGAs are composed of many logic elements (e.g., logic cells), each representing primitive logic operations such as AND, OR, and NOT. These logic cells are arranged in a matrix with programmable interconnects. To deliver functionality, these logic cells may need to operate together to achieve circuit definition and synchronize clock timing constraints. Placing each logic cell and routing the wires between them can be an algorithmic challenge. On lightly utilized FPGAs, compile times can be less than ten minutes. However, as FPGAs are increasingly utilized for a variety of applications, the placement and routing challenges can increase, increasing the time it takes to compile a given function onto the FPGA. Consequently, adding additional logic to an FPGA that already consumes a significant portion of its routing resources can require hours of compile time.
[0281] One approach is to design hardware using specific processing primitives, such as parsing, matching, and action primitives. These can be used to construct a processing pipeline in which all packets undergo each of three steps. First, the packet is parsed to construct a metadata representation of the protocol header. Second, the packet is flexibly matched against rules stored in a table. Finally, when a match is found, the packet is acted upon based on the table entry selected in the match operation.
[0282] To implement functionality using the parsing / matching / action model, the P4 programming language (or similar languages) can be used. The P4 programming language is target-independent, meaning that programs written in P4 can be compiled to run on different types of hardware (e.g., CPU, FPGA, ASIC, NPU, etc.). Each different type of target provides its own compiler to map the P4 source code to the appropriate target switch model.
[0283] P4 can be used to provide a programming model that allows high-level programs to express packet processing operations for a packet processing pipeline. This approach is suitable for expressing operations naturally in a declarative style. In the P4 language, programmers express the parsing, matching, and action phases as operations to be performed on received packets. These operations are aggregated to enable efficient execution by specialized hardware. However, this declarative style may not be suitable for expressing programs that are imperative in nature, such as eBPF programs.
[0284] In a network interface device, a series of eBPF programs may need to be executed serially. In this case, a chain of eBPF programs is generated, with each program calling another. Each program can modify state, and the output is as if the entire chain had been executed sequentially. For a compiler, collecting all the parsing, matching, and manipulation steps can be challenging. Even after an eBPF program chain has been installed, it may be necessary to install, remove, or modify the chain, which can present further challenges.
[0285] To provide an example of such a procedure that needs to be repeated, see Figure 10 , Figure 10 An example of a sequence of programs e1, e2, and e3 configured to process a packet is shown. For example, each program can be an eBPF program. Each program is configured to parse a received packet, perform a lookup of table 1010 to determine an action in a matching entry in table 1010, and then perform the action for the packet. The action may include modifying the packet. Each eBPF program may also perform operations based on local and shared state. Packet P0 is initially processed by eBPF program e1 and then passed and modified to the next program e2 in the pipeline. The output of the program sequence is the output of the final program in the pipeline, i.e., e3.
[0286] It may be complex for a compiler to combine the effects of each of n such programs into a single P4 program. Additionally, some programming models (e.g., XDP) may require that programs be inserted and removed dynamically and quickly at any point in the program sequence in response to changing circumstances.
[0287] According to some embodiments of the present application, a network interface device including a plurality of processing units is provided. Each processing unit is configured to perform at least one predetermined operation in hardware. Each processing unit includes a memory for storing its own local state. Each processing unit includes a digital circuit for modifying the state. The digital circuit may be a dedicated integrated circuit. Each processing unit is configured to run a program including configurable parameters to perform a corresponding plurality of operations. Each processing unit may be an atom. Atoms are defined by specific programming and routing of a predetermined template. This defines their specific operational behavior and their logical position in the process provided by the connected plurality of processing units. Where the term "atom" is used in the specification, this may be understood to refer to a data processing unit configured to perform its operation in a single step. In other words, an atom performs its operation as an atomic operation.
[0288] An atom can be thought of as a collection of hardware structures that can be configured to repeatedly perform one of a series of computations, taking one or more inputs and producing one or more outputs.
[0289] Atoms are provided by the hardware. Atoms can be configured by the compiler. Atoms can be configured to perform computations.
[0290] During compilation, at least some of the plurality of processing units are configured to execute operations to perform functions on data packets received by at least some of the plurality of processing units at the network interface device. Each of at least some of the plurality of processing units is configured to perform at least one predetermined operation to perform functions with respect to the data packets. In other words, the operations that the connected processing units are configured to perform are performed on the received data packets. The operations are performed sequentially by at least some of the plurality of processing units. Collectively, execution of each of the plurality of operations provides functions, such as filtering the received data packets.
[0291] By arranging each atom to perform at least one of their respective predetermined operations to perform a function, as above with respect to Figure 5 Compile time can be reduced compared to the FPGA application example described above. Moreover, for functions that use processing units dedicated to performing specific operations in hardware, such as the one above Figure 4 As discussed, the speed at which functions are performed can be increased compared to using a CPU in a network interface device to execute software to perform the functions for each data packet.
[0292] refer to Figure 6 , which shows an example of a network interface device 600 according to an embodiment of the present application. The network interface device includes a hardware module 610 configured to perform processing of data packets received at an interface of the network interface device 600. Although Figure 6The hardware module 610 is shown performing functions (eg, filtering) on packets on a receive path, but the hardware module 610 may also be used to perform functions (eg, load balancing or firewall) on packets on a transmit path received from a host.
[0293] The network interface device 600 includes a host interface 620 for transmitting and receiving data packets with a host and a network MAC interface 630 for transmitting and receiving data packets with a network.
[0294] The network interface device 600 includes a hardware module 610, which includes multiple processing units 640a, 640b, 640c, and 640d. Each processing unit can be an atomic processing unit. The term atomic is used in this specification to refer to a processing unit. Each processing unit is configured to perform at least one operation in hardware. Each processing unit includes a digital circuit 645 configured to perform the at least one operation. The digital circuit 645 can be an application-specific integrated circuit. Each processing unit also includes a memory 650 for storing state information. When executing the corresponding multiple operations, the digital circuit 645 updates the state information. In addition to local memory, each processing unit can also access a shared memory 660, which can also store state information accessible to each of the multiple processing units.
[0295] The state information in the shared memory 660 and / or the state information in the memory 650 in the processing units may include at least one of the following: metadata passed between processing units, temporary variables, content of data packets, content of one or more shared mapping tables.
[0296] The plurality of processing units, taken together, can provide functions to be performed on data packets received at the network interface device 600. The compiler outputs instructions to configure the hardware module 610 to perform the functions on incoming data packets by arranging at least some of the plurality of processing units to perform at least one predetermined operation on each incoming data packet. This can be achieved by linking (i.e., connecting) at least some of the processing units 640a, 640b, 640c, and 640d together so that each connected processing unit performs at least one operation on each incoming data packet. Each processing unit performs its at least one operation in a specific order to perform the functions. This order can allow two or more processing units to execute in parallel (i.e., simultaneously) with each other. For example, one processing unit can read from a data packet during a period of time (defined by a periodic signal (e.g., a clock signal) of the hardware module 610) while a second processing unit also reads data from a different location in the same data packet during that period of time.
[0297] In some embodiments, the data packet is passed to each stage represented by a processing unit in sequence. In this case, each processing unit completes its processing before passing the data packet to the next processing unit to perform its processing.
[0298] exist Figure 6 In the example shown, processing units 640a, 640b, and 640d are connected together at compile time so that each performs at least one of their respective operations, thereby performing a function, such as filtering received data packets. Processing units 640a, 640b, and 640d form a pipeline for processing data packets. Data packets can move through the pipeline in stages, with each stage having an equal time period. The time period can be defined by a periodic signal or tick. The time period can be defined by a clock signal. Several cycles of the clock can define a time period for each stage in the pipeline. At the end of each repeating time period, the data packet moves through a stage in the pipeline. The time period can be a fixed interval. Alternatively, each time period in a stage in the pipeline can take a variable amount of time. When the previous processing stage completes its operation, a signal may be generated indicating the next stage in the pipeline, which may take a variable amount of time. By delaying the signal by a predetermined amount of time, a pause can be introduced at any stage in the pipeline.
[0299] Each processing unit 640a, 640b, 640d can be configured to access shared memory 660 as part of at least one of their respective operations. Each of the processing units 640a, 640b, 640d can be configured to pass metadata between each other as part of at least one of their respective operations. Each of the processing units 640a, 640b, 640d can be configured to access data packets received from the network as part of at least one of their respective operations.
[0300] In this example, processing unit 640c is not used to perform processing of received data packets to provide functionality and is omitted from the pipeline.
[0301] Data packets received at the network MAC layer 630 may be passed to the hardware module 610 for processing. Figure 6 6. Not shown, but the processing performed by hardware module 610 may be part of a larger processing pipeline that provides additional functionality with respect to data packets beyond that provided by hardware module 610. Figure 14 shown and will be described in more detail below.
[0302] The first processing unit 640a is configured to perform at least one first operation on the data packet. This first at least one operation may include at least one of: reading from the data packet, reading and writing shared state in memory 660, and / or performing a table lookup to determine an action. The first processing unit 640a is then configured to generate a result from its at least one operation. The result may be in the form of metadata. The result may include a modification to the data packet. The result may include a modification to the shared state in memory 660. The second processing unit 640b is configured to perform at least one operation on the first data packet based on the result of the operation performed by the first processing unit 640a. The second processing unit 640b generates a result from its at least one operation and passes the result to the third processing unit 640d, which is configured to perform at least one operation on the first data packet. The first, second, and third processing units 640a, 640b, and 640d are collectively configured to provide functionality related to the data packet. The data packet may then be passed to the host interface 620, from which it is delivered to the host system.
[0303] Thus, it can be seen that the connected processing units form a pipeline for processing packets received at the network interface device. The pipeline can provide processing of eBPF programs. The pipeline can provide processing of multiple eBPF programs. The pipeline can provide processing of multiple modules executed in sequence.
[0304] The connection of the processing units in the hardware module 610 can be performed by programming the routing function of the pre-synthesized interconnect structure of the hardware module 610. The interconnect structure provides connections between the various processing units of the hardware module 610. The interconnect structure can be programmed according to the topology supported by the structure. Possible example topologies are shown below. Figure 15 Have a discussion.
[0305] Hardware module 610 supports at least one bus interface. The at least one bus interface receives data packets at hardware module 610 (e.g., from a host or a network). The at least one bus interface outputs data packets from hardware module 610 (e.g., to a host or a network). The at least one bus interface receives control messages at hardware module 610. The control messages can be used to configure hardware module 610.
[0306] Relative to Figure 5 The FPGA application 515 shown, Figure 6 The example shown has the advantage of reduced compilation time. For example, Figure 6 The hardware module 610 may take less than 10 seconds to compile the filter function. Figure 4 Compared to the example of the CPU array shown, Figure 6 The example shown has the advantage of improved processing speed.
[0307] An application can be executed in such a hardware module 610 by mapping a general purpose program (or programs) to a pre-synthesized datapath. The compiler builds the datapath by linking any number of processing stage instances, each of which is built from one of the pre-synthesized processing stage atoms.
[0308] Each atom is built from circuits. Each circuit can be defined using RTL (Register Transfer Language) or a high-level language. Each circuit is synthesized using a compiler or tool chain. Atoms can be synthesized as hard logic and can therefore be used as hard (ASIC) resources in the hardware modules of a network interface device. Atoms can be synthesized as soft logic. Constraints can be provided to atoms in soft logic, which assign and maintain the location and routing information of the synthesized logic on the physical device. Atoms can be designed using configurable parameters that specify the behavior of the atoms. Each parameter can be a variable, or even a sequence of operations (microprogram), which can specify at least one operation to be performed by the processing unit during a clock cycle of the processing pipeline. The logic that implements the atoms can be clocked synchronously or asynchronously.
[0309] The atomic processing pipeline itself can be configured to operate based on a periodic signal. In this case, each packet and metadata advances one stage along the pipeline in response to each occurrence of the signal. The processing pipeline can also operate asynchronously. In this case, a high level of backpressure in the pipeline will cause each downstream stage to begin processing only when it has data from the upstream stage available to it.
[0310] When compiling a function to be executed by multiple such atoms, the sequence of computer code instructions is broken into multiple operations, each of which is mapped to a single atom. Each operation can represent a single line of disassembled instructions in the computer code instruction. Each operation is assigned to one of the atoms to be executed. There may be one atom per expression in the computer code instruction. Each atom is associated with an operation type and is selected to perform at least one operation in the computer code instruction based on its associated operation type. For example, an atom can be preconfigured to perform a load operation from a data packet. Thus, such an atom is assigned to execute an instruction representing a load operation from a data packet in the computer code.
[0311] In computer code instructions, one atom may be selected per line. Thus, when a function is implemented in a hardware module containing such atoms, there may be 100 such atoms, each performing its own operation to execute the function for that data packet.
[0312] Each atom can be constructed according to one of a set of processing stage templates that determine its associated operation type. The compilation process is configured to generate instructions based on its associated type to control each atom to perform a specific at least one operation. For example, if an atom is pre-configured to perform a packet access operation, the compilation process can assign an operation to the atom to load certain information from the header of the packet (e.g., the source ID of the packet). The compilation process is configured to send instructions to the hardware module, where the atoms are configured to perform the operations assigned to them by the compilation process.
[0313] The processing stage templates that specify atomic behavior are logic stage templates (e.g., providing operations on registers, registers and stacks, as well as branches), packet access state templates (e.g., providing packet data loading and / or packet data storage), and mapping table access stage templates (e.g., mapping table lookup algorithm, mapping table size).
[0314] The data packet access phase may include at least one of: reading a sequence of bytes from the data packet, replacing a sequence of bytes with a different sequence of bytes in the data packet, inserting bytes into the data packet, and deleting bytes from the data packet.
[0315] The mapping table access stage can be used to access different types of mapping tables (e.g., lookup tables), including direct index arrays and associative arrays. The mapping table access stage can include at least one of the following: reading a value from a location, writing a value to a location, and replacing the value at a location in the mapping table with another value. The mapping table access stage can include a comparison operation in which a value is read from a location in the mapping table and compared to a different value. If the value read from the location is less than the different value, a first action can be performed (e.g., no operation is performed, the value at the location is exchanged for a different value, or the values are added). Otherwise, a second operation may be performed (e.g., no operation is performed, a value is exchanged, or a value is added). In either case, the value read from the location can be provided to the next processing stage.
[0316] Each mapping table access stage can be implemented in a stateful processing unit. Figure 17 , which shows an example of a circuit 1700 that can be included in an atom configured to perform processing of a mapping table access phase. Circuit 1700 may include a hash function 1710 configured to perform a hash of an input value used as input to a lookup table. Circuit 1700 includes a memory 1720 configured to store a state associated with the operation of the atom. Circuit 1700 includes an arithmetic logic unit 1730 configured to perform the operation.
[0317] The logic stage can perform calculations on values provided by the previous stage. The processing units configured to implement the logic stage can be stateless processing units. Each stateless processing unit can perform simple arithmetic operations. Each processing unit can perform, for example, 8-bit operations.
[0318] Each logical stage can be implemented in a stateless processing unit. Figure 18 , which shows an example of a circuit 1800 that can be included in an atom configured to perform logic stage processing. Circuit 1800 includes an array of arithmetic logic units (ALUs) and multiplexers. The ALUs and multiplexers are arranged in layers, and the output of one layer of processing performed by the ALUs is used by the multiplexers to provide input to the next layer of ALUs.
[0319] The stage pipeline implemented in the hardware module may include the first packet access stage (pkt0), then the first logic stage (logic0), then the first map access stage (map0), then the second logic stage (logic1), then the second packet access stage (pkt1), etc. Therefore, it can take the following form:
[0320] pkt0->logic0->map0->logic1->pktl
[0321] In some examples, stage pkt0 extracts the required information from the packet. Stage pkt0 passes this information to logic0. Logic0 determines whether the packet is a valid IP packet. In some cases, logic0 forms a mapping request and sends it to map0, which performs the mapping operation. Stage map0 may perform updates to the lookup table. Stage logic1 then collects the results from the mapping operation and decides whether to discard the packet.
[0322] In some cases, the map request is disabled to cover situations where a map operation should not be performed on this packet. In the case where a map operation is not performed, logic 0 indicates to logic 1 whether the packet should be dropped based on whether it is a valid IP packet. In some examples, the lookup table contains 256 entries, each of which is an 8-bit value.
[0323] The example described includes only five stages. However, as mentioned above, more can be used. Furthermore, all operations do not have to be performed sequentially, but certain operations on the same data packet can be performed simultaneously by different processing units.
[0324] Figure 6The illustrated hardware module 610 shows a single pipeline of atoms for performing functions related to data packets. However, the hardware module 610 may include multiple pipelines for processing data packets. Each of the multiple pipelines may perform a different function with respect to the data packets. The hardware module 610 may be configured to interconnect a first group of atoms of the hardware module 610 to form a first data processing pipeline. The hardware module 610 may also be configured to interconnect a second group of atoms of the hardware module 610 to form a second data processing pipeline.
[0325] In order to compile the functions to be implemented in a hardware module including multiple processing units, a series of steps starting from a series of computer codes may be performed. A compiler that may be run on a processor on a host device or a network interface device may access the disassembled computer code sequence.
[0326] First, the compiler is configured to divide a sequence of computer code instructions into separate phases. Each phase may include operations according to one of the aforementioned processing phase templates. For example, one phase may read a data packet. Another phase may update mapping table data. Another phase may make a pass / abandon decision. The compiler assigns each of the multiple operations represented by the code to each of the multiple phases.
[0327] Secondly, the compiler is configured to assign each processing stage, as determined by the code, to be executed by a different processing unit. This means that each of the at least one operation of the processing stage is executed by a different processing stage. The compiler's output can then be used to cause the processing units to execute the operations of each stage in a specific order to perform the function.
[0328] The output of the compiler includes generated instructions for causing the processing units of the hardware module to perform the operations associated with each processing stage.
[0329] The output of the compiler can also be used to generate logic in the hardware module to respond to control messages used to configure the hardware module 610. Figure 14 Such control messages are described in more detail.
[0330] The compilation process for compiling the functions to be executed on the network interface device 600 may be performed based on a determination that the process for providing the functions is safe for execution in the kernel of the host device. The determination of the safety of the program may be made by referring to the above. Figure 3 Once it has been determined that the process is safe for execution in the kernel, the process can be compiled for execution in the network interface device.
[0331] refer to Figure 15The figure shows a representation of at least some of a plurality of processing units performing at least one operation to perform a function on a data packet. Such a representation can be generated by a compiler and used to configure a hardware module to perform the function. The representation indicates the order in which the operations can be performed and how some processing units can perform their operations in parallel.
[0332] Representation 1500 is in the form of a list with rows and columns. Certain entries in the list show atoms, such as atom 1510a, configured to perform their respective operations. The row to which a processing unit belongs indicates the timing of the operations performed by that processing unit for a particular data packet. Each row may correspond to a single time period represented by one or more cycles of a clock signal. Processing units belonging to the same row perform their operations in parallel.
[0333] The input to the logic stage is provided in row 0, and the computation flows forward to the following rows. By default, an atom receives the results of its processing in the same column as the atom but in the previous row. For example, atom 1510b receives the results of its processing from atom 1510a and performs its own processing based on these results.
[0334] When using local routing resources, an atom can also access the output of a previous row of atoms whose column numbers differ by no more than two. For example, atom 1510d can receive the results of processing performed by atom 1510c.
[0335] When using global routing resources, atoms can also access the outputs of atoms in the first two rows and any columns. This can be performed by using global routing resources. For example, atom 1510f can receive the result of processing performed by atom 1510e.
[0336] These constraints on routing between atoms are given as examples, and other constraints can be applied. Imposing more restrictive constraints can make it easier to route information between atoms. Imposing less restrictive constraints can make scheduling easier. If the number of atoms of a given type (e.g., mapping, logic, or packet access) is exhausted or routing between atoms is impossible, compiling the function into a hardware module will fail.
[0337] The specific constraints are determined by the topology supported by the interconnect structure that the hardware blocks support. The interconnect structure is programmed so that the atoms of the hardware blocks perform their operations in a specific order and provide data to each other within the constraints. Figure 15 One specific example of how the interconnect structure may be programmed is shown.
[0338] When synthesizing the FPGA application 515 into the FPGA (such as Figure 5In the process shown in Figure 1, a place and route algorithm is used. However, in this case, the solution space is restricted, so the algorithm has a limited execution time.
[0339] There is a trade-off between processing speed or efficiency and compilation time. According to an embodiment of the present application, it may be desirable to first perform the processing in at least one processing unit (which may be as described above with reference to Figure 6 The at least one processing unit may then execute the function with respect to the received data packet during a first time period. During operation of the network interface device, the function is executed. The second at least one processing unit (which may be a processor as described above) may execute the function with respect to the received data packet. Figure 6 The FPGA application or template type processing unit can be configured to execute a function for a data packet. The function can then be migrated from the first at least one processing unit to the second at least one processing unit, so that the second at least one processing unit then executes the function for a data packet subsequently received at the network interface device. Therefore, the slower compile time of the second at least one processing unit does not prevent the network interface device from executing the function for the data packet before compiling the function for the second at least one processing unit, because the first at least one processing unit can be compiled more quickly and can be used to execute the function for the data packet while the function for the second at least one processing unit is being compiled. Since the second at least one processing unit generally has a faster processing time, migrating to the second at least one processing unit at compile time allows for faster processing of data packets received at the network interface device.
[0340] According to an embodiment of the present application, the compilation process can be configured to run on at least one processor of a data processing system, wherein the at least one processor is configured to send instructions for the first at least one processing unit and the second at least one processing unit to perform at least one function on the data packet at the appropriate time. The at least one processor may include a host CPU. The at least one processor may include a control processor on a network interface device. The at least one processor may include a combination of one or more processors on the host system and one or more processors on the network interface device.
[0341] Thus, at least one processor is configured to execute a first compilation process to compile a function to be executed by a first at least one processing unit of a network interface device. The at least one processing unit is further configured to execute a second compilation process to compile a function to be executed by a second at least one processing unit of the network interface device. Before the second compilation process is completed, the at least one processing unit instructs the first at least one processing unit to execute the function on a data packet received from the network. Subsequently, after the second compilation process is completed, the at least one processing unit instructs the second at least one processing unit to begin executing the function on the data packet received from the network.
[0342] Performing these steps enables the network interface device to use the first at least one processing unit (which may have a shorter compile time but slower processing and / or less efficient) to perform functions while waiting for the second compile process to complete. When the second compile process is complete, the network interface device can then use the second at least one processing unit (which may have a longer compile time but faster processing and / or more efficient) to perform functions in addition to or instead of the first at least one processing unit.
[0343] refer to Figure 7 , which shows an exemplary network interface device 700 according to an embodiment of the present application. Reference elements identical to those shown in previous figures are denoted by the same reference numerals.
[0344] The network interface device includes a first at least one processing unit 710. The first at least one processing unit 710 may include Figure 6 The hardware module 610 includes multiple processing units. The first at least one processing unit 710 may include one or more CPUs, such as Figure 4 shown.
[0345] The function is compiled to run on the first at least one processing unit 710, so that within a first time period, the function is executed by the first at least one processing unit 710 with respect to a data packet received from the network. Before a second compilation process for the second at least one processing unit is completed, the first at least one processing unit 710 is instructed by the at least one processor to execute the function with respect to the data packet received from the network.
[0346] The network interface device includes a second at least one processing unit 720. The second at least one processing unit 720 may include an FPGA with an FPGA application (e.g. Figure 5 shown), or may include Figure 6 The hardware module 610 shown includes multiple processing units.
[0347] During the first time period, a second compilation process is performed to compile the functionality for running on the second at least one processing unit. That is, the network interface device is configured to dynamically compile the FPGA application 515 .
[0348] After the first time period (ie, after completing the second compiling process), the second at least one processing unit 720 is configured to start executing functions on data packets received from the network.
[0349] After the first time period, the first at least one processing unit 710 may stop performing functions with respect to packets received from the network. In some embodiments, the first at least one processing unit 710 may partially stop performing functions with respect to packets. For example, if the first at least one processing unit includes multiple CPUs, after the first time period, one or more of the CPUs may stop processing packets received from the network, while the remaining CPUs of the multiple CPUs continue to perform such processing.
[0350] The first at least one processing unit 710 may be configured to perform functions on packets of the first data stream. When the second compilation process is completed, the second at least one processing unit 720 may start performing functions on packets of the first data stream. When the second compilation process is completed, the first at least one processing unit may stop performing functions on packets of the first data stream.
[0351] Different combinations are possible for the first at least one processing unit and the second at least one processing unit. For example, in some embodiments, the first at least one processing unit 710 includes multiple CPUs (e.g., Figure 4 As shown), the second at least one processing unit 720 includes a hardware module having multiple processing units (such as Figure 6 In some embodiments, the first at least one processing unit 710 includes multiple CPUs (such as Figure 4 As shown), the second at least one processing unit 720 includes an FPGA (as shown Figure 5 In some embodiments, the first at least one processing unit 710 includes a hardware module having multiple processing units (such as Figure 6 As shown), the second at least one processing unit 720 includes an FPGA (as shown Figure 5 shown).
[0352] refer to Figure 11 , which shows how the connected multiple processing units 640a, 640b, 640d can perform their respective at least one operation on the data packet. Each processing unit is configured to perform its respective at least one operation on the received data packet.
[0353] At least one operation of each processing unit can represent a logical stage in a function (e.g., a function of an eBPF program). The at least one operation of each processing unit can be expressed by an instruction executed by the processing unit. The instruction can determine the behavior of the atom.
[0354] Figure 11 It shows how a data packet (P0) proceeds along the processing stages implemented by each processing unit.
[0355] Each processing unit performs processing on the data packet in a specific order specified by the compiler. This order may allow some processing units to be configured to perform their processing in parallel. The processing may include accessing at least a portion of the data packet stored in memory. Additionally or alternatively, the processing may include performing a lookup in a lookup table to determine the action to perform on the data packet. Additionally or alternatively, the processing may include modifying state 1110.
[0356] The processing units exchange metadata M0, M1, M2, and M3 with each other. The first processing unit 640a is configured to perform at least one predetermined operation of each of them and generate metadata M1 in response. The first processing unit 640a is configured to pass the metadata M1 to the second processing unit 640b.
[0357] At least some of the processing units perform at least one operation based on at least one of the following: the contents of the data packet, its own memory state, the global shared state, and metadata associated with the data packet (e.g., M0, M1, M2, M3). Some processing units may be stateless.
[0358] Each processing unit can perform its associated operation type on the data packet (P0) within at least one clock cycle. In some embodiments, each processing unit can perform its associated operation type within a single clock cycle. Each processing unit can be individually timed to perform its operations. This timing can be in addition to the timing of the processing unit's processing pipeline.
[0359] Examining the operation of second processing unit 640b in more detail, second processing unit 640b is configured to connect to first processing unit 640a, which is configured to perform a first at least one predetermined operation on a first data packet. Second processing unit 640b is configured to receive the result of the first at least one predetermined operation from the first additional processing unit. Second processing unit 640b is configured to perform a second at least one predetermined operation based on the result of the first at least one predetermined operation. Second processing unit 640b is configured to connect to third processing unit 640d, which is configured to perform a third at least one predetermined operation on the first data packet. Second processing unit 640b is configured to send the result of the second at least one predetermined operation to third processing unit 640d for processing in the third at least one predetermined operation.
[0360] The processing unit may similarly operate to provide functionality with respect to each of the plurality of data packets.
[0361] The embodiments of the present application enable pipeline processing of multiple data packets simultaneously when functionality permits.
[0362] refer to Figure 12 , which illustrates the data packet pipeline. As shown in the figure, different data packets can be processed simultaneously by different processing units. First processing unit 640a performs at least one operation corresponding to third data packet (P2) at a first time (t0). Second processing unit 640b performs at least one operation corresponding to second data packet (P1) at a first time (t0). Third processing unit 640d performs at least one operation corresponding to first data packet (P0) at a first time (t0).
[0363] After each processing unit has performed at least one corresponding operation, each data packet moves along a stage in the sequence. For example, at a subsequent second time (t1), the first processing unit 640a performs its corresponding at least one operation on the fourth data packet (P3) at the first time (t0). The second processing unit 640b performs its corresponding at least one operation on the third data packet (P2) at the first time (t0). The third processing unit 640d performs its corresponding at least one operation on the first data packet (P1) at the first time (t0).
[0364] It should be understood that in some embodiments, there may be multiple packets in a given phase.
[0365] In some embodiments, packets can move from one stage to the next without necessarily being in lock step.
[0366] A pipeline running at a fixed clock can have a constant bandwidth as long as there are no pipeline hazards. This can reduce jitter in the system.
[0367] To avoid hazards when executing instructions (eg, conflicts when accessing shared state), each processing unit may be configured to execute no-op (ie, processing unit stall) instructions when necessary.
[0368] In some embodiments, operations (e.g., simple arithmetic, increments, adding / subtracting constant values, shifts, adding / subtracting values from a data packet or from metadata) require one clock cycle to execute by a processing unit. This may mean that a shared state value needed by one processing unit has not yet been updated by another processing unit. As a result, outdated values in shared state 1110 can be read by processing units that need them. Consequently, hazards may occur when reading and writing values from shared state. On the other hand, operations on intermediate values can be passed as metadata without incurring hazards.
[0369] An example of hazards that can be avoided when reading and writing shared state 1110 can be given in the context of an increment operation. Such an increment operation can be an operation that increments a packet counter in shared state 1110. In one implementation of the increment operation, during a first time slot in the pipeline, second processing unit 640b is configured to read the value of the counter from shared state 1110 and provide the output of the read operation (e.g., as metadata M2) to third processing unit 640d. Third processing unit 640d is configured to receive the value of the counter from second processing unit 640b. During a second time slot, third processing unit 640d increments the value and writes the new incremented value to shared state 1110.
[0370] A problem may arise when performing such an increment operation, namely, if the second processing unit 640b attempts to access the counter stored in the shared state 1110 during the second time slot, the second processing unit 640b may read the previous value of the counter before the counter value in the shared state 1110 is updated by the third processing unit 640d.
[0371] Therefore, to address this issue, the second processing unit 640b can be stalled during the second time slot (by the second processing unit 640b executing a no-operation instruction or a pipeline bubble). A stall can be understood as a delay in the execution of the next instruction. This delay can be achieved by executing a "no-operation" instruction instead of the next instruction. Then, during the subsequent third time slot, the second processing unit 640b reads the counter value from the shared state 1110. During the third time slot, the counter in the shared state 1110 has already been updated, thereby ensuring that the second processing unit 640b reads the updated value.
[0372] In some embodiments, each atom is configured to read from state, update state, and write the updated state during a single pipeline time slot. In this case, the above-described stalling of the processing unit may not be used. However, stalling the processing unit can reduce the cost of the required memory interface.
[0373] In some embodiments, to avoid hazards, a processing unit in a pipeline may wait until other processing units in the pipeline have completed their processing before performing its own operations.
[0374] As mentioned previously, the compiler builds the datapath by chaining any number of processing stage instances, each of which is built from one of a predetermined number of pre-synthesized processing stage templates (three in the given example). The processing stage templates are the logic stage template (e.g., providing arithmetic operations on registers, scratchpads, and metadata), the packet access state template (e.g., providing packet data loads and / or packet data stores), and the map access stage template (e.g., map lookup algorithms, map table sizes).
[0375] Each processing stage instance can be implemented by a single processing unit. That is, each processing stage includes at least one corresponding operation performed by the processing unit.
[0376] Figure 13 13 shows an example of how processing stages can be connected together in pipeline 1300 to process received data packets. Figure 13 As shown, a first data packet is received and stored at FIFO 1305. One or more call parameters are received at first logic stage 1310. The call parameters may include a program selector that identifies a function to be executed for the received data packet. The call parameters may include an indication of the packet length of the received data packet. First logic stage 1310 is configured to process the call parameters and provide an output to first data packet access stage 1315.
[0377] The first packet access stage 1315 loads data from the first packet at the network tap 1320. The first packet access stage 1315 may also write data to the first packet based on the output of the first logic stage 1310. The first packet access stage 1315 may write data to the front of the first packet. The first packet access stage 1315 may also overwrite data in the packet.
[0378] The loaded data and any other metadata and / or parameters are then provided to the second logic stage 1325, which performs processing on the first packet and provides the output parameters to the first map access stage 1330. The first map access stage 1330 uses the output from the second logic stage 1325 to perform a lookup in a lookup table to determine the action to be performed for the first packet. The output is then passed to the third logic stage 1335, which processes the output and passes the result to the second packet access stage 1340.
[0379] The second data packet access stage 1340 may read data from the first data packet and / or write data to the first data packet based on the output of the third logic stage 1335. The results of the second data packet access stage 1340 are then passed to the fourth logic stage 1345, which is configured to perform processing on the inputs it receives.
[0380] The pipeline may include multiple packet access stages, logic stages, and mapping access stages. The final logic stage 1350 is configured to output a return parameter. The return parameter may include a pointer identifying the beginning of a packet. The return parameter may include an indication of an action to be performed on the packet. The indication of the action may indicate whether to discard the packet. The indication of the action may indicate whether to forward the packet to a host system. The network interface device may include at least one processing unit configured to discard the corresponding packet in response to the indication that the packet is to be discarded.
[0381] Pipeline 1300 may further include one or more bypass FIFOs 1355a, 1355b, and 1355c. The bypass FIFOs may be used to pass processing data, such as data from the first packet, around the map access stage and / or the packet access stage. In some embodiments, the map access stage and / or the packet access stage do not require data from the first packet to perform their respective at least one operation. The map access stage and / or the packet access stage may perform their respective at least one operation based on input parameters.
[0382] refer to Figure 8 , which shows a method 800 performed by the network interface device 600, 700 according to an embodiment of the present application.
[0383] At S810, a hardware module of a network interface device is arranged to perform a function. The hardware module includes a plurality of processing units, each of which is configured to perform a type of operation in hardware with respect to a data packet. S810 includes arranging at least some of the plurality of processing units to perform their respective predetermined types of operations in a specific order to provide the function for each received data packet. Arranging the hardware module in this manner includes connecting at least some of the plurality of processing units such that a received data packet is processed by each of a plurality of operations of at least some of the plurality of processing units. The connection can be achieved by configuring routing hardware of the hardware module to route data packets and associated metadata between the processing units.
[0384] At S820 , a first data packet is received from a network at a first interface of a network interface device.
[0385] At S830, each of the at least some processing units connected during the compilation process at S810 processes the first data packet. Each of the at least some processing units performs the type of operation it is pre-configured to perform with respect to the at least one data packet. Thus, a function is executed with respect to the first data packet.
[0386] At S840, the processed first data packet is transmitted to its destination. This may include sending the data packet to a host or sending the data packet over a network.
[0387] refer to Figure 9 , which shows a method 900 that can be executed in the network interface device 700 according to an embodiment of the present application.
[0388] At S910, a first at least one processing unit (ie, a first circuit) of a network interface device is configured to receive and process a data packet received from a network. The processing includes performing a function on the data packet. The processing is performed during a first time period.
[0389] At S920 , a second compiling process is performed during the first time period to compile functions for execution on the second at least one processing unit (ie, the second circuit).
[0390] At S930 , it is determined whether the second compilation process is completed. If not, the method returns to S910 and S920 , wherein the first at least one processing unit continues to process the data packets received from the network, and the second compilation process continues.
[0391] At S940, in response to determining that the second compilation is complete, the first at least one processing unit stops performing functions on the received data packets. In some embodiments, the first at least one processing unit may stop performing functions only for certain data flows. The second at least one processing unit may then perform functions for those specific data flows instead (at S950).
[0392] At S950 , when the second compiling process is completed, the second at least one processing unit is configured to start executing a function on a data packet received from the network.
[0393] refer to Figure 16 , which shows a method 1600 according to an embodiment of the present application. The method 1600 can be executed in a network interface device or a host device.
[0394] At S1610 , a compile process is performed to compile a function to be executed by the first at least one processing unit.
[0395] At S1620, a compilation process is performed to compile a function to be executed by the second at least one processing unit. This process includes assigning each of the plurality of processing units of the second at least one processing unit to perform at least one operation associated with one of the plurality of stages for processing a data packet to provide the first function. Each of the plurality of processing units is configured for a type of processing, and the assignment is performed based on a determination of the type of processing that the processing unit is configured to perform that is suitable for performing the corresponding at least one operation. In other words, the processing units are selected based on their templates.
[0396] In step 1630 , before the compilation process in step S1620 is completed, an instruction is sent to enable the first at least one processing unit to execute the function. The instruction may be sent before the compilation process in step S1620 is started.
[0397] At S1640, after the compilation process at S1620 is completed, an instruction is sent to the second circuit to enable the second circuit to execute a function on the data packet. The instruction may include the compilation instruction generated at S1620.
[0398] Functionality according to embodiments of the present application may be provided as pluggable components of processing slices in a network interface. Figure 14 , which shows an example of how a slice 1425 may be used in the network interface device 600. The slice 1425 may be referred to as a processing pipeline.
[0399] The network interface device 600 includes a transmit queue 1405 for storing packets received from the host to be processed by the slice 1425 and then sent over the network. The network interface device 600 includes a receive queue 1410 for storing packets received from the network 1410 to be processed by the slice 1425 and then delivered to the host. The network interface device 600 includes a receive queue 1415 for storing packets received from the network that have been processed by the slice 1425 and are to be delivered to the host. The network interface device 600 includes a transmit queue for storing packets received from the host that have been processed by the slice 1425 and are to be delivered to the network.
[0400] Slice 1425 of network interface device 600 includes multiple processing functions for processing data packets on the receive and transmit paths. Slice 1425 may include a protocol stack configured to perform protocol processing on data packets on the receive and transmit paths. In some embodiments, network interface device 600 may include multiple slices. At least one of the multiple slices may be configured to process receive data packets received from the network. At least one of the multiple slices may be configured to process transmit data packets for transmission on the network. Slices may be implemented using hardware processing devices such as at least one FPGA and / or at least one ASIC.
[0401] Accelerator components 1430a, 1430b, 1430c, 1430d can be inserted into a slice at different stages, as shown. Each accelerator component provides functionality related to packets traversing the slice. Accelerator components can be inserted or removed dynamically (i.e., during operation of the network interface device). Therefore, accelerator components are pluggable components. Accelerator components are logical areas allocated to slice 1425. Each of them supports a streaming packet interface that allows packets flowing through the slice to flow into and out of the component.
[0402] For example, one type of accelerator component may be configured to provide encryption of data packets on a receive or transmit path, while another type of accelerator component may be configured to provide decryption of data packets on a receive or transmit path.
[0403] The functionality discussed above is provided by executing operations performed by a plurality of connected processing units (as described above with reference to Figure 6 Similarly, an array of network processing CPUs (as discussed above) may be provided by an accelerator component. Figure 4 discussed) and / or FPGA applications (as discussed above Figure 5 The functionality provided by the accelerator component may be provided by the accelerator component.
[0404] As described, during operation of the network interface device, processing performed by a first at least one processing unit (e.g., a plurality of connected processing units) can be migrated from a second at least one processing unit. To accomplish this migration, components of the slice 1425 that are processed by the first at least one processing unit can be replaced with components that are processed by the second at least one processing unit.
[0405] The network interface device may include a control processor configured to insert and remove components from the slice 1425. During the first time period discussed above, a component whose function is performed by the first at least one processing unit may exist in the slice 1425. The control processor may be configured to, after the first time period, remove the pluggable component whose function is provided by the first at least one processing unit from the slice 1425 and insert the pluggable component whose function is provided by the second at least one processing unit into the slice 1425.
[0406] In addition to or in lieu of inserting or removing components from a slice, the control processor may also load programs into the component and issue control plane commands to control the flow of frames into the component. In this case, it may cause the component to run or not run without being inserted into or removed from the pipeline.
[0407] In some embodiments, control plane or configuration information is carried on the datapath, eliminating the need for a separate control bus. In some embodiments, requests to update the configuration of datapath components are encoded as messages that are carried on the same bus as network packets. Thus, the datapath can carry two types of packets: network packets and control packets.
[0408] Control packets are formed by the control processor and injected into slice 1425 using the same mechanism used to send or receive data packets using slice 1425. This same mechanism can be a transmit queue or a receive queue. Control packets can be distinguished from network packets in any suitable manner. In some embodiments, different types of packets can be distinguished by one or more bits in a metadata word.
[0409] In some embodiments, the control packet contains a routing field in the metadata word that determines the path the control packet takes through slice 1425. The control packet can carry a series of control commands. Each control command can target one or more components of slice 1425. The corresponding datapath component is identified by the component ID field. Each control command encodes a request for the corresponding identified component. The request can be to change the configuration of the component. The request can control whether the component is activated, that is, whether the component performs its function for packets traversing the slice.
[0410] Thus, in some embodiments, the control processor of the network interface device 600 is configured to send a message causing one of the components of the slice to begin executing a function for packets received at the network interface device. This message is a control plane message that is sent via the pluggable component and causes the atomic switching of frames to the component to execute the function. This component then executes on all received packets traversing the slice until it is removed from the slice. The control processor is configured to send a message causing another component of the slice to stop executing the function for packets received at the network interface device 600.
[0411] To switch components in and out of data slice 1425, sockets may exist at various points in the ingress and egress data paths. The control processor may enable additional logic to enter and exit slice 1425. This additional logic may take the form of FIFOs placed between components.
[0412] The control processor may send control plane messages through slice 1425 to the configured components of slice 1425. The configuration may determine the functions performed by the components of slice 1425. For example, a control message sent through slice 1425 may cause a hardware module to be configured to perform a function on a data packet. Such a control message may interconnect the atoms of the hardware module into a pipeline of the hardware module to provide a certain function. Such a control message may cause the atoms of the hardware module to be configured so as to select an operation to be performed by each selected atom. Since each atom is pre-configured to perform a certain operation, the selection of the operation for each atom depends on the type of operation that each atom is pre-configured to perform.
[0413] Now refer to Figures 19 to 21 Some other embodiments are described. In this embodiment, a packet processing program or feedforward pipeline runs in an FPGA. A method for implementing the packet processing program or feedforward pipeline in a subunit of an FPGA will be described. The packet processing program or feedforward pipeline can be an eBPF program, a P4 program, or any other suitable program.
[0414] The FPGA may be provided in the network interface device.In some embodiments, the packet processing program is deployed or run only after the network interface device is installed relative to its host.
[0415] A packet handler or feed-forward pipeline can implement a logic flow without loops.
[0416] In some embodiments, the program can be written in a non-privileged domain or a less privileged domain, such as in userland. The program can be run in a privileged domain or a more privileged domain (such as a kernel). The hardware running the program may require no arbitrary loops.
[0417] In the following embodiments, reference is made to an eBPF program example. However, it will be appreciated that other embodiments may be used with any other suitable program.
[0418] It should be appreciated that one or more of the following embodiments may be used in combination with one or more of the previous embodiments.
[0419] Some embodiments may be provided in the context of an FPGA, an ASIC, or any other suitable hardware device. Some embodiments utilize subunits of an FPGA or ASIC, etc. The following examples are described with reference to an FPGA. It should be understood that similar processing may be performed using an ASIC or any other suitable hardware device.
[0420] A subunit may be an atom. Some examples of atoms have been previously described. It should be understood that any of those previously described examples of atoms may alternatively or additionally be used as a subunit. Alternatively or additionally, these subunits may be referred to as "slices" or configurable logic blocks.
[0421] Each of these subunits can be configured to execute a single instruction or multiple related instructions. In the latter case, the related instructions can provide a single output (which can be defined by one or more bits).
[0422] The subunits can be considered as computing units. The subunits can be arranged in a pipeline where packets are processed sequentially. In some embodiments, the subunits can be dynamically assigned to execute individual instructions in a program.
[0423] In some embodiments, a subunit can be all or part of a unit used to define, for example, a block of an FPGA. In some FPGAs, a block of an FPGA is called a slice. In some embodiments, a subunit or atom is equivalent to a slice.
[0424] By mapping the corresponding atoms or subunits to the corresponding blocks or slices of the FPGA, improved resource utilization can be achieved compared to an approach that maps RTL atoms to FPGA resources. Such a latter approach may result in the RTL atoms requiring a relatively large number of individual blocks or slices of the FPGA.
[0425] In some embodiments, compilation can be atomic. This may have the advantage of pipeline processing. Data packets can be processed sequentially. The compilation process can be performed relatively quickly.
[0426] In some embodiments, arithmetic operations may require one slice per byte. Logical operations may require half a slice per byte. Depending on the width of the shift operation, a shift operation may require a set of slices. Comparison operations may require one slice per byte. Selection operations may require half a slice per byte.
[0427] As part of the compilation process, placement and routing are performed. Placement is the assignment of specific physical subunits to execute one or more specific instructions. Routing ensures that one or more outputs of a specific subunit are routed to the correct destination, which can be, for example, another subunit or multiple subunits.
[0428] Placement and routing can use a process that assigns operations to specific subunits starting at one end of the pipeline. In some embodiments, the most critical operations can be placed before less critical operations. In some embodiments, routes can be assigned while placing specific operations. In some embodiments, routes can be selected from a limited set of pre-computed routes. This will be described in more detail later.
[0429] In some embodiments, if a route cannot be assigned, the operation is reserved for later use.
[0430] In some embodiments, the pre-computed routes may be byte-wide routes. However, this is merely exemplary, and in other embodiments, different route widths may be defined. In some embodiments, multiple routes of different sizes may be provided.
[0431] In some embodiments, routing may be limited to routing between nearby subunits.
[0432] In some embodiments, the sub-units may be physically arranged on the FPGA in a regular structure.
[0433] In some embodiments, to facilitate routing, rules can be established about how subunits can communicate. For example, a subunit can only provide output to the subunit next to it, above it, or below it.
[0434] Alternatively or additionally, for routing purposes, a limit may be placed on the distance to the next subunit. For example, a subunit may only output data to adjacent subunits or subunits within a limited distance (e.g., no more than one intermediate subunit).
[0435] refer to Figure 19 , which illustrates the methods of some embodiments.
[0436] In some embodiments, an FPGA may have one or more "static" regions and one or more "dynamic" regions. The static regions provide standard configurations, and the dynamic regions may provide functionality that meets the requirements of the end user. For example, the static portion may be defined before the end user receives the network interface device, such as before the network interface device is installed relative to a host. For example, the static regions may be configured to cause the network interface device to provide certain functionality. The static regions will provide pre-computed paths between atoms. As will be discussed in more detail below, routing may be performed between one or more static regions passing through one or more dynamic regions. When the network interface device is deployed relative to a host, the dynamic regions may be configured by the end user according to their needs. The dynamic regions may be configured to perform different functions for the end user over a period of time.
[0437] In step S1, a first compilation process is performed to provide a first bitfile, referred to as a master bitfile 50 and a tool checkpoint 52. In some embodiments, this is a bitfile for at least a portion of the static region. When downloaded to the FPGA, the bitfile will cause the FPGA to function as specified by the program from which the bitfile was compiled. In some embodiments, the program used in the first compilation process can be any one or more programs, or can be a test program specifically designed to help determine routing within a portion of the FPGA. In some embodiments, a series of simple programs can be used instead or in addition.
[0438] The program can be modified or have a reconfigurable partition that can be used by the compiler. By moving the network out of the reconfigurable partition, the program can be modified to make the compiler's job easier.
[0439] Step S1 can be performed in a design tool. By way of example only, the Vivado tool can be used with Xilinx FPGAs. A checkpoint file can be provided by the design tool. The checkpoint file represents a snapshot of the design at the time the bitfile is generated. The checkpoint file can include one or more synthesized netlists, design constraints, placement information, and routing information.
[0440] In step S2 the bitfile is analysed taking into account the checkpoint file to provide a bitfile description 54. The analysis may be one or more of: detecting resources, generating routing, checking timing, generating one or more partial bitfiles and generating a bitfile description.
[0441] The analysis can be configured to extract routing information from the bit file. The analysis can be configured to determine which wires or routes a signal has traveled.
[0442] The analysis phase can be performed at least in part in a synthesis or design tool. In some embodiments, a scripting tool such as Vivado can be used. The scripting tool can be TCL (Tool Command Language). TCL can be used to add or modify Vivado functionality. Vivado functionality can be called and controlled by TCL scripts.
[0443] The bitfile description 54 defines how a given portion of the FPGA is used. For example, the bitfile description will indicate which atoms can be routed to other atoms, as well as one or more routes that can be routed between these atoms. For example, for each atom, the bitfile description will indicate where the atom's input can come from and where the atom's output can be routed, along with one or more data output routes. The bitfile description is independent of any program.
[0444] The bitfile description may contain one or more routing information, an indication of which routing conflicts occur, and a description of how to generate the bitfile from the desired atomic configuration.
[0445] A bitfile description may provide a set of routes between a set of atoms, but available before any particular instruction is executed by a given atom.
[0446] A bit file description may be for a portion of the FPGA. A bit file description may be for a dynamic portion of the FPGA. The bit file description will include which routes are available and / or which routes are not available. For example, the bit file may allow for any routing required across the dynamic portion of the FPGA, such as through the static portion of the FPGA, and indicate available routes for the dynamic portion of the FPGA.
[0447] It should be understood that in some embodiments, the bit file description may be obtained in any suitable manner. For example, the bit file description may be provided by a provider of an FPGA or an ASIC.
[0448] In some embodiments, a bit file description may be provided by a design tool. In this embodiment, the analysis step may be omitted. The design tool may output a bit file description. The bit file description may be for the static portion of the FPGA, including any required routing across the dynamic portion of the FPGA.
[0449] It will be appreciated that any other suitable technique may be used to generate the bit file description.In the example described above, a tool for designing an FPGA is used to provide analysis for generating a bit file.
[0450] It should be understood that different tools can be used in other embodiments. In some embodiments, the tool can be specific to a product or a series of products. For example, the provider of the FPGA can provide relevant tools for managing the FPGA.
[0451] In other embodiments, a general scripting tool may be used.
[0452] In some embodiments, different tools or different techniques can be used to determine the partial bit files. For example, the master bit file can be analyzed to determine which features correspond to which features. This may require generating multiple partial bit files.
[0453] It should be understood that step S3 is performed when the network interface device is installed relative to the host and executed on the physical FPGA device. Steps S1 and S2 can be performed as part of a design synthesis process to generate a bitfile image that implements the network interface device. In some embodiments, steps S1 and / or S2 are used to characterize the behavior of the FPGA. Once the FPGA characteristics are determined, the bitfile description is stored in the memory of all physical network interface devices, which will operate in a given defined manner.
[0454] In step S3, compilation is performed using the bitfile description and the eBPF program. The output of the compilation is a partial bitfile of the eBPF program. Compilation adds routing to the partial bitfile and adds the programming to be executed by each slice in the slice.
[0455] It should be understood that a bit file description can be provided in the deployed system. The bit file description can be stored in a memory. The bit file description can be stored on an FPGA, a network interface device, or a host device. In some embodiments, the bit file description is stored in a flash memory connected to an FPGA on a network interface device, or the like. The flash memory may also contain the master bit file.
[0456] eBPF programs can be stored with the bitfile description or separately. They can be stored on the FPGA, on the network interface device, or on the host. In the case of eBPF, the program can be transferred from a user-mode program to the kernel, both running on the host. The kernel transfers the program to a device driver, which then transfers it to a compiler running on the host or network interface device. In some embodiments, the eBPF program can be stored on the network interface device so that it can be run before booting the host OS.
[0457] The compiler may be provided at any suitable location on the network interface device, FPGA, or host. By way of example only, the compiler may run on a CPU on the network interface device.
[0458] The compiler flow will now be described. The compiler front end receives an eBPF program. The eBPF program can be written in any suitable language. For example, an eBPF program can be written in a C-type language. The compiler front end is configured to convert the program into an intermediate representation (IR). In some embodiments, the IR can be LLVM-IR or any other suitable IR.
[0459] In some embodiments, pointer analysis may be performed to create packet / mapping access primitives.
[0460] It should be understood that in some embodiments, optimization of the IR may be performed by a compiler. This may be optional in some embodiments.
[0461] The high-level synthesis backend of the compiler is configured to divide the program pipeline into multiple stages, generate data packet access taps and emit C code. In some embodiments, the design tool and / or the HLS part of the design tool used can be called to synthesize the output of the HLS stage.
[0462] The FPGA Atom compiler backend breaks the pipeline into stages and generates packet access taps. If-if transformations are performed to convert control dependencies into data dependencies. The design is placed and routed. The partial bitfile for the eBPF program is emitted.
[0463] There may be routing issues, such as Figure 20a As shown, there is a routing conflict. For example, slice A can communicate with slice C, and slice B can communicate with slice D. Figure 20a In the layout of , the common routing portion 60 has been allocated to the communication between slice A and slice C and the communication between slice A and slice C. In some embodiments, this can avoid the routing conflict. In this regard, reference is made to Figure 20b As can be seen, a separate route 62 is provided between slice A and slice C, compared to the route 64 between slice B and slice D.
[0464] In some embodiments, the bit file description may include multiple different routes for at least some subunit pairs. The compilation process will check for routing conflicts, such as Figure 20a In case of routing conflicts, the compiler can resolve or avoid such conflicts by selecting one of the appropriate alternative routes.
[0465] Figure 21 A partition 66 in the FPGA for executing eBPF programs is shown. This partition will interface with the static portion of the FPGA, for example, via a series of input flip-flops 68 and a series of output flip-flops. In some embodiments, as previously discussed, routing 70 may be present in the design.
[0466] The compiler may need to handle routing across the FPGA region being configured by the compiler. The compiler needs to generate a partial bitfile that fits within the reconfigurable partition in the main bitfile. When generating a main bitfile using a reconfigurable partition, the design tools will avoid using logic resources within the reconfigurable partition so that the partial bitfile can use those resources. However, the design tools may not be able to avoid using routing resources within the reconfigurable partition.
[0467] Therefore, the analysis tool will need to avoid using routing resources in the master bitfile that are already being used by the design tool. The analysis tool may need to ensure that its list of available routes in the bitfile description does not include any used resources that are being used by the master bitfile. The available routes can be defined in terms of routing templates, and because FPGAs are very regular, these routing templates can be used in many places within the FPGA. The routing resources used by the master bitfile break this regularity, which means that the analysis tool avoids using these templates in locations that conflict with the master bitfile. The analysis tool may need to generate new routing templates that can be used in those places and / or prevent certain routing templates from being used in specific locations.
[0468] We will now describe some examples of the functionality provided by the compiler when converting some example eBPF program slices into instructions to be executed by atomics.
[0469] Some embodiments may use any suitable synthesis tool to generate the bitfile description. By way of example only, some embodiments may use the Bluespec tool based on a model that uses atomic transactions for hardware.
[0470] In the first example, the eBPF program slice has two instructions:
[0471] Instruction 1: r1 += r2
[0472] Instruction 2: r1 += r3
[0473] The first instruction adds the number in register 1 (r1) to the number in register 2 (r2) and places the result in r1. The second instruction adds r1 to r3 and places the result in r1. Both instructions in this example use 64-bit registers, but only the lowest 32 bits are used. The upper 32 bits of the result are padded with zeros.
[0474] The compiler converts these into atomically executed instructions. A 32-bit addition instruction requires 32 pairs of lookup tables (LUTs), a 32-bit carry chain, and 32 flip-flops.
[0475] Each pair of lookup tables will add two bits to produce a 2-bit result. The structure of the carry chain allows a bit to be carried from the digit column to the next column during addition and allows a bit to be borrowed from the next column during subtraction.
[0476] The 32 flip-flops are storage elements that accept a value in one clock cycle and reproduce that value in the next clock cycle. These can be used to limit the amount of work completed per clock cycle and simplify timing analysis.
[0477] In some embodiments, the FPGA may include multiple slices. In some example slices, the carry chain propagates from the bottom (CIN) of the slice to the top (COUT) of the slice, which is then connected to the CIN input of the next slice.
[0478] In the example with a 4-bit carry chain per slice, eight slices are used to perform a 32-bit addition. In this embodiment, the atomicity can be considered to be provided by a pair of slices. This is because in some embodiments, it may be convenient to operate on 8-bit values.
[0479] In the example where each slice has an 8-bit carry chain, four slices are used to perform a 32-bit addition. In this embodiment, the atoms can be considered to be provided by the slices.
[0480] It will be appreciated that this is exemplary only and, as previously stated, atoms may be defined in any suitable manner.
[0481] In this example, the case where the FPGA has a slice that supports 8-bit carry chaining will now be used for the compilation of the first example eBPF program slice.
[0482] There are three 32-bit wide input values and one 32-bit wide output value. There may be other earlier instructions that produce these three input values. In the following, some arbitrary positions of slices (atoms) will be assumed.
[0483] The following numbering convention will be used. Slices (atoms) are arranged in regular rows and columns. XnYm represents the position of the atom in the arrangement. Xn represents the column and Ym represents the row. X6Y0 indicates that the slice is in column 6 and row 0. It should be understood that any other suitable numbering scheme can be used in other embodiments.
[0484] Assume that initial values are generated simultaneously at the following locations:
[0485] r1: slices X6Y0, X6Y1, X6Y2, and X6Y3
[0486] r2: slices X6Y4, X6Y5, X6Y6, and X6Y7
[0487] r3: slices X6Y8, X6Y9, X6Y10, and X6Y11
[0488] The result of the first instruction needs to be computed from four adjacent slices in the same column in order for the carry chain to connect correctly. The compiler may choose to compute this result in slices X7Y0, X7Y1, X7Y2, and X7Y3. To do this, the inputs need to be connected. There will be a connection from X6Y0 to X7Y0, from X6Y1 to X7Y1, from X6Y2 to X7Y2, and from X6Y3 to X7Y3. Corresponding connections are also required from X6Y4-X6Y7 to X7Y0-X7Y3.
[0489] These will be full-byte connections, meaning each of the 8 input bits is connected to the corresponding output bit. For example:
[0490] The output of slice X6Y0 flip-flop 0 is connected to input 0 of slice X7Y0 LUT 0.
[0491] The output of slice X6Y0 flip-flop 1 is connected to input 0 of slice X7Y0 LUT 1.
[0492] And so on, until
[0493] The output of slice X6Y0 flip-flop 7 is connected to input 0 of slice X7Y0 LUT 7.
[0494] During the first clock cycle, the r1 and r2 values from slices X6Y0-X6Y7 will be transferred to the inputs of slices X7Y0-X7Y3, will be processed by the LUT and carry chain, and the results will be stored in the flip-flops of these slices (X7Y0-X7Y3), ready to be used in the next cycle.
[0495] Move on to instruction 2. The compiler needs to choose a location to compute the result of instruction 2. It might choose to slice X7Y4 to X7Y7. Similarly, the result of instruction 1 (X7Y0 to X7Y3) is concatenated into the input of instruction 2 (X7Y4 to X7Y7).
[0496] The value of r3 is also needed. If r1, r2, and r3 are generated in cycle 0, then r1 + r2 will be generated in cycle 1. The value of r3 needs to be delayed one clock cycle so that it can be generated in cycle 1. The compiler can choose to generate r3 in cycle 1 by using slices X7Y8 through X7Y11. Then, connections need to be made from the original slice that generated r3 in cycle 0 (X6Y8 through X6Y11) to the new slice that generated the same value in cycle 1 (X7Y8 through X7Y11). Once this is done, connections are now needed from these new slices to the slice used for instruction 2. Therefore, the output of slice X7Y8 will be connected to the input of slice X7Y4, and so on.
[0497] The FPGA bitfile will contain the following characteristics:
[0498] - Full byte connection from X6Y0 to X7Y0 input 0 (initial r1 byte 0)
[0499] - Full byte connection from X6Y1 to X7Y1 input 0 (initial r1 byte 1)
[0500] - Full byte connection from X6Y2 to X7Y2 input 0 (initial r1 byte 2)
[0501] - Full byte connection from X6Y3 to X7Y3 input 0 (initial r1 byte 3)
[0502] - Full byte connection from X6Y4 to X7Y0 input 1 (initial r2 byte 0)
[0503] - Full byte connection from X6Y5 to X7Y1 input 1 (initial r2 byte 1)
[0504] - Full byte connection from X6Y6 to X7Y2 input 1 (initial r2 byte 2)
[0505] - Full byte connection from X6Y7 to X7Y3 input 1 (initial r2 byte 3)
[0506] - Full byte connection from X6Y8 to X7Y8 input 0 (initial r3 byte 0)
[0507] - Full byte connection from X6Y9 to X7Y9 input 0 (initial r3 byte 1)
[0508] - Full byte connection from X6Y10 to X7Y10 input 0 (initial r3 byte 2)
[0509] - Full byte connection from X6Y11 to X7Y11 input 0 (initial r3 byte 3)
[0510] - Slice X7Y0 is configured to add input 0 to input 1 (instruction 1 byte 0)
[0511] - Slice X7Y1 is configured to add input 0 to input 1 (instruction 1 byte 1)
[0512] - Slice X7Y2 is configured to add input 0 to input 1 (instruction 1 byte 2)
[0513] - Slice X7Y3 is configured to add input 0 to input 1 (instruction 1 byte 3)
[0514] - Slice X7Y8 is configured to copy input 0 to output (r3 delay byte 0)
[0515] - Slice X7Y9 is configured to copy input 0 to output (r3 delay byte 1)
[0516] - Slice X7Y10 is configured to copy input 0 to output (r3 delay byte 2)
[0517] - Slice X7Y11 is configured to copy input 0 to output (r3 delay byte 3)
[0518] - Full byte connection from X7Y0 to X7Y4 input 0 (instruction 1 byte 0)
[0519] - Full byte connection from X7Y1 to X7Y5 input 0 (instruction 1 byte 1)
[0520] - Full byte connection from X7Y2 to X7Y6 input 0 (instruction 1 byte 2)
[0521] - Full byte connection from X7Y3 to X7Y7 input 0 (instruction 1 byte 3)
[0522] - Full byte connection from X7Y8 to X7Y4 input 1 (r3 delay byte 0)
[0523] - Full byte connection from X7Y9 to X7Y5 input 1 (r3 delay byte 1)
[0524] - Full byte connection from X7Y10 to X7Y6 input 1 (r3 delay byte 2)
[0525] - Full byte connection from X7Y11 to X7Y7 input 1 (r3 delay byte 3)
[0526] - Slice X7Y4 is configured to add input 0 to input 1 (instruction 2 byte 0)
[0527] - Slice X7Y5 is configured to add input 0 to input 1 (instruction 2 byte 1)
[0528] - Slice X7Y6 is configured to add input 0 to input 1 (instruction 2 byte 2)
[0529] - Slice X7Y7 is configured to add input 0 to input 1 (instruction 2 byte 3)
[0530] The compiler does not need to generate the upper 32 bits of the result of instruction 2, since they are known to be zero. It can just record that fact and use zero whenever it needs to use them.
[0531] A second example of compilation of an eBPF slice will now be described.
[0532] Instruction 1: r1&=0xff
[0533] Instruction 2: r2&=0xff
[0534] Instruction 3: if r1 <r2 goto L1
[0535] Instruction 4: r1 = r2
[0536] Label L1.
[0537] The first instruction performs a bitwise AND operation on r1 with the constant 0xff and places the result in r1. If the corresponding bit was originally set to 1 in r1 and the corresponding bit is set, the given bit in the result will be set to the constant 1. Otherwise, it will be set to zero. Bits 0 through 7 of the constant 0xff are set to 1, and bits 8 through 63 are cleared, so the result will be bits 0 through 7 of r1 unchanged, while bits 8 through 63 will be set to zero. This simplifies the compiler's work because it knows that bits 8 through 63 are zero and does not need to generate them. The second instruction performs the same operation on r2.
[0538] Instruction 3 checks if r1 is less than r2 and jumps to label L1 if so. This skips instruction 4, which simply copies the value from r2 into r1. This instruction sequence finds the minimum of byte 0 of r1 and byte 0 of r2, placing the result in byte 0 of r1.
[0539] The compiler can use a technique called "if conversion" to convert a conditional jump into a select instruction:
[0540] Instruction 1: r1&=0xff
[0541] Instruction 2: r2&=0xff
[0542] Instruction 5: c1 = (r1 <r2)
[0543] Instruction 6: r1=c1? r1:r2
[0544] Instruction 5 compares r1 to r2 and sets c1 to 1 if r1 is less than r2, otherwise sets c1 to zero. Instruction 6 is a select instruction that copies r1 to r1 if c1 is set (this will have no effect), otherwise copies r2 to r1. If c1 is equal to 1, instruction 3 will skip instruction 4, which means r1 will retain its value from instruction 1. In this case, the select instruction also leaves r1 unchanged. If c1 is equal to 0, instruction 3 will not skip instruction 4, so r2 will be copied into r1 by instruction 4. Again, the select instruction will copy r2 into r1, so the new sequence has the same effect as the old one.
[0545] Instruction 6 is not a valid eBPF instruction. However, when the compiler processes it, the instruction is represented in LLVM-IR. Instruction 6 will be a valid instruction in LLVM-IR.
[0546] Now, these instructions need to be assigned to atomics. Assume that the input r1 is available in slices X0Y0 to X0Y7, and r2 is available in slices X0Y8 to X0Y15. Instructions 1 and 2 cause the compiler to note that the first 7 bytes of r1 and r2 are set to zero.
[0547] The compiler might then choose to compute the result of instruction 5 in slice X1Y0. A full-byte connection is required from the output of slice X0Y0 to input 0 of slice X1Y0, and a full-byte connection is required from the output of slice X0Y8 to input 1 of slice X1Y0. The two values are compared by subtracting one from the other and then checking if the calculation overflows by attempting to borrow from the next bit. The result of this comparison is then stored in flip-flop 7 of slice X1Y1.
[0548] Like the first example, r1 and r2 will need to be delayed one cycle to present their values at the correct time to instruction 6. The compiler might use slices X1Y1 and X1Y2 for r1 and r2 respectively.
[0549] The select instruction requires three inputs: c1, r1, and r2. Note that r1 and r2 are one byte wide, while c1 is only one bit wide. Suppose the compiler computes the result of the select instruction slice X2Y0. The selection is performed bit by bit, with each LUT in slice X2Y0 processing one bit:
[0550] If c1 is set, bit 0 of the result is r1 bit 0 and r2 bit 0
[0551] otherwise
[0552] If c1 is set, bit 1 of the result is r1 bit 1 and r2 bit 1
[0553] otherwise
[0554] ...and so on, until
[0555] If c1 is set, bit 7 of the result is r1 bit 7 and r2 bit 7
[0556] otherwise.
[0557] Each LUT may need to access the corresponding bit from r1 and the corresponding bit from r2, but all LUTs need to access c1. This means that cl needs to be copied between the bits of input 0 of the slice. Therefore, the connection of the input of instruction 6 will be: copy bit 7 of the output of slice X1Y0 to input 0 of slice X2Y0.
[0558] A full byte connection from the output of slice X1Y1 to input 1 of slice X2Y0.
[0559] A full byte connection from the output of slice X1Y2 to input 2 of slice X2Y0.
[0560] Another problem that needs to be solved is related to shift instructions. Consider the following example:
[0561] Shifting a 16-bit word left by 5 bits requires:
[0562] Set output bit 0 to zero
[0563] Set output bit 1 to zero
[0564] Set output bit 2 to zero
[0565] Set output bit 3 to zero
[0566] Set output bit 4 to zero
[0567] Copies input bit 0 to output bit 5
[0568] Copies input bit 1 to output bit 6
[0569] …
[0570] Copies input bit 10 to output bit 15
[0571] It should be noted that the input and output here are connected. The connected input is the output from the first slice. The connected output will go into the input of the second slice.
[0572] This connection may not be possible within a slice, but rather through the interconnect between slices. The compiler can assume that the 16-bit input value is produced by two adjacent slices in the same column, because the compiler can guarantee that the value will be produced there.
[0573] For example, suppose the input is produced by slices X0Y4 and X0Y5, and the output is going into slices X1Y4 and X1Y5. In this case, the following connections are required:
[0574] Bit 0 of slice X1Y4 is known to be zero and therefore does not need to be
[0575] Bit 1 of slice X1Y4 is known to be zero, so it is not necessary
[0576] Bit 2 of slice X1Y4 is known to be zero, so it is not needed.
[0577] Bit 3 of slice X1Y4 is known to be zero, so it is not needed.
[0578] Bit 4 of slice X1Y4 is known to be zero, so it is not needed.
[0579] Bit 5 of slice X1Y4 comes from bit 0 of slice X0Y4
[0580] Bit 6 of slice X1Y4 comes from bit 1 of slice X0Y4
[0581] Bit 7 of slice X1Y4 comes from bit 2 of slice X0Y4
[0582] Bit 0 of slice X1Y5 comes from bit 3 of slice X0Y4
[0583] Bit 1 of slice X1Y5 comes from bit 4 of slice X0Y4
[0584] Bit 2 of slice X1Y5 comes from bit 5 of slice X0Y4
[0585] Bit 3 of slice X1Y5 comes from bit 6 of slice X0Y4
[0586] Bit 4 of slice XIY5 comes from bit 7 of slice X0Y4
[0587] Bit 5 of slice X1Y5 comes from bit 0 of slice X0Y5
[0588] Bit 6 of slice XIY5 comes from bit 1 of slice X0Y5
[0589] Bit 7 of slice X1Y5 comes from bit 2 of slice X0Y5
[0590] The 8 connections to the inputs of slice X1Y5 can be thought of as shift connections or routing. Slice X1Y4 can use the same structure, but for the inputs to X1Y3 and X1Y4, it doesn't matter what inputs appear there because bits 5-7 match and the slice can ignore bits 0-4.
[0591] You might want to be able to shift by any amount between 1 and 7 bits. Concatenation by shifting 0 or 8 bits is the same as concatenation by a full byte, since in that case each bit is concatenated to the corresponding bit of the other slice.
[0592] Depending on the width of the values being shifted, a variable amount of shifting can be done in two or three stages. These stages are:
[0593] Stage 1: Shift by 0, 1, 2, or 3.
[0594] Stage 2: Shift by 0, 4, 8, or 12.
[0595] Stage 3: Shift by 0, 16, 32, or 48 (32-bit or 64-bit words only).
[0596] As another example, assume that the arithmetic right shift of a byte is a variable amount, the value to be shifted is produced by the slice X3Y2, and the shift amount is produced by X3Y3.
[0597] Arithmetic right shifts require an "Arithmetic Right Shift" type connection. This type of connection takes the outputs of one slice and connects it to the inputs of another slice, but in the process shifts them right by a constant amount, copying the sign bit as needed.
[0598] For example, an "arithmetic right shift by 3" connection would have:
[0599] Output bit 0 comes from input bit 3
[0600] Output bit 1 comes from input bit 4
[0601] Output bit 2 comes from input bit 5
[0602] Output bit 3 comes from input bit 6
[0603] Output bit 4 comes from input bit 7
[0604] Output bit 5 comes from input bit 7 (sign bit)
[0605] Output bit 6 comes from input bit 7 (sign bit)
[0606] Output bit 7 comes from input bit 7 (sign bit)
[0607] Stage 1 might be computed in slice X4Y2, in which case it requires the following connections:
[0608] From slice X3Y2 full byte to slice X4Y2 input 0
[0609] Arithmetic right shift 1 from slice X3Y2 to subslice X4Y2 input 1
[0610] Arithmetic right shift 2 from slice X3Y2 to subslice X4Y2 input 2
[0611] Arithmetic right shift 3 from slice X3Y2 to subslice X4Y2 input 3
[0612] Copy slice X3Y3 bit 0 to slice X4Y2 input 4
[0613] Copy slice X3Y3 bit 1 to slice X4Y2 input 5
[0614] Slice X4Y2 would then be configured to select one of the first four inputs based on input 4 and input 5 as follows:
[0615] Input 4 is 0, Input 5 is 0: Select Input 0
[0616] Input 4 is 1, input 5 is 0: Select input 1
[0617] Input 4 is 0, Input 5 is 1: Select Input 2
[0618] Input 4 is 1 and Input 5 is 1: Select Input 3
[0619] The offset can be copied from slice X3Y3 to slice X4Y3 to provide a delayed version.
[0620] Phase 2 can be calculated in slice X5Y2, in which case it requires the following connections:
[0621] Input 0 from the full byte of slice X4Y2 to slice X5Y2
[0622] Arithmetic right shift 4 from slice X4Y2 to subslice X5Y2 input 1
[0623] Copy slice X4Y3 bit 2 to slice X5Y2 input 2
[0624] Slice X5Y2 would then be configured to select either Input 0 or Input 1 based on Input 2 as follows:
[0625] Input 2 is 0: Select input 0
[0626] Input 2 is 1: Select input 1
[0627] The output of the slice X5Y2 will be the result of the variable arithmetic right shift operation.
[0628] The bitfile for a given atom might look like this:
[0629] Atom identity information
[0630] A list of other atoms from which a given atom can receive input, and the available routes for that input
[0631] A given atom is a list of other atoms that can provide an output and the available routes for that output.
[0632] It will be appreciated that, because FPGAs are regular structures, there may be a general template that can be used for multiple atoms and modified for individual atoms as necessary.
[0633] For example, a bitfile description of slice X7Y1 might specify the following possible inputs and outputs:
[0634] From the input of X6Y1 via route A or route B
[0635] From the input of X6Y5 via route C or route D
[0636] From the input of X7Y0 via route E or route F
[0637] Via route G or route H to the output of X8Y1
[0638] Via route I or route J to the output of X7Y2
[0639] Output via Routing K or Routing L to X7Y5.
[0640] The compiler will use this bitfile description to provide partial bitfiles for the input and output of slice X7Y1 for the first eBPF example described previously,
[0641] From the input of X6Y1 via route A
[0642] From the input of X6Y5 via route C
[0643] Output via Routing K or Routing L to X7Y5.
[0644] For example, a bitfile description of a slice XnYm could specify the following possible inputs and outputs:
[0645] From the input of Xn-lYm via route A or route B
[0646] From the input of Xn-lYm+4 via routing C or routing D
[0647] From the input of XnYm-1 via route E or route F
[0648] Output via routing G or routing H to Xn+1Ym
[0649] Output to XnYm+1 via route I or route J
[0650] Output via route K or route L to XnYm+4.
[0651] This bitfile description can be modified to remove one or more routes that the compiler cannot use, as described previously. This could be because the route is used by another atom or is being used for cross-partition routing.
[0652] It should be understood that the compiler can be implemented by a computer program comprising computer executable instructions that can be executed by one or more computer processors. The compiler can run on hardware such as at least one processor operating in conjunction with one or more memories.
[0653] It should be noted that while the above describes exemplifying embodiments, there are numerous variations and modifications which may be made to the disclosed solution without departing from the scope of the present invention.
[0654] Therefore, the embodiments may vary within the scope of the appended claims. In general, some embodiments may be implemented in hardware or dedicated circuits, software, logic, or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software executed by a controller, microprocessor, or other computing device, but the embodiments are not limited thereto.
[0655] The embodiments may be implemented by computer software stored in a memory and executed by at least one data processor of the entity involved, or may be executed by hardware or by a combination of software and hardware.
[0656] The software may be stored on physical media such as memory chips or memory blocks implemented within the processor, magnetic media such as hard disk or floppy disks, and optical media such as DVDs and their data variants, CDs.
[0657] The memory may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory.
[0658] The data processor may be of any type suitable to the local technical environment and may include, as non-limiting examples, one or more of a general purpose computer, a special purpose computer, a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a gate level circuit, and a processor based on a multi-core processor architecture.
[0659] Various modifications and variations will become apparent to those skilled in the relevant arts in view of the foregoing description, when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings will still fall within the scope defined by the appended claims.
Claims
1. A network interface device for connecting a host to a network, characterized in that The network interface device includes: a first interface configured to receive a plurality of data packets; a configurable hardware module comprising a plurality of processing units, each processing unit being associated with a predetermined type of operation that can be performed in a single step, wherein at least some of the plurality of processing units are associated with different predetermined types of operations, wherein the hardware module is configurable to interconnect at least some of the plurality of processing units to provide a first data processing pipeline for processing one or more of the plurality of data packets so as to perform a first function on the one or more of the plurality of data packets, wherein for a first data packet among one or more data packets among the plurality of data packets, two or more of at least some of the plurality of processing units in the first data processing pipeline are configured to perform their associated at least one predetermined operation on the first data packet in parallel; and a shared memory accessible by two or more of at least some of the plurality of processing units, wherein the shared memory is configured to store a state associated with the first data packet, wherein during execution of the first function by the hardware module, the two or more of the plurality of processing units are configured to access and modify the state associated with the first data packet, wherein a first processing unit among two or more of at least some of the plurality of processing units is configured to cease execution during access by a second processing unit among two or more of at least some of the plurality of processing units to a value of a state associated with the first data packet.
2. The network interface device according to claim 1, wherein: Two or more of at least some of the plurality of processing units are configured to: performing its associated predetermined type of operation within a predetermined length of time specified by the clock signal; and In response to the end of the predetermined time period, the result of the respective at least one operation is transmitted to the next processing unit.
3. The network interface device according to claim 1, wherein: Each of the plurality of processing units includes an application specific integrated circuit configured to perform the at least one operation associated with the respective processing unit.
4. The network interface device according to claim 1, wherein: Each of the plurality of processing units includes digital circuitry and memory for storing state associated with processing performed by the digital circuitry, wherein the digital circuitry is configured to perform a predetermined type of operation associated with the respective processing unit in communication with the memory.
5. The network interface device according to claim 1, wherein: One or more of the plurality of processing units are individually configured to perform specific operations for respective pipelines based on their associated predetermined operation types.
6. The network interface device according to claim 1, wherein: The hardware module is configured to receive an instruction and, in response to the instruction, perform at least one of the following operations: interconnecting at least some of the plurality of processing units to provide a data processing pipeline for processing one or more packets of the plurality of packets; causing one or more processing units of the plurality of processing units to perform their associated predetermined operation types for the one or more data packets; adding one or more processing units of the plurality of processing units to a data processing pipeline; as well as One or more processing units of the plurality of processing units are removed from the data processing pipeline.
7. The network interface device according to claim 1, wherein: The predetermined operation includes at least one of the following operations: loading at least one value of the first data packet from a memory; storing at least one value of the data packet in a memory; as well as A lookup is performed in a lookup table to determine the action to be performed on the packet.
8. The network interface device according to claim 1, wherein: One or more of at least some of the plurality of processing units are configured to pass at least one result of its associated at least one predetermined operation to a next processing unit in the first processing pipeline, and the next processing unit is configured to perform a next predetermined operation based on the at least one result.
9. The network interface device according to claim 1, wherein: Each of the different predetermined operation types is defined by a different template.
10. The network interface device according to claim 1, wherein: The predetermined operation type includes at least one of the following operations: Access data packets; accessing a lookup table stored in a memory of the hardware module; performing logical operations on data loaded from the data packet; and Perform logical operations on data loaded from a lookup table.
11. The network interface device according to claim 1, wherein: The hardware module includes routing hardware, wherein the hardware module is configured to: route data packets between the plurality of processing units in a specific order specified by the first data processing pipeline by configuring the routing hardware, thereby interconnecting at least some of the plurality of processing units to provide a first data processing pipeline.
12. The network interface device according to claim 1, wherein: The hardware module may be configured to interconnect at least some of the plurality of processing units to provide a second data processing pipeline for processing one or more of the plurality of data packets to perform a second function different from the first function.
13. The network interface device according to claim 1, wherein: The hardware module may be configured to interconnect at least some of the plurality of processing units to provide a second data processing pipeline after interconnecting at least some of the plurality of processing units to provide the first data processing pipeline.
14. The network interface device according to claim 1, wherein: The network interface device includes additional circuitry separate from the hardware module and configured to perform a first function on one or more packets of the plurality of packets.
15. The network interface device according to claim 14, wherein: The further circuitry includes at least one of the following: Field Programmable Gate Arrays; and Multiple central processing units.
16. The network interface device according to claim 14 or 15, characterized in that: The network interface device includes at least one controller, the further circuitry being configured to execute the first function on a data packet during a compiling process for the first function to be executed in the hardware module, and the at least one controller being configured to, in response to completion of the compiling process, control the hardware module to begin executing the first function on the data packet.
17. The network interface device according to claim 16, wherein: The at least one controller is configured to, in response to determining that a compile process for the first function to be executed in the hardware module has been completed, control the further circuit to stop executing the first function on the data packet.
18. The network interface device according to claim 14 or 15, characterized in that: The network interface device includes at least one controller, the hardware module is configured to perform the first function on the data packet during the compilation process for the first function to be performed in the additional circuit, and the at least one controller is configured to: determine that the compilation process for the first function to be performed in the additional circuit is completed, and in response to the determination, control the additional circuit to start performing the first function on the data packet.
19. The network interface device according to claim 18, wherein: The at least one controller is configured to, in response to determining that a compile process for the first function to be executed in the further circuit is completed, control the hardware module to stop executing the first function on the data packet.
20. The network interface device according to claim 1, wherein: The network interface device includes at least one controller configured to perform a compilation process to provide the first functionality to be executed in the hardware module.
21. A data processing system, characterized in that: The data processing system includes a host device and the network interface device according to claim 1, wherein the data processing system includes at least one controller configured to perform a compilation process to provide the first function to be executed in the hardware module.
22. The data processing system according to claim 21, wherein: The at least one controller is provided by one or more of: the network interface device; and The host device.
23. The data processing system according to claim 21, wherein: The compiling process is performed in response to a determination by the at least one controller that the computer program representing the first function is safe for execution in kernel mode of the host device.
24. The data processing system according to claim 21, wherein: The at least one controller is configured to perform a compilation process by directing each of at least some of the plurality of processing units to perform at least one operation represented by a series of computer code instructions in a specific order of the first data processing pipeline, wherein the plurality of operations provide the first functionality for the one or more data packets of the plurality of data packets.
25. The data processing system according to claim 21, wherein: The at least one controller is configured to: Before the compilation process is completed, sending a first instruction to cause further circuitry of the network interface device to perform the first function on the data packet; and A second instruction is sent, so that the hardware module starts to execute the first function on the data packet after the compiling process is completed.
26. A method for implementation in a network interface device, characterized in that, The method comprises: receiving a plurality of data packets at the first interface; and configuring a hardware module to interconnect at least some of the plurality of processing units of the hardware module to provide a first data processing pipeline for processing one or more packets of the plurality of packets to perform a first function on the one or more packets of the plurality of packets, wherein each processing unit is associated with a predetermined operation type that can be performed in a single step, wherein at least some of the plurality of processing units are associated with different predetermined operation types, and wherein for a first data packet among one or more data packets among the plurality of data packets, two or more of at least some of the plurality of processing units in the first data processing pipeline are configured to perform at least one predetermined operation associated therewith on the first data packet in parallel, wherein during execution of the first function, two or more of at least some of the plurality of processing units are configured to access and modify a state associated with the first data packet within a shared memory of the network interface device, the shared memory being accessible by two or more of at least some of the plurality of processing units in the first data processing pipeline, Wherein, during execution of the first function, a first processing unit among two or more of at least some of the plurality of processing units is configured to cease execution during access by a second processing unit among two or more of at least some of the plurality of processing units in the first data processing pipeline to a value of a state associated with the first data packet.
27. A non-transitory computer-readable medium, characterized in that The non-transitory computer-readable medium includes program instructions for causing a network interface device to perform a method, the method comprising: receiving a plurality of data packets at the first interface; and configuring a hardware module to interconnect at least some of the plurality of processing units of the hardware module to provide a first data processing pipeline for processing one or more packets of the plurality of packets to perform a first function on the one or more packets of the plurality of packets, wherein each processing unit is associated with a predetermined operation type that can be performed in a single step, wherein at least some of the plurality of processing units are associated with different predetermined operation types, and wherein for a first data packet among one or more data packets among the plurality of data packets, two or more of at least some of the plurality of processing units in the first data processing pipeline are configured to perform at least one predetermined operation associated therewith on the first data packet in parallel, wherein during execution of the first function, two or more of at least some of the plurality of processing units are configured to access and modify a state associated with the first data packet within a shared memory of the network interface device, the shared memory being accessible by two or more of at least some of the plurality of processing units in the first data processing pipeline, Wherein, during execution of the first function, a first processing unit among two or more of at least some of the plurality of processing units is configured to cease execution during access by a second processing unit among two or more of at least some of the plurality of processing units in the first data processing pipeline to a value of a state associated with the first data packet.
Citation Information
Patent Citations
Data processing system and method for task scheduling in a data processing system
CN103765384A
Data processing device with mechanism for controlling bus priority of multiple processors
US20070214302A1
Adjustable cycle pipeline system and method
US7861067B1
Streaming interconnect architecture
US9940284B1