System and method for implementing machine learning network algorithms in data plane
Through the system architecture of integrated circuit devices and the use of feature extraction and inference circuits, the problems of high latency and high resource consumption caused by feature extraction and model inference in integrated circuit devices are solved, efficient feature generation and group or stream processing are achieved, and the flexibility and accuracy of network problem solving are improved.
Patent Information
- Application Number
- CN202510246220.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-04-10
- Filing Date
- 2025-03-04
- Publication Date
- 2025-10-17
AI Technical Summary
When deploying algorithms in the data plane or control plane of integrated circuit devices, feature extraction and model inference may result in high latency and high computing resource consumption, while making it difficult to maintain high-throughput algorithm mapping.
The system architecture adopts integrated circuit devices, including feature extraction circuits and inference circuits. Groups are replicated through a replication engine, feature types are determined using programmable data plane circuits, and features are generated and normalized through extraction circuits. The inference circuits perform group-by-group or stream-by-stream inference, achieving flexible feature selection and data pipeline processing.
It improves the accuracy of feature extraction and the efficiency of inference, reduces computing resources and power consumption, achieves high-quality feature generation and group or stream processing, and solves network problems.
Smart Images

Figure HDA0005295569740000011 
Figure HDA0005295569740000021 
Figure HDA0005295569740000031
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure generally relates to a system architecture for performing feature extraction, stream processing, machine learning (ML) model selection, and / or artificial intelligence (AI) ML inference on an integrated circuit device. BACKGROUND
[0002] This section is intended to introduce the reader to various aspects of art that can be related to various aspects of the present disclosure and is not intended to limit the scope of the claimed subject matter. Believing that this discussion helps the reader to better understand the context of the various aspects of the present disclosure as they can be presented for the claims that follow thereafter. Accordingly, it should be understood that these statements are to be read in this light and not as admissions of prior art.
[0003] Many network functions can be implemented using algorithms (e.g., congestion control, classification, anomaly detection, quality of service (QoS) policy adjustment, etc.) in a data plane or a control plane of an integrated circuit device. However, deploying algorithms in the data plane or the control plane can involve implementation of feature extraction, stream classification, model selection, and / or model inference (e.g., AI / ML model inference) for a large number of flows. Feature extraction and model inference can be implemented in a central processing unit (CPU), which can result in high latency due to data transfer to and from the CPU. Alternatively, feature extraction can be implemented in fixed hardware blocks, which can result in higher consumption of computational resources and power. Moreover, features for solving network problems can vary, and high throughput of mapping algorithms to models can be difficult to maintain. Therefore, it can be desirable to enable flexibility in feature extraction, feature selection based on a particular workload, and / or per-packet or per-flow inference in a data pipeline to solve network problems. BRIEF DESCRIPTION OF DRAWINGS
[0004] Various aspects of the present disclosure can be better understood when read in conjunction with the following detailed description and with reference to the drawings, in which:
[0005] Figure 1 is a block diagram of a system for programming an integrated circuit device according to embodiments of the present disclosure;
[0006] Figure 2 is a block diagram of an integrated circuit device according to embodiments of the present disclosure; Figure 1
[0007] Figure 3 is a block diagram of a programmable fabric of an integrated circuit device according to embodiments of the present disclosure; Figure 1
[0008] Figure 4 is a block diagram of a programmable fabric of an integrated circuit device according to embodiments of the present disclosure; Figure 1 Examples of circuitry employed by the integrated circuit device of
[0009] Figure 5 is performed by the integrated circuit device of Figure 1 a block diagram of a first processing flow or a second processing flow performed by example circuitry employed by the integrated circuit device of
[0010] Figure 6 is performed by the integrated circuit device of Figure 1 a packet processing pipeline of the integrated circuit device of
[0011] Figure 7 is an example method for feature extraction via data plane circuitry according to embodiments of the disclosure;
[0012] Figure 8 is an example illustration of a packet processing pipeline and feature extraction circuitry using on-chip memory or off-chip memory according to embodiments of the disclosure;
[0013] Figure 9 is an example illustration of data plane memory partitioning according to embodiments of the disclosure; and
[0014] Figure 10 is an example illustration of a data processing system that can contain Figure 1 a block diagram of a data processing system of the integrated circuit device of DETAILED DESCRIPTION
[0015] One or more specific embodiments will be described below. In an effort to provide a concise description of these embodiments, all features of an actual implementation can not be described in the specification. It should be appreciated that in the development of any such actual implementation, as in any engineering or design project, numerous implementation-specific decisions must be made to achieve the developers’ specific goals, such as compliance with system-related and business-related constraints, which can vary from one implementation to another. Moreover, it should be appreciated that such a development effort might be complex and time consuming, but would nevertheless be a routine undertaking of design, fabrication, and manufacture for those of ordinary skill having the benefit of this disclosure.
[0016] When introducing elements of various embodiments of the present disclosure, the articles "a," "an," and "the" are intended to mean that there are one or more of the elements. The terms "comprising," "including," and "having" are intended to be inclusive and mean that there can be additional elements other than the listed elements. Additionally, it should be understood that references to "one embodiment" or "an embodiment” of the present disclosure are not intended to be interpreted as excluding the existence of additional embodiments that also incorporate the recited features.
[0017] The present systems and techniques relate to embodiments for generating features from an incoming packet stream, normalizing the features, and interfacing with an inference engine to perform packet-by-packet or flow-by-flow inference. For example, an integrated circuit device can employ a system architecture that includes feature extraction circuitry and inference circuitry (e.g., AI / ML inference circuitry). In one embodiment, the integrated circuit device can receive a packet (or a stream of packets) at a network interface of the integrated circuit device. The integrated circuit device can replicate the packet or create a copy of the packet via a replication engine and provide the copy of the packet and associated metadata to the extraction circuitry. In another embodiment, the integrated circuit device can provide the packet to data plane circuitry (e.g., programmable data plane, programmable data plane fabric) to determine a type of AI model to implement and / or a type of feature to extract. The data plane circuitry can then provide the packet and one or more hints to the extraction circuitry to enable the extraction circuitry to extract the determined type of feature.
[0018] The extraction circuitry can be implemented as offload processing circuitry or as a co-processing system with a packet processing pipeline of the data plane circuitry. The extraction circuitry can extract (e.g., generate) features from the packet, normalize the features, and / or transform the features. For example, the extraction circuitry can extract the features via the packet processing pipeline or via dedicated circuitry. Additionally, the extraction circuitry can include feature vector transformation circuitry to pre-process and / or normalize data associated with the packet for the inference circuitry. The extraction circuitry can be programmable, which can enable a user to customize the normalization and / or transformation of the data. Accordingly, based on the customized normalization and / or transformation of the data, accuracy of the inference circuitry can be improved.
[0019] In one embodiment, the extraction circuitry can provide (e.g., send, transmit) the extracted features directly to the inference circuitry. The inference circuitry can then perform packet-by-packet or flow-by-flow inference based on the extracted features. For example, the inference circuitry can utilize various types of AI models and / or various instances of a single AI model to identify a solution (e.g., generate a result) to a particular network problem. In another embodiment, the extraction circuitry can provide the extracted features to the data plane circuitry of the integrated circuit device. The data plane circuitry can enable analysis of the extracted features to determine a list of AI models to employ on the extracted features. The data plane circuitry can then provide the list of AI models and the extracted features to the inference circuitry, which performs packet-by-packet or flow-by-flow inference based on the list of AI models and the extracted features. As such, the system architecture of the integrated circuit device described herein can enable the integrated circuit device to programmably generate high quality features, normalize the features, and interface with the inference circuitry to perform packet processing (e.g., packet-by-packet or flow-by-flow processing) to resolve network problems.
[0020] With the above in mind, Figure 1 is a block diagram of a system 10 that can implement one or more functions. For example, a designer can wish to implement a function such as the operations of the present disclosure on an integrated circuit device 12 (e.g., a programmable logic device such as a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC), an integrated circuit system). In some cases, the designer can specify a high-level program to be implemented, such as program or This can enable the designer to more efficiently and easily provide programming instructions to configure a set of programmable logic cells of the integrated circuit device 12 without having specific knowledge of a low-level hardware description language (e.g., Verilog or VHDL). For example, because Similar to other high-level programming languages such as C++, a programmable logic designer familiar with such programming languages can have a shorter learning curve than a designer who is required to learn an unfamiliar low-level hardware description language to implement a new function in the integrated circuit device 12. Additionally or alternatively, a subset of the high-level program can be implemented using a lower-level language such as a register transfer language (RTL) and / or translated into a lower-level language.
[0021] The designer can implement the high-level design using design software 14 such as INTEL CORPORATION’s version). The design software 14 can use a compiler 16 to convert the high-level program into a low-level description. In some embodiments, the compiler 16 and the design software 14 can be packaged into a single software application. The compiler 16 can provide machine-readable instructions representative of the high-level program to a host 18 and to the integrated circuit device 12. The host 18 can receive a host program 22 (which can be implemented by the kernel program 20). To implement the host program 22, the host 18 can pass instructions from the host program 22 to the integrated circuit device 12 via a communication link 24, which can be, for example, a direct memory access (DMA) communication or a peripheral component interconnect express (PCIe) communication. In some embodiments, the kernel program 20 and the host 18 can enable configuration of logic blocks 26 on the integrated circuit device 12. The logic blocks 26 can include circuitry and / or other logic elements and can be configured to implement arithmetic operations such as addition and multiplication.
[0022] The designer can use the design software 14 to generate and / or specify a low-level program, such as the low-level hardware description language described above. For example, the design software 14 can be used to map a workload to one or more routing resources of the integrated circuit device 12 based on timing, wire usage, logic utilization, and / or routability. Additionally or alternatively, the design software 14 can be used to route first data to a portion of the integrated circuit device 12 and second data, power, and clock signals to a second portion of the integrated circuit device 12. Furthermore, in some embodiments, the system 10 can be implemented without the host program 22 and / or without a separate host program 22. Furthermore, in some embodiments, the techniques described herein can be implemented in a circuit according to a non-programmable circuit design. Therefore, the embodiments described herein are intended to be illustrative and not restrictive.
[0023] Turning now to a more detailed discussion of the integrated circuit device 12, Figure 2 1 is a block diagram of an example of an integrated circuit device 12 as a programmable logic device, such as a field programmable gate array (FPGA). Furthermore, it should be understood that the integrated circuit device 12 may be any other suitable type of programmable logic device, such as a structured ASIC (eASIC) such as Intel Corporation's ASIC implementation. TM ) and / or application specific standard products). The integrated circuit device 12 may have input / output circuitry 42 for driving signals out of the device via input / output pins 44 and for receiving signals from other devices. Interconnect resources 46 (such as global and local vertical and horizontal wires and buses) and / or configuration resources (e.g., hardwired couplings, logical couplings not implemented by designer logic) may be used to route signals on the integrated circuit device 12. In addition, the interconnect resources 46 may include fixed interconnects (wires) and programmable interconnects (i.e., programmable connections between various fixed interconnects). For example, the interconnect resources 46 may be used to route signals (such as clock or data signals) throughout the integrated circuit device 12. Additionally or alternatively, the interconnect resources 46 may be used to route power (e.g., voltage) throughout the integrated circuit device 12. Programmable logic 48 may include combinational and sequential logic circuits. For example, the programmable logic 48 may include lookup tables, registers, and multiplexers. In various embodiments, the programmable logic 48 may be configured to perform custom logic functions. The programmable interconnects associated with the interconnect resources may be considered part of the programmable logic 48.
[0024] Programmable logic devices, such as integrated circuit device 12, can include programmable elements 50 having programmable logic 48. In some embodiments, at least some of the programmable elements 50 can be grouped into logic array blocks (LABs). As noted above, a designer (e.g., user, customer) can (re)program (e.g., (re)configure) the programmable logic 48 to perform one or more desired functions. For example, some programmable logic devices can be programmed or reprogrammed by configuring the programmable elements 50 using a mask-programmed arrangement, which is performed during semiconductor fabrication. Other programmable logic devices are configured after the semiconductor fabrication operation has been completed, such as by programming the programmable elements 50 using electrical programming or laser programming. In general, the programmable elements 50 can be based on any suitable programmable technology, such as fuse, antifuse, electrically programmable read-only memory technology, random access memory cells, mask-programmed elements, etc.
[0025] Many programmable logic devices are electrically programmed. With an electrically programmed arrangement, the programmable elements 50 can be formed from one or more memory cells. For example, during programming, configuration data is loaded into the memory cells using the input / output pins 44 and the input / output circuitry 42. In one embodiment, the memory cells can be implemented as random access memory (RAM) cells. The use of RAM technology-based memory cells described herein is intended to be merely an example. Moreover, since these RAM cells are loaded with configuration data during programming, they are sometimes referred to as configuration RAM cells (CRAM). The memory cells can each provide a respective static control output signal that controls the state of an associated logic component in the programmable logic 48. In some embodiments, the output signal can be applied to the gate of a metal-oxide semiconductor (MOS) transistor within the programmable logic 48.
[0026] The integrated circuit device 12 can include any programmable logic device, such as a field programmable gate array (FPGA) 70, as Figure 3FPGA 70 is referred to as an FPGA for purposes of this example, but it should be understood that the device can be any suitable type of programmable logic device (e.g., an application specific integrated circuit and / or an application specific standard product). In one example, FPGA 70 is a sectorized FPGA of the type described in U.S. Patent Application Publication No. 2016 / 0049941, entitled “Programmable Circuit Having Multiple Sectors,” which is incorporated by reference herein in its entirety for all purposes. FPGA 70 can be formed on a single plane. Additionally or alternatively, FPGA 70 can be a three-dimensional FPGA having a base die and a fabric die of the type described in U.S. Patent No. 10833679, entitled “Multi-Purpose Interface for Configuration Data and Designer Fabric Data,” which is incorporated by reference herein in its entirety for all purposes.
[0027] In Figure 3 example, FPGA 70 can include a transceiver 72, which can include and / or use input / output circuitry, such as input / output circuitry 42 in Figure 2 to drive signals off of FPGA 70 and to receive signals from other devices. Interconnect resources 46 can be used to route signals, such as clock or data signals, throughout FPGA 70. FPGA 70 is sectorized, meaning that programmable logic resources can be distributed through a plurality of discrete programmable logic sectors 74. Programmable logic sectors 74 can include a plurality of programmable elements 50 having operations defined by configuration memory 76 (e.g., CRAM). Power supply 78 can provide voltage (e.g., supply voltage) and current sources to a power distribution network (PDN) 80, which distributes power to various components of FPGA 70. Circuits operating FPGA 70 cause power to be drawn from power distribution network 80.
[0028] There can be any suitable number of programmable logic sectors 74 on the FPGA 70. Indeed, while 29 programmable logic sectors 74 are shown here, it will be appreciated that more or fewer sectors can occur in actual implementations (e.g., in some cases, approximately 50, 100, 500, 1000, 5000, 10000, 50000, or 100000 sectors or more). The programmable logic sectors 74 can include a sector controller (SC) 82 for controlling operation of the programmable logic sectors 74. The sector controller 82 can communicate with a device controller (DC) 84.
[0029] The sector controller 82 can accept commands and data from the device controller 84 and can read data from and write data to its configuration memory 76 based on control signals from the device controller 84. In addition to these operations, the sector controller 82 can be augmented with a number of additional capabilities. For example, such capabilities can include local ordering of reads and writes to implement error detection and correction on the configuration memory 76, and ordering of test control signals to implement various test modes.
[0030] The sector controller 82 and the device controller 84 can be implemented as state machines and / or processors. For example, the operations of the sector controller 82 or the device controller 84 can be implemented as separate routines in a memory containing a control program. Such a control program memory can be fixed in read-only memory (ROM) or stored in writable memory such as random access memory (RAM). The size of the ROM can be larger than the size to store only one copy of each routine. This can allow for multiple variants of the routines according to the “modes” that the local controller can enter. When the control program memory is implemented as RAM, the RAM can be written with new routines to implement new operations and functionality into the programmable logic sector 74. This can provide scalability in an efficient and easily understood manner. This can be useful because new commands can bring a large amount of local activity within the sector, while only a small amount of communication between the device controller 84 and the sector controller 82 is required.
[0031] Accordingly, the sector controller 82 can communicate with the device controller 84, which can coordinate the operations of the sector controller 82 and communicate commands initiated from outside the FPGA 70. To support this communication, the interconnect resources 46 can act as a network between the device controller 84 and the sector controller 82. The interconnect resources 46 can support various signals between the device controller 84 and the sector controller 82. In one example, these signals can be transmitted as communication packets.
[0032] The use of the configuration memory 76 based on RAM technology described herein is intended to be merely an example. Moreover, the configuration memory 76 can be distributed among the various programmable logic sectors 74 of the FPGA 70 (e.g., as RAM cells). The configuration memory 76 can provide respective static control output signals that control the state of the associated programmable elements 50 or programmable components of the interconnect resources 46. The output signals of the configuration memory 76 can be applied to the gates of metal-oxide-semiconductor (MOS) transistors that control the state of the programmable elements 50 or programmable components of the interconnect resources 46.
[0033] The programmable elements 50 of the FPGA 40 can also include some signal metal (e.g., communication wires) to transmit signals. In one embodiment, the programmable logic sectors 74 can be provided in the form of vertical routing channels (e.g., interconnects formed along the y-axis of the FPGA 70) and horizontal routing channels (e.g., interconnects formed along the x-axis of the FPGA 70), and each routing channel can include at least one trace to route at least one communication wire. If desired, the communication wires can be shorter than the entire length of the routing channels. That is, the communication wires can be shorter than the first die region or the second die region. A wire of length L can span L routing channels. As such, the length of four wires in a horizontal routing channel can be referred to as “H4” wires, and the length of four wires in a vertical routing channel can be referred to as “V4” wires.
[0034] As mentioned above, some embodiments of programmable logic fabric can be configured using indirect configuration techniques. For example, an external host device can pass configuration data packets to the configuration management hardware of the FPGA 70. The data packets can be passed internally using data paths and specific firmware that are typically custom-made for passing configuration packets and can be based on a specific host device driver (e.g., for compatibility). The customization can also be associated with a specific device tape out, which typically results in a high cost for the specific tape out and / or reduced scalability of the FPGA 70.
[0035] With the above in mind, Figure 4 is in accordance with embodiments of the present disclosure Figure 1FIG. 1 illustrates an example of a system 10 that includes an integrated circuit device 12 that employs circuitry (including extraction circuitry 100 (e.g., extraction blocks, extraction components) and inference circuitry 102 (e.g., inference blocks, inference components)). As shown, the integrated circuit device 12 can include the extraction circuitry 100, the inference circuitry 102, a replication engine 106 (e.g., replication circuitry), a data plane circuitry 108 (e.g., programmable data plane), and / or an interface 110 (e.g., host interface). Further, the integrated circuit device 12 can be in communication with a central processing unit 104 (e.g., CPU circuitry, CPU configuration), which can be the host 18 or at least a portion of the host 18.
[0036] The extraction circuitry 100 can include feature extraction circuitry 112 and feature vector transform (FVT) circuitry 114. The inference circuitry 102 can include an AI / ML inference engine 116. It should be noted that the interconnection between the extraction circuitry 100, the inference circuitry 102, and / or the data plane circuitry 108 can be included in a monolithic system or a disaggregated system. In this way, the packet processing pipeline of the data plane circuitry 108 can operate with the extraction circuitry 100 and / or the inference circuitry 102 as co-processors.
[0037] The CPU 104 can operate or function the same or similar as processing circuitry (e.g., processor, processing system). Indeed, the CPU 104 can fetch (e.g., receive, extract) instructions from memory (e.g., memory of the integrated circuit device 12 and / or the host 18) and execute the instructions. For example, the instructions can include computations, logical operations, data movement, and / or control operations. Further, the CPU 104 can configure, control, and / or coordinate the operation of the extraction circuitry 100, the inference circuitry 102, the replication engine 106, and / or the data plane circuitry 108. As an example, the CPU 104 can configure and / or control the extraction circuitry 100, the inference circuitry 102, and / or the data plane circuitry 108 via software instructions and / or hardware configurations that set communication protocols, data flows, hardware circuit states, and the like.
[0038] The extraction circuitry 100 can include the feature extraction circuitry 112 as a standalone circuit (e.g., operating independently), which can operate or function as an offload processing circuit (e.g., feature extraction offload processing circuit) or as a co-processor to the data plane circuitry 108. The extraction circuitry 100 can extract raw features via the feature extraction circuitry 112 as a standalone circuit or through the packet processing pipeline of the data plane circuitry 108 (e.g., via one or more stages of the packet processing pipeline).
[0039] The extraction circuit 100 can then normalize and / or transform one or more features of one or more packets from a workload via the FVT circuit 114. In effect, the FVT circuit 114 can enable derivation of AI / ML easy-to-control features for the data plane circuit 108 and / or the inference circuit 102. In some embodiments, the integrated circuit device can replicate the FVT circuit 114 to enable load balancing across the replicas by performing a hash function on a unique flow ID. Additionally, the FVT circuit 114 can perform a floating point analysis to determine a minimum and / or maximum standard deviation, a minimum and / or maximum normalization, a mean square sum, or any other suitable floating point analysis of the extracted features.
[0040] The extraction circuit 100 can be programmable to control implementation or execution of a particular set of features. Further, the extraction circuit 100 can extract features for each packet at line rate and / or per flow for a large number of flows (e.g., one million flows, three million flows, eight million flows, etc.). Additionally or alternatively, the extraction circuit 100 can maintain a respective history for a group (e.g., batch) of packets for individual flows of the large number of flows. The extraction circuit 100 can consume in-band network telemetry device features to consolidate packets and features (e.g., flow features).
[0041] The extraction circuit 100 can be embedded as an offload processing circuit on any suitable network device associated with the integrated circuit device 12 (e.g., a network interface card (NIC), a switch, an infrastructure processing unit (IPU), a CPU, a system on a chip (SoC), etc.). In one embodiment, the extraction circuit 100 can be integrated as an offload processing circuit with a packet processing pipeline (e.g., a packet processing system) included in the data plane circuit 108 of the integrated circuit device 12. For example, the packet processing system can include a monolithic system that includes the extraction circuit 100. Further, the monolithic system can be connected with off-chip components (such as a memory, a CPU, etc.) through a high-speed interconnect.
[0042] In another embodiment, the extraction circuit 100 can operate as a standalone circuit connected to any suitable component of the integrated circuit device 12 via a high-speed interconnect. In another embodiment, the extraction circuit 100 can be integrated as hardware on the integrated circuit device 12 (e.g., an FPGA-based system) and connected with programmable logic fabric and any other suitable component of the integrated circuit device 12 using any suitable interconnect. In another embodiment, the extraction circuit 100 can be integrated as software on the integrated circuit device 12 and implementable on programmable logic fabric of the integrated circuit device 12.
[0043] The extraction circuit 100 can enable reuse of hardware resources to map features of each packet of each workload. That is, various workloads can share hardware resources to independently map features of each packet. The packet processing system can dynamically select features of each packet depending on the workload. For packet processing, the extraction circuit 100 can directly receive packets from the replication engine 106 and implement packet parsing, flow classification, feature extraction, pre-processing, and / or normalization of the packets.
[0044] Alternatively, the extraction circuit 100 can work as a co-processing system with a packet processing pipeline that performs parsing and / or classification and pipes data flows to the extraction circuit 100 to implement feature extraction, pre-processing, and normalization of the data. Accordingly, at least some packet processing pipeline stages can be reused to perform feature extraction, which can result in reduced computational resources for processing and / or extraction. Additional details regarding the extraction circuit 100 are described below. Figures 4 to 8
[0045] The inference circuit 102 can be integrated with a suitable network device of the integrated circuit device 12 and can enable mapping of network algorithms to AI / ML models via the AI / ML inference engine 116. To provide a solution, the inference circuit 102 can fetch (e.g., call, request) one or more models from the data plane circuit 108. In practice, the inference circuit 102 can host one or more types of models and / or one or more instances of a single model. Programmability and / or reusability of the inference circuit 102 can be based on the type of models hosted and / or the multiple instances of a single model. Further, the inference circuit 102 can dynamically update one or more weights and architectures of the AI / ML models based on varying network conditions via the AI / ML inference engine 116.
[0046] The inference circuit 102 can receive a set of features and packet classification data from the extraction circuit 100 and / or the data plane circuit 108 and perform inference using the set of features and packet classification data. Further, the inference circuit 102 can perform load balancing across multiple instances of the AI / ML model based on connection or flow identification (ID). The inference circuit 102 can link the AI / ML model for a particular workload or run the AI / ML model in parallel. Additionally, the inference circuit 102 can enable multiple workloads to concurrently use a set of AI / ML models via orchestration (e.g., computation) from the data plane circuit 108.
[0047] The integrated circuit device 12 can receive incoming one or more packets at the interface 110. In one embodiment, the integrated circuit device 12 can provide the incoming packets to the replication engine 106 for replication and / or routing. In effect, the replication engine 106 can receive the packets and make a copy (e.g., replicate) of each packet and its associated metadata. Further, the replication engine 106 can provide the copy of each packet to the extraction circuit 100.
[0048] In another embodiment, the integrated circuit device 12 can receive incoming packets at the interface 110 and provide the packets to the data plane circuit 108. The data plane circuit 108 can support dynamic AI / ML models at line rate. The data plane circuit 108 can employ the extraction circuit 100 and / or the inference circuit 102 in tandem with the data plane circuit or integrated with the packet processing pipeline of the data plane circuit 108. For example, the integrated circuit device 12 can employ a shim layer to enable the data plane circuit 108 to communicate with the extraction circuit 100 and / or the inference circuit 102. In some embodiments, the shim layer can also collect result data from the inference circuit 102 and write the result data to a pre-pended header (PPH). Additional details regarding the data flow via the interface 110, the replication engine 106, the data plane circuit 108, the extraction circuit 100, the inference circuit 102, and / or the CPU 104 are described below with respect to FIGS. 2-5. Figure 5
[0049] Figure 5 is performed by the integrated circuit device 12 of Figure 1 a first processing flow or a second processing flow performed by example circuitry employed by the integrated circuit device 12 in accordance with embodiments of the present disclosure. It should be noted that the integrated circuit device 12 can perform the first processing flow or the second processing flow. Further, for the first processing flow, the extraction circuit 100 and the inference circuit 102 can be implemented as independent circuitry independent of the data plane circuit 108. Further, for the second processing flow, the extraction circuit 100 and the inference circuit 102 can be implemented in tandem with the data plane circuit 108 interconnected (e.g., high speed interconnect) or integrated with the data plane circuit 108 to enable a hybrid system (where the data plane circuit 108 (e.g., packet processing pipeline) employs the extraction circuit 100 and the inference circuit 102 as co-processors) and / or send packets with or without replication. Further, it should be noted that the integrated circuit device 12 can instruct the example circuitry described herein to perform operations via any suitable circuitry, such as the CPU 104.
[0050] For the first processing flow, at process 120, interface 110 can receive one or more packets (e.g., a workload). At process 122, interface 110 can provide the packets to replication engine 106. At process 124, replication engine 106 can receive the packets, replicate the packets and their associated metadata (e.g., packet length, timestamp, inter-packet arrival time, average size, etc.), and provide the replicated packets and their associated metadata to extraction circuit 100. Extraction circuit 100 can include feature extraction circuit 112, which can operate as a standalone circuit. Extraction circuit 100 can employ feature extraction circuit 112 to extract (e.g., generate) a set of features (e.g., raw features) from the packets. For example, the set of features can include a size of each packet, a direction of travel of each packet, a timestamp associated with each packet, a maximum and minimum inter-packet arrival time, etc. Additionally, extraction circuit 100 can employ FVT circuit 114 to normalize and / or transform the set of features, such as performing any suitable pre-processing on the packets for use by inference circuit 102.
[0051] At process 126, extraction circuit 100 can provide the set of features to inference circuit 102. Inference circuit 102 can employ AI / ML inference engine 116 to apply one or more AI / ML models to the set of features to make predictions and / or decisions based on the data associated with the set of features. In this way, inference circuit 102 can perform inference by analyzing and / or predicting network behavior, detecting anomalies, optimizing performance, etc. using AI / ML models based on the set of features to generate a result that can provide a solution to a network problem. It should be noted that, in some embodiments, AI / ML inference engine 116 can be implemented in programmable logic fabric of integrated circuit device 12. At process 128, the result generated by inference circuit 102 (e.g., a solution to a network problem) can be provided to data plane circuit 108. In some embodiments, based on a configuration of inference circuit 102, the result can be provided to CPU 104 in addition to or instead of being provided to data plane circuit 108.
[0052] As described herein, for the second processing flow, extraction circuit 100, inference circuit 102, and data plane circuit 108 can be interconnected to enable a hybrid system. Thus, at process 130, interface 110 can provide packets to data plane circuit 108. Data plane circuit 108 can enable classification of the packets (e.g., by determining a flow identification (ID) and / or other parameters of the packets), determination or selection of a type of features to extract, generation of one or more hints to be provided to extraction circuit 100 based on the determined type of features to extract, determination of a type of AI / ML model to run, etc.
[0053] At process 132, the data plane circuit 108 can provide the hint to the extraction circuit 100. In this way, the extraction circuit 100 can extract a set of features via the feature extraction circuit 112 and perform normalization and / or transformation via the FVT circuit 114 based on the hint. At process 134, after extracting the set of features, the extraction circuit 100 can provide the set of features to the data plane circuit 108. Further, the data plane circuit 108 can enable selection of a list of AI / ML models to be run for inference based on a quality of the set of features or features included therein. At process 136, the data plane circuit 108 can provide the set of features and the list of AI / ML models to the inference circuit 102. The inference circuit 102 can employ the AI / ML inference engine 116 to apply the list of selected AI / ML models to the set of features to generate a result. At process 138, the inference circuit 102 can provide the result to the data plane circuit 108. Thus, the system architecture of the integrated circuit device 12 described herein can enable generation of high quality features, normalization of features, and inference to perform per-packet or per-flow processing to solve network problems.
[0054] With the above in mind, Figure 6 is an integrated circuit device 12 according to embodiments of the present disclosure Figure 1 packet processing pipeline 150 of the integrated circuit device 12. That is, the packet processing pipeline 150 can be used for feature extraction. For example, a packet processor (P4) program independent of a programming protocol can enable definition of network packet processing behavior by implementing a first program 152 (e.g., a first tenant) associated with the packet processing pipeline 150 and a second program 154 (e.g., a second tenant) associated with the extraction circuit 100. Thus, the first program 152 can enable packet processing and the second program 154 can enable feature extraction. Further, the first program 152 and the second program 154 can be loaded onto the integrated circuit device 12 during compilation of the integrated circuit device 12 (e.g., alternatively, at runtime). A control plane interface (e.g., a management interface of the data plane circuit 108) can enable fusion of the first program 152 and the second program 154 into a single packet processing pipeline 150. In effect, the first program 152 and the second program 154 can be fused to work using defined headers (e.g., a uniformly defined common header) and satisfy hardware constraints.
[0055] The first program 152 and / or the second program 154 can enable analysis of a header of a packet received by the packet processing pipeline 150 to determine the liveliness of the header. That is, the first program 152 and / or the second program 154 can determine whether a particular header in the header is considered to be live (e.g., active or relevant to a current processing stage of the packet processing pipeline 150). In this way, the first program 152 and / or the second program 154 can enable intelligent placement of the extraction circuit 100 within the packet processing pipeline 150. As a result, utilization of existing hardware resources can be improved.
[0056] The packet processing pipeline 150 (e.g., the first program 152) can include a plurality of stages, which can include a parser circuit 156, a local termination circuit 158, a classification circuit 160, an access control list (ACL) or firewall circuit 162, a forwarding circuit 164, a routing circuit 166 (e.g., next hop tracking (NEXTHOP) protocol, address resolution protocol (ARP), media access control (MAC) protocol), and a de-parser circuit 168. The parser circuit 156 can be a programmable parser programmed to identify a plurality of headers present in a packet (or each packet of a plurality of packets) received by the integrated circuit device 12. Further, the plurality of identified headers can be combined to provide data indicative of the headers and header offsets present in the packet. The local termination circuit 158 can determine whether a packet terminates at a certain level (e.g., level 2 (L2) or level 3 (L3)). In effect, the local termination circuit 158 can determine which layer last processes the packet before the packet is delivered to its destination or discarded.
[0057] The classification circuit 160 can determine a flow identification (ID) of a flow (e.g., a five-tuple flow) utilizing the features to be extracted and / or a flow classification table. The flow classification table can be associated with a feature extraction batch or a timeout parameter to enable feature extraction per flow, per packet, per a plurality of packets, or per a timeout. In this way, the classification circuit 160 can enable a hybrid system in which certain flows of a plurality of packets are feature extracted while other packets wait for a timeout to occur.
[0058] The ACL or firewall circuit 162 can include a set of rules to permit or deny network traffic based on various criteria (e.g., IP address, protocol type, port, etc.). The forwarding circuit 164 can enable a network device to decide how to forward a packet. For example, the forwarding circuit 164 can be associated with a forwarding information base (FIB) table, which can enable mapping of a destination address to an outgoing interface and a next hop address. The routing circuit 166 can determine a destination of a packet. Further, the de-parser circuit 168 can reconstruct data associated with a packet into its original or designated format for transmission or storage.
[0059] Extraction circuit 100 (e.g., second program 154) can include header build circuit 170 and feature extraction circuit 112. Extraction circuit 100 can enable extraction of packet features and accumulation in local memory of integrated circuit device 12. Header build circuit 170 can determine whether a current packet can initiate feature extraction and / or build of a feature extraction header. It should be noted that the example circuits described herein are illustrative only, and packet processing pipeline 150 and extraction circuit 100 can include any other suitable circuits for performing packet processing and / or extraction. Additional details regarding operation of extraction circuit 100 collocated with packet processing pipeline 150 in data plane circuit 108 will be described below with respect to Figure 6 Figure 7
[0060] Figure 7 is a flowchart of an example method 190 for feature extraction via data plane circuit 108 according to embodiments of the present disclosure. While method 190 is described using steps in a particular order, it should be understood that the present disclosure contemplates that the described steps can be performed in a different order than shown and that certain described steps can be skipped or not performed.
[0061] At block 192, integrated circuit device 12 can receive one or more packets (e.g., workload, packet stream) at data plane circuit 108. Integrated circuit device 12 can then determine whether the packet is to be forwarded to another device (e.g., using local termination circuit 158). If the packet is to be used / consumed locally in integrated circuit device 12, at block 194, integrated circuit device 12 can determine a flow ID and feature extraction parameters associated with the packet (e.g., using classification circuit 160). For example, the feature extraction parameters can include types of features to be extracted from the packet, such as header information (e.g., source and / or destination Internet Protocol (IP) addresses or ports, protocol, etc.), timing information (e.g., time of arrival, time to live, etc.), flow size (e.g., number of packets or bytes), flow duration (e.g., time between first and last packet of the flow), inter-arrival time (IAT) between packets in the flow (e.g., time between each packet in the flow), protocol information, and / or any other suitable features. Further, integrated circuit device 12 can determine whether the packet can initiate feature extraction, and if it is determined that the packet can initiate feature extraction, integrated circuit device 12 can proceed to block 196.
[0062] At block 196, the integrated circuit device 12 can construct a feature extraction header based on the feature extraction parameters (e.g., using the header construction circuit 170 of the extraction circuit 100). Indeed, the integrated circuit device 12 can determine which features are applicable to the packet. Moreover, in some embodiments, the integrated circuit device 12 can organize and / or prepare the data associated with the packet in a manner that facilitates extraction of information associated with the feature extraction parameters. For example, the integrated circuit device 12 can identify header information (e.g., data) related to the extraction from the packet, aggregate the header information, and so on.
[0063] At block 198, the integrated circuit device 12 can extract the packet number and the packet size based on the feature extraction header. Moreover, at block 200, the integrated circuit device 12 can extract the IAT based on the feature extraction header. It should be noted that blocks 198 and 200 are illustrative only, and any other suitable features can be extracted based on the feature extraction parameters associated with the respective packet. Moreover, it should be noted that any of the features can be extracted in any order.
[0064] Additionally, fusing feature extraction with packet processing can enable splitting multiple read-modify-write operations among several stages of the packet processing pipeline 150 by creating multiple match-action tables. Indeed, each stage of the packet processing pipeline can be associated with a respective match-action table. For example, a first match-action table associated with feature extraction of the first header, the second header, the third header, and the fourth header can be placed at a first stage (e.g., at the local termination circuit 158). As another example, based on availability of a parallel lookup budget, a second match-action table associated with feature extraction of the second header and the fourth header can be placed at a second stage (e.g., at the classification circuit 160) instead of at the first stage. As another example, a third match-action table can be placed at a third stage (e.g., at the ACL or firewall circuit 162) associated with feature extraction of the fifth header, the sixth header, and the seventh header.
[0065] Accordingly, throughput within the system architecture of the integrated circuit device 12 is improved. The match-action tables can enable reducing accelerator functional unit (AFU) resource utilization by reducing and / or limiting data movement. Indeed, a compiler of the integrated circuit device 12 can perform a header and / or metadata liveliness analysis based on the match-action tables for reusability and minimization of metadata. The metadata liveliness analysis can also be applied to metadata received from the first program and / or the second program. Accordingly, metadata can be reused among multiple programs.
[0066] The integrated circuit device 12 can include on-chip (e.g., fabric memory) and / or off-chip memory, which can be used to vary the latency, size, and / or bandwidth of different applications. However, on-chip memory can be limited, while off-chip memory can incur latency. Further, feature extraction can be stateful and involve storing historical data associated with a large number of flows (e.g., millions of flows). Thus, the match-action table and / or the flow register array can be placed based on total memory utilization. In one embodiment, a user can adjust the placement of the match-action table and / or the flow register array by indicating the placement within a program (e.g., a P4 program) implemented on the integrated circuit device 12. In another embodiment, a compiler of the integrated circuit device 12 can automatically determine the placement of the match-action table and / or the flow register array. Thus, if memory usage exceeds a memory budget of the integrated circuit device 12, the flow register array can be placed in off-chip memory, while on-chip memory serves as a cache.
[0067] With the above in mind, Figure 8 is an example illustration of the packet processing pipeline 150 and the extraction circuit 100 employing on-chip memory and off-chip memory according to embodiments of the present disclosure. The integrated circuit device 12 can include on-chip memory 230 (e.g., 230A, 230B, 230C, 230D, 230E) and can utilize off-chip memory 232 (e.g., 232A and 232B). For example, the on-chip memory 230 can include block random access memory (BRAM) and the off-chip memory 232 can include double data rate (DDR) memory.
[0068] The on-chip memory 230 can enable storage of smaller match-action table entries. The off-chip memory 232 can enable storage of larger match-action table entries (e.g., ACLs or firewall tables). Additionally or alternatively, the flow register array can be placed on the on-chip memory 230. However, when storage is insufficient, the flow register array can be placed on the off-chip memory 232 and the on-chip memory 230 can serve as a cache. Thus, a hybrid memory system including the integrated circuit device 12 utilizing the on-chip memory 230 and / or the off-chip memory 232 can enable additional storage of historical data associated with feature extraction of a large number of flows.
[0069] Sometimes, when the CPU 104 adds new flow entries in the data plane circuit 108, latency can be involved, which can result in loss of feature data. Thus, it can be desirable to employ local memory to store flow data by partitioning the local memory (e.g., the on-chip memory 230). With the above in mind, Figure 9is an example illustration of a memory partition 258 of the data plane circuit 108. For example, the memory partition 258 can include CPU-managed memory 264 and / or scratch memory 266. Further, the memory partition 258 can be configured by the control plane circuit 262 over the interface 260.
[0070] The CPU-managed memory 264 can be managed by the CPU 104. For example, the CPU 104 can perform management of flow installation and / or mapping of flow tuples to flow IDs. Additionally or alternatively, the packet processing pipeline 150 can perform flow table lookups via a flow lookup table 268 and / or perform flow tuple hash functions (e.g., stateless) via a hash and digest circuit 270. The flow lookup table 268 can enable allocation of a unique flow ID for each flow in the data plane circuit 108. When the packet processing pipeline 150 encounters a large number of flows, the packet processing pipeline 150 can employ the scratch memory 266 to map data to flow IDs. Thus, the scratch memory 266 can enable storage of flow data for any number of new flows until a flow entry is added.
[0071] With respect to Figure 4 The integrated circuit device 12 described can be included in a data processing system such as a computer or a similar device that includes memory and / or storage. The memory and / or storage can include the memory partition 258 described. The data processing system can include the CPU 104 and / or the control plane circuit 262 described. Figure 10Components of a data processing system 300 (shown in FIG. 1 ). Data processing system 300 may include integrated circuit device 12, host processor 302 (e.g., CPU 104), memory and / or storage circuitry 304, and network interface 306. Data processing system 300 may include more or fewer components (e.g., an electronic display, user interface structures, application-specific integrated circuits (ASICs)). Integrated circuit device 12 can be efficiently programmed to listen for requests from the host and pre-fill caches with data based on the requests to reduce memory access times. In other words, integrated circuit device 12 can accelerate functions of a host (such as host processor 302). Host processor 302 may include any of the aforementioned processors that can manage data processing requests for data processing system 300 (e.g., to perform encryption, decryption, machine learning, video processing, speech recognition, image recognition, data compression, database search ranking, bioinformatics, network security pattern recognition, spatial navigation, cryptocurrency operations, etc.). Memory and / or storage circuitry 304 may include random access memory (RAM), read-only memory (ROM), one or more hard drives, flash memory, etc. Memory and / or storage circuit 304 can store data to be processed by data processing system 300. In some cases, memory and / or storage circuit 304 can also store configuration programs (e.g., bitstreams, mapping functions) for programming FPGA 70. Network interface 306 can allow data processing system 300 to communicate with other electronic devices. Data processing system 300 can include several different packages, or can be contained in a single package on a single package substrate. For example, the components of data processing system 300 can be located on several different packages at one location (e.g., a data center) or at multiple locations. For example, the components of data processing system 300 can be located in separate geographic locations or regions, such as cities, states, or countries.
[0072] Data processing system 300 may be part of a data center that processes a variety of different requests. For example, data processing system 300 may receive data processing requests via network interface 306 to perform encryption, decryption, machine learning, video processing, speech recognition, image recognition, data compression, database search ranking, bioinformatics, network security pattern recognition, spatial navigation, digital signal processing, or other specialized tasks.
[0073] Although the embodiments described in this disclosure may be susceptible to various modifications and alternative forms, specific embodiments have been shown by way of example in the drawings and are described in detail herein. However, it should be understood that this disclosure is not intended to be limited to the particular forms disclosed. This disclosure is intended to cover all modifications, equivalents, and alternatives that fall within the spirit and scope of this disclosure as defined by the appended claims.
[0074] The technology presented and claimed herein was made in view of the real-world realities and particular examples that have been explicitly improved by the technology, and thus is not abstract, intangible or purely theoretical. Further, if any of the claims appended to this specification contain one or more elements designated as “means for [performing] [functioning]…” or “step for [performing] [functioning]…”, it is intended that such elements be interpreted under 35 U.S.C. 112(f). However, for any claim containing elements designated in any other manner, it is intended that such elements not be interpreted under 35 U.S.C. 112(f).
[0075] Example Embodiments
[0076] Example Embodiment 1. Integrated circuitry comprising a data plane circuit comprising a packet processing pipeline and an extraction circuit, wherein the extraction circuit is configurable to extract one or more features of a packet via the packet processing pipeline. The integrated circuitry further comprises an inference circuit configurable to perform machine learning inference based on the one or more features.
[0077] Example Embodiment 2. The integrated circuitry of example embodiment 1, wherein the packet processing pipeline is configurable to determine a flow identification and one or more extraction parameters.
[0078] Example Embodiment 3. The integrated circuitry of example embodiment 2, wherein the extraction circuit is configurable to construct one or more extraction headers based on the one or more extraction parameters.
[0079] Example Embodiment 4. The integrated circuitry of example embodiment 3, wherein the extraction circuit is configurable to extract the one or more features based on the one or more extraction headers.
[0080] Example Embodiment 5. The integrated circuitry of example embodiment 1, wherein the inference circuit is configurable to apply one or more artificial intelligence (AI) machine learning (ML) models to the one or more features to generate a result.
[0081] Example Embodiment 6. The integrated circuitry of example embodiment 1, wherein the extraction circuit comprises feature extraction circuitry to extract the one or more features and feature vector transformation circuitry to perform pre-processing or normalization on the one or more features.
[0082] Example Embodiment 7. The integrated circuitry of example embodiment 1, comprising an interface configurable to receive a packet and provide the packet to the data plane circuit or the extraction circuit.
[0083] Example Embodiment 8. The integrated circuitry of example embodiment 1, comprising a hybrid memory system, wherein the hybrid memory system comprises on-chip memory and off-chip memory.
[0084] Example Embodiment 9. The integrated circuitry of example embodiment 8, wherein the on-chip memory, the off-chip memory, or both are configurable to store at least one of one or more match-action table entries and one or more flow register arrays.
[0085] Example Embodiment 10. The integrated circuitry of example embodiment 1, comprising a local memory configurable to be partitioned into central processing unit (CPU) managed memory and staging memory.
[0086] Example Embodiment 11. The integrated circuitry of example embodiment 1, wherein the packet processing pipeline comprises a first stage and a second stage, wherein the first stage is associated with a first match-action table, and wherein the second stage is associated with a second match-action table.
[0087] Example Embodiment 12. The integrated circuitry of example embodiment 1, wherein the one or more features comprise header information of a packet, timing information of a packet, flow size of a packet, flow duration of a packet, inter-packet arrival time of a packet, protocol information of a packet, or any combination thereof.
[0088] Example Embodiment 13. A data plane circuit comprising: a packet processing pipeline comprising one or more stages, wherein the packet processing pipeline is configurable to determine one or more feature extraction parameters; an extraction circuit configurable to extract one or more features of one or more packets via a plurality of stages of the one or more stages based on the one or more feature extraction parameters.
[0089] Example Embodiment 14. The data plane circuit of example embodiment 13, wherein the extraction circuit comprises a feature extraction circuit and a feature vector transformation circuit.
[0090] Example Embodiment 15. The data plane circuit of example embodiment 14, wherein the feature vector transformation circuit is configurable to derive artificial intelligence (AI) machine learning (ML) features for the data plane circuit, an inference circuit, or both.
[0091] Example Embodiment 16. The data plane circuit of example embodiment 14, wherein the feature vector transformation circuit is configurable to perform floating point analysis to determine a minimum standard deviation, a maximum standard deviation, a minimum normalization, a maximum normalization, a mean square sum, or any combination thereof.
[0092] Example embodiment 17. A method for performing packet processing, comprising: receiving a plurality of packets using an extraction circuit; extracting one or more features of the plurality of packets using the extraction circuit and a packet processing pipeline; transmitting the one or more features to an inference circuit using the extraction circuit; and generating a result based on the one or more features using the inference circuit.
[0093] Example Embodiment 18. The method of Example Embodiment 17, wherein generating a result based on the one or more features comprises applying one or more models to the one or more features using inference circuitry.
[0094] Example Embodiment 19. The method of Example Embodiment 17, comprising: normalizing data associated with the plurality of packets using an extraction circuit.
[0095] Example Embodiment 20. The method of Example Embodiment 17, comprising: using the inference circuit to reason about the plurality of packets per packet or per flow of the plurality of packets.
Claims
1. An integrated circuit system for implementing a machine learning network algorithm in a data plane, comprising: data plane circuitry comprising a packet processing pipeline and extraction circuitry, wherein the extraction circuitry is configurable to extract one or more features of a packet via the packet processing pipeline; and Inference circuitry can be configured to perform machine learning inference based on the one or more features.
2. The integrated circuit system according to claim 1, wherein: The packet processing pipeline can be configured to determine a flow identification and one or more extraction parameters.
3. The integrated circuit system according to claim 2, wherein: The extraction circuitry can be configured to construct one or more extraction headers based on the one or more extraction parameters.
4. The integrated circuit system according to claim 3, wherein: The extraction circuitry can be configured to extract the one or more features based on the one or more extraction headers.
5. The integrated circuit system according to claim 1, wherein: The inference circuitry can be configured to apply one or more artificial intelligence (AI) machine learning (ML) models to the one or more features to generate a result.
6. The integrated circuit system according to claim 1, wherein: The extraction circuit includes a feature extraction circuit for extracting the one or more features and a feature vector transformation circuit for performing preprocessing or normalization on the one or more features.
7. The integrated circuit system of claim 1, comprising an interface configurable to receive the packet and provide the packet to the data plane circuitry or the extraction circuitry.
8. The integrated circuit system of claim 1, comprising a hybrid memory system, wherein the hybrid memory system comprises on-chip memory and off-chip memory.
9. The integrated circuit system according to claim 8, wherein: The on-chip memory, the off-chip memory, or both can be configured to store at least one of one or more match action table entries and one or more flow register arrays.
10. The integrated circuit system according to any one of claims 1 to 9, comprising a local memory configurable to be partitioned into a memory managed by a central processing unit (CPU) and a scratch memory.
11. The integrated circuit system according to any one of claims 1 to 9, wherein: The packet processing pipeline includes a first stage and a second stage, wherein the first stage is associated with a first match-action table, and wherein the second stage is associated with a second match-action table.
12. The integrated circuit system according to any one of claims 1 to 9, wherein: The one or more characteristics include header information of the packet, timing information of the packet, flow size of the packet, flow duration of the packet, inter-packet arrival time of the packet, protocol information of the packet, or any combination thereof.
13. A data plane circuit for implementing a machine learning network algorithm in a data plane, the data plane circuit comprising: a packet processing pipeline comprising one or more stages, wherein the packet processing pipeline is configurable to determine one or more feature extraction parameters; and Extraction circuitry is configurable to extract one or more features of the one or more packets via a plurality of the one or more stages based on the one or more feature extraction parameters.
14. The data plane circuit according to claim 13, wherein: The extraction circuit includes a feature extraction circuit and a feature vector conversion circuit.
15. The data plane circuit according to claim 14, wherein: The feature vector transformation circuitry can be configured to derive artificial intelligence (AI) machine learning (ML) features for the data plane circuitry, the inference circuitry, or both.
16. The data plane circuit of claim 14, wherein: The feature vector conversion circuit can be configured to perform floating point analysis to determine minimum standard deviation, maximum standard deviation, minimum normalization, maximum normalization, mean sum of squares, or any combination thereof.
17. A method for performing packet processing, comprising: receiving a plurality of packets using an extraction circuit; extracting one or more features of the plurality of packets using the extraction circuitry and the packet processing pipeline; transmitting the one or more features to inference circuitry using the extraction circuitry; as well as A result is generated based on the one or more features using the inference circuitry.
18. The method according to claim 17, wherein Generating the result based on the one or more features includes applying one or more models to the one or more features using the inference circuitry.
19. The method according to claim 17, comprising: The data associated with the plurality of packets is normalized using the extraction circuitry.
20. The method according to any one of claims 17 to 19, comprising: Reasoning is performed on the plurality of packets using the inference circuitry on a per-packet or per-flow basis.
Citation Information
Patent Citations
Multi-purpose interface for configuration data and user fabric data
US10833679B2
Programmable circuit having multiple sectors
US20160049941A1