Feature extraction for in-line network analysis

By using feature extraction and reconfigurable neural network circuits to process packet streams, the problem of non-real-time network analysis in existing technologies is solved, enabling rapid detection of network anomalies and intrusions and improving the timeliness of network protection.

CN116506314BActive Publication Date: 2025-10-24AVAGO TECHNOLOGIES INTERNATIONAL SALES PTE LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211657132.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-01-27
Filing Date
2022-12-22
Publication Date
2025-10-24
Estimated Expiration
2042-12-22

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve real-time detection when performing network analysis based on the end of a large number of packets, resulting in untimely network intrusion detection and an inability to provide timely and adequate network protection.

Method used

By employing feature extraction circuits and reconfigurable neural network circuits, and receiving raw packet streams, time statistics and feature data processing are performed. Combined with hash tables and multiplexers, rapid identification of packet attributes and stream attributes and neural network calculations are achieved to determine network characteristics.

Benefits of technology

It enables real-time or near-line-speed network analysis, allowing for rapid detection of network anomalies and intrusions, thus improving the timeliness and effectiveness of network protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116506314B_ABST
    Figure CN116506314B_ABST
Patent Text Reader

Abstract

An apparatus and method for performing feature extraction for in-line network analysis is described herein. In one aspect, the apparatus includes feature extraction circuitry, input processing circuitry, and reconfigurable neural network circuitry. In one aspect, the feature extraction circuitry receives a raw packet stream and obtains time statistics for a flow from first packet attributes or first flow attributes of the raw packet stream. In one aspect, the feature extraction circuitry generates feature data including one or more statistical features based on the time statistics for the flow. In one aspect, the input processing circuitry scales the feature data to generate adjusted feature data. In one aspect, the reconfigurable neural network circuitry performs computations corresponding to a neural network on the adjusted feature data to determine a predicted network characteristic.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates generally to systems and methods for performing network analysis. BACKGROUND

[0002] Various approaches have been proposed for implementing deep learning techniques based on neural networks to determine network characteristics or network conditions. For example, neural networks can be applied to application classification, anomaly detection, congestion handling, and intrusion detection (DDoS detection, etc.) based on flow statistics. In one implementation, a large number of packets at the end of a flow (e.g., more than 10,000 packets) can be applied to a neural network to determine network characteristics or network conditions. However, this implementation based on the end of a flow with a large number of packets can be inappropriate for real-time analysis. For example, detecting network intrusion based on a large number of packets can be too late and can not allow timely adequate network protection. SUMMARY

[0003] In one aspect, the present application provides a network device comprising: a feature extraction circuit configured to receive a raw packet stream, obtain time statistics of a flow from first packet attributes or first flow attributes of the raw packet stream, and generate feature data including one or more statistical features based on the time statistics of the flow; an input processing circuit configured to scale the feature data to generate adjusted feature data; and a reconfigurable neural network circuit configured to perform computations corresponding to a neural network on the adjusted feature data to determine predicted network characteristics.

[0004] In another aspect, the present application provides a network device comprising: a control circuit configured to determine an index based on first packet attributes or first flow attributes of a raw packet stream; a multiplexer configured to select second packet attributes or second flow attributes of the raw packet stream corresponding to a hash key according to the index; a hash table circuit configured to identify a flow from the hash key; and one or more computation circuits configured to: obtain time statistics of the flow, and generate feature data including one or more statistical features based on the time statistics of the flow according to configuration settings corresponding to the first packet attributes or the first flow attributes of the raw packet stream determined by the control circuit.

[0005] In another aspect, the present application provides a method comprising: determining, by a control circuit, an index based on first packet properties or first flow properties of an original packet stream; selecting, by a multiplexer, second packet properties or second flow properties of the original packet stream corresponding to a hash key according to the index; determining, by the control circuit, configuration settings corresponding to the first packet properties or the first flow properties of the original packet stream; identifying, by a hash table circuit, a flow according to the hash key; obtaining, by one or more computation circuits, time statistics of the flow; and generating, by the one or more computation circuits, feature data including one or more statistical features based on the time statistics of the flow according to the configuration settings. BRIEF DESCRIPTION OF DRAWINGS

[0006] Various objects, aspects, features, and advantages of the present disclosure will be better understood and appreciated with reference to the following detailed description taken in conjunction with the accompanying drawings, in which like reference characters identify corresponding elements throughout the drawings. In the figures, like reference numbers indicate similar or corresponding elements unless otherwise noted.

[0007] Figure 1A is a block diagram depicting a network environment including one or more access points in communication with one or more devices or stations, in accordance with some embodiments.

[0008] Figure 1B and 1C is a block diagram depicting a computing device that can be used in connection with the methods and systems described herein, in accordance with some embodiments.

[0009] Figure 2 illustrates an apparatus for performing in-line network analysis based on a reconfigurable neural network, in accordance with an embodiment.

[0010] Figure 3 illustrates a flow diagram showing a process for determining predicted network characteristics based on a reconfigurable neural network, in accordance with an embodiment.

[0011] Figure 4 illustrates a schematic diagram of a feature computation circuit, in accordance with an embodiment.

[0012] Figure 5 illustrates a flow diagram showing a process for adaptively generating feature data, in accordance with an embodiment.

[0013] Figure 6 illustrates a schematic diagram of an input processing circuit, in accordance with an embodiment.

[0014] Figure 7 illustrates a flow diagram showing a process for adaptively adjusting feature data to obtain adjusted feature data, in accordance with an embodiment.

[0015] Figure 8FIG. illustrates a schematic diagram of a reconfigurable neural network circuit, in accordance with an embodiment.

[0016] Figure 9 FIG. illustrates a flow diagram showing a process of adaptively performing computations of a neural network by a reconfigurable neural network circuit, in accordance with an embodiment.

[0017] Figure 10 FIG. illustrates a schematic diagram of an output processing circuit, in accordance with an embodiment.

[0018] Figure 11 FIG. illustrates a flow diagram showing a process of adaptively generating output data including predicted network characteristics, in accordance with an embodiment.

[0019] Details of various embodiments of the method and system are described in the drawings and in the following detailed description. DETAILED DESCRIPTION

[0020] For the purpose of reading the description of various embodiments of the present disclosure set forth below, the following description of the sections of the specification and their respective contents can be helpful:

[0021] - Section A describes network and computing environments that can be used in the practice of the embodiments described herein; and

[0022] - Section B describes embodiments of systems and methods for in-line network analysis.

[0023] A. Computing and Network Environment

[0024] Before discussing specific embodiments of the present solution, aspects of operating environments and related system components (e.g., hardware elements) in connection with the methods and systems described herein can be helpful. Referring to Figure 1A , an embodiment of a network environment is depicted. In brief overview, the network environment includes a wireless communication system that comprises one or more access points (APs) 106, one or more wireless communication devices 102, and network hardware components 192. For example, the wireless communication devices 102 can include a laptop computer 102, a tablet computer 102, a personal computer 102, and / or a cellular telephone device 102. Referring to Figure 1B and 1CDetails of embodiments of each wireless communication device 102 and / or AP 106 are described in greater detail. In one embodiment, the network environment can be an ad hoc network environment, an infrastructure wireless network environment, a sub-network environment, etc. The APs 106 can be operatively coupled to network hardware 192 via a local area network connection. The network hardware 192, which can include routers, gateways, switches, bridges, modems, system controllers, appliances, etc., can provide a local area network connection for the communication system. Each of the APs 106 can have an associated antenna or array of antennas to communicate with the wireless communication devices within its area. The wireless communication devices 102 can register with a particular AP 106 to receive service from the communication system (e.g., via a SU-MIMO or MU-MIMO configuration). Some of the wireless communication devices can communicate directly via an assigned channel and communication protocol for direct connection (e.g., point-to-point communication). Some of the wireless communication devices 102 can be mobile or relatively static with respect to the APs 106.

[0025] In some embodiments, the APs 106 include a device or module (including a combination of hardware and software) that allows the wireless communication devices 102 to connect to a wired network using wireless fidelity (WiFi) or other standards. The APs 106 can sometimes be referred to as wireless access points (WAPs). The APs 106 can be implemented (e.g., configured, designed, and / or built) to operate in a wireless local area network (WLAN). In some embodiments, the APs 106 can connect to a router (e.g., via a wired network) as a standalone device. In other embodiments, the APs 106 can be a component of a router. The APs 106 can provide access to a network for multiple devices. For example, the APs 106 can connect to a wired Ethernet connection and provide a wireless connection using radio frequency links to enable other devices 102 to utilize the wired connection. The APs 106 can be implemented to support standards for sending and receiving data using one or more radio frequencies. Those standards and frequencies used can be defined by the IEEE (e.g., the IEEE 802.11 standards). The APs 106 can be configured and / or used to support public Internet hotspots and / or extend the range of Wi-Fi signals on a network.

[0026] In some embodiments, the access points 106 can be used for (e.g., in-home or in-building) wireless networks (e.g., IEEE 802.11, Bluetooth, ZigBee, any other type of radio frequency-based network protocol, and / or variations thereof). Each of the wireless communication devices 102 can include a built-in radio and / or be coupled to a radio. Such wireless communication devices 102 and / or access points 106 can operate in accordance with the various aspects of the disclosure presented herein to improve performance, reduce cost and / or size, and / or enhance broadband applications. Each wireless communication device 102 can have the ability to act as a client node seeking access to resources (e.g., data and connectivity to networked nodes (e.g., servers)) via one or more access points 106.

[0027] A network connection can include any type and / or form of network and can include any one or combination of the following: a point-to-point network, a broadcast network, a telecommunication network, a data network and a computer network. The topology of the network can be a bus, star, or ring network topology. The network can have any such network topology as known to those having ordinary skill in the art capable of supporting the operations described herein. In some embodiments, different types of data can be transmitted via different protocols. In other embodiments, the same types of data can be transmitted via different protocols.

[0028] The communication devices 102 and access points 106 can be deployed as any type and form of computing device (e.g., a computer, a network appliance or an appliance capable of communicating on any type and form of network and performing the operations described herein) and / or executing on. Figure 1B and 1C A block diagram of a computing device 100 suitable for implementing embodiments of the wireless communication devices 102 or APs 106 is depicted. As shown in Figure 1B and 1C As shown in Figure 1B As shown in Figure 1C As shown in

[0029] The central processing unit 121 is any logic circuitry that responds to and processes instructions fetched from the main memory unit 122. In many implementations, the central processing unit 121 is provided by a microprocessor unit, such as: a microprocessor unit manufactured by the Intel Corporation of Santa Clara, California; a microprocessor unit manufactured by the International Business Machines of White Plains, New York; or a microprocessor unit manufactured by the Advanced Micro Devices of Sunnyvale, California. The computing device 100 can be based on any of these processors or any other processor capable of operating as described herein.

[0030] The main memory unit 122 can be one or more memory chips capable of storing data and allowing direct access to any location by the microprocessor 121, such as any type or variation of static random access memory (SRAM), dynamic random access memory (DRAM), ferroelectric RAM (FRAM), NAND flash memory, NOR flash memory, and solid state drive (SSD). The main memory 122 can be based on any of the memory chips described above or any other available memory chip capable of operating as described herein. In Figure 1B In the embodiment shown in FIG. 1, the processor 121 communicates with the main memory 122 via a system bus 150 (described in greater detail below). Figure 1C An embodiment of the computing device 100 is depicted in which the processor communicates directly with the main memory 122 via a memory port 103. For instance, in Figure 1C In the embodiment shown in FIG. 1, the main memory 122 can be a DRDRAM.

[0031] Figure 1C An embodiment is depicted in which the main processor 121 communicates with the cache memory 140 via an auxiliary bus (sometimes called a backside bus). In other embodiments, the main processor 121 communicates with the cache memory 140 using the system bus 150. The cache memory 140 is typically larger than the cache memory 130 and is a lower performance, lower cost memory, such as one or more of: SRAM, BSRAM, or EDRAM. Figure 1CIn the embodiment shown in the figure, the processor 121 communicates with various I / O devices 130 via a local system bus 150. Various buses can be used to connect the central processing unit 121 to any of the I / O devices 130, for example, a VESA VL bus, an ISA bus, an EISA bus, a MicroChannel Architecture (MCA) bus, a PCI bus, a PCI-X bus, a QuickPCI bus, or a NuBus. For embodiments in which the I / O device is a video display 124, the processor 121 can use an advanced graphics port (AGP) to communicate with the display 124. Figure 1C An embodiment of the computer 100 is depicted in which the host processor 121 can communicate directly with I / O device 130b via HYPERTRANSPORT, RAPIDIO, or INFINIBAND communication technology. Figure 1C An embodiment is also depicted in which a local bus and a direct communication are mixed: the processor 121 communicates with I / O device 130a using a local interconnect bus while communicating with I / O device 130b directly.

[0032] A wide variety of I / O devices 130a-130n can be present in the computing device 100. Input devices include keyboards, mice, trackpads, trackballs, microphones, dials, touch pads, touch screens, and drawing tablets. Output devices include video displays, speakers, inkjet printers, laser printers, projection devices, and dye-sublimation printers. An I / O device can be a storage device such as a disk drive, floppy disk drive, hard disk drive, or optical drive. The I / O devices can be controlled by an I / O controller 123 as shown in the figure. The I / O controller can control one or more I / O devices such as a keyboard 126 and a pointing device 127, e.g., a mouse or optical pen. In addition, an I / O device can be a storage device such as a disk drive, floppy disk drive, hard disk drive, or optical drive. The computing device 100 can provide Figure 1B

[0033] Again, referring to the figure Figure 1B ​The computing device 100 can support any suitable installation device 116, such as a disk drive, a CD-ROM drive, a CD-R / RW drive, a DVD- ROM drive, a flash memory device, various formats of tape drives, a USB device, a hard drive, a network interface, or any other device suitable for installing software and programs. The computing device 100 can further include a storage device, such as one or more hard disk drives or redundant arrays of independent disks, for storing an operating system and other related software, and for storing application software programs, such as any program or software 120 for implementing (e.g., configured and / or designed for) the systems and methods described herein. Optionally, any of the installation devices 116 can also be used as the storage device. Additionally, the operating system and the software can be run from the removable media.

[0034] Moreover, the computing device 100 can include a network interface 118 to interface to a network 104 through a variety of connections including, but not limited to, standard telephone line, LAN, or WAN link (e.g., 802.11, Tl, T3, 56kb, X.25, SNA, DECNET), broadband connections (e.g., ISDN, Frame Relay, ATM, Gigabit Ethernet, Ethernet over SONET), wireless connections, or some combination thereof. Connections can be established using a variety of communication protocols, such as TCP / IP, IPX, SPX, NetBIOS, Ethernet, ARCNET, SONET, SDH, Fiber Distributed Data Interface (FDDI), RS232, IEEE 802.11, IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, IEEE 802.11n, IEEE 802.11ac, IEEE 802.11ad, CDMA, GSM, WiMax, and direct asynchronous. In one embodiment, the computing device 100 communicates with other computing devices 100' via any type and / or form of gateway or tunneling protocol, such as a Secure Socket Layer (SSL) or Transport Layer Security (TLS). The network interface 118 can include a built-in network adapter, network interface card, PCMCIA network card, card bus network adapter, wireless network adapter, USB network adapter, modem, or any other device suitable for interfacing the computing device 100 to any type of network capable of communication and performing the operations described herein.

[0035] In some embodiments, the computing device 100 may include or be connected to one or more display devices 124a to 124n. Thus, any of the I / O devices 130a to 130n and / or the I / O controller 123 may include any type and / or form of suitable hardware, software, or a combination of hardware and software to support, enable, or provide for the computing device 100 to connect to and use the display devices 124a to 124n. For example, the computing device 100 may include any type and / or form of video adapter, video card, driver, and / or library to interface, communicate, connect, or otherwise use the display devices 124a to 124n. In one embodiment, the video adapter may include multiple connectors to interface to the display devices 124a to 124n. In other embodiments, the computing device 100 may include multiple video adapters, each of which is connected to a display device 124a to 124n. In some embodiments, any portion of the operating system of the computing device 100 may be configured to use multiple displays 124a to 124n. In other embodiments, the I / O device 130 may be a bridge between the system bus 150 and an external communication bus, such as a USB bus, an Apple desktop bus, an RS-232 serial connection, a SCSI bus, a FireWire bus, a FireWire 800 bus, an Ethernet bus, an AppleTalk bus, a Gigabit Ethernet bus, an Asynchronous Transfer Mode bus, a FibreChannel bus, a Serial Attached Small Computer System Interface bus, a USB connection, or an HDMI bus.

[0036] Figure 1B and 1CThe computing device 100 of the sort depicted in FIG. 1A can operate under the control of an operating system, which controls scheduling of tasks and access to system resources. The computing device 100 can be running any operating system, such as one of the versions of the MICROSOFT WINDOWS operating systems, the different versions of the Unix and Linux operating systems, the MAC OS for Macintosh computers, any embedded operating system, any real-time operating system, any open source operating system, any proprietary operating system, any operating systems for mobile computing devices, or any other operating system capable of running on the computing device and performing the operations described herein. Typical operating systems include, but are not limited to: the Android operating system produced by Google, the WINDOWS 7-11 operating system produced by the Microsoft Corporation of Redmond, Washington, the MAC OS produced by the Apple Computer of Cupertino, California, the WebOS produced by Research In Motion (RIM) of Waterloo, Ontario, the OS / 2 produced by International Business Machines of Armonk, New York, the Linux operating system produced by Caldera International of Salt Lake City, Utah, or any type and / or form of Unix operating system, among others.

[0037] The computer system 100 can be any workstation, telephone, desktop computer, laptop or notebook computer, server, handheld computer, mobile telephone, other portable telecommunication device, media playing device, gaming system, mobile computing device, or any other type and / or form of computing, telecommunications or media device that is capable of communication and that has sufficient processor power and memory capacity to perform the operations described herein. In some embodiments, the computing device 100 can have different processors, operating systems and input devices consistent with the device.

[0038] Aspects of the operations environments and components described above will become apparent in the context of the systems and methods disclosed herein.

[0039] B. System and Method for In-line Network Analysis

[0040] Systems (or apparatuses) and methods are described herein for performing or predicting network analysis at line rate. In one aspect, a network apparatus includes a reconfigurable neural network circuit to determine an indication of a predicted network characteristic. The predicted network characteristic can be a network anomaly, a network intrusion, a predicted congestion, or a configuration value for a traffic manager to improve QoS (e.g., reduce dropped packets). In one aspect, the reconfigurable neural network circuit includes a set of compute circuits configured to perform computations according to neural network parameters of a neural network to determine the indication of the predicted network characteristic. The neural network parameters can be weights, biases, quantization parameters, stride sizes, pooling sizes, types of activation functions, etc. In one aspect, the reconfigurable neural network circuit includes a controller to determine configuration settings corresponding to packet attributes or flow attributes of an original packet stream. Examples of packet attributes include a packet source, a packet destination, a traffic class or flag. Examples of flow attributes include an identification protocol, a total number of bytes in a flow up to a current packet, a flag count within a flow, a table lookup result, an indication of whether a flow is an elephant flow, or any flow attribute computed by a stateful pipe component. The configuration settings can indicate a configuration for the reconfigurable neural network circuit to implement the neural network. The reconfigurable neural network circuit can include a storage device to provide the neural network parameters of the neural network to the set of compute circuits according to the configuration settings. Thus, different neural networks can be adaptively implemented for different packets to perform different network analysis.

[0041] In one aspect, a network apparatus can be implemented as a linear feedforward packet processing pipeline that is capable of performing per-packet inference (or computation) on packets processed by the pipeline. Packets can be streamed into and out of the network apparatus or reconfigurable neural network circuit at a rate of one packet per clock cycle. Recirculation can be provided locally, with a corresponding non-linear degradation of processing bandwidth.

[0042] In one aspect, the network apparatus can perform per-packet feature computation, scaling and interpolation, and post-processing according to packet attributes or flow attributes. In one aspect, the feature computation, scaling and interpolation, neural network computation, and post-processing can be performed based on one or more tables. Each table can include a list of indices to be applied to a subsequent table or to identify configuration settings for a corresponding packet attribute or flow attribute. For example, a packet attribute of a first packet can correspond to a first entry of a table, and a second attribute of a second packet can correspond to a second entry of the table. Thus, different configuration settings of various components of the network apparatus can be selected for different packets, thereby allowing different analysis or computation to be performed adaptively for different packets. For example, different analysis or computation for application classification, intrusion detection, or congestion prediction can be performed for different packets.

[0043] Advantageously, the disclosed apparatus can perform computations in real-time or at line rate based on one or more neural networks. For example, the disclosed apparatus can implement different neural networks and perform computations at a rate of approximately billions of packets per second (Bpps) and achieve a processing capability of approximately trillions of operations per second (TOPS) for different neural networks in a pipeline.

[0044] Figure 2 FIG. illustrates an apparatus 200 (or system) for performing in-line network analysis based on a reconfigurable neural network circuit 240, according to an embodiment. The apparatus 200 can be an access point 106, a network hardware 192, a network switch, or any network apparatus. In some embodiments, the apparatus 200 includes a feature computation circuit 220, an input processing circuit 230, a reconfigurable neural network circuit 240, and an output processing circuit 250. The apparatus 200 can also include a raw packet data bus 202 and a pipeline bus 205, multiplexers 218, 238, 258, and demultiplexers 228, 248, 268. The pipeline bus 205 can also provide packet metadata, flow metadata, packet attributes, and flow attributes. The multiplexers 218, 238, 258 can selectively provide data between the raw packet data bus 202 and the pipeline bus 205 to the feature computation circuit 220, the input processing circuit 230, and the reconfigurable neural network circuit 240. The demultiplexers 228, 248, 268 can selectively provide data output from the feature computation circuit 220, the input processing circuit 230, and the output processing circuit 250 to the pipeline bus 205. These components can collectively operate to adaptively perform computations of one or more neural networks to generate or determine predicted network characteristics. In some embodiments, the apparatus 200 includes more, fewer, or different components than shown in Figure 2 For example, some of the multiplexers 218, 238, 258 and / or demultiplexers 228, 248, 268 can be omitted or disposed at locations different than shown in Figure 2

[0045] ​In some embodiments, the feature computation circuit 220 is a circuit or component that receives input attribute data 215 including packet attributes or flow attributes and generates feature data 225 including one or more statistical features of one or more packets or a flow based on the packet attributes or flow attributes. Examples of statistical features include flow start time, last packet time, total packet count, total packet length, minimum packet length, maximum packet length, average packet length, average packet length difference, median packet length, minimum packet inter-arrival time (IAT), maximum IAT, average IAT, average IAT difference, median IAT, flow duration, packet rate, number of flags, etc. The average value can be an exponential moving average over a time period. The median value can be an approximate median value. The feature computation circuit 220 can receive the input attribute data 215 from other components (e.g., a processor, a counter, or a stateful table of the apparatus 200) through the pipeline bus 205. From the packet attributes or flow attributes in the input attribute data 215, the feature computation circuit 220 can obtain or collect the statistical features. The feature computation circuit 220 can perform computation on the stored statistical features according to the packet attributes or flow attributes to obtain derived statistical features. The feature computation circuit 220 can generate the feature data 225 including one or more statistical features and provide the one or more statistical features to the input processing circuit 230 through the pipeline bus 205. Details of embodiments and operations of the feature computation circuit 220 are provided below with respect to Figure 4 and 5 Details of embodiments and operations of the input processing circuit 230 are provided below with respect to

[0046] In some embodiments, the input processing circuit 230 is a circuit or component that receives feature data 225’ corresponding to the feature data 225 through the pipeline bus 205 and multiplexer 238 and adjusts the feature data 225’ to generate or obtain adjusted feature data 235. In some embodiments, the reconfigurable neural network circuit 240 can implement a quantized neural network. In one aspect, the input processing circuit 230 can adaptively adjust the feature data 225’ such that the adjusted feature data 235 can be adequately processed by the reconfigurable neural network circuit 240 implementing the quantized neural network. For example, the input processing circuit 230 can adaptively perform scaling and interpolation on the feature data 225’ according to the packet attributes or flow attributes. Details of embodiments and operations of the input processing circuit 230 are provided below with respect to Figure 6 and 7 Details of embodiments and operations of the input processing circuit 230 are provided below with respect to

[0047] In some embodiments, the reconfigurable neural network circuit 240 is a circuit or component that receives adjusted feature data 235' corresponding to the adjusted feature data 235 through the pipeline bus 205 and performs computations of a neural network on the adjusted feature data 235' to obtain an indication 245 of a predicted network characteristic. In some embodiments, the reconfigurable neural network circuit 240 can alternatively receive the raw packet stream from the raw packet data bus 202 and perform computations on the raw packet stream. In one aspect, the reconfigurable neural network circuit 240 includes a set of computation circuits configured to perform computations according to neural network parameters (e.g., weights, biases, quantization parameters, stride sizes, pooling sizes, types of activation functions, etc.) of a neural network to produce or determine an indication 245 of a predicted network characteristic. The indication 245 can be an output of the computations of the neural network. The set of computation circuits can implement different neural networks or perform computations of different neural networks according to packet attributes or stream attributes of the raw packet stream. For example, input signal selection of the set of computation circuits can be changed according to packet attributes or stream attributes of the raw packet stream. For example, different neural network parameters can be applied to the set of computation circuits according to packet attributes or stream attributes of the raw packet stream. The reconfigurable neural network circuit 240 can implement different neural networks for different packets in a pipeline configuration. Details regarding implementations and operations of the reconfigurable neural network circuit 240 are provided below with respect to Figure 8 and 9 Details regarding implementations and operations of the reconfigurable neural network circuit 240 are provided below with respect to

[0048] In some embodiments, the output processing circuit 250 is a circuit or component that receives the indication 245 of a predicted network characteristic from the reconfigurable neural network circuit 240 that implements a quantized neural network and performs post-processing on the indication 245 of a predicted network characteristic to produce output data 255 that includes the predicted network characteristic. In one approach, the output processing circuit 250 can adaptively produce the output data 255 according to packet attributes or stream attributes through regression analysis or classification analysis. Details regarding implementations and operations of the output processing circuit 250 are provided below with respect to Figure 10 and 11 Details regarding implementations and operations of the output processing circuit 250 are provided below with respect to

[0049] Figure 3 A flowchart illustrating a process 300 of determining a predicted network characteristic based on a reconfigurable neural network circuit 240 according to an embodiment is shown. In some embodiments, the process 300 is performed by the device 200. In some embodiments, the process 300 is performed by other entities. In some embodiments, the process 300 includes more, fewer, or different steps than shown in Figure 3

[0050] ​In one approach, the feature computation circuit 220 generates 310 feature data 225 including one or more statistical features of one or more packets. The feature computation circuit 220 can obtain or collect time statistics from packet attributes or flow attributes, and perform computations on the stored time statistics to generate the feature data 225 including one or more statistical features or derived statistical features.

[0051] In one approach, the input processing circuit 230 generates 320 adjusted feature data 235 based on the feature data 225 (or the feature data 225'). For example, the input processing circuit 230 can apply scaling and interpolation to the feature data 225 (or the feature data 225') to obtain the adjusted feature data 235. In one aspect, different neural networks can be set or trained to perform computations for input data values having different ranges or precisions. The input processing circuit 230 can adjust the feature data 225 (or the feature data 225') so that the adjusted feature data 235 can have appropriate value ranges for computations of a neural network (e.g., a quantized neural network).

[0052] In one approach, the reconfigurable neural network circuit 240 performs 330 computations of a neural network based on the adjusted feature data 235 (or the adjusted feature data 235') to obtain an indication 245 of a predicted network characteristic. In one aspect, the reconfigurable neural network circuit 240 includes a set of computation circuits that perform computations on the adjusted feature data 235 (or the adjusted feature data 235') according to neural network parameters of the neural network to determine the indication 245 of the predicted network characteristic. In one example, different inputs or signals of the set of computation circuits can be set or selected for different packet attributes or flow attributes. In one example, different neural network parameters can be applied to the set of computation circuits for different packet attributes or flow attributes.

[0053] In one approach, the output processing circuit 250 generates 340 output data 255 including a predicted network characteristic based on the indication 245. In one approach, the output processing circuit 250 can determine whether the indication 245 is for regression or classification, and compute or determine a value or a decision vector as the output data 255 based on the determination of whether the indication 245 is for regression or classification.

[0054] In one aspect, the feature computation circuit 220, the input processing circuit 230, the reconfigurable neural network circuit 240, and the output processing circuit 250 can operate in a linear pipeline. By adaptively configuring each of the feature computation circuit 220, the input processing circuit 230, the reconfigurable neural network circuit 240, and the output processing circuit 250 differently for different packets according to packet attributes or flow attributes, the feature computation circuit 220, the input processing circuit 230, the reconfigurable neural network circuit 240, and the output processing circuit 250 can operate in a linear pipeline. For example, the feature computation circuit 220 can generate feature data 225 for a first packet of a raw packet flow during a first clock cycle. Then, the feature computation circuit 220 can generate feature data 225 for a second packet of the raw packet flow while the input processing circuit 230 generates adjusted feature data 235 for the first packet of the raw packet flow during a second clock cycle. In one aspect, each of the feature computation circuit 220, the input processing circuit 230, the reconfigurable neural network circuit 240, and the output processing circuit 250 can operate in a linear pipeline. Thus, the apparatus 200 can perform computations at a rate of approximately billions of packets per second (Bpps) such that complex network analysis can be performed in real-time or at line rate.

[0055] Figure 4 A diagram illustrates a schematic of a feature computation circuit 220 according to embodiments. In some embodiments, the feature computation circuit 220 includes ternary content addressable memory (TCAM) match circuit 410, profile table storage 420, multiplexers (MUXs) 402A, 404A, 430, de-multiplexer (deMUX) 402B, 490, hash table circuit 440, configuration table storage 450, first level feature computation circuit 460, second level feature computation circuit 470, precision adjustment circuit 480. These components can be implemented as field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), or any logic circuits. These components can collectively operate to receive input attribute data 405 and generate feature data 485 from the input attribute data 405. In one aspect, the input attribute data 405 corresponds to the input attribute data 215 and the feature data 485 corresponds to the feature data 225. In some embodiments, the feature computation circuit 220 includes more, fewer, or different components than shown in Figure 4 more, fewer, or different components than shown in

[0056] In one aspect, the TCAM matching circuit 410 and the profile table storage 420 can constitute or operate as a controller or a decoder. In one aspect, the TCAM matching circuit 410 receives the input attribute data 405 including the packet attributes or the flow attributes of the original packet stream through the multiplexer 402A. The TCAM matching circuit 410 can store a plurality of values associated with corresponding indices. The TCAM matching circuit 410 can utilize the packet attributes, the flow attributes, or a combination thereof as a key, and perform an AND operation between the key and a mask. Then, the TCAM matching circuit 410 can determine or search for a result of an AND operation between one of the stored values and the mask that matches a result of the AND operation between the key and the mask. The TCAM matching circuit 410 can provide the profile index 415 associated with the value to the profile table storage 420 through the demultiplexer 402B, the pipeline bus 205, and the multiplexer 404A. The profile table storage 420 can store a table including a list of configuration indices corresponding to the profile index. The profile table storage 420 can provide the configuration index 428 corresponding to the received profile index 415 to the configuration table storage 450 through the pipeline bus 205 and the MUX 430. The profile table storage 420 can also store a table including a list of MUX control configuration settings corresponding to the profile index. The profile table storage 420 can provide the MUX control configuration setting 425 corresponding to the received profile index 415 to the MUX 430. According to the MUX control configuration setting 425, the MUX 430 can provide one or more fields in the pipeline bus 205 corresponding to the hash key 435 to the hash table circuit 440, and provide one or more fields in the pipeline bus 205 corresponding to the configuration index 428 to the configuration table storage 450. The profile table storage 420 can also provide the deMUX control configuration setting 492 corresponding to the received profile index 415 to the demultiplexer 490. According to the deMUX control configuration setting 492, the demultiplexer 490 can provide the feature data 485 to the pipeline bus 205.

[0057] In one aspect, the configuration table storage 450 can store a table including a list of configuration settings corresponding to the configuration index. The configuration table storage 450 can be embodied as a static random access memory (SRAM) or any storage device. In some embodiments, the configuration table storage 450 can be implemented as a controller or a decoder. The configuration table storage 450 can provide the configuration settings 452, 455, 458 corresponding to the received configuration index 428 to the first level feature calculation circuit 460, the second level feature calculation circuit 470, and the precision adjustment circuit 480, respectively. The configuration settings 452, 455, 458 can indicate configurations of the first level feature calculation circuit 460, the second level feature calculation circuit 470, and the precision adjustment circuit 480, respectively.

[0058] In one aspect, the TCAM matching circuit 410, the profile table storage 420, and the configuration table storage 450 can identify configuration settings for particular packet attributes or particular flow attributes. Rather than implementing a single component or table to determine configuration settings, implementing multiple components or tables can help improve storage and computational efficiency. For example, the feature computation circuit 220 can support a large number of permutations of different configuration settings (e.g., 10,000 to 300,000). Implementing a single table to store a list of such a large number of configuration settings can consume a large amount of storage resources. By implementing the TCAM matching circuit 410, the profile table storage 420, and the configuration table storage 450 as disclosed herein, each of the TCAM matching circuit 410, the profile table storage 420, and the configuration table storage 450 can be implemented with less storage resources (e.g., 100 kb). Thus, the feature computation circuit 220 can be implemented in a small form factor. In some embodiments, the feature computation circuit 220 can include an array or multiple TCAM matching circuits, profile table storages, and configuration table storages to support a larger number of permutations.

[0059] In one aspect, the MUX 430 receives a set of fields 422 of flow or packet attributes from the pipeline bus 205 and selectively provides one or more fields corresponding to packet attributes or flow attributes 438 to the first level feature computation circuit 460 and one or more fields corresponding to a hash key 435 to the hash table circuit 440 according to a MUX control configuration setting 425. In some embodiments, the MUX 430 is embodied as an array of multiplexers. Examples of packet attributes include packet length, time stamp, etc. The hash key 435 can be formed from a source address, a destination address, a protocol, a source port, a destination port, any subset of packet or flow attributes, or any combination thereof.

[0060] In one aspect, the hash table circuit 440 receives one or more fields corresponding to the hash key 435 from the MUX 430 and identifies the input stream according to the hash key 435. The hash table circuit 440 can include or can be embodied as a memory, a flip-flop, or a digital logic circuit. The hash table circuit 440 can determine whether the database 462 has a corresponding entry for the stream by searching a hash table. The hash table circuit 440 can store a set of database indexes for corresponding hash keys. The database indexes can include indexes of the database 462. In one example, the hash table circuit 440 can determine whether the hash table circuit 440 stores an entry that matches the received hash key 435. If the hash table circuit 440 stores an entry that matches the hash key 435, the hash table circuit 440 can determine that the database 462 has a corresponding entry for the stream. If the hash table circuit 440 does not store an entry that matches the hash key 435, the hash table circuit 440 can determine that the database 462 does not have a corresponding entry for the stream. If the hash table circuit 440 determines that the database 462 does not have a corresponding entry for the stream, the hash table circuit 440 can store the hash key 435 and provide a database index 445 of the database 462 corresponding to the hash key 435 to the database 462 and / or the first level feature calculation circuit 460. If the hash table circuit 440 determines that the database 462 has a corresponding entry for the stream, the hash table circuit 440 can provide the database index 445 to the database 462 and / or the first level feature calculation circuit 460.

[0061] In one aspect, the first level feature computation circuit 460 can perform computations on one or more fields corresponding to the packet attributes or flow attributes 438 according to the configuration settings 452 to obtain time statistics for a flow, and store the time statistics or update a database 462 with the time statistics. Examples of time statistics include a packet count in a flow, a total packet byte count, a minimum packet length, a maximum packet length, an average packet length, a minimum inter-arrival time, a maximum inter-arrival time, an average inter-arrival time, a flag count, etc. For example, the first level feature computation circuit 460 can compare the stored time statistics with the packet attributes or flow attributes 438 selected by the MUX 430 or compare the packet or flow attributes with a constant according to the configuration settings 452 to return a Boolean result. For example, the first level feature computation circuit 460 can update the time statistics according to the configuration settings 452 by adjusting the stored time statistics by subtracting or adding an attribute value, obtaining a minimum or maximum between the stored time statistics and an attribute value, obtaining an approximate median or an exponential moving average. The first level feature computation circuit 460 can provide the time statistics as a statistical feature or perform computations on the time statistics to obtain derived time statistics as a statistical feature. The first level feature computation circuit 460 can provide a first result 465' including the statistical features (time statistics and / or derived time statistics) to the precision adjustment circuit 480 according to the configuration settings 452 or provide the first result 465 to the second level feature computation circuit 470.

[0062] In one aspect, the packet attributes or flow attributes utilized by the TCAM match circuit 410 to determine the profile index 415, the packet attributes or flow attributes utilized by the hash table circuit 440 to obtain the hash key 435, and the packet attributes or flow attributes utilized by the first level feature computation circuit 460 to obtain the time statistics can be different.

[0063] In one aspect, aging control can be provided according to the configuration settings 452. For example, entries for original packet flows in the database 462 that have not been accessed or updated for a predetermined number of clock cycles can be removed from the database 462. The hash table circuit 440 can remove entries of the hash table that include corresponding database indices of the removed entries in the database 462. Thus, the database 462 can not be overloaded due to infrequent packet flows. For example, an entry can be aged out if there is no hit for a preconfigured number of clock cycles or a preconfigured number of packets in a packet flow indicated by the configuration settings 452. This allows the database 462 and the hash table circuit 440 to efficiently release entries for flows for which flow end conditions cannot be easily detected or packets indicating flow end conditions are dropped.

[0064] In one aspect, saturation control can be provided according to configuration setting 452. For example, if a value of an entry in database 462 exceeds an allowable range of values, database 462 can indicate according to configuration setting 452 that the value is invalid, return a predetermined value (e.g., a threshold value), or maintain the value. For example, if database 462 has 16 bits to store a packet count for a flow, the packet count can go from 0 to 65535. At 65535, an accumulator can saturate. Also, certain features can depend on a flow duration, such as a packet rate, a byte rate, etc. When a flow duration is too long, a timer that tracks the flow duration can saturate. When a particular accumulator or counter saturates, an affected feature can be calculated based on the saturated value (e.g., 65535 and beyond for the 65536th packet), an invalid signal can be provided to a downstream pipeline component to ignore an output of reconfigurable neural network circuit 240 or cause input processing circuit 230 to interpolate a value.

[0065] In one aspect, second level feature calculation circuit 470 performs calculations on first results 465 according to configuration setting 455 to obtain second results 475. For example, second level feature calculation circuit 470 can perform a minimum selection between two features, a maximum selection between two features, an average value calculation, adding or subtracting two features, etc. Second level feature calculation circuit 470 can provide second results 475 to precision adjustment circuit 480.

[0066] In one aspect, precision adjustment circuit 480 receives first results 465' and / or second results 475 and generates feature data 485 according to configuration setting 458. In one example, precision adjustment circuit 480 can determine according to configuration setting 458 whether feature data 485 should be provided starting from a first packet in a flow or after a particular number of packets. In one example, precision adjustment circuit 480 can apply a shift operation to first results 465' or second results 475 according to configuration setting 458 to quantize. For example, a 24-bit feature in first results 465' or second results 475 can be right shifted by 8 bits to obtain feature data 485 that includes two 8-bit features to apply to a quantized neural network.

[0067] In one aspect, demultiplexer 490 is an N-bit demultiplexer. In some embodiments, demultiplexer 490 corresponds to or is implemented as demultiplexer 228. In some embodiments, demultiplexer 490 is embodied as a demultiplexer array. Demultiplexer 490 can receive feature data 485 from precision adjustment circuit 480 and selectively provide feature data 485 to input processing circuit 230 over pipeline bus 205 according to deMUX control configuration setting 492. For example, demultiplexer 490 can provide feature data 485 at a corresponding field according to deMUX control configuration setting 492.

[0068] In one aspect, the feature computation circuit 220 can be adaptively configured or arranged to obtain different statistical features of different raw packet streams. The feature computation circuit 220 can be configured differently for each packet. By adaptively configuring the first-level feature computation circuit 460, the second-level feature computation circuit 470, and the precision adjustment circuit 480, the first-level feature computation circuit 460, the second-level feature computation circuit 470, and the precision adjustment circuit 480 can operate in a linear pipeline. For example, the first-level feature computation circuit 460 can generate a first result 465 of a first packet of a raw packet stream during a first clock cycle. Then, while the second-level feature computation circuit 470 generates a second result 475 of the first packet of the raw packet stream based on the first result 465 obtained during the first clock cycle during a second clock cycle, the first-level feature computation circuit 460 can generate a first result 465 of a second packet of the raw packet stream. In some embodiments, each of the TCAM matching circuit 410, the profile table storage 420, the first-level feature computation circuit 460, the second-level feature computation circuit 470, and the precision adjustment circuit 480 can be internally pipelined and execute over multiple clock cycles. Thus, the feature computation circuit 220 can perform computation in real-time or at line rate speed to obtain different statistical features of different packets.

[0069] Figure 5 A flowchart illustrating a process 310 of adaptively generating feature data according to embodiments is shown. In one approach, the process 310 is performed by the feature computation circuit 220. In some embodiments, the process 310 is performed by other entities. In some embodiments, the process 310 includes more, fewer, or different steps than shown in Figure 5

[0070] In one approach, the feature computation circuit 220 receives 510 input attribute data 405 including packet attributes or stream attributes of a raw packet stream.

[0071] In one approach, the feature computation circuit 220 determines 520 a hash key based on the packet attributes or the stream attributes through a first table (e.g., a profile table). For example, the TCAM matching circuit 410 can determine a profile index 415 with the matched packet attributes or stream attributes and provide the profile index 415 to the profile table storage 420. In response to the profile index 415, the profile table stored by the storage 420 can determine, identify, or provide a corresponding MUX control configuration setting. According to the MUX control configuration setting, the MUX 430 can select and provide one or more fields of one or more attributes of a stream in the pipeline bus 205 corresponding to the hash key 435.

[0072] ​In one approach, the feature computation circuit 220 determines 530 the configuration index 428 based on the packet attributes or the flow attributes through a first table (e.g., the profile table stored by the storage device 420). For example, in response to the profile index 415, the profile table stored by the storage device 420 can determine, identify, or provide the corresponding configuration index 428.

[0073] In one approach, the feature computation circuit 220 determines 540 the configuration settings (e.g., configuration settings 452, 455, 458) based on the configuration index 428 through a second table (e.g., the configuration table stored by the storage device 450). The configuration settings can indicate the configuration of the computation circuits (e.g., the first-level feature computation circuit 460, the second-level feature computation circuit 470, and the precision adjustment circuit 480).

[0074] In one approach, the feature computation circuit 220 identifies 550 the flow according to the hash key 435. The hash table circuit 440 can determine whether the database 462 has a corresponding entry for the flow according to the hash key 435. For example, the hash table circuit 440 can determine whether the hash table circuit 440 stores an entry that matches the received hash key 435. If the hash table circuit 440 stores an entry that matches the hash key 435, the hash table circuit 440 can determine that the database 462 has a corresponding entry for the flow. If the hash table circuit 440 does not store an entry that matches the hash key 435, the hash table circuit 440 can determine that the database 462 does not have an entry corresponding to the flow. The hash table circuit 440 can send an indication to the first-level feature computation circuit 460 indicating whether the database has a corresponding entry for the flow. In response to determining that the database 462 does not have a corresponding entry for the flow, the hash table circuit 440 can send the database index 445 of the database 462 corresponding to the hash key 435 to the first-level feature computation circuit 460 and the database 462. In response to determining that the database 462 has a corresponding entry for the flow, the hash table circuit 440 can send the database index 445 corresponding to the hash key 435 to the first-level feature computation circuit 460 and the database 462. The entry of the database 462 associated with the database index 445 can be updated.

[0075] In one approach, the feature computation circuitry 220 obtains 560 temporal statistics for the identified flow. For example, the first-level feature computation circuitry 460 may perform computations on one or more fields corresponding to the packet attributes or flow attributes 438 according to the configuration settings 452 to obtain the temporal statistics for the flow. If the hash table circuitry 440 determines that the database 462 does not have a corresponding entry for the flow (as indicated by the indicator), the first-level feature computation circuitry 460 may cause the database 462 to create a new entry using the database index and store the temporal statistics for the flow in the new entry. If the hash table circuitry 440 determines that the database 462 has a corresponding entry for the flow (as indicated by the indicator), the first-level feature computation circuitry 460 may cause the database 462 to update the corresponding entry at the database index 445 with the temporal statistics for the flow.

[0076] In one approach, the feature computation circuitry 220 performs 570 computations on the temporal statistics according to the configuration settings. For example, the first-level feature computation circuitry 460 may perform a first-level computation 572 on the temporal statistics (or statistical features) according to the configuration settings 452 to obtain a first result 465. For example, the second-level feature computation circuitry 470 may perform a second-level computation 575 on the first result 465 according to the configuration settings 455 to obtain a second result 475.

[0077] In one approach, feature computation circuitry 220 generates 580 feature data 485 according to configuration settings. For example, precision adjustment circuitry 480 may generate feature data 485 having a specific number of bits or based on first result 465 or second result 475 according to configuration settings 458 for computation by the quantized neural network.

[0078] Figure 6 Illustrated is a schematic diagram of the input processing circuit 230 according to an embodiment. In some embodiments, the input processing circuit 230 includes MUX 605A...605D, demultiplexer 605E, TCAM matching circuit 610, policy table storage 620, MUX control circuit 630, configuration table storage 640, and scaling and interpolation circuit 650. These components can be implemented as a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), or any logic circuit. These components can operate together to receive input attribute data 608 and generate adjusted feature data 698. In one aspect, the adjusted feature data 698 corresponds to the adjusted feature data 235. In one aspect, the input processing circuit 230 generates the adjusted feature data 698 for application to a quantized neural network. In some embodiments, the input processing circuit 230 includes Figure 6more, fewer, or different components than those shown in the figures. In some embodiments, TCAM matching circuit 610, policy table storage 620, MUX control circuit 630, and configuration table storage 640 are embodied as a single component or a single controller.

[0079] In one aspect, TCAM matching circuit 610 and policy table storage 620 can constitute or operate as a controller 615 or a decoder. In one aspect, TCAM matching circuit 610 receives input attribute data 608 including packet attributes or flow attributes of a raw packet stream through multiplexer 605 A. TCAM matching circuit 610 can utilize the packet attributes, the flow attributes, or a combination thereof as a key, and perform an "and" operation between the key and a mask. TCAM matching circuit 610 can then determine or search for a result of an "and" operation between one of the stored values and the mask that matches the result of the "and" operation between the key and the mask. TCAM matching circuit 610 can provide a policy index 612 associated with the value to policy table storage 620. Policy table storage 620 can store a table including a list of configuration indexes corresponding to the policy index. In some embodiments, policy table storage 620 is embodied as a SRAM or any storage device. Policy table storage 620 can provide a configuration index 628 corresponding to the received policy index 612 to configuration table storage 640. Policy table storage 620 can provide a control index 625 associated with configuration index 628 to MUX control circuit 630.

[0080] In one aspect, configuration table storage 640 can store a table including a list of configuration settings corresponding to the configuration index. Configuration table storage 640 can be embodied as a static random access memory (SRAM) or any storage device. In some embodiments, configuration table storage 640 can be implemented as a controller or a decoder. Configuration table storage 640 can provide configuration settings 645A to 645F corresponding to the received configuration index 628 to various components of scaling and interpolation circuit 650. Configuration settings 645A to 645F can indicate configurations of the components of scaling and interpolation circuit 650.

[0081] In one aspect, the TCAM matching circuit 610, the policy table storage 620, and the configuration table storage 640 can identify configuration settings 645 for a particular packet attribute or a particular flow attribute. Rather than implementing a single component or table to determine the configuration settings 645, implementing multiple components or tables can help improve storage and computational efficiency. For example, the input processing circuit 230 can support a large number of permutations of different configuration settings 645 (e.g., 10,000 to 300,000). Implementing a single table to store a list of such a large number of configuration settings 645 can consume a large amount of storage resources. By implementing the TCAM matching circuit 610, the policy table storage 620, the configuration table storage 640, and the MUX control circuit 630 as disclosed herein, each of the TCAM matching circuit 610, the policy table storage 620, the configuration table storage 640, and the MUX control circuit 630 can be implemented with less storage resources (e.g., 100 kb). Thus, the input processing circuit 230 can be implemented in a small form factor.

[0082] The MUX control circuit 630 can be a circuit to control the MUXs 605B, 605C, 605D, and the demultiplexer 605E according to the control index 625. The MUX control circuit 630 can include a table of different configurations or control signals of the corresponding control index of the MUXs 605B, 605C, 605D, and the demultiplexer 605E. The MUX control circuit 630 can receive the control index 625 and can generate the control signals 635 corresponding to the control index 625. The MUX control circuit 630 can apply the control signals 635 to the MUXs 605B, 605C, 605D, and the demultiplexer 605E.

[0083] In some embodiments, MUX 605B is an N-bit multiplexer (e.g., 16 bits), and MUX 605C is a 1-bit multiplexer. In some embodiments, each of MUX 605B and MUX 605C can be embodied as an array of multiplexers. MUX 605B can select an ordinal feature 655A of the feature data (e.g., feature data 225') to be processed, where MUX 605C can select a categorical feature 655B of the feature data (e.g., feature data 225') to be processed. Ordinal feature 655A can be a feature represented by a bit-width up to twice the quantization precision, where categorical feature 655B can be a binary feature (e.g., a flag) or a feature represented by one bit (e.g., a one-hot encoded feature). For example, for an 8-bit quantization precision, ordinal feature 655A can be represented by up to 16 bits. MUX 605B can selectively provide the N-bit ordinal feature 655A from the pipeline bus 205 or the raw packet data bus 202 to MUX 660 of the scaling and interpolation circuit 650 according to control signal 635. Similarly, MUX 605C can selectively provide the 1-bit categorical feature 655B from the pipeline bus 205 or the raw packet data bus 202 to MUX 660 according to control signal 635.

[0084] In some embodiments, MUX 605D is a 1-bit multiplexer. In some embodiments, MUX 605D can be embodied as an array of multiplexers. MUX 605D can receive a 1-bit valid feature indicator 658 from the pipeline bus 205 and selectively provide the 1-bit valid feature indicator to OR gate 695 of the scaling and interpolation circuit 650 according to control signal 635. The 1-bit valid feature indicator 658 can indicate whether an ordinal feature 655A or a categorical feature 655B has a valid value. According to valid feature indicator 658, the scaling and interpolation circuit 650 can perform interpolation.

[0085] In some embodiments, the scaling and interpolation circuit 650 receives feature data (e.g., ordinal feature 655A or categorical feature 655B) and generates adjusted feature data 698 according to configuration settings 645A-F from the configuration table storage 640. In some embodiments, the scaling and interpolation circuit 650 includes an N-bit multiplexer 660, a left shifter 665, a right shifter 670, a mask operator 675, an adder 680, a clamping circuit 685, an N-bit multiplexer 690, and an OR gate 695. The multiplexer 660 can be embodied as a multiplexer array. In some embodiments, the shifter 665 can be embodied as a left shifter or a left shifter array. In some embodiments, the shifter 670 can be embodied as a right shifter or a right shifter array. In some embodiments, the mask operator 675 is embodied as a mask operator array. In some embodiments, the adder 680 is embodied as an adder array. The adder 680 can be a signed adder. In some embodiments, the clamping circuit 685 is embodied as a clamping circuit array. These components can collectively operate to perform scaling and interpolation on received feature data to generate adjusted feature data 698. In some embodiments, the scaling and interpolation circuit 650 includes fewer, more, or different components than shown in FIG. 6. Figure 6 In some embodiments, the scaling and interpolation circuit 650 includes fewer, more, or different components than shown in FIG. 6.

[0086] The multiplexer 660 can provide the ordinal feature 655A or the categorical feature 655B to the shifter 665 according to the configuration setting 645A.

[0087] Shifter 665, 670, mask operator 675, adder 680, and clamp circuit 685 can constitute a scaling circuit to perform scaling and masking operations. In one aspect, left shifter 665 can perform a left shift operation and right shifter 670 can perform a right shift operation to scale, according to configuration setting 645B. In one aspect, mask operator 675 can perform a masking operation on the shifted value from shifter 670 according to configuration setting 645C from configuration table storage 640. In one example, mask operator 675 is implemented as an N-bit AND gate to perform an AND logical operation between the shifted output from shifter 670 and a reference value in configuration setting 645C from configuration table storage 640. Adder 680 can add an offset to the output of mask operator 675 according to an offset value in configuration setting 645D from configuration table storage 640. Clamp circuit 685 can clamp the output of adder 680. In one aspect, clamp circuit 685 clamps the value to a predefined range based on the quantization precision of reconfigurable neural network circuit 240. For example, for an 8b quantized neural network, clamp circuit 685 can clamp the value between -128 and +127. For example, if the input value is less than -128, clamp circuit 685 can set the output value to -128. For example, if the input value is greater than 127, clamp circuit 685 can set the output value to 127. For example, if the input value is between -128 and 127, clamp circuit 685 can set the output value to the input value. Thus, shifter 665, 670, mask operator 675, adder 680, and clamp circuit 685 can perform scaling and masking operations with simple components (e.g., shifters 665, 670) without the need for complex circuits (e.g., multipliers, dividers, etc.). Thus, scaling and interpolation circuit 650 can be implemented in a simple architecture and save computational resources.

[0088] In one aspect, scaling and interpolation circuit 650 can discard unwanted lower bits and mask out lower bits as needed in order to bucket the values. Scaling and interpolation circuit 650 can then convert to signed integers and have the adjusted values uniformly spread in the quantization range (e.g., between -128 and 127 for 8b). Scaling and interpolation circuit 650 can be implemented with a simple architecture as shown in Figure 6 In one aspect, scaling and interpolation circuit 650 can discard unwanted lower bits and mask out lower bits as needed in order to bucket the values. Scaling and interpolation circuit 650 can then convert to signed integers and have the adjusted values uniformly spread in the quantization range (e.g., between -128 and 127 for 8b). Scaling and interpolation circuit 650 can be implemented with a simple architecture as shown in

[0089] In one aspect, multiplexer 690 and OR gate 695 may constitute an interpolation circuit to perform interpolation. In some embodiments, multiplexer 690 is embodied as a multiplexer array, and OR gate 695 is embodied as an OR gate array. In one aspect, the interpolation circuit may detect invalid values ​​in the characteristic data and replace the invalid values ​​with configured values. For example, multiplexer 690 receives the output of clamp circuit 685 and the assigned or configured value to be applied in configuration setting 645E. Based on a 1-bit valid characteristic indicator 658 or configuration setting 645F (e.g., force interpolation control), multiplexer 690 may select or provide the output of clamp circuit 685 or the assigned value (or configured value) as the adjusted characteristic data 698.

[0090] In some embodiments, demultiplexer 605E is an N-bit demultiplexer. In some embodiments, demultiplexer 605E corresponds to or is implemented as demultiplexer 248. In some embodiments, demultiplexer 605E may be embodied as a demultiplexer array. Demultiplexer 605E may be coupled to the output of multiplexer 690 of scaling and interpolation circuit 650. Demultiplexer 605E may receive adjusted feature data 698 from the N-bit output of MUX 690 and selectively provide the adjusted feature data 698 to reconfigurable neural network circuit 240 via pipeline bus 205 based on control signal 635. For example, demultiplexer 605E may provide the adjusted feature data 698 at the corresponding field based on control signal 635.

[0091] Figure 7 The diagram illustrates a flow chart showing a process 320 of adaptively adjusting feature data to obtain adjusted feature data according to an embodiment. In one approach, process 320 is performed by input processing circuitry 230. In some embodiments, process 320 is performed by other entities. In some embodiments, process 320 includes Figure 7 More, fewer, or different steps than shown in .

[0092] In one approach, the input processing circuitry 230 receives 710 input attribute data 608 comprising packet attributes or flow attributes.

[0093] In one approach, the input processing circuitry 230 determines 720 a configuration index from a first table (e.g., a policy table) based on the packet attributes or flow attributes. For example, the TCAM matching circuitry 610 may determine a policy index 612 using the matched packet attributes or flow attributes and provide the policy index 612 to the policy table storage 620. In response to the policy index 612, the policy table stored by the storage 620 may determine, identify, or provide a corresponding configuration index 628.

[0094] In one approach, the input processing circuit 230 determines 730 configuration settings based on the configuration index 628 through a second table (e.g., a configuration table stored by storage 640). The configuration settings can indicate configurations of various circuits or components of the scaling and interpolation circuit 650 (e.g., N-bit multiplexers 660, left shifter 665, right shifter 670, mask operator 675, adder 680, clamping circuit 685, N-bit multiplexers 690, OR gates 695, etc.).

[0095] In one approach, the input processing circuit 230 applies 740 scaling and interpolation to the feature data (e.g., feature data 225’ or 655) according to the configuration settings to obtain adjusted feature data (e.g., adjusted feature data 235 or 698). In one aspect, scaling is performed by simple components (e.g., shifters) and interpolation is performed by multiplexers (without the need to use complex logic circuits, such as multipliers, dividers, or other complex circuits) for quantization. Thus, the input processing circuit 230 can be implemented in a small form factor and perform scaling and interpolation in a fast manner with reduced power consumption.

[0096] Figure 8 A schematic diagram illustrating a reconfigurable neural network circuit 240 according to embodiments is shown. In some embodiments, the reconfigurable neural network circuit 240 includes multiplexers 805, TCAM matching circuit 810, policy table storage 820, MUX control circuit 840, QNN control circuit 830, QNN parameter profile table storage 890, and a set of compute circuits 850. These components can be implemented as field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), or any logic circuits. These components can collectively operate to receive feature data 808 from the pipeline bus 205 or one or more packets 818 from the raw packet data bus 202 and generate an output of predicted network characteristics or an indication 845. In one aspect, the feature data 808 corresponds to adjusted feature data 235 (or adjusted feature data 235’), and the indication 845 of predicted network characteristics corresponds to the indication 245 of predicted network characteristics. In some embodiments, the reconfigurable neural network circuit 240 includes more, fewer, or different components than shown in Figure 8 In some embodiments, the TCAM matching circuit 810, policy table storage 820, MUX control circuit 840, and QNN control circuit 830 are embodied as a single component or a single controller.

[0097] In one aspect, the TCAM matching circuit 810 and the policy table storage 820 can constitute or operate as a controller 815 or a decoder. In one aspect, the TCAM matching circuit 810 receives input attribute data 802 including packet attributes or flow attributes of the original packet stream through the multiplexer 805. The TCAM matching circuit 810 can utilize the packet attributes, the flow attributes, or a combination thereof as a key and perform an AND operation between the key and a mask. Then, the TCAM matching circuit 810 can determine or search for a result of an AND operation between one of the stored values and the mask that matches the result of the AND operation between the key and the mask. The TCAM matching circuit 810 can provide a policy index 812 associated with the value to the policy table storage 820. The policy table storage 820 can store a table including a list of configuration indexes corresponding to the policy index. In some embodiments, the policy table storage 820 is embodied as an SRAM or any storage device. The policy table storage 820 can provide a configuration index 828A corresponding to the received policy index 812 to the QNN control circuit 830. The policy table storage 820 can provide a MUX control index 828B associated with the configuration index 828A to the MUX control circuit 840.

[0098] In some embodiments, the TCAM matching circuit 810 and the policy table storage 820 can determine different configuration settings for different packets. For example, the TCAM matching circuit 810 and the policy table storage 820 can determine a first configuration setting corresponding to the packet attributes or the flow attributes of a first packet of the original packet stream during a first clock cycle. Then, the TCAM matching circuit 810 and the policy table storage 820 can determine a second configuration setting corresponding to the packet attributes or the flow attributes of a second packet of the original packet stream during a second clock cycle that is immediately after or after the first clock cycle. Thus, the reconfigurable neural network circuit 240 can set, configure, or operate differently for different packets of different neural networks in a pipelined manner.

[0099] In one aspect, QNN control circuit 830 can set, configure, or control the operation of the set of compute circuits 850 and QNN parameter profile table storage 890 according to configuration index 828A. QNN control circuit 830 can include storage that stores a table that includes a list of configuration settings for corresponding configuration indices. QNN control circuit 830 can provide a first configuration setting that includes control signals 838A-838C corresponding to the received configuration index 828A to various components of the set of compute circuits 850. The first configuration setting that includes control signals 838A-838C can indicate which components (e.g., MAC circuits, convolution layers) to enable or select. QNN control circuit 830 can also provide a second configuration setting that includes a QNN profile index 838D corresponding to the received configuration index 828A to QNN parameter profile table storage 890. The second configuration setting that includes QNN profile index 838D can indicate which neural network parameters of which neural network to apply to which layer or sub-group of the set of compute circuits 850. In some embodiments, QNN control circuit 830 can also perform permission control and generate a busy signal. In some embodiments, if there is a predicted resource conflict to use or reuse the same compute resource during a clock cycle, QNN control circuit 830 can not permit a packet or a set of features on the same clock cycle for inference.

[0100] MUX control circuit 840 can be a circuit to control MUX 848 according to MUX control index 828B. MUX control circuit 840 can include a table of different configurations or control signals of MUX 848 for corresponding MUX control indices. MUX control circuit 840 can receive MUX control index 828B and can generate MUX control signals 835 corresponding to MUX control index 828B. MUX control circuit 840 can apply MUX control signals 835 to MUX 848.

[0101] In one aspect, the TCAM matching circuit 810, the policy table storage 820, and the QNN control circuit 830 can identify configuration settings for a particular packet attribute or a particular flow attribute. Rather than implementing a single component or table to determine the configuration settings, implementing multiple components or tables can help improve storage and computing efficiency. For example, the reconfigurable neural network circuit 240 can support a large number of permutations (e.g., 10,000 to 300,000) of different configuration settings for different neural networks. Implementing a single table to store a list of such a large number of configuration settings can consume a large amount of storage resources. By implementing the TCAM matching circuit 810, the policy table storage 820, the QNN control circuit 830, and the MUX control circuit 840 as disclosed herein, each of the TCAM matching circuit 810, the policy table storage 820, the QNN control circuit 830, and the MUX control circuit 840 can be implemented with less storage resources (e.g., 100 kb). Thus, the reconfigurable neural network circuit 240 can be implemented in a small form factor.

[0102] In one aspect, the QNN parameter profile table storage 890 can include a plurality of bins, where each bin can store neural network parameters for a corresponding layer of a neural network. Examples of neural network parameters include weights, biases, quantization parameters, stride size, pooling size. In one aspect, each bit can be identified by a corresponding QNN profile index 838D. For example, a first bin can store neural network parameters for a first layer of a neural network, and a second bin can store neural network parameters for a second layer of a neural network. In one aspect, the neural network parameters can be trained such that a neural network implemented according to the neural network parameters can generate an indication of a predicted network characteristic (e.g., network anomaly, intrusion detection, predicted congestion, etc.) for input feature data 808 or raw packets 818. The neural network parameters can be trained prior to deployment of the apparatus 200. The QNN parameter profile table storage 890 can receive a QNN profile index 838D and apply a signal corresponding to the neural network parameters (e.g., weights, bias values, activation functions) stored by the bin corresponding to the QNN profile index 838D to the corresponding computing circuit 850. In some embodiments, the QNN parameter profile table storage 890 can receive a different QNN profile index 838D every clock cycle or for every packet, and accordingly provide a different signal corresponding to different neural network parameters to the set of computing circuits 850 every clock cycle or for every packet.

[0103] In some embodiments, the set of compute circuits 850 includes multiplexers 848, 870A, 870B, 870C, 852, neurons 855, 885, pooling circuits 858, and controllable delay lines 860A...860C. These components can collectively operate to perform computations of one or more neural networks on the feature data 808 or the one or more packets 818 to produce the indication 845 of the predicted network property. In some embodiments, the set of compute circuits 850 includes more, fewer, or different components than shown in Figure 8

[0104] In one aspect, the set of compute circuits 850 includes a first portion 865 and a second portion 868 to implement, for example, two types of layers: convolutional layers (CNNs) and dense layers. For example, a convolutional layer can receive the feature data 808 or the one or more packets 818 and perform computations to identify spatial features in the feature data 808 or the one or more packets 818. Then, a dense layer can perform computations on the identified spatial features to produce the indication 845 of the predicted network property.

[0105] In one aspect, the first portion 865 of the set of compute circuits 850 includes multiplexers 848, 870A, 870B, 852, delay lines 860A, 860B, 860C, neurons 855, and pooling circuits 858. In one example, the first portion 865 of the set of compute circuits 850 can implement a convolutional layer. In some embodiments, the first portion 865 of the set of compute circuits 850 can implement other types of neural network layers.

[0106] ​In one aspect, the multiplexers 852, neurons 855, and pooling circuits 858 are arranged in layers or stacks, where each layer or each stack includes a corresponding set of multiplexers 852, a corresponding set of neurons 855, and a corresponding pooling circuit 858. In some embodiments, a certain layer or stack can omit the pooling circuit 858. Each neuron 855 can be embodied as a multiply-accumulate (MAC) circuit with quantization or any reconfigurable computing circuit. Each pooling circuit 858 can be a max-pooling circuit to perform a max-pooling function or an average-pooling circuit to perform an average-pooling. The set of multiplexers 852, the set of neurons 855, and the pooling circuit 858 in a layer can implement a corresponding layer of a neural network. In one aspect, the set of multiplexers 852, the set of neurons 855, and the pooling circuit 858 in a layer can be set, controlled, or configured according to the neural network parameters (e.g., weights, bias values) stored in the QNN parameter profile table storage 890 by a corresponding bin. For example, each multiplexer 852 can be individually controlled or configured according to a neural network parameter to provide convolution striding. For example, each neuron 855 can perform a multiplication or multiply-accumulate operation according to a corresponding set of weights and bias values in a neural network parameter.

[0107] In one aspect, the multiplexers 848 apply the feature data 808 or the raw data in the one or more packets 818 as input according to the MUX control signals 835 from the MUX control circuit 840. In some embodiments, the multiplexers 848 correspond to or are implemented as the multiplexers 258. In one aspect, the multiplexers 870A can be set, controlled, or configured according to the control signals 838A from the QNN control circuit 830 to support recirculation. In one aspect, the multiplexers 870B can be set, controlled, or configured according to the control signals 838B from the QNN control circuit 830 to bypass a certain layer. In one aspect, the delay lines 860A...860C ensure that each pass can go through the same number of cycles of a corresponding layer of the first portion 865 of the set of computing circuits to facilitate the design of the QNN control circuit 830.

[0108] In one aspect, the second portion 868 of the set of computing circuits 850 includes the multiplexers 870C and the neurons 885. In one example, the second portion 868 of the set of computing circuits 850 can implement a dense layer.

[0109] In one aspect, the neurons 885 are arranged in layers or stacks, where each layer or stack includes a set of corresponding neurons 885. Each neuron 885 can be embodied as a quantized multiplication and accumulation circuit or any reconfigurable computational circuit. A set of neurons 885 in a layer can implement a corresponding layer of a neural network. In one aspect, the set of neurons 885 in a layer can be set, controlled, or configured according to neural network parameters (e.g., weights, bias values) stored in the QNN parameter profile table storage device 890 by corresponding bin. For example, each neuron 885 can perform a multiplication or multiplication and accumulation operation according to a set of corresponding weights and bias values ​​in the neural network parameters.

[0110] In one aspect, the reconfigurable neural network circuit 240 can operate in a linear pipeline. The reconfigurable neural network circuit 240 can operate in a linear pipeline to perform computations for the same neural network or different neural networks. For example, the QNN control circuit 830 can provide the QNN profile index 838D to the QNN parameter profile table storage 890 so that the QNN parameter profile table storage 890 can apply a signal corresponding to the first neural network parameter of the first layer of the first neural network in the first bin to the first layer (e.g., CI0 . ... i-1 ). Thus, the first layer of the set of computing circuits 850 (eg, CI0 . . . CI i-1 ) may perform computations for the first packet according to the first neural network parameters in the first bin during the first clock cycle. The QNN control circuitry 830 may then provide the QNN profile index 838D to the QNN parameter profile table storage 890 so that the QNN parameter profile table storage 890 may apply a signal corresponding to the second neural network parameters of the second layer of the first neural network in the second bin to the second layer (e.g., CHO . . . CH1 . ) of the set of computation circuits 850. c0-1 ) and applying a signal corresponding to a third neural network parameter of a first layer of a second neural network in a third bin to a first layer (e.g., CI0 . . . CI0 ) of the set of computational circuits 850 during a second clock cycle immediately following the first clock cycle or after the first clock cycle. i-1 ). Therefore, the second layer of the set of computing circuits 850 (eg, CHO . . . CH c0-1 ) may perform calculations on the output of the first packet in the first clock cycle based on the second neural network parameters of the second layer of the first neural network in the second partition during the second clock cycle based on the first layer of the group of computing circuits 850, and the first layer of the group of computing circuits 850 (e.g., CI0 . . . CI i-1) can perform computations for the second packet according to third neural network parameters of the first layer of the second neural network in the third partition during a second clock cycle. By applying the neural network parameters of layers of different neural networks to different layers or different subsets of the set of computational circuits 850 for each clock cycle, the reconfigurable neural network circuit 240 can perform computations for different neural networks in a pipeline at a rate of approximately billions of packets per second (Bpps) and achieve a processing capability of approximately trillions of operations per second (TOPS).

[0111] Figure 9 The diagram illustrates a flow chart showing a process 330 for adaptively performing computations of a neural network by reconfigurable neural network circuit 240 according to an embodiment. In one approach, process 330 is performed by reconfigurable neural network circuit 240. In some embodiments, process 330 is performed by other entities. In some embodiments, process 330 includes Figure 9 More, fewer, or different steps than shown in .

[0112] In one approach, the reconfigurable neural network circuit 240 receives 910 input attribute data 802 comprising packet attributes or flow attributes.

[0113] In one approach, the reconfigurable neural network circuit 240 determines 920 a configuration index from a first table (e.g., a policy table) based on packet attributes or flow attributes. For example, the TCAM matching circuit 810 may determine a policy index 812 using the matched packet attributes or flow attributes and provide the policy index 812 to the policy table storage device 820. In response to the policy index 812, the policy table stored by the storage device 820 may determine, identify, or provide a corresponding configuration index 828A.

[0114] In one approach, the reconfigurable neural network circuit 240 determines 930 a first configuration setting including control signals 838A-838C and a second configuration setting including a QNN profile index 838D based on the configuration index 828A, for example, from a second table (e.g., a configuration table stored by the QNN control circuit 830). The first configuration setting including the control signals 838A-838C may indicate how to set, control, or configure one or more components (e.g., multiplexers) of the set of computational circuits 850. The second configuration setting including the QNN profile index 838D may indicate which neural network parameters of which neural network are to be applied to which layer or subset of the set of computational circuits 850.

[0115] In one approach, the reconfigurable neural network circuit 240 configures 940 the set of compute circuits 850 according to the configuration settings. For example, the QNN control circuit 830 can provide the QNN profile index 838D to the QNN parameter profile table storage 890 to apply the neural network parameters of the corresponding layer of the corresponding neural network corresponding to the QNN profile index 838D to the corresponding subset or layer of the set of compute circuits 850. The reconfigurable neural network circuit 240 can also set, control, or configure one or more multiplexers (e.g., 848, 870A...870C) to support recirculation or bypass capabilities.

[0116] In one approach, the reconfigurable neural network circuit 240 applies 950 the feature data 808 (or adjusted feature data 235, 235', 698) to the set of compute circuits to obtain an indication 845 of a predicted network characteristic. In one aspect, the reconfigurable neural network circuit 240 can be set, controlled, or configured differently for different neural networks, such that different analyses can be performed for different packets or feature data.

[0117] Figure 10 FIG. illustrates an output processing circuit 250 according to embodiments. In some embodiments, the output processing circuit 250 includes MUXs 1005, 1030A, 1030B, 1085 and demultiplexers 1090A, 1090B, TCAM matching circuit 1010, policy table storage 1020, MUX control circuit 1060, configuration table storage 1040, classification analysis processor 1050, and regression analysis processor 1068. These components can be implemented as field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), or any logic circuits. These components can collectively operate to receive an output of the reconfigurable neural network circuit 240 or an indication 1045 of a predicted network characteristic, and generate output data 1095 indicating the predicted network characteristic. In one aspect, the indication 1045 corresponds to the indication 245 or 845, and the output data 1095 corresponds to the output data 255. In one aspect, the output processing circuit 250 generates the output data 1095 for an output of a quantized neural network. In some embodiments, the output processing circuit 250 includes more, fewer, or different components than shown in FIG. 10. Figure 10 In some embodiments, the TCAM matching circuit 1010, the policy table storage 1020, the MUX control circuit 1060, and the configuration table storage 1040 are embodied as a single component or a single controller.

[0118] In one aspect, the TCAM matching circuit 1010 and the policy table storage 1020 can constitute or operate as a controller 1015 or a decoder. In one aspect, the TCAM matching circuit 1010 receives the input attribute data 1008 including the packet attributes or the flow attributes of the original packet stream through the multiplexer 1005. The TCAM matching circuit 1010 can utilize the packet attributes, the flow attributes, a combination thereof as a key, and perform an "and" operation between the key and a mask. Then, the TCAM matching circuit 1010 can determine or search for a result of an "and" operation between one of the stored values and the mask that matches the result of the "and" operation between the key and the mask. The TCAM matching circuit 1010 can provide the policy index 1012 associated with the value to the policy table storage 1020. The policy table storage 1020 can store a table including a list of configuration indexes corresponding to the policy index. In some embodiments, the policy table storage 1020 is embodied as a SRAM or any storage device. The policy table storage 1020 can provide the configuration index 1028 corresponding to the received policy index 1012 to the configuration table storage 1040 through the multiplexer 1030B. The policy table storage 1020 can provide the control index 1025 associated with the configuration index 1028 to the MUX control circuit 1060 through the multiplexer 1030A. In one aspect, the multiplexers 1030A, 1030B can be coupled to Figure 8 The de-multiplexer 1090B can provide the configuration settings 1002 or the recirculation control to the pipeline bus 205 to inform the downstream pipeline components that the reconfigurable neural network circuit 240 is not capable of producing a valid indication of the predicted network characteristics at the clock cycle due to recirculation.

[0119] In one aspect, the configuration table storage 1040 can store a table including a list of configuration settings corresponding to the configuration index. The configuration table storage 640 can be embodied as a static random access memory (SRAM) or any storage device. In some embodiments, the configuration table storage 1040 can be implemented as a controller or a decoder. The configuration table storage 1040 can provide the configuration settings corresponding to the received configuration index 1028 to various components of the output processing circuit 250. The configuration settings can indicate configurations of the components of the output processing circuit 250 (e.g., the classification analysis processor 1050, the regression analysis processor 1068, etc.). In one aspect, the configuration settings indicate or correspond to a type of the output data 1095 (e.g., apply classification, network anomaly detection, network intrusion detection, predicted congestion, or a configuration value for the traffic manager to improve QoS, etc.).

[0120] In one aspect, the TCAM matching circuit 1010, the policy table storage 1020, and the configuration table storage 1040 can identify configuration settings for particular packet attributes or particular flow attributes. Rather than implementing a single component or table to determine configuration settings, implementing multiple components or tables can help improve storage and computational efficiency. For example, the output processing circuit 250 can support a large number of permutations of different configuration settings (e.g., 10,000 to 300,000). Implementing a single table to store a list of such a large number of configuration settings can consume a large amount of storage resources. By implementing the TCAM matching circuit 1010, the policy table storage 1020, and the configuration table storage 1040 as disclosed herein, each of the TCAM matching circuit 1010, the policy table storage 1020, and the configuration table storage 1040 can be implemented with less storage resources (e.g., 100 kb). Thus, the output processing circuit 250 can be implemented in a small form factor.

[0121] In one aspect, the indication 1045 of the output or predicted network characteristics of the reconfigurable neural network circuit 240 can be processed by a classification analysis processor 1050 or a regression analysis processor 1068. The classification analysis processor 1050 can use the indication 1045 of the predicted network characteristics to compute a one-hot decision vector for a multi-label or multi-class classification problem, where the regression analysis processor 1068 can use the indication 1045 of the predicted network characteristics to predict values for a multivariate regression problem. Classification analysis can involve converting the indication 1045 or neural network output into a 1-bit decision value (e.g., a one-hot classification / decision vector). For example, the classification analysis processor 1050 can compare the indication 1045 or neural network output to a threshold value indicated by a configuration setting from the configuration table storage 1040 and generate a 1-bit indication according to the comparison (e.g., above the threshold value or below the threshold value). In some embodiments, the regression analysis processor 1068 includes conversion circuit 1070, multiplexer 1075, left shifter 1078, and right shifter 1080. Regression analysis can involve converting the indication 1045 from a signed integer to an unsigned integer by the conversion circuit 1070. The output of the conversion circuit 1070 can be scaled or adjusted by the shifters 1078, 1080. The multiplexer 1075 can be implemented to bypass the conversion circuit 1070. The multiplexer 1085 can select the output of the regression analysis processor 1068 or the output of the classification analysis processor 1050 and provide the selected output as output data 1095 to the pipeline bus 205 through the demultiplexer 1090A.

[0122] In one aspect, the output processing circuit 250 interprets the indication of the predicted network characteristics as a solution to a regression problem or a classification problem. The classification problem can include both multi-class and multi-label classification problems. For a regression problem, the output processing circuit 250 can optionally convert the output from a signed integer to an unsigned integer, and then shift by a preprogrammed value, which can then drive the output on a global bus. For a classification problem, the raw output of each QNN output layer (DO) neuron can be translated into a lb decision, thereby forming a decision vector, which can then drive the output of the decision vector on the pipeline bus 205. The neurons in the output layer can be separated into groups. For each group, if a neuron has the highest activation value of all neurons in the group, hardware can set the output to 1 for the neuron if the highest activation value is above a preprogrammed threshold or confidence threshold, and otherwise can set the output to 0. The preprogrammed threshold or confidence threshold can be changed based on the application or flow's tolerance for false positives or false negatives. If two neurons have the same raw activation value, a static priority can be enforced and the neuron with the lower index can be set to 1 and the other neuron to 0. The maximum number of neuron groups can be equal to the number of neurons of the output layer of the QNN that is implemented in hardware. Typically, for multi-class classification networks, neurons belonging to the same network can be placed in the same group, while for multi-label classification problems, neurons from the same network can be placed into separate groups with one neuron in each group.

[0123] In one aspect, the output data 1095 can be employed for various network applications. For example, the output data 1095 can be utilized for application classification. In one example, network characteristics of one or more packets can be obtained to determine or identify whether the one or more packets are for video streaming, email, browsing a website, etc. For example, the output data 1095 can be utilized for intrusion detection. In one example, network characteristics of one or more packets can be obtained to determine different types of DoS attacks. For example, the output data 1095 can be utilized for congestion prediction. In one example, network characteristics of one or more packets can be obtained to determine a particular traffic pattern that can indicate short-term congestion in a traffic manager.

[0124] Figure 11 A flow diagram illustrating a process 340 that adaptively generates output data 1095 (or output data 255) including predicted network characteristics, in accordance with an embodiment, is shown. In some embodiments, the process 340 is performed by the output processing circuit 250. In some embodiments, the process 340 is performed by other entities. In some embodiments, the process 340 includes more, fewer, or different steps than shown in Figure 11

[0125] ​In one approach, the output processing circuit 250 receives 1110 input attribute data 1008 containing packet attributes or flow attributes.

[0126] In one approach, the output processing circuit 250 determines 1120 a configuration index based on the packet attributes or flow attributes through a first table (e.g., a policy table). For example, the TCAM matching circuit 1010 can determine a policy index with the matching packet attributes or flow attributes, and provide the policy index to the policy table storage 1020. In response to the policy index, the policy table stored by the storage 1020 can determine, identify, or provide a corresponding configuration index 1028.

[0127] In one approach, the output processing circuit 1030 determines 1130 a configuration setting based on the configuration index 1028 through a second table (e.g., a configuration table stored by the storage 1040). The configuration setting can indicate a configuration of various circuits or components of the output processing circuit 1030 (e.g., the classification analysis processor 1050, the shifters 1078, 1080, the multiplexers 1075, 1085, etc.). In one aspect, the configuration setting can indicate or correspond to a type of output data 1095 (e.g., apply classification, network anomaly detection, network intrusion detection, predicted congestion, or a configuration value for a traffic manager to improve QoS, etc.).

[0128] In one approach, the output processing circuit 1030 determines 1140 whether to perform a classification analysis or a regression analysis. For example, the configuration table storage 1040 determines whether to perform a classification analysis or a regression analysis from the table stored by the configuration table storage 1040 according to the configuration index, and determines or obtains a configuration setting of the classification analysis processor 1050 or the regression analysis processor 1068.

[0129] In response to determining to apply a classification analysis, the classification analysis processor 1050 can generate 1150 output data 1095 through a classification. The classification analysis processor 1050 can use the indication 1045 of predicted network characteristics to compute a one-hot decision vector for a multi-label or multi-class classification problem. In response to determining to apply a regression analysis, the regression analysis processor 1068 can generate 1160 output data 1095 through a regression analysis. For example, the regression analysis processor 1068 can use the indication 1045 of predicted network characteristics to predict values for a multivariate regression problem. The multiplexer 1085 can select the output of the regression or the classification, and provide the selected output as output data 1095 through the demultiplexer 1090A to the pipeline bus 205.

[0130] In one aspect, by implementing TCAM matching circuitry, profile table storage or policy table storage and configuration table storage for different components (e.g., feature computation circuitry 220, input processing circuitry 230, reconfigurable neural network circuitry 240, output processing circuitry 250) can help the device implement a large number of quantized neural networks in order to obtain a large number of statistical features and compute a large number of predicted network characteristics in an efficient manner. For example, device 200 can support a large number of permutations (e.g., more than a million) of different configuration settings for different components (e.g., feature computation circuitry 220, input processing circuitry 230, reconfigurable neural network circuitry 240, output processing circuitry 250), where each component (e.g., feature computation circuitry 220, input processing circuitry 230, reconfigurable neural network circuitry 240, output processing circuitry 250) can implement a set of storage devices with less storage resources (e.g., 100kb each). Thus, device 400 can achieve area efficiency while supporting a large number of different computations for different neural networks.

[0131] It should be noted that certain paragraphs of the disclosure can refer to terms such as "first" and "second" for the purpose of identifying or differentiating one from another or from others, in conjunction with transmitting spatial streams, probe frames, responses, and subgroups of devices. These terms are not intended to relate to entities (e.g., first device and second device) only in time or according to an order, although in some cases these entities can include such a relationship. These terms also do not limit the number of possible entities that can operate within a system or environment. It should be understood that the systems described above can provide multiple components of any or each of those components and that these components can be provided on independent machines or on multiple machines in a distributed system in some embodiments. In addition, the systems and methods described above can be provided as one or more computer-readable programs or executable instructions embodied on or in one or more articles of manufacture, such as a soft disc, a hard disc, a CD-ROM, a flash memory, a PROM, a RAM, a ROM, or a magnetic tape. The programs can be implemented in any programming language, such as LISP, PERL, C, C++, C#, or in any byte code language such as JAVA. The software programs or executable instructions can be stored on or in one or more articles of manufacture as object code.

[0132] While the foregoing written description of the methods and systems enables a person skilled in the art to make and use embodiments, those skilled in the art will understand and appreciate the many modifications, combinations, and equivalents that exist within the scope of the specific embodiments, methods, and examples described herein. Thus, the present methods and systems should not be limited by the embodiments, methods, and examples described above, but rather should be defined by the scope and spirit of the disclosure.

Claims

1. A network device comprising: a feature extraction circuit configured to: receive a raw packet stream; obtain time statistics of a flow from a first packet attribute or a first flow attribute of the raw packet stream; and generate feature data including one or more statistical features based on the time statistics of the flow; an input processing circuit configured to scale the feature data to generate adjusted feature data; a control circuit configured to determine a configuration setting of a plurality of configuration settings for different packet attributes or flow attributes corresponding to the first packet attribute or the first flow attribute, the configuration setting configuring one or more components of a reconfigurable neural network circuit to adaptively implement different neural networks to perform different computations based on packet attributes or flow attributes; and wherein the reconfigurable neural network circuit is configured to adaptively implement the neural network as a different neural network based at least in part on the configuration setting determined to correspond to the first packet attribute or the first flow attribute and configured to perform computations corresponding to the neural network on the adjusted feature data to determine predicted network characteristics including network anomalies, network intrusions, predicted congestion, or configuration values for a traffic manager to improve quality of service.

2. The network device of claim 1, wherein the raw packet stream includes a first packet and a second packet, and wherein the feature extraction circuit is further configured to: obtain the time statistics of the flow from the first packet attribute or the first flow attribute of the first packet, generate the feature data including the one or more statistical features based on the time statistics of the flow, obtain additional time statistics of another flow from a second packet attribute or a second flow attribute of the second packet, and generate additional feature data including one or more additional statistical features based on the additional time statistics of the another flow.

3. The network device of claim 2, wherein the feature extraction circuit is implemented as a linear pipeline.

4. The network device of claim 1, wherein the first packet attribute is one of a packet source address, a packet destination address, a traffic class, or a flag.

5. The network device of claim 1, wherein the feature extraction circuit includes: a first control circuit including: a matching circuit to determine an index from a second packet attribute or a second flow attribute of the raw packet stream, a first storage device storing a first table of a set of configuration indices, the first storage device configured to provide a configuration index corresponding to the index, and a second storage device storing a second table of a plurality of configuration settings, the second storage device configured to provide a configuration setting corresponding to the configuration index.

6. The network device of claim 5, wherein the feature extraction circuit includes: a hash table circuit configured to: determine whether a database storing time statistics has a corresponding entry for the flow based on a hash key; send an indication to a first computation circuit indicating whether the database has a corresponding entry for the flow; in response to determining that the database does not have the corresponding entry for the flow, storing the hash key corresponding to a third packet attribute or a third flow attribute of the original packet flow; in response to determining that the database does not have the corresponding entry for the flow, sending an index of the database corresponding to the hash key to the first computing circuit and the database; and in response to determining that the database has the corresponding entry for the flow, sending the index of the database corresponding to the hash key to the first computing circuit and the database.

7. The network device of claim 6, wherein the first computing circuit is further configured to: in response to the indication indicating that the database does not have the corresponding entry for the flow, obtain the time statistics based on the first packet attribute or the first flow attribute; in response to the indication indicating that the database does not have the corresponding entry for the flow, store the time statistics through the corresponding entry associated with the index; perform a computation on the time statistics stored by the database in the corresponding entry associated with the index and the first packet attribute or the first flow attribute to obtain updated time statistics according to the configuration setting; store the updated time statistics at the corresponding entry of the database; and wherein the feature extraction circuit includes a second computing circuit configured to perform a computation on the updated time statistics to determine derived time statistics according to the configuration setting.

8. The network device of claim 7, wherein the feature extraction circuit includes: a precision adjustment circuit configured to perform a shift operation on the updated time statistics or the derived time statistics to produce the feature data including the one or more statistical features according to the configuration setting.

9. The network device of claim 6, wherein the feature extraction circuit is further configured to: identify an entry of the hash table circuit that is not accessed for a predetermined number of clock cycles or a predetermined number of packets; and remove the entry from the hash table circuit.

10. The network device of claim 6, wherein the feature extraction circuit is configured to: identify a set of time statistics stored by the database that have values that are outside of a predetermined range of values; and invalidate feature data computed based on the set of time statistics.

11. A network device comprising: a control circuit configured to determine an index based on a first packet attribute or a first flow attribute of an original packet flow; a multiplexer configured to select a second packet attribute or a second flow attribute of the original packet flow corresponding to a hash key according to the index; a hash table circuit configured to identify a flow according to the hash key; and one or more computing circuits configured to: obtain time statistics of the flow, and produce feature data including one or more statistical features based on the time statistics of the flow, wherein the control circuit is configured to determine a configuration setting of a plurality of configuration settings for different packet attributes or flow attributes that corresponds to the first packet attribute or the first flow attribute of the original packet flow, ​ The configuration setting configures one or more components of a reconfigurable neural network circuit to adaptively implement different neural networks to perform different computations based on packet attributes or flow attributes; and wherein the reconfigurable neural network circuit is configured to adaptively implement the neural network as a different neural network based at least in part on the configuration setting corresponding to the first packet attribute or the first flow attribute of the original packet flow and is configured to perform computations on the feature data corresponding to the neural network to predict network characteristics including network anomalies, network intrusions, predicted congestion, or configuration values for a traffic manager to improve quality of service.

12. The network device of claim 11, wherein the control circuitry includes: a first control circuit configured to determine a configuration index through a first table based on the first packet attribute or the first flow attribute of the original packet flow, and a second control circuit configured to determine the configuration setting corresponding to the configuration index through a second table.

13. The network device of claim 11, wherein the one or more computation circuits include a first computation circuit and a second computation circuit; and wherein the hash table circuitry is further configured to: determine whether a database storing time statistics has a corresponding entry for the flow based on a hash key, send an indication to the first computation circuit indicating whether the database has the corresponding entry for the flow, and in response to determining that the database (i) does have the corresponding entry for the flow or (ii) does not have the corresponding entry for the flow, respectively: (i) send the index of the database corresponding to the hash key to the first computation circuit and the database; or (ii-a) store the hash key in the database, and (ii-b) send an index of the database corresponding to the stored hash key to the first computation circuit and the database.

14. The network device of claim 13, wherein the first computation circuit is further configured to: in response to the indication indicating that the database does not have the corresponding entry for the flow, obtain the time statistics based on a third packet attribute or a third flow attribute, in response to the indication indicating that the database does not have the corresponding entry for the flow, store the time statistics through the corresponding entry associated with the index, perform computations on the time statistics stored by the database in the corresponding entry associated with the index and the third packet attribute or the third flow attribute according to the configuration setting to obtain updated time statistics, and store the updated time statistics at the corresponding entry of the database, and wherein the second computation circuit is to perform computations on the updated time statistics according to the configuration setting to determine derived time statistics.

15. The network device of claim 14, wherein the one or more computation circuits include: a first computation circuit configured to perform computations on the time statistics stored in the corresponding entry of the database according to the configuration setting to determine derived time statistics, and a second computation circuit configured to perform computations on the derived time statistics according to the configuration setting to determine predicted network characteristics. a precision adjustment circuit configured to perform a shift operation on the updated time statistics or the derived time statistics according to the configuration setting to generate the feature data including the one or more statistical features.

16. A method comprising: determining, by control circuitry, an index based on a first packet attribute or a first flow attribute of an original packet stream; selecting, by a multiplexer, a second packet attribute or a second flow attribute of the original packet stream corresponding to a hash key according to the index; determining, by the control circuitry, a configuration setting among a plurality of configuration settings for different packet attributes or flow attributes corresponding to the first packet attribute or the first flow attribute of the original packet stream, the configuration setting configuring one or more components of a reconfigurable neural network circuit to adaptively implement different neural networks to perform different computations based on packet attributes or flow attributes; identifying, by a hash table circuitry, a flow according to the hash key; obtaining, by one or more computation circuitries, time statistics of the flow; generating, by the one or more computation circuitries, feature data including one or more statistical features based on the time statistics of the flow according to the configuration setting; and adaptively implementing, by the reconfigurable neural network circuit, the neural network as different neural networks based at least in part on the configuration setting determined to correspond to the first packet attribute or the first flow attribute of the original packet stream and performing computations on the feature data corresponding to the neural network to predict network characteristics, the predicted network characteristics including network anomalies, network intrusions, predicted congestion, or configuration values for a traffic manager to improve quality of service.

17. The method of claim 16, wherein determining the configuration setting corresponding to the first packet attribute or the first flow attribute of the original packet stream includes: determining, by the control circuitry, a configuration index by a first table based on the first packet attribute or the first flow attribute of the original packet stream; and determining, by the control circuitry, the configuration setting corresponding to the configuration index by a second table.

18. The method of claim 16, wherein the one or more computation circuitries include a first computation circuitry and a second computation circuitry, and wherein identifying the flow includes: determining, by the hash table circuitry, whether a database storing time statistics has a corresponding entry for the flow based on the hash key, sending, by the hash table circuitry, an indication to the first computation circuitry indicating whether the database has the corresponding entry for the flow, and in response to determining that the database (i) does have the corresponding entry for the flow or (ii) does not have the corresponding entry for the flow, respectively: (i) sending, by the hash table circuitry, the index of the database corresponding to the hash key to the first computation circuitry and the database; or (ii-a) storing, by the hash table circuitry, the hash key in the database, and (ii-b) sending, by the hash table circuitry, an index of the database corresponding to the stored hash key to the first computation circuitry and the database.

19. The method of claim 18, ​ wherein obtaining, by the one or more computing circuits, the time statistics for the flow includes: in response to the indication indicating that the database does not have the corresponding entry for the flow, obtaining, by the first computing circuit, the time statistics based on a third packet attribute or a third flow attribute, and in response to the indication indicating that the database does not have the corresponding entry for the flow, storing, by the database, the time statistics through the corresponding entry associated with the index; and wherein generating, by the one or more computing circuits, feature data including one or more statistical features based on the time statistics for the flow includes: performing, by the first computing circuit according to the configuration settings, a computation on the time statistics stored by the database in the corresponding entry associated with the index and the third packet attribute or the third flow attribute to obtain updated time statistics, performing, by the second computing circuit according to the configuration settings, a computation on the updated time statistics to determine derived time statistics, and performing, by a precision adjustment circuit according to the configuration settings, a shift operation on the updated time statistics or the derived time statistics to generate the feature data including the one or more statistical features.

Citation Information

Patent Citations

  • Machine learning runtime library for neural network acceleration

    CN111247533A

  • System and method for assessing streaming video quality of experience in the presence of end-to-end encryption

    US20170093648A1