An inference engine system on a data plane bypass
By setting up an inference engine system on the data plane bypass and using RTC accelerator for group inference, the high latency and resource occupation problems of traditional network devices in NN inference are solved, and the NN inference effect with high throughput and low latency is achieved.
Patent Information
- Application Number
- CN202510329317.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-03-20
AI Technical Summary
Traditional network data plane devices have difficulty supporting neural network (NN) inference, especially with challenges in high throughput and low latency, and existing high-performance hardware has high latency and excessive resource utilization when deploying NNs in the network.
Set up an inference engine system on the data plane bypass, including input module, inference engine module and output module, use RTC accelerator to perform group inference, support multiple NN models, and store inference result sets through storage modules.
It realizes high throughput and low latency NN inference, is independent of the data plane, does not affect the original network performance, and has high flexibility and extensive NN model support.
Smart Images

Figure CN119849642B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data plane architectures, and particularly to an inference engine system on a data plane bypass. Background Art
[0002] In recent years, the Neural Network (NN) algorithm has become increasingly important in the network field and has become the core technology for mining data features, simulating complex relationships, and many key applications. The NN can deeply analyze various features in network data packets (such as packet length, protocol type, and port number, etc.), and use a multi-layer NN structure to abstract and extract features from the data layer by layer, so as to accurately identify complex traffic patterns. The NN algorithm has a very wide range of applications in the network field, such as traffic classification, intrusion detection, encrypted traffic classification, network faults, etc. Therefore, deploying the NN algorithm inside the network will bring many improvements and optimizations to network tasks.
[0003] However, the packet forwarding pipeline structure of the traditional network data plane does not support NN inference. First of all, NN inference requires a large amount of computing resources and memory, which are often difficult to meet in traditional network data plane devices. Secondly, the architecture of the data plane packet forwarding pipeline lacks flexibility in data processing logic, making it difficult to adapt to the complex and flexible computing steps required for NN inference. Therefore, the pipeline architecture of the traditional data plane has significant limitations in supporting NN inference and requires new hardware and architectures to make up for these deficiencies. On the other hand, current high-performance hardware devices (CPU / GPU) face challenges in supporting the deployment of NN in the network, especially in meeting the requirements of high throughput and low latency. When deploying network-oriented NN algorithms on traditional CPUs / GPUs, there are mainly the following two problems: one is that the per-packet inference latency is relatively high: when performing per-packet processing on CPUs and GPUs, the inference latency is usually in the millisecond level, which is difficult to match the microsecond-level latency of traffic forwarding; the other is that 99% of the latency is wasted on traffic migration and data batching: specifically, the processes of transferring data packets from the network interface to the CPU or GPU memory and organizing the data into batches suitable for NN processing occupy most of the latency time, thus further exacerbating the overall latency of the system.
[0004] In view of this, it is an urgent technical problem for those skilled in the art to provide an inference engine system with high throughput, low latency, high flexibility, and that does not affect the functions and performance of the original network data plane, which is set on the data plane bypass. Summary of the Invention
[0005] To solve the above technical problems, the object of the present invention is to provide an inference engine system on a data plane bypass, which supports NN inference and simultaneously has the characteristics of high throughput, low latency, high flexibility, and does not affect the functions and performance of the original network data plane.
[0006] The object of the present invention is to provide an inference engine system on a data plane bypass;
[0007] The technical solution provided by the present invention is as follows:
[0008] An inference engine system on a data plane bypass, comprising: an input module, an inference engine module, and an output module;
[0009] The input module and the output module are respectively connected to the data plane;
[0010] The inference engine module is respectively connected to the input module and the output module;
[0011] The input module is used to copy and transmit the traffic mirror of the data plane;
[0012] The inference engine module is used to perform inference based on the traffic mirror to obtain an inference result set;
[0013] The output module is used to output the inference result set to the data plane.
[0014] Preferably, the inference engine module includes a plurality of RTC accelerators;
[0015] The RTC accelerator is used to perform grouped inference on the traffic mirror to obtain inference results, and integrate the inference results to obtain the inference result set.
[0016] Preferably, the RTC accelerator includes: a grouping module, an inference instruction module, an inference processing module, and an integration module;
[0017] The grouping module is used to group the traffic mirror according to a preset rule;
[0018] The inference instruction module is used to send a preset inference instruction;
[0019] The inference processing module is used to perform corresponding inference operations on each group of traffic mirrors according to the inference instruction to obtain grouped inference results;
[0020] The integration module is used to integrate the grouped inference results to obtain the inference result set.
[0021] Preferably, the RTC accelerator further includes: a model construction module and a model training module;
[0022] A model construction module for constructing an initial neural network model;
[0023] A model training module for pre-training the initial neural network model to obtain the weights of the pre-trained neural network model;
[0024] The inference processing module is specifically configured to perform corresponding inference operations on each group of traffic mirrors through the weights of the pre-trained neural network model to obtain the grouped inference results.
[0025] Preferably, the inference processing module includes: a vector operation unit, a vector accumulator, and a non-linear activation component;
[0026] The vector accumulator is respectively connected to the vector operation unit and the non-linear activation component;
[0027] The vector operation unit is configured to perform matrix-vector multiplication operations on the traffic mirror to obtain a first intermediate calculation result;
[0028] The vector accumulator is configured to accumulate the first intermediate calculation result to obtain a second intermediate calculation result;
[0029] The non-linear activation component is configured to perform non-linear activation operations on the second intermediate calculation result to obtain the grouped inference results.
[0030] Preferably, the RTC accelerator further includes: an instruction cache module, a weight cache module, and a register cache module;
[0031] The instruction cache module is connected to the inference instruction module;
[0032] The weight cache module is respectively connected to the inference processing module and the model training module;
[0033] The register cache module is respectively connected to the vector operation unit, the vector accumulator, and the non-linear activation component;
[0034] The instruction cache module is configured to store the inference instructions;
[0035] The weight cache module is configured to store the weights of the pre-trained neural network model;
[0036] The register cache module is configured to store the first intermediate calculation result, the second intermediate calculation result, and the grouped inference results.
[0037] Preferably, the inference instructions include vector operation instructions, weight cache instructions, and register file read / write instructions;
[0038] The vector operation instruction is used to control the vector operation unit to perform calculations;
[0039] The weight cache instruction is used to store and read the weights of the pre-trained neural network model;
[0040] The register file read / write instruction is used to control the register cache module to perform read and write operations.
[0041] Preferably, the vector operation unit includes multiple SIMD channels;
[0042] The SIMD channel includes multiple dot product units;
[0043] The dot product unit includes multiple multipliers and an adder tree;
[0044] The adder tree includes multiple adders, which are used to accumulate the calculation results of the multipliers to obtain the first intermediate calculation result.
[0045] Preferably, an inference engine system on a data plane bypass further includes: a storage module;
[0046] The storage module is respectively connected to the inference engine module and the output module;
[0047] The storage module is used to store the inference result set.
[0048] Preferably, the storage module is a flow lookup table;
[0049] The flow lookup table is specifically used to store the inference result set in the form of a five-tuple hash value.
[0050] The present invention discloses an inference engine system on a data plane bypass, including: an input module, an inference engine module, and an output module; during operation, the input module copies the data plane traffic mirror image and transmits the traffic mirror image to the inference engine module; the inference engine module performs inference on the data plane traffic mirror image to obtain an inference result set; the output module outputs the inference result set to the data plane; by setting up an inference engine system on the data plane bypass and adopting a working mode of performing inference independently of the data plane, the present invention achieves the beneficial effects of being able to flexibly support multiple NN inference models, as well as achieving high throughput, low latency, and not affecting the functions and performance of the original network data plane. Description of the Drawings
[0051] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0052] Figure 1 It is a schematic structural diagram of an inference engine system on a data plane bypass in an embodiment of the present invention;
[0053] Figure 2 It is a schematic architecture diagram of an inference engine system deployed on a network card data plane bypass in an embodiment of the present invention;
[0054] Figure 3 It is a schematic structural diagram of an RTC accelerator in an embodiment of the present invention. Detailed implementation manners
[0055] To enable those skilled in the art to better understand the technical solutions in the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by the present application.
[0056] As Figure 1 shown, an embodiment of the present invention provides an inference engine system on a data plane bypass, including: an input module, an inference engine module, and an output module;
[0057] The input module and the output module are respectively connected to the data plane;
[0058] The inference engine module is respectively connected to the input module and the output module;
[0059] The input module is used to copy and transmit the traffic mirror of the data plane;
[0060] The inference engine module is used to perform inference based on the traffic mirror to obtain an inference result set;
[0061] The output module is used to output the inference result set to the data plane.
[0062] During the actual working process, the input module is connected to the traffic inlet of the data plane. The input module copies the traffic mirror of the data plane and transmits it to the inference engine module. The inference engine module performs inferences based on the traffic mirror of the data plane to obtain an inference result set, and the output module outputs the inference result set to the traffic outlet of the data plane. Based on the design architecture of the data plane bypass, the inference engine module is independent of the data plane and uses the traffic mirror of the data plane for inferences. The inference process will not cause additional latency to the operation of the data plane and will not affect the functions and performance of the data plane. At the same time, the inference engine module is independent of the data plane and is no longer restricted by the data plane pipeline workflow, having higher design flexibility to support multiple NN inference models and realize a hardware design space with high throughput.
[0063] As Figure 2 shown, as an embodiment, the embodiment of the present invention provides an inference engine system deployed on the bypass of the network card data plane. Network card data plane setting and working principle: The network card is set to the simple forwarding mode, that is, when traffic enters from port 0 of the network card and enters the receive queue of the network card through the inlet, it will be immediately transferred to the transmit queue and outlet of the network card and output from port 1 of the network card. This operation realizes the most common packet forwarding operation in the data plane. The inference engine system is integrated between the inlet and outlet of the network card, receives the mirrored traffic input by the data plane for inferences, and outputs the inference result set to the data plane.
[0064] Preferably, the inference engine module includes multiple RTC accelerators;
[0065] The RTC accelerator is used to group, infer, and integrate the traffic mirror in sequence to obtain an inference result set.
[0066] It should be noted that the RTC (Run-to-complete) accelerator is a programmable RTC accelerator; the RTC accelerator supports multiple NN model operators, such as general matrix-vector multiplication (GEMV) and non-linear activation functions; the RTC accelerator in the present invention supports inferences based on the original bytes of the data plane traffic; the RTC accelerator also supports cyclic calculations on the grouped inference results.
[0067] Preferably, the RTC accelerator includes: an inference instruction module, an inference processing module, and a register cache module;
[0068] The inference instruction module is used to send inference instructions;
[0069] The inference processing module is used to execute corresponding inference operations according to the inference instructions to obtain inference results;
[0070] The register cache module is used to store the intermediate calculation results generated by the inference processing module.
[0071] It should be noted that the inference instruction module may include multiple instruction slots, each of which can store an instruction, and the inference instruction module can send multiple instructions simultaneously.
[0072] Preferably, the RTC accelerator further includes: a model construction module and a model training module;
[0073] The model construction module is used to construct an initial neural network model;
[0074] The model training module is used to pre-train the initial neural network model to obtain the weights of the pre-trained neural network model;
[0075] The inference processing module is specifically configured to perform corresponding inference operations on each group of traffic mirrors through the weights of the pre-trained neural network model to obtain grouped inference results.
[0076] It should be noted that the initial neural network model can be a perceptron model, a BP neural network, and a convolutional neural network (CNN). The model training module pre-trains the initial neural network model according to a preset scenario to obtain the corresponding weights of the pre-trained neural network model.
[0077] In the actual application process, the model construction module and the model training module can also be set outside the system of the present invention, that is, the neural network model construction and the neural network model pre-training are performed outside the system, and the weights of the pre-trained model are written into this system. This implementation mode has the same effect as setting the model construction module and the model training module inside the system of the present invention.
[0078] As an embodiment, a three-layer perceptron model (MLP) is constructed and pre-trained. The specific construction and training steps are as follows:
[0079] Step 1.1, prepare a traffic data set in pcap format. Open-source traffic data sets such as ISCX-VPN-NonVPN-2016 and USTC-TFC can be selected;
[0080] Step 1.2, divide the data set into a training set and a test set. The ratio of the training set to the test set is generally 7:3 or 8:2;
[0081] Step 1.3, clean the data set to remove traffic with non-TCP / UDP protocols (remove protocols such as ARP and ICMP);
[0082] Step 1.4, perform fine-grained preprocessing on the cleaned data, remove the source / destination addresses of the Ethernet header and IP header of the packet, retain the remaining fields of the IP header, the TCP / UDP header, and the front part of the payload. Each traffic data packet finally retains 32 bytes of content;
[0083] Step 1.5, the model construction module implements a three-layer MLP model using an artificial intelligence framework such as Pytorch. The model takes a (1×32)-dimensional vector as input, corresponding to 32 bytes of raw bytes intercepted from the packet, and outputs a (1×c)-dimensional vector, where c is the number of classification results of the inference. Taking the inference results of malicious traffic / benign traffic as an example, in this embodiment, c = 2. The output dimensions of each layer of the three-layer MLP model are 64, 32, and c respectively, and the ReLU function is used as the non-linear activation function between the three MLP layers. Step 1.5 can be further split into the following steps: Step 1.5.1, intercept 20 bytes of content (excluding source and destination addresses) in the IP header of the traffic data packet and the first 12 bytes in the payload;
[0084] Step 1.5.2, send the intercepted raw bytes of the data packet into the neural network model for inference of the neural network model;
[0085] Step 1.5.3, classify according to the inference results of the neural network, such as classifying into malicious traffic / benign traffic, or further classifying the benign traffic, such as multimedia traffic, email traffic, etc.;
[0086] Step 1.6, use the preprocessed dataset to train the MLP model on the Pytorch framework. The loss function can be the cross-entropy function, mean square error function, mean absolute error, or mean absolute percentage error, and the AdamW optimizer in the Pytorch framework is used as the optimizer to obtain the pre-trained neural network model;
[0087] Step 1.7, after the training is completed, perform inference on the test set, and the model training module obtains the final neural network model weights of the MLP model;
[0088] Step 1.8, export and save the neural network model weights in the final MLP model.
[0089] Preferably, the inference processing module includes: a vector operation unit, a vector accumulator, and a non-linear activation component;
[0090] The vector accumulator is respectively connected to the vector operation unit and the non-linear activation component;
[0091] The vector operation unit is used to perform matrix-vector multiplication on the traffic mirror to obtain a first intermediate calculation result;
[0092] The vector accumulator is used to accumulate the first intermediate calculation result to obtain a second intermediate calculation result;
[0093] The non-linear activation component is used to perform a non-linear activation operation on the second intermediate calculation result to obtain a grouped inference result.
[0094] It should be noted that the nonlinear activation component can use ReLU function, Sigmoid function, Tanh, LeakyReLU function, ELU function, GELU function or Softmax function as the nonlinear activation function to perform nonlinear activation operation; the nonlinear activation module can be implemented by hardware lookup table and software lookup table, etc.
[0095] Preferably, the RTC accelerator further includes: an instruction cache module, a weight cache module and a register cache module;
[0096] The instruction cache module is connected to the inference instruction module;
[0097] The weight cache module is connected to the inference processing module and the model training module respectively;
[0098] The register cache module is respectively connected with the vector operation unit, the vector accumulator and the nonlinear activation component;
[0099] An instruction cache module, used to store inference instructions;
[0100] Weight cache module, used to store the pre-trained neural network model weights;
[0101] The register cache module is used to store the first intermediate calculation result, the second intermediate calculation result and the grouping reasoning result.
[0102] In the actual working process, before the inference engine module performs inference, the weights of the neural network model pre-trained in the model training module are written into the weight cache module; the register cache module includes multiple data registers.
[0103] Preferably, the inference instructions include vector operation instructions, weight cache instructions, and register file read and write instructions;
[0104] Vector operation instructions are used to control the vector operation unit to perform calculations;
[0105] Weight cache instructions, used to store and read pre-trained neural network model weights;
[0106] Register file read and write instructions are used to control the register cache module to read and write.
[0107] In the actual working process, the inference engine system on the data plane bypass of the present invention also includes a corresponding instruction set, taking the inference instruction module including three instruction slots as an example, specifically including the following instructions:
[0108] Table 1 Inference instruction table
[0109]
[0110] Preferably, the vector operation unit includes multiple SIMD channels;
[0111] The SIMD channels include multiple dot product units;
[0112] The dot product unit includes multiple multipliers and an adder tree;
[0113] The adder tree includes multiple adders for accumulating the calculation results of the multipliers to obtain a first intermediate calculation result.
[0114] It should be noted that the vector operation unit may be composed of multiple SIMD (Single-instruction -multiple-data) channels, and the SIMD channels are stacked by m dot product units. There are n multipliers placed in parallel in the dot product unit, which can perform multiple multiplication operations simultaneously; the dot product unit is connected to the adder tree for accumulating the operation results of the multipliers into the output result of the dot product operation, and the adder tree contains multiple levels of parallel adders; the dot product units in the SIMD channels share the same input vector and operate on the relevant columns of the matrix; the vector operation unit can complete the product of a vector with a dimension of (1, n) and a matrix with a dimension of (n, m) in one startup; the number of SIMD channels, the number of dot product units, and the number of multipliers in the vector operation unit can be determined according to the actual inference scenario.
[0115] It should be noted that the first intermediate calculation result calculated by the vector operation unit is the product of the input vector and each column (or sub-column) of the matrix. The vector accumulator accumulates the first intermediate calculation result calculated by the vector operation unit to obtain the final result. Through this operation, the first intermediate calculation results of the operations of the vector and the matrix sub-columns can be combined to support the block vector-matrix multiplication operation.
[0116] As an implementation, in the vector operation unit, there are 4 SIMD channels, each channel has 8 vector dot product units, and each vector dot product unit has 8 multipliers. That is, the vector operation unit in the embodiment of the present invention can complete the multiplication of a vector with a dimension of (1×32) and a matrix with a dimension of (32×8) at one time; the register cache module includes 32 data registers, and the bit width of each register is 32 bytes.
[0117] Preferably, an inference engine system on a data plane bypass further includes: a storage module;
[0118] The storage module is respectively connected to the inference engine module and the output module;
[0119] The storage module is used to store the inference result set.
[0120] Preferably, the storage module is a traffic query table;
[0121] The traffic query table is specifically used to store the inference result set in the form of a five-tuple hash value.
[0122] It should be noted that the five-tuple hash value is obtained by performing a hash operation on five key elements in network communication: source IP address, destination IP address, source port number, destination port number, and protocol type, and is used for quickly searching and matching network connections.
[0123] During the actual working process, when the data plane queries the inference result, the output module reads the corresponding inference result set from the traffic query table through the five-tuple hash value, and then returns the inference result set to the data plane through the output module. The data plane performs corresponding subsequent operations according to the inference result set.
[0124] As a specific implementation manner, the inference engine system of the present invention is an inference engine system deployed on the bypass of the network card data plane. The inference engine is based on a three-layer MLP model, and the weight cache module has written the weights of the pre-trained three-layer MLP model. The pre-training process has been described above and will not be elaborated here. In this embodiment, in the vector operation unit, there are 4 SIMD channels, each channel has 8 vector dot product units, and each vector dot product unit has 8 multipliers, that is, the vector operation unit in the embodiment of the present invention can complete the multiplication of a (1×32)-dimensional vector and a (32×8)-dimensional matrix at one time; the register cache module includes 32 data registers, and the bit width of each register is 32 bytes. The inference process of the inference engine system of the present invention is specifically as follows:
[0125] Step 2.1, the input module inputs a 32-byte traffic mirroring original packet intercepted according to a preset rule to the inference engine module;
[0126] Step 2.2, the inference engine module performs the first-layer inference in the MLP model. The inference processing module receives the 32-byte input from the input module, fetches the corresponding MLP model weights from the weight cache module, performs a matrix-vector multiplication operation, and outputs a (1×64)-dimensional vector (this operation corresponds to a (1×32)×(32×64) matrix-vector multiplication), and then stores it in the register cache module after passing through the ReLU non-linear activation operation of the non-linear activation module;
[0127] Step 2.3, the inference engine module performs the second-layer inference in the MLP model. The inference processing module reads the inference result of the first layer of the MLP model stored in the register cache module, fetches the corresponding MLP model weights from the weight cache module, performs a matrix-vector multiplication operation, and outputs a vector of dimension (1×32) (this operation corresponds to a matrix-vector multiplication of (1×64)×(64×32)), and then stores it in the register cache module after passing through the ReLU non-linear activation operation of the non-linear activation module;
[0128] Step 2.4, the inference engine module performs the third-layer inference in the MLP model. The inference processing module reads the inference result of the second layer of the MLP model stored in the register cache module, fetches the corresponding MLP model weights from the weight cache module, performs a matrix-vector multiplication operation, and outputs a vector of dimension (1×2) (this operation corresponds to a matrix-vector multiplication of (1×32)×(32×2)), and then obtains the grouped inference result after passing through the ReLU non-linear activation operation of the non-linear activation module and stores it in the register cache module;
[0129] Step 2.5, the integration module integrates the grouped inference results in the register cache module and writes them into the traffic query table.
[0130] In the embodiments provided in the present application, it should be understood that the disclosed method and system can be implemented in other ways. The system embodiments described above are only illustrative. For example, the division of modules is only a logical function division, and there can be other division methods in actual implementation. For example, multiple modules or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed with each other can be through some interfaces, and the indirect coupling or communication connection of devices or modules can be electrical, mechanical, or other forms.
[0131] In addition, each functional module in the embodiments of the present invention can be all integrated in one processor, or each module can be separately used as a device, or two or more modules can be integrated in one device; each functional module in the embodiments of the present invention can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.
[0132] Those of ordinary skill in the art will understand that all or part of the steps to implement the above method embodiments can be completed through program instructions and related hardware. The foregoing program instructions can be stored in a computer-readable storage medium. When the program instructions are executed, they perform the steps including the above method embodiments; and the foregoing storage medium includes: various media such as removable storage devices, read-only memory (ROM), magnetic disks, or optical discs that can store program code.
[0133] It should be understood that in this application, if the terms "system", "device", "unit", and / or "module" are used, they are only a way to distinguish different components, elements, parts, portions, or assemblies at different levels. However, if other words can achieve the same purpose, the term can be replaced by other expressions.
[0134] As shown in this application and the claims, unless the context clearly indicates an exception, words such as "a", "an", "one", and / or "the" are not specifically singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of the clearly identified steps and elements, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements. An element defined by the statement "comprising one..." does not exclude the existence of another identical element in the process, method, article, or device including the element.
[0135] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of this application, the meanings of "a plurality of" and "several" are two or more, unless otherwise specifically and clearly defined.
[0136] If a flowchart is used in this application, the flowchart is used to illustrate the operations performed by the system according to the embodiments of this application. It should be understood that the operations before or after may not be executed precisely in sequence. On the contrary, the operations can be executed in reverse order or simultaneously. At the same time, other operations can also be added to these processes, or one or several operations can be removed from these processes.
[0137] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An inference engine system on a data plane bypass, characterized in that It includes: an input module, an inference engine module, and an output module; The input module and the output module are respectively connected to the data plane; The inference engine module is respectively connected to the input module and the output module; The input module is used to copy and transmit the traffic mirror of the data plane; The inference engine module includes multiple RTC accelerators, which are used to perform inference based on the traffic mirror to obtain an inference result set; wherein, the RTC accelerator is a programmable accelerator that supports multiple neural network model operators, and is used to group, infer, and integrate the traffic mirror in sequence to obtain the inference result set; The output module is used to output the inference result set to the data plane; The RTC accelerator includes: a grouping module, an inference instruction module, an inference processing module, and an integration module; The grouping module is used to group the traffic mirror according to a preset rule; The inference instruction module is used to send a preset inference instruction; The inference processing module is used to perform corresponding inference operations on each group of traffic mirrors through the weights of the pre-trained neural network model to obtain grouped inference results; The integration module is used to integrate the grouped inference results to obtain the inference result set; The inference processing module includes: a vector operation unit, a vector accumulator, and a non-linear activation component; The vector accumulator is respectively connected to the vector operation unit and the non-linear activation component; The vector operation unit is used to perform matrix-vector multiplication on the traffic mirror to obtain a first intermediate calculation result; The vector accumulator is used to accumulate the first intermediate calculation result to obtain a second intermediate calculation result; The non-linear activation component is used to perform a non-linear activation operation on the second intermediate calculation result to obtain the grouped inference result.
2. The inference engine system on the data plane bypass as described in claim 1, wherein The RTC accelerator further includes: a model construction module and a model training module; The model construction module is used to construct an initial neural network model; The model training module is used to pre-train the initial neural network model to obtain the weights of the pre-trained neural network model.
3. The inference engine system on the data plane bypass as described in claim 2, wherein The RTC accelerator further includes: an instruction cache module, a weight cache module, and a register cache module; The instruction cache module is connected to the inference instruction module; The weight cache module is respectively connected to the inference processing module and the model training module; The register cache module is respectively connected to the vector operation unit, the vector accumulator, and the non-linear activation component; The instruction cache module is used to store the inference instruction; The weight cache module is used to store the weights of the pre-trained neural network model; The register cache module is used to store the first intermediate calculation result, the second intermediate calculation result, and the grouped inference result.
4. The inference engine system on the data plane bypass as claimed in claim 3, wherein The inference instruction includes a vector operation instruction, a weight cache instruction, and a register file read / write instruction; The vector operation instruction is used to control the vector operation unit to perform calculations; The weight cache instruction is used to store and read the weights of the pre-trained neural network model; The register file read / write instruction is used to control the register cache module to perform read / write operations.
5. The inference engine system on the data plane bypass as described in claim 1, wherein The vector operation unit includes multiple SIMD channels; The SIMD channel includes multiple dot product units; The dot product unit includes multiple multipliers and an adder tree; The adder tree includes multiple adders for accumulating the calculation results of the multipliers to obtain the first intermediate calculation result.
6. The inference engine system on the data plane bypass as claimed in claim 1, wherein, It further includes: a storage module; The storage module is respectively connected to the inference engine module and the output module; The storage module is used to store the inference result set.
7. The inference engine system on the data plane bypass according to claim 6, characterized in that The storage module is a traffic query table; The traffic query table is specifically used to store the inference result set in the form of a five-tuple hash value.
Citation Information
Patent Citations
Real-time flow detection system and method compatible with multiple inference engines
CN114189368A