Efficient aggregation method and device for programmable switch
By efficiently deploying lookup tables in programmable switches, supporting larger-scale neural networks and abstracting them into efficient aggregation operations, the problems of limited scale and long aggregation time in the prior art are solved, and high-speed inference tasks are achieved.
Patent Information
- Application Number
- CN202411937200.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-23
AI Technical Summary
The neural networks supported by the prior art on programmable switches are limited in scale and have a small scope of application, which makes it a long time to aggregate data and makes it difficult to achieve high-speed inference tasks.
By efficiently deploying at least one first lookup table in a programmable switch, a larger-scale neural network is supported, abstract matrix multiplication and convolution operations are efficient aggregation operations, reducing model inference time.
It realizes supporting larger-scale neural networks on programmable switches, with a wide range of application, while reducing model inference time and achieving high-speed inference tasks.
Smart Images

Figure CN120030370A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of next generation Internet architecture, network intelligence, and ultra-high-speed programmable data plane technology, and in particular to a high-efficiency aggregation method and device for programmable switches. Background Art
[0002] In recent years, with the rapid development of artificial intelligence, more and more fields are trying to use artificial intelligence to give new capabilities to related fields, such as more accurate and intelligent identification of malicious traffic based on artificial intelligence technology. With the diversification of Internet application types and the growth of dynamic demand for cyberspace security, more and more workers are trying to use artificial intelligence to empower the network field, or use the network to support artificial intelligence. However, the existing solutions for implementing neural network models in the data plane are limited by the scarce resources on programmable switches, the scale of neural networks that can be supported is limited, the scope of application is small, and it takes more time to aggregate data in programmable switches. Summary of the invention
[0003] The present invention aims to solve one of the technical problems in the related art at least to a certain extent.
[0004] To this end, an object of the present invention is to propose an efficient aggregation method for programmable switches. The method efficiently deploys at least one first lookup table in a programmable switch so that the programmable switch can support larger-scale neural networks and has a wide range of applications. At the same time, efficient aggregation is performed based on the programmable switch to obtain inference results, so that matrix multiplication and convolution operations in the neural network model are abstracted into efficient aggregation operations in the programmable switch, thereby reducing the model inference time and achieving high-speed inference tasks.
[0005] Another object of the present invention is to provide an efficient aggregation device for programmable switches.
[0006] To achieve the above object, an embodiment of the present invention provides an efficient aggregation method for a programmable switch, including:
[0007] Acquire a trained neural network model and a training data set used to train the neural network model;
[0008] Based on the training data in the training data set, generating a first binary clustering tree corresponding to the neural network model;
[0009] Determining convolution kernel parameters / fully connected layer parameters in the trained neural network model;
[0010] Based on the first binary clustering tree and the convolution kernel parameters / fully connected layer parameters, obtaining at least one corresponding weighted value vector;
[0011] The at least one weighted value vector is stored in the corresponding at least one first lookup table, and the at least one first lookup table is efficiently deployed in a programmable switch, and an inference result is obtained by performing efficient aggregation based on the programmable switch.
[0012] The efficient aggregation method for programmable switches according to the embodiment of the present invention may also have the following additional technical features:
[0013] Furthermore, the generating a first binary clustering tree corresponding to the neural network model based on the training data in the training data set includes:
[0014] determining clustering parameters for clustering;
[0015] Based on the clustering parameters and the minimum square error, clustering the training data to obtain a clustering result, wherein the clustering result includes average values of different dimensions of the training data in each cluster;
[0016] Based on the average values of different dimensions of the training data in each cluster, a first binary clustering tree corresponding to the neural network model is generated.
[0017] Furthermore, the obtaining of at least one corresponding weighted value vector based on the first binary clustering tree and the convolution kernel parameters / fully connected layer parameters includes:
[0018] Multiplying the node parameters in the first binary clustering tree by the convolution kernel parameters / fully connected layer parameters to obtain at least one corresponding second binary clustering tree;
[0019] Based on the at least one second binary clustering tree, at least one corresponding weight value vector is obtained.
[0020] Further, storing the at least one weighted value vector into the corresponding at least one first lookup table, and efficiently deploying the at least one first lookup table in a programmable switch, and performing efficient aggregation based on the programmable switch to obtain an inference result, includes:
[0021] Storing the at least one weighted value vector in a corresponding first lookup table based on a target data format, wherein all weighted value vectors in the second binary clustering tree are stored in the same first lookup table;
[0022] Processing the plurality of first lookup tables based on the dimension of the input data, and deploying the processed second lookup tables and / or index tables in the programmable switch;
[0023] The inference result is obtained by performing efficient aggregation based on the programmable switch.
[0024] Further, the processing of the plurality of first lookup tables based on the dimension of the input data and deploying the processed second lookup tables and / or index tables in the programmable switch includes:
[0025] If the dimension of the input data is less than or equal to a preset threshold, aggregating the data in a first preset number table among the plurality of first lookup tables, and deploying a second lookup table after aggregating the data in the programmable switch;
[0026] If the dimension of the input data is greater than a preset threshold, an index table is generated based on the multiple first lookup tables, wherein the index in the index table is the position of the data in the first lookup table in the corresponding second binary clustering tree;
[0027] Aggregate the data in the second preset number table of the multiple first lookup tables, and deploy the second lookup table after the data aggregation and the index table in the programmable switch.
[0028] Furthermore, efficient aggregation based on the programmable switch is performed to obtain inference results, including:
[0029] The programmable switch receives a data packet from a network port, dynamically parses the data packet according to a specific protocol through a parser of a programmable data plane, and matches the data packet with the second lookup table and / or the index table through a ternary content addressable memory to obtain a corresponding plurality of output vectors;
[0030] The multiple output vectors are added and aggregated to obtain an inference result corresponding to the data packet.
[0031] Further, the matching of the data packet with the second lookup table and / or the index table through a ternary content addressable memory to obtain a corresponding plurality of output vectors includes:
[0032] If the dimension of the input data is less than or equal to a preset threshold, the ternary content addressable memory matches the data packet with the TCAM table and the second lookup table to obtain a corresponding plurality of output vectors;
[0033] If the dimension of the input data is greater than a preset threshold, the ternary content addressable memory matches the data packet with the TCAM table and the index table to obtain a corresponding index code;
[0034] The index code is matched with the second lookup table to obtain corresponding multiple output vectors.
[0035] To achieve the above object, another embodiment of the present invention provides an efficient aggregation device for a programmable switch, the device comprising:
[0036] An acquisition module, used to acquire a trained neural network model and a training data set used to train the neural network model;
[0037] A generating module, configured to generate a first binary clustering tree corresponding to the neural network model based on the training data in the training data set;
[0038] A determination module, used to determine the convolution kernel parameters / fully connected layer parameters in the trained neural network model;
[0039] A processing module, configured to obtain at least one corresponding weighted value vector based on the first binary clustering tree and the convolution kernel parameters / fully connected layer parameters;
[0040] An aggregation module is used to store the at least one weighted value vector into the corresponding at least one first lookup table, and efficiently deploy the at least one first lookup table in a programmable switch, and obtain an inference result by performing efficient aggregation based on the programmable switch.
[0041] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The above and / or additional aspects and advantages of the present invention will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0043] Figure 1 A flowchart of an efficient aggregation method for a programmable switch according to an embodiment of the present invention;
[0044] Figure 2 A schematic diagram of performing addition on different data formats according to an embodiment of the present invention;
[0045] Figure 3 The figure is a schematic diagram of the structure of an efficient aggregation device for a programmable switch according to an embodiment of the present invention. DETAILED DESCRIPTION
[0046] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and should not be construed as limiting the present invention.
[0047] The following describes an efficient aggregation method and device for a programmable switch according to an embodiment of the present invention with reference to the accompanying drawings.
[0048] First, an efficient aggregation method for programmable switches proposed according to an embodiment of the present invention will be described with reference to the accompanying drawings.
[0049] Figure 1 The figure is a flow chart of an efficient aggregation method for a programmable switch according to an embodiment of the present invention.
[0050] like Figure 1 As shown, the efficient aggregation method for programmable switches includes the following steps:
[0051] Step S1, obtaining a trained neural network model and a training data set for training the neural network model;
[0052] In one embodiment of the present invention, the trained neural network model includes matrix multiplication and convolutional layers and / or fully connected layers.
[0053] Step S2, generating a first binary clustering tree corresponding to the neural network model based on the training data in the training data set;
[0054] Among them, in one embodiment of the present invention, after obtaining the training data set for training the neural network model through the above steps, a first binary clustering tree corresponding to the neural network model can be generated based on the training data in the training data set, so as to subsequently determine the weighted value vector of the neural network model.
[0055] And, in one embodiment of the present invention, unlike the traditional clustering method, the clustering method adopted by the present invention is to use feature thresholds to divide data points without relying on traditional geometric distances. Specifically, the similarity can be determined by checking whether the feature values of the data points are on the same side of the predefined threshold, and efficient processing can be achieved through simple comparison or range matching operations, thereby avoiding complex distance calculations.
[0056] Specifically, in one embodiment of the present invention, the method for generating a first binary clustering tree corresponding to a neural network model based on training data in a training data set may include the following steps:
[0057] Step S21, determining clustering parameters for clustering;
[0058] Step S22, clustering the training data based on the clustering parameters and the minimum square error to obtain a clustering result, wherein the clustering result includes the average values of different dimensions of the training data in each cluster;
[0059] Step S23, generating a first binary clustering tree corresponding to the neural network model based on the average values of different dimensions of the training data in each cluster.
[0060] In one embodiment of the present invention, the training data includes multiple feature dimensions, and the clustering parameters may include a target feature dimension for clustering and a corresponding segmentation threshold for segmenting the data.
[0061] And, in one embodiment of the present invention, the least square error (SSE) can be used as a division criterion, and the training data is clustered by the above clustering parameters and the least square error to obtain a clustering result, thereby ensuring that the data points in each cluster are as close to their centroid as possible and maximizing the similarity within the cluster. The clustering result includes the average values of different dimensions of the training data in each cluster.
[0062] Furthermore, in one embodiment of the present invention, the centroid of each cluster in the above clustering result includes the average value corresponding to the target dimension in the training data in the cluster, and based on the average values of different dimensions of the training data in the centroid, a first binary clustering tree corresponding to the neural network model is generated, and each node of the first binary clustering tree may also include a partition threshold of the cluster corresponding to the centroid. Specifically, the corresponding first binary clustering tree can be generated according to the average value in the centroid of each cluster in the clustering result, wherein the method for generating the first binary clustering tree is the same as the method for generating a binary tree in the prior art, and the embodiments of the present disclosure will not be described in detail here.
[0063] Step S3, determining the convolution kernel parameters / fully connected layer parameters in the trained neural network model;
[0064] In one embodiment of the present invention, if the trained neural network model includes a convolutional layer, the convolution kernel parameters in the trained neural network model are determined; if the trained neural network model includes a fully connected layer, the fully connected layer parameters in the trained neural network model are determined.
[0065] Step S4, obtaining at least one corresponding weighted value vector based on the first binary clustering tree and the convolution kernel parameters / fully connected layer parameters;
[0066] Among them, in one embodiment of the present invention, after determining the first binary clustering tree and convolution kernel parameters / fully connected layer parameters through the above steps, at least one corresponding weighted value vector can be obtained based on the first binary clustering tree and convolution kernel parameters / fully connected layer parameters.
[0067] Specifically, in one embodiment of the present invention, the method for obtaining at least one corresponding weighted value vector based on the first binary clustering tree and the convolution kernel parameters / fully connected layer parameters may include the following steps:
[0068] Step S41, multiplying the node parameters in the first binary clustering tree by the convolution kernel parameters / fully connected layer parameters to obtain at least one corresponding second binary clustering tree;
[0069] Step S42: obtaining at least one corresponding weighted value vector based on at least one second binary clustering tree.
[0070] In one embodiment of the present invention, if the convolution layer in the trained neural network includes multiple convolution kernels, the node parameters in the first binary clustering tree are multiplied by the parameters of each convolution kernel to obtain at least one corresponding second binary clustering tree, wherein the parameters of each convolution kernel correspond to a second binary clustering tree. The processing method of the fully connected layer parameters is the same as that of the above-mentioned convolution kernel parameters, and the detailed introduction of the fully connected layer parameters is not repeated in this embodiment of the present disclosure.
[0071] Among them, in one embodiment of the present invention, after obtaining at least one second binary clustering tree through the above steps, at least one corresponding weighted value vector can be obtained based on the at least one second binary clustering tree. Specifically, in one embodiment of the present invention, the second binary clustering trees can be traversed based on the existing binary tree traversal method to obtain the weighted value vector corresponding to each second binary clustering tree.
[0072] Step S5: store at least one weighted value vector into at least one corresponding first lookup table, and efficiently deploy the at least one first lookup table in a programmable switch, and obtain an inference result by performing efficient aggregation based on the programmable switch.
[0073] In one embodiment of the present invention, the convolution operation has the characteristic of sparse connection. For a certain input I Convolution,i The corresponding output quantity Relatively few, among which n i is the number of outputs corresponding to a certain input. For example, for a convolution with padding 0, for each convolution kernel, the first element I of its input Convolution,1 Only corresponds to the first element O of the output Convolution,1 Based on this, the output result connected to the input element can be stored in the corresponding lookup table T Convolution After querying the corresponding lookup table, the corresponding output vector is retrieved at one time. Where k∈{1,2,…,n i}.
[0074] In one embodiment of the present invention, after obtaining at least one weighted value vector through the above steps, the at least one weighted value vector can be stored in the corresponding at least one first lookup table, and the at least one first lookup table can be efficiently deployed in a programmable switch to obtain an inference result by efficient aggregation based on the programmable switch.
[0075] Specifically, in one embodiment of the present invention, the method of storing at least one weighted value vector into the corresponding at least one first lookup table, and efficiently deploying the at least one first lookup table in a programmable switch, and efficiently aggregating and obtaining an inference result based on the programmable switch may include the following steps:
[0076] Step S51, storing at least one weighted value vector in a corresponding first lookup table based on a target data format, wherein all weighted value vectors in a second binary clustering tree are stored in the same first lookup table;
[0077] Step S52, processing the plurality of first lookup tables based on the dimension of the input data, and deploying the processed second lookup tables and / or index tables in the programmable switch;
[0078] Step S53, performing efficient aggregation based on the programmable switch to obtain the inference result.
[0079] In one embodiment of the present invention, since matrix multiplication requires addition aggregation of large-scale elements, and programmable switches generally lack sufficient hardware to support this aggregation operation, the present invention designs a novel target data format to achieve efficient aggregation.
[0080] Specifically, in one embodiment of the present invention, it is difficult to achieve efficient aggregation on the above-mentioned programmable switch, mainly because the programmable switch only supports limited 8-bit, 16-bit, and 32-bit adders. Figure 2 Schematic diagram of adding different data formats proposed in the embodiment of the present disclosure. Figure 2 As shown, 8-bit addition may cause overflow and produce unpredictable errors; the number of 16-bit adders is difficult to meet the requirements, so it is necessary to use 32-bit adders as much as possible to accurately implement as many aggregation operations as possible.
[0081] Also, in one embodiment of the present invention, simply using a 32-bit adder to implement two 16-bit additions expanded from 8-bit numbers may also cause problems. Figure 2 As shown, the addition of two negative numbers in the lower 16 bits may cause underflow, affecting the result of the upper 16 bits. Based on this, the present invention proposes a target data format, specifically a 13-6-13 data format, that is, 13 bits are used to implement 8-bit fixed-point addition, wherein the upper 5 bits are used to avoid errors that may be caused by addition overflow, and the middle 6 bits are used to absorb the underflow generated by the lower bits to avoid affecting the calculation results of the higher bits.
[0082] Furthermore, in one embodiment of the present invention, after obtaining at least one lookup table through the above steps, multiple first lookup tables can be processed based on the dimension of the input data, and the processed second lookup tables and / or index tables can be deployed in the programmable switch to efficiently utilize resources on limited resources.
[0083] Specifically, in one embodiment of the present invention, the method for processing a plurality of first lookup tables based on the dimension of input data and deploying the processed second lookup tables and / or index tables in a programmable switch may include the following steps:
[0084] Step 1: If the dimension of the input data is less than or equal to a preset threshold, the data in the first preset number table of the plurality of first lookup tables are aggregated, and the second lookup table after the aggregated data is deployed in the programmable switch;
[0085] Step 2: if the dimension of the input data is greater than a preset threshold, an index table is generated based on the multiple first lookup tables, wherein the index in the index table is the position of the data in the first lookup table in the corresponding second binary clustering tree;
[0086] Step 3: Aggregate the data in the second preset number table in the multiple first lookup tables, and deploy the second lookup table and index table after the aggregated data in the programmable switch.
[0087] The input data may be data whose dimensions need to be extracted, which is determined in advance, that is, data whose dimensions are required for processing.
[0088] And, in one embodiment of the present invention, if the dimension of the input data is less than or equal to the preset threshold, the resource occupation of the ternary content addressable memory (TCAM) is small at this time, and it does not reach the minimum table unit in the programmable switch, so the data in the first preset number table (for example, 2) can be aggregated. After re-aggregation, although the occupied ternary content addressable memory (TCAM) resources are greatly improved compared with before, it still only needs to occupy two minimum table units, which is no different from the resource overhead before aggregation, but the number of queries is reduced by one half. In addition, the above method can pre-aggregate the query results of two times, reducing the bandwidth occupation of the data path by one half.
[0089] For example, assuming that the dimension of the weighted value vector in the first lookup table 1 is m, the dimension of the weighted value vector in the first lookup table 2 is m, and each of the weighted value vectors in the first lookup table 1 and the first lookup table 2 is 1-dimensional, then when the first lookup table 1 is aggregated with the second lookup table 2, the value of each dimension in the first lookup table 1 is aggregated with the value of each dimension in the first lookup table 2, and the dimension of the second lookup table is m×m, and the aggregation result corresponding to each dimension is 2-dimensional.
[0090] Further, in one embodiment of the present invention, if the dimension of the input data is greater than a preset threshold, less coding is required to achieve efficient deployment. Specifically, in one embodiment of the present invention, the data in the second preset number table (e.g., 3) can be aggregated, and after re-aggregation, corresponding index coding can be generated based on the aggregated data, thereby reducing the occupancy of the data path bandwidth. In one embodiment of the present invention, the above-mentioned preset threshold can be set as needed, for example, the above-mentioned preset threshold is 6 dimensions.
[0091] Furthermore, in one embodiment of the present invention, after the processed second lookup table and / or index table is deployed on a programmable switch through the above steps, an inference result can be obtained by efficient aggregation based on the programmable switch.
[0092] Specifically, in one embodiment of the present invention, the method for obtaining inference results by efficient aggregation based on a programmable switch may include the following steps:
[0093] Step a, the programmable switch receives a data packet from a network port, dynamically parses the data packet according to a specific protocol through a parser of a programmable data plane, and matches the data packet with a second lookup table and / or an index table through a ternary content addressable memory to obtain a corresponding plurality of output vectors;
[0094] Step b: Add and aggregate the multiple output vectors to obtain the inference result corresponding to the data packet.
[0095] In one embodiment of the present invention, each data packet is assembled according to a specific protocol. After the programmable switch receives the data packet from the network port, the data packet can be dynamically parsed according to the specific protocol through the parser of the programmable data plane, and the feature vector required by the parsed neural network model is pre-saved in the metadata, and then the metadata is matched with the position in the second lookup table and / or index table using the ternary content addressable memory, and multiple pre-cached output vectors are taken out from the second lookup table, and the multiple output vectors are added and aggregated to obtain the neural network inference result.
[0096] And, in one embodiment of the present invention, the method of matching the data packet with the second lookup table and / or the index table through the ternary content addressable memory to obtain the corresponding multiple output vectors may include: if the dimension of the input data is less than or equal to the preset threshold, the ternary content addressable memory matches the data packet with the TCAM table and the second lookup table to obtain the corresponding multiple output vectors; if the dimension of the input data is greater than the preset threshold, the ternary content addressable memory matches the data packet with the TCAM table and the index table to obtain the corresponding index code, and matches the index code with the second lookup table to obtain the corresponding multiple output vectors. Wherein, the TCAM table can be obtained according to the threshold on the first binary clustering tree and the selected dimension.
[0097] In one embodiment of the present invention, if the dimension of the input data is less than or equal to a preset threshold, the TCAM table and the second lookup table may be matched to obtain corresponding multiple output vectors, thereby reducing the number of queries of the TCAM table.
[0098] Also, in one embodiment of the present invention, assuming that the dimension of the input data is n, each dimension of the input data can obtain a d-bit index code through the index table. Based on this, the input data is queried through the TCAM table and the index table to obtain d-bit index codes corresponding to the n dimensions, and the corresponding nd-bit index codes are generated after combination, and the index codes are matched with the second lookup table to obtain corresponding multiple output vectors, thereby reducing the occupancy of the data path bandwidth and the query overhead of the TCAM table entries.
[0099] According to the efficient aggregation method for programmable switches proposed in an embodiment of the present invention, by efficiently deploying at least one first lookup table in a programmable switch, the programmable switch can support larger-scale neural networks and has a wide range of applications. At the same time, efficient aggregation is performed based on the programmable switch to obtain inference results, so that the matrix multiplication and convolution operations in the neural network model are abstracted into efficient aggregation operations in the programmable switch, thereby reducing the model inference time and achieving high-speed inference tasks.
[0100] Next, a high-efficiency aggregation device for programmable switches proposed according to an embodiment of the present invention will be described with reference to the accompanying drawings.
[0101] Figure 3 The figure is a schematic diagram of the structure of an efficient aggregation device for a programmable switch according to an embodiment of the present invention.
[0102] like Figure 3 As shown, the efficient aggregation device 10 for programmable switches includes: an acquisition module 301, a generation module 302, a determination module 303, a processing module 304 and an aggregation module 305, wherein
[0103] An acquisition module 301 is used to acquire a trained neural network model and a training data set for training the neural network model;
[0104] A generating module 302, for generating a first binary clustering tree corresponding to the neural network model based on the training data in the training data set;
[0105] A determination module 303 is used to determine the convolution kernel parameters / fully connected layer parameters in the trained neural network model;
[0106] The processing module 304 is used to obtain at least one corresponding weight value vector based on the first binary clustering tree and the convolution kernel parameters / fully connected layer parameters;
[0107] The aggregation module 305 is used to store at least one weighted value vector into the corresponding at least one first lookup table, and efficiently deploy the at least one first lookup table in the programmable switch, and obtain the inference result by performing efficient aggregation based on the programmable switch.
[0108] Furthermore, the generation module 302 is specifically used for:
[0109] determining clustering parameters for clustering;
[0110] Based on the clustering parameters and the minimum square error, the training data is clustered to obtain the clustering results, wherein the clustering results include the average values of different dimensions of the training data in each cluster;
[0111] Based on the average values of different dimensions of the training data in each cluster, the first binary clustering tree corresponding to the neural network model is generated.
[0112] Furthermore, the determination module 202 is further configured to:
[0113] Multiplying the node parameters in the first binary clustering tree by the convolution kernel parameters / fully connected layer parameters to obtain at least one corresponding second binary clustering tree;
[0114] Based on the at least one second binary clustering tree, at least one corresponding weight value vector is obtained.
[0115] Furthermore, the processing module 304 is specifically used to:
[0116] Multiplying the node parameters in the first binary clustering tree by the convolution kernel parameters / fully connected layer parameters to obtain at least one corresponding second binary clustering tree;
[0117] Based on at least one second binary clustering tree, at least one corresponding weight value vector is obtained.
[0118] Furthermore, the aggregation module 305 is specifically used for:
[0119] storing at least one weighted value vector in a corresponding first lookup table based on a target data format, wherein all weighted value vectors in a second binary clustering tree are stored in the same first lookup table;
[0120] Processing the plurality of first lookup tables based on the dimension of the input data, and deploying the processed second lookup tables and / or index tables in the programmable switch;
[0121] The inference results are obtained by efficient aggregation based on programmable switches.
[0122] Furthermore, the above-mentioned reasoning module 205 is also used for:
[0123] If the dimension of the input data is less than or equal to a preset threshold, the data in the first preset number table in the plurality of first lookup tables are aggregated, and the second lookup table after the aggregated data is deployed in the programmable switch;
[0124] If the dimension of the input data is greater than a preset threshold, an index table is generated based on the multiple first lookup tables, wherein the index in the index table is the position of the data in the first lookup table in the corresponding second binary clustering tree;
[0125] Aggregate the data in the second preset number table in the plurality of first lookup tables, and deploy the second lookup table after the data aggregation and the index table in the programmable switch.
[0126] Furthermore, the above device is also used for:
[0127] The programmable switch receives a data packet from the network port, dynamically parses the data packet according to a specific protocol through a parser of the programmable data plane, and matches the data packet with the second lookup table and / or the index table through a ternary content addressable memory to obtain a corresponding plurality of output vectors;
[0128] Multiple output vectors are added and aggregated to obtain the inference result corresponding to the data packet.
[0129] Furthermore, the above device is also used for:
[0130] If the dimension of the input data is less than or equal to a preset threshold, the ternary content addressable memory matches the data packet with the TCAM table and the second lookup table to obtain a corresponding plurality of output vectors;
[0131] If the dimension of the input data is greater than a preset threshold, the ternary content addressable memory matches the data packet with the TCAM table and the index table to obtain a corresponding index code;
[0132] The index code is matched with the second lookup table to obtain corresponding multiple output vectors.
[0133] According to the efficient aggregation device for programmable switches proposed in an embodiment of the present invention, by efficiently deploying at least one first lookup table in a programmable switch, the programmable switch can support larger-scale neural networks and has a wide range of applications. At the same time, efficient aggregation is performed based on the programmable switch to obtain inference results, so that the matrix multiplication and convolution operations in the neural network model are abstracted into efficient aggregation operations in the programmable switch, thereby reducing the model inference time and achieving high-speed inference tasks.
[0134] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0135] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0136] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations of the present invention. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present invention.
Claims
1. An efficient aggregation method for programmable switches, characterized in that: The method comprises: Acquire a trained neural network model and a training data set used to train the neural network model; Based on the training data in the training data set, generating a first binary clustering tree corresponding to the neural network model; Determining convolution kernel parameters / fully connected layer parameters in the trained neural network model; Based on the first binary clustering tree and the convolution kernel parameters / fully connected layer parameters, obtaining at least one corresponding weighted value vector; The at least one weighted value vector is stored in the corresponding at least one first lookup table, and the at least one first lookup table is efficiently deployed in a programmable switch, and an inference result is obtained by performing efficient aggregation based on the programmable switch.
2. The method according to claim 1, characterized in that The step of generating a first binary clustering tree corresponding to the neural network model based on the training data in the training data set includes: determining clustering parameters for clustering; Based on the clustering parameters and the minimum square error, clustering the training data to obtain a clustering result, wherein the clustering result includes average values of different dimensions of the training data in each cluster; Based on the average values of different dimensions of the training data in each cluster, a first binary clustering tree corresponding to the neural network model is generated.
3. The method according to claim 1, characterized in that The obtaining, based on the first binary clustering tree and the convolution kernel parameters / fully connected layer parameters, corresponding to at least one weighted value vector, comprises: Multiplying the node parameters in the first binary clustering tree by the convolution kernel parameters / fully connected layer parameters to obtain at least one corresponding second binary clustering tree; Based on the at least one second binary clustering tree, at least one corresponding weight value vector is obtained.
4. The method according to claim 3, characterized in that The storing of the at least one weighted value vector into the corresponding at least one first lookup table, and efficiently deploying the at least one first lookup table in a programmable switch, and obtaining the inference result by performing efficient aggregation based on the programmable switch, comprises: Storing the at least one weighted value vector in a corresponding first lookup table based on a target data format, wherein all weighted value vectors in the second binary clustering tree are stored in the same first lookup table; Processing the plurality of first lookup tables based on the dimension of the input data, and deploying the processed second lookup tables and / or index tables in the programmable switch; The inference result is obtained by performing efficient aggregation based on the programmable switch.
5. The method according to claim 4, characterized in that The processing of the plurality of first lookup tables based on the dimension of the input data and deploying the processed second lookup tables and / or index tables in the programmable switch includes: If the dimension of the input data is less than or equal to a preset threshold, aggregating the data in a first preset number table among the plurality of first lookup tables, and deploying a second lookup table after aggregating the data in the programmable switch; If the dimension of the input data is greater than a preset threshold, an index table is generated based on the multiple first lookup tables, wherein the index in the index table is the position of the data in the first lookup table in the corresponding second binary clustering tree; Aggregate the data in the second preset number table of the multiple first lookup tables, and deploy the second lookup table after the data aggregation and the index table in the programmable switch.
6. The method according to claim 4, characterized in that The method of obtaining the inference result by performing efficient aggregation based on the programmable switch includes: The programmable switch receives a data packet from a network port, dynamically parses the data packet according to a specific protocol through a parser of a programmable data plane, and matches the data packet with the second lookup table and / or the index table through a ternary content addressable memory to obtain a corresponding plurality of output vectors; The multiple output vectors are added and aggregated to obtain an inference result corresponding to the data packet.
7. The method according to claim 6, characterized in that The step of matching the data packet with the second lookup table and / or the index table through a ternary content addressable memory to obtain a corresponding plurality of output vectors includes: If the dimension of the input data is less than or equal to a preset threshold, the ternary content addressable memory matches the data packet with the TCAM table and the second lookup table to obtain a corresponding plurality of output vectors; If the dimension of the input data is greater than a preset threshold, the ternary content addressable memory matches the data packet with the TCAM table and the index table to obtain a corresponding index code; The index code is matched with the second lookup table to obtain corresponding multiple output vectors.
8. An efficient aggregation device for a programmable switch, characterized in that: The device comprises: An acquisition module, used to acquire a trained neural network model and a training data set used to train the neural network model; A generating module, configured to generate a first binary clustering tree corresponding to the neural network model based on the training data in the training data set; A determination module, used to determine the convolution kernel parameters / fully connected layer parameters in the trained neural network model; A processing module, configured to obtain at least one corresponding weighted value vector based on the first binary clustering tree and the convolution kernel parameters / fully connected layer parameters; An aggregation module is used to store the at least one weighted value vector into the corresponding at least one first lookup table, and efficiently deploy the at least one first lookup table in a programmable switch, and obtain an inference result by performing efficient aggregation based on the programmable switch.
9. An electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.