A method, system, device and medium for estimating flow frequency
By using the parameterized sketch model in stream frequency estimation, the problem of poor generalization in the prior art is solved, and stronger adaptability and prediction capabilities are achieved.
Patent Information
- Application Number
- CN202510142255.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-10
AI Technical Summary
The existing sketch-based data structure for stream frequency estimation has the problem of poor generalization and it is difficult to adapt to changes in network traffic distribution in different scenarios.
The parameterized sketch model is adopted, including the encoder module, hash module and decoder module, and the stream tags and their frequency are stored by writing operations, the frequency is estimated by reading operations, and the model is optimized by training to improve universality.
The generalization of the parameterized sketch model and global prediction capabilities are improved, and it can more effectively adapt to changes in network traffic distribution in different scenarios.
Smart Images

Figure CN119603203B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of network flow measurement, and in particular to a flow frequency estimation method, system, equipment and medium. Background Art
[0002] Flow frequency estimation is a basic research direction in the field of network traffic measurement. It refers to estimating the number of times each flow label corresponds to a flow within a certain measurement period. It is extremely important for many applications in the network field, such as network traffic monitoring and distributed denial of service attack (DDoS) detection.
[0003] In the prior art, sketch-based data structures are widely used in flow frequency estimation tasks. However, the performance of traditional sketches will be reduced due to changes in network traffic characteristics. In order to adapt to the network traffic distribution in different scenarios, researchers will try to assume scenarios and develop targeted statistical models. However, it is complicated to develop a model for each scenario and it is difficult to reuse across scenarios. In order to deeply explore the potential knowledge of flow data distribution, learning-based methods to fit flow data distribution have become an optional solution.
[0004] In addition, in the prior art, RLSketch, which detects large flows based on reinforcement learning, uses statistical data of network flows to predict potential large flows, achieving high accuracy with a small memory. However, RLSketch is only optimized for LDSketch and is not universal. The researchers also designed a general network measurement framework MLSketch based on machine learning, which uses a linear regression model to correct the estimation results of traditional sketch for error-prone flows; and MLSketch can be applied to multiple measurement tasks by simply changing the training process. However, MLSketch requires manual setting of thresholds to determine whether the measurement results are error-prone, resulting in a low degree of automation. TalentSketch adds an error-prone flow determination model based on the measurement model of MLSketch. First, the error-prone flow determination model is used to adaptively determine whether the flow to be queried is error-prone, and then the measurement model is used to correct the estimation results of traditional sketch for error-prone flows. Although both MLSketch and TalentSketch can adapt to multiple measurement tasks, in order to adapt to the dynamically changing short-term network traffic distribution, the linear regression model needs to sample data frequently for offline training, and frequent data sampling makes the communication cost too high. In order to avoid the communication overhead caused by offline training, ALSketch updates the measurement rules in real time through online learning. Although ALSketch can reduce communication costs through online learning, the traffic data distribution learned by ALSketch, MLSketch and TalentSketch is short-term, and short-term features can only maintain short-term measurement tasks. In order to maintain the measurement performance continuously and effectively, a large amount of training overhead is inevitable. In response to the above training overhead problem, the researchers also designed a meta-sketch based on meta-learning to enhance the adaptability of the model by pre-training on different tasks. When faced with a new data distribution, meta-sketch can adapt to distribution changes stably in the long term through one-time processing. However, in order to adapt to the data distribution of the training set, meta-sketch will have the problem of overfitting, so it is not general-purpose.
[0005] In summary, the existing methods for estimating flow frequency based on sketch data structures still have various problems, especially the problem of poor versatility, which needs to be improved. Summary of the invention
[0006] Therefore, the technical problem to be solved by the present invention is to overcome the problem of poor versatility of the method for estimating flow frequency based on sketch data structure in the prior art.
[0007] In order to solve the above technical problems, the present invention provides a flow frequency estimation method, comprising:
[0008] Step S1: Obtain a packet set from the network, and parse the packet set to obtain a flow label set;
[0009] Step S2: forming a key-value pair set of flow labels and their corresponding frequencies according to the frequency of occurrence of each flow label in the flow label set, and caching the key-value pair set;
[0010] Step S3: storing the key-value pair set of the cached flow label and its corresponding frequency through the write operation of the parameterized sketch model;
[0011] Step S4: when querying the frequency of the flow label, the estimated frequency is obtained through a read operation of a parameterized sketch model, wherein the parameterized sketch model includes an encoder module, a hash module, and a decoder module connected in sequence.
[0012] In one embodiment of the present invention, the writing operation in step S3 includes encoding the flow label, specifically:
[0013] The encoder module in the parameterized sketch model performs two encodings on the flow label f, where:
[0014] The first encoding is: through the encoder module Encode the input flow label f to obtain the embedding vector used to represent the flow label f , since the frequency of the flow label f is equal to the embedding vector The number of occurrences will be embedded in the vector As the write word w of the write counter C;
[0015] The second encoding is: through the encoder module Embedding vector Encode and generate query word q, and find the embedding vector through query word q The storage location in the counter C;
[0016] Said and They are all fully connected neural networks with 3-4 layers, and the counter C is located in the hash module.
[0017] In one embodiment of the present invention, the write operation in step S3 further includes: generating a sparse addressing word according to the query word q , specifically:
[0018] Transpose the query word q to get ,Will And the attention matrix in the hashing module Multiply to generate an address word with preliminary addressing capability ;
[0019] Use the activation function SparseMax in the hash module to perform a sparse operation on the address word a, and filter out the sparse address words with fewer slots in the corresponding counter C. .
[0020] In one embodiment of the present invention, the write operation in step S3 further includes: writing the sparse address word Matrix multiplication with the word w is memorized , and then the memory Then add it to counter C In the corresponding slot, the counter C is updated with the formula:
[0021] ;
[0022] in, , They are matrix multiplication and matrix addition, respectively.
[0023] In one embodiment of the present invention, the method of obtaining the estimated frequency by reading the parameterized sketch model in step S4 includes:
[0024] According to the counter C and the sparse address word The memory r is calculated as follows:
[0025] ;
[0026] Where C is a counter, is a sparse address word, r is the C The corresponding memory, is matrix multiplication, z is the embedding vector corresponding to the flow label f, is the Hadamard product operation;
[0027] By inputting the memory r into the decoder module ,pass Decoding frequency ,in, It is a fully connected neural network with 3-4 layers.
[0028] In one embodiment of the present invention, the parameterized sketch model is trained by constructing a loss function, specifically:
[0029] The training data set with a total number of flow labels k obtained by extracting a subset of packets from the packet set and parsing is , Indicates flow labels, Indicates flow labels, Indicates that in a measurement cycle, the label is The number of times the flow appears;
[0030] Suppose the occurrence count of the flow label obtained by frequency estimation is , Indicates that in one measurement cycle, the estimated label is The number of times the flow appears; the loss function of the prediction performance is set to the first optimization objective of the parameterized sketch model, the formula is:
[0031] ;
[0032] in, and Represents the parameters of the model to be learned, and the mean absolute error AAE and the mean relative error ARE satisfy: ;
[0033] Set the second optimization goal of the parameterized sketch model. The formula is:
[0034] ;
[0035] in, Used to calculate the variance of a series; Indicates that within a measurement cycle The total number of non-sparse items in each column is equivalent to the number of times each slot in the counter C is mapped;
[0036] The loss function of the parameterized sketch model is constructed according to the first optimization objective and the second optimization objective. The formula is:
[0037] ;
[0038] Among them, loss represents the loss function of the parameterized sketch model. is an adjustable hyperparameter.
[0039] In order to solve the above technical problems, the present invention provides a base flow frequency estimation system, comprising:
[0040] Acquisition module: used to acquire a packet set from the network, and parse the packet set to obtain a flow label set;
[0041] A cache module: used to form a key-value pair set of flow labels and their corresponding frequencies according to the frequency of occurrence of each flow label in the flow label set, and cache the key-value pairs;
[0042] Storage module: used to store the key-value pair set of cached flow labels and their corresponding frequencies through the write operation of the parameterized sketch model;
[0043] Query module: used to obtain an estimated frequency through a read operation of a parameterized sketch model when querying the frequency of a flow label, wherein the parameterized sketch model includes an encoder module, a hash module, and a decoder module connected in sequence.
[0044] In order to solve the above technical problems, the present invention provides a network flow measurement device, including the above flow frequency estimation system.
[0045] To solve the above technical problems, the present invention provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above stream frequency estimation method when executing the computer program.
[0046] In order to solve the above technical problem, the present invention provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the above stream frequency estimation method are implemented.
[0047] The above technical solution of the present invention has the following advantages compared with the prior art:
[0048] The present invention finds the address that makes the result optimal under the premise of making the mapping result of the hash module relatively uniform, and the game process between optimizing the hash module and optimizing the prediction performance is explainable: optimizing the prediction performance will make the local performance better, but the versatility of the hash module will be weakened; optimizing the hash module will make the global hash conflict relatively fair, the local prediction performance will be weakened, but the versatility of the hash module will be enhanced; the present invention makes the parameterized sketch model more versatile through this game training method, and the global prediction ability will also be more robust. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to make the contents of the present invention more clearly understood, the present invention is further described in detail below based on specific embodiments of the present invention in conjunction with the accompanying drawings.
[0050] Figure 1 is a flow chart of the method of the present invention;
[0051] Figure 2 It is a schematic diagram of the overall framework of the parameterized sketch model in an embodiment of the present invention;
[0052] Figure 3 It is an overall flow chart of frequency estimation in an embodiment of the present invention. DETAILED DESCRIPTION
[0053] The present invention is further described below in conjunction with the accompanying drawings and specific embodiments so that those skilled in the art can better understand the present invention and implement it, but the embodiments are not intended to limit the present invention. Embodiment 1
[0054] Reference Figure 1 As shown, the present invention relates to a flow frequency estimation method, comprising:
[0055] Step S1: Obtain a packet set from the network, and parse the packet set to obtain a flow label set;
[0056] Step S2: forming a key-value pair set of flow labels and their corresponding frequencies according to the frequency of occurrence of each flow label in the flow label set, and caching the key-value pair set;
[0057] Step S3: storing the key-value pair set of the cached flow label and its corresponding frequency through the write operation of the parameterized sketch model;
[0058] Step S4: when querying the frequency of the flow label, the estimated frequency is obtained through a read operation of a parameterized sketch model, wherein the parameterized sketch model includes an encoder module, a hash module, and a decoder module connected in sequence.
[0059] The following is a detailed introduction to this embodiment:
[0060] The overall framework of the parametric sketch model of the present invention is as follows Figure 2 As shown in the figure, there are two rows of block diagrams, the first row of block diagrams represents the frequency estimation process, and the second row of block diagrams represents the parameterized sketch model training process. After the parameterized sketch model training is completed, only the first row of block diagrams of the parameterized sketch model is needed in actual use. After a measurement cycle or a training is completed, all the counters of the parameterized sketch model must be set to 0.
[0061] In the frequency estimation process, first, the present invention parses the packet set from the network to obtain the flow label set, and then uses a small amount of memory for preliminary caching to obtain the <flow label, frequency> key-value pair set. Secondly, the present invention stores the cached <flow label, frequency> key-value pair set through a forward propagation write operation. Finally, when a query task is received, the present invention performs a forward propagation read operation to obtain the estimated frequency. .
[0062] In the model training process, first, the present invention will periodically sample to form a data packet set (i.e., a packet subset) required for training. Second, the packet subset is stored in the parameterized sketch model in the same way as in the frequency estimation process. Finally, the address generated in the forward propagation process is used to And the predicted frequency read last Perform back propagation to update the model parameters required for forward propagation. Can be compared with the real flow frequency Obtaining the error (such as the average relative error, the average absolute error) and reducing the error is the first training goal of the present invention. The first training goal can enhance the ability of the model to fit the current flow data distribution, thereby improving the prediction performance of the model. It can be used to calculate the number of hash conflicts in each slot of the memory matrix, and reducing the difference between the number of hash conflicts in each slot is the second training goal of the present invention. The second training goal makes the address mapped by the hash module relatively more uniform, reducing the possibility of serious hash conflicts. Uniform hashing will reduce the local prediction performance on the training set to a certain extent, but fairer hash mapping makes the global performance of the present invention better. The following will introduce the two parts of the present invention in detail.
[0063] Part I: Frequency Estimation Process
[0064] The overall process of frequency estimation is as follows Figure 3 As shown in the figure, it specifically includes three modules: encoder module, hash module, and decoder module. First, the controller in the encoder module encodes the stream label to obtain the write word w written into the parameterized counter C and the query word q used to interact with the hash module. These vectors are machine-understandable codes. Secondly, the attention matrix A in the hash module dynamically focuses on the more important features in q to generate the addressing word a. On this basis, the hash module uses sparse SparseMax (an optimized Softmax function) to sparse a to obtain the sparse addressing word. Finally, using , The operation can write w to the corresponding position in C. The operation can read the storage information of the corresponding position in C, which can be decoded by the decoder module to form a frequency that humans can understand. The specific implementation of each module and operation is as follows:
[0065] Encoder module: This module needs to complete two types of encoding: (1) The controller in the encoder Encode the input flow label f so that it can be embedded in a machine-understandable vector represents the flow item, rather than the flow label f which cannot measure the distance. Since the frequency is equal to the number of times the embedding vector z appears, the embedding vector z will also be used as the write word w to the parameterized counter C. The subsequent decoder will fit the number of times the embedding vector z is superimposed in the counter C. (2) The controller in the encoder module is Encode z to generate query word q. It should be noted that q and w can be directly obtained by encoding f, but since f only contains one piece of information, it is difficult to mine a more meaningful query word q from f.
[0066] The above two encodings can be summarized as the controller part shown in formula (1):
[0067] Formula (1)
[0068] Where f represents the flow label. Map the input variables from the human-understandable feature space to the potential feature space to form vectors q and w. Because all operations in the parameterized sketch model, including addressing operations, read operations, and write operations, are parameterized, the variables that interact with the parameterized sketch model are also parameterized. Therefore, It is a necessary module to transform Sketch from non-parametric to parametric.
[0069] In order to increase the versatility of the present invention, , and the following decoder All use only fully connected neural networks with 3-4 layers.
[0070] Hash module: In the traditional sketch, the goal of the hash function is to calculate the address of the stream write counter, and the process of mapping the address meets two criteria. Criteria 1: The address corresponding to any stream label is unique. Criteria 2: The addresses corresponding to all stream labels are evenly distributed. The present invention uses differentiable read and write operations to implement the hash function, and adds the variance of the number of hash conflicts occurring in all slots to the loss function as an additional loss term. Differentiable read and write make the hash operation of the present invention capable of being optimized by back propagation, and the added additional loss term makes the present invention explainable.
[0071] The specific steps of the hash module of the present invention are as follows: (1) First, Multiplying the batch matrix of A can produce an address word a, where A is an attention matrix (the attention matrix Based on the attention mechanism, it selectively interacts with certain features of q. batch represents the number of input batches and represents the maximum output length of the hash function, which corresponds to the number of storage units of counter C. Any storage unit in counter C (i.e. slots, each slot corresponds to a vector) can store a vector of length w. At this time, after the feature screening of A, a has a certain addressing ability. However, the addresses in the counter corresponding to a that has not been sparsed are multiple, and multiple addresses will increase the probability of hash conflicts. (2) Therefore, in this embodiment, the activation function SparseMax is used to perform sparse operation on a by referring to the neural Bloom filter, and a more prominent sparse addressing word is screened out. . Corresponding to fewer slots This can make parameterized hashing closer to the principle 1 of traditional hashing. These two steps can be summarized as an address mapping process, as shown in formula (2):
[0072] Formula (2)
[0073] Among them, q represents the query word, represents a parameterized hashing process, Represents the sparse addressing word mapped by the hash operation.
[0074] From an intuitive point of view, The fewer the corresponding slots, the fewer hash conflicts will occur, and the accuracy of model prediction will also increase. The most ideal number of mapping addresses is the number of traditional hash mapping results, that is, one. Reducing the number of hash mapping addresses will greatly reduce the number of model parameters. However, since the parameterized sketch model has an optimal model parameter, reducing the number of model parameters may not necessarily be beneficial to the model fitting data distribution. Specifically, when the number of model parameters is higher than the optimal model parameter, the learned model is often more complex, which is more likely to cause overfitting. Therefore, reducing the number of parameters can increase the versatility of the model. At this time, reducing the number of hash mapping addresses can not only reduce the performance loss caused by hash conflicts, but also reduce the interference of addresses with little correlation on reading and writing. When the number of model parameters is less than the optimal model parameter, the learned model is too simple, making the model lose its measurement ability. When the number of hash mapping addresses is already very small, further reducing the number of hash mapping addresses will only reduce hash conflicts to a limited extent, and the model's ability to fit data distribution will be weakened. Therefore, when training a parameterized sketch model, measurement accuracy should be used as the standard, rather than blindly using the low sparsity of the mapping results as the standard.
[0075] In addition, in order to adapt to the data distribution of the training set, the parameterized hash will arrange a mapping result that is more suitable for the distribution of the training set for any stream as much as possible. Therefore, the parameterized hash has an overfitting phenomenon. The present invention draws on the criterion 2 of the traditional hash as an additional training target of the present invention, and the specific method will be described in detail later.
[0076] Write operation: Write operation contains two abstract operation symbols , In specific applications, the most basic instantiation method is shown in formula (3).
[0077] Formula (3)
[0078] Among them, C is the counter, w is the write word, is a sparse addressing word, , for matrix multiplication and matrix addition.
[0079] First, w and When performing matrix multiplication, a vector with the same shape as C is obtained, which can be understood as a temporary counter for storing w. In this temporary counter, the value of w is concentrated in The values of the slots corresponding to the non-sparse items of will be close to zero. That is, w will only be stored at address Finally, let the total counter storing all memories and the counter temporarily storing w's memory perform matrix addition to get the latest total counter.
[0080] Read operation: The read operation contains an operation symbol In specific applications, the most basic instantiation method is shown in formula (4).
[0081] Formula (4)
[0082] Where C is a counter, is a sparse address word, r is the C The corresponding memory, is matrix multiplication. The principle of reading is similar to that of writing, and they both rely on To find the slot corresponding to the read and write operations.
[0083] In the specific implementation process, since the counter stores the encoding of the stream, the storage of smaller words (i.e. words less than the preset value) will be interfered by larger words (i.e. words greater than the preset value). When multiple larger words are superimposed in the counter, the influence of smaller words on the counter is severely weakened. Therefore, in order to enhance the weight of smaller words, When instantiating, it is necessary to add the weights used for constraints, as shown in formula (5).
[0084] Formula (5)
[0085] Among them, z is the embedding vector of the flow label, , that is, the corresponding embedding vector is used to constrain writing. is the Hadamard product operation. The difference from formula (4) is that when the counter reads After the memory of the corresponding position, let the memory read Perform a Hadamard product with 1 / z. This weakens the memory corresponding to the larger letter and strengthens the memory corresponding to the smaller letter.
[0086] Decoder module: This module Used to construct parameterized memory (i.e., memory r read in the counter) and frequency The mapping relationship between them is that by inputting r into , which can be decoded into frequencies that humans can understand Without considering hash conflicts, the memory r read in the counter corresponds to the written word w one-to-one, w corresponds to the embedded vector z of the flow one-to-one, and z corresponds to the flow label f one-to-one. In other words, every time the frequency of the flow label f corresponding to the written word w increases once, the memory corresponding to w will be superimposed once more in the counter C. Therefore, the decoder fits r and The mapping relationship between them corresponds to the real logical relationship: the number of times the memory r corresponding to the writing w is stacked is equal to the frequency of occurrence of the flow with the same flow label f corresponding to the writing w.
[0087] Part 2: Parametric Sketch Model Training Process
[0088] The training process of the parameterized sketch model consists of two stages: pre-training and deployment. The dataset used in the pre-training stage comes from historical network data streams. The huge amount of historical stream data enables the model to learn , 、A、 The initial parameters (the training process is equivalent to , 、A、 Parameters are optimized), and the parameterized sketch obtains basic flow frequency estimation capabilities. In the deployment phase, the model only needs to sample small batches of flow data on-chip, so that the model can adapt to the ever-changing distribution of network flow data. The present invention divides CAIDA 2019 in a ratio of 7:3 to simulate the training of these two stages respectively. Among them, in the pre-training stage, the present invention uses 7 / 10 of CAIDA 2019 to train the parameterized sketch. In the deployment phase, the present invention extracts 500 to 5000 flow items from 3 / 10 of CAIDA 2019 for adaptive training. Although the data used in the pre-training phase comes from historical data, the data in the deployment phase comes from regular sampling. However, the data format is consistent, and the training scheme will also be consistent. The specific training process of the present invention is as follows:
[0089] Define the training data set with a total number of flow items k obtained by parsing the current package set as , Indicates flow labels, Indicates flow labels, Indicates that within a certain measurement period, the label is The number of occurrences of the flow. Suppose the number of occurrences of the flow label f obtained by frequency estimation is , Indicates that within a certain measurement period, the model estimates the label to be The two training objectives based on the present invention are shown in formula (6):
[0090] Formula (6)
[0091] Among them, AAE is the mean absolute error and ARE is the average relative error. The two can evaluate the prediction performance of the model. The lower the two are, the better the model performance is. Therefore, the loss function of the prediction performance is set as the first optimization objective of the parameterized sketch model, as shown in formula (7):
[0092] Formula (7)
[0093] in, and Represents the parameters to be learned by the model. Setting the loss function in combination with the parameters to be learned by the model enables the loss function to adaptively mix multiple training indicators.
[0094] However, if too much attention is paid to the loss of prediction performance, parameterized hashing will suffer from overfitting, that is, each flow is assigned to a slot that is more suitable for the data distribution of the current training data. To this end, the present invention draws on the characteristics of traditional hashing fairness and sets the second optimization goal of the parameterized sketch model as shown in formula (8):
[0095] Formula (8)
[0096] in, Can calculate the variance of a series; Indicates that within a measurement cycle The total number of non-sparse items in each column is equivalent to the number of times each slot in the counter C is mapped. Different flow labels may be mapped to the same slot of the counter, and the slots can be shared, but the present invention hopes that each slot shares a little more evenly. Therefore, combined with formula (7), the loss function of the parameterized sketch model is shown in formula (9):
[0097] Formula (9)
[0098] Among them, loss represents the loss function of the parameterized sketch model, and the model parameters are updated through back propagation of loss; is a tunable hyperparameter that controls the degree of parameterized hashing uniformity. The larger it is, the greater the model's efforts to optimize the hash function, while the efforts to predict accuracy will become smaller. The game process between optimizing the hash function and optimizing the prediction performance is explainable: optimizing the prediction performance will make the local performance better, but the versatility of the hash module will be weakened. Optimizing the hash module will make global hash conflicts relatively fair, the versatility of the hash module will be enhanced, but the performance of local predictions will be weakened.
[0099] It should be noted that is a sparse addressing code, that is, only a small number of items are non-sparse items, which makes the final mapping slots smaller. Assume that there are two , among which the first is (0.5, 0.5, 0.5, 0.0000009, 0.0000009), the second is (0.5, 0.0000009, 0.0000009, 0.0000009, 0.0000009). When mapping slots, items greater than 0.000001 are called non-sparse items and mapped to 1, and items less than 0.000001 are called sparse items and mapped to 0 (items less than 0.000001 can be ignored in the result of matrix multiplication, so they can be regarded as not mapped to the corresponding slot). Then the first The contribution to the number of slot mappings in counter C is (1, 1, 1, 0, 0). The second The contribution to the number of slot mappings in counter C is (1, 0, 0, 0, 0), so the first and the second The total contribution of to the number of slot mappings is (2, 1, 1, 0, 0), which is equivalent to the above formula (8) . This is equivalent to calculating the variance of (2, 1, 1, 0, 0). The smaller the variance, the better, indicating that the hash mapping is uniform. Among them, (2, 1, 1, 0, 0) means that after the insertion of two streams, the first slot in counter C is mapped twice, the second and third slots are mapped once, and the fourth and fifth slots are mapped 0 times. Embodiment 2
[0100] This embodiment provides a flow frequency estimation system, including:
[0101] Acquisition module: used to acquire a packet set from the network, and parse the packet set to obtain a flow label set;
[0102] A cache module: used to form a key-value pair set of flow labels and their corresponding frequencies according to the frequency of occurrence of each flow label in the flow label set, and cache the key-value pairs;
[0103] Storage module: used to store the key-value pair set of cached flow labels and their corresponding frequencies through the write operation of the parameterized sketch model;
[0104] Query module: used to obtain an estimated frequency through a read operation of a parameterized sketch model when querying the frequency of a flow label, wherein the parameterized sketch model includes an encoder module, a hash module, and a decoder module connected in sequence. Embodiment 3
[0105] This embodiment provides a network flow measurement device, including the flow frequency estimation system described in the second embodiment. Embodiment 4
[0106] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method for estimating the number of stream frequencies described in the first embodiment are implemented. Embodiment 5
[0107] This embodiment provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the method for estimating the flow frequency number described in the first embodiment are implemented.
[0108] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes. The schemes in the embodiments of the present application may be implemented in various computer languages, for example, object-oriented programming language Java and literal scripting language JavaScript, etc.
[0109] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0110] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0111] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0112] Although the preferred embodiments of the present application have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.
[0113] Obviously, the above embodiments are merely examples for clear explanation and are not intended to limit the implementation methods. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation methods here. The obvious changes or modifications derived from these are still within the protection scope of the invention.
Claims
1. A method for estimating flow frequency, characterized in that: include: Step S1: Obtain a packet set from the network, and parse the packet set to obtain a flow label set; Step S2: forming a key-value pair set of flow labels and their corresponding frequencies according to the frequency of occurrence of each flow label in the flow label set, and caching the key-value pair set; Step S3: storing the key-value pair set of the cached flow label and its corresponding frequency through the write operation of the parameterized sketch model; The writing operation in step S3 includes encoding the flow label, specifically: The encoder module in the parameterized sketch model performs two encodings on the flow label f, where: The first encoding is: through the encoder module Encode the input flow label f to obtain the embedding vector used to represent the flow label f , since the frequency of the flow label f is equal to the embedding vector The number of occurrences will be embedded in the vector As the write word w of the write counter C; The second encoding is: through the encoder module Embedding vector Encode and generate query word q, and find the embedding vector through query word q The storage location in the counter C; Said and They are all fully connected neural networks with 3-4 layers, and the counter C is located in the hash module; The writing operation in step S3 further includes: generating a sparse addressing word according to the query word q , specifically: Transpose the query word q to get ,Will And the attention matrix in the hash module Multiply to generate an address word with preliminary addressing capability ; Use the activation function SparseMax in the hash module to perform a sparse operation on the address word a, and filter out the sparse address words with fewer slots in the corresponding counter C. ; The write operation in step S3 further includes: Matrix multiplication with the word w is memorized , and then the memory Then add it to counter C In the corresponding slot, the counter C is updated with the formula: ; in, , They are matrix multiplication and matrix addition respectively; Step S4: when querying the frequency of the flow label, the estimated frequency is obtained through a read operation of a parameterized sketch model, wherein the parameterized sketch model includes an encoder module, a hash module, and a decoder module connected in sequence.
2. The method for estimating flow frequency according to claim 1, characterized in that: The method for obtaining the estimated frequency by reading the parameterized sketch model in step S4 includes: According to the counter C and the sparse address word The memory r is calculated as follows: ; Where C is a counter, is a sparse address word, r is the C The corresponding memory, is matrix multiplication, z is the embedding vector corresponding to the flow label f, is the Hadamard product operation; By inputting the memory r into the decoder module ,pass Decoding frequency ,in, It is a fully connected neural network with 3-4 layers.
3. The method for estimating flow frequency according to claim 1, characterized in that: The parameterized sketch model is trained by constructing a loss function, specifically: The training data set with a total number of flow labels k obtained by extracting a subset of packets from the packet set and parsing is , Indicates flow labels, Indicates flow labels, Indicates that in a measurement cycle, the label is The number of times the flow appears; Suppose the occurrence count of the flow label obtained by frequency estimation is , Indicates that in one measurement cycle, the estimated label is The number of times the flow appears; the loss function of the prediction performance is set to the first optimization objective of the parameterized sketch model, the formula is: ; in, and Represents the parameters of the model to be learned, and the mean absolute error AAE and the mean relative error ARE satisfy: ; Set the second optimization goal of the parameterized sketch model. The formula is: ; in, Used to calculate the variance of a series; Indicates that within a measurement cycle The total number of non-sparse items in each column is equivalent to the number of times each slot in the counter C is mapped; The loss function of the parameterized sketch model is constructed according to the first optimization objective and the second optimization objective. The formula is: ; Among them, loss represents the loss function of the parameterized sketch model. is an adjustable hyperparameter.
4. A stream frequency estimation system, using the stream frequency estimation method according to claim 1, characterized in that: include: Acquisition module: used to acquire a packet set from the network, and parse the packet set to obtain a flow label set; A cache module: used to form a key-value pair set of flow labels and their corresponding frequencies according to the frequency of occurrence of each flow label in the flow label set, and cache the key-value pairs; Storage module: used to store the key-value pair set of cached flow labels and their corresponding frequencies through the write operation of the parameterized sketch model; Query module: used to obtain an estimated frequency through a read operation of a parameterized sketch model when querying the frequency of a flow label, wherein the parameterized sketch model includes an encoder module, a hash module, and a decoder module connected in sequence.
5. A network flow measurement device, characterized in that: It comprises the stream frequency estimation system as claimed in claim 4.
6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the stream frequency estimation method according to any one of claims 1 to 3 are implemented.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the flow frequency estimation method according to any one of claims 1 to 3 are implemented.
Citation Information
Patent Citations
Data stream counting method and device based on Flag flag bit and storage medium
CN116303585A