High-precision flow size measurement system based on machine learning

Through a machine learning-based stream size measurement system, combining size and flow separation, misstream sampling and off-chip learning unit, the problems of insufficient storage space and hash noise error in high-speed network traffic are solved, and high-precision measurement of small and medium-sized streams are achieved.

CN120342903APending Publication Date: 2025-07-18STATE GRID ANHUI ELECTRIC POWER CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510478906.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the high-speed network traffic measurement, insufficient storage space leads to difficulty in measuring high-precision traffic, hash noise introduces overestimation errors, especially on the measurement of small and medium-sized streams, and the low sampling frequency cannot meet the real-time requirements.

Method used

Using a machine learning-based stream size measurement system, through the combination of size and flow separation, easily misstream sampling, small flow recording and under-chip learning units, machine learning is used to dynamically fit hash noise distribution to improve the measurement accuracy of small and medium-sized streams.

Benefits of technology

It significantly improves the measurement accuracy of small and medium-sized streams, corrects the impact of hash noise, and meets the real-time and high-precision measurement requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120342903A_ABST
    Figure CN120342903A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a high-precision flow size measurement system based on machine learning, and belongs to the technical field of electric digital signal processing. The measurement system comprises a large and small stream separation recording unit used for carrying out large and small stream separation operation on an accessed data stream; the cross-prone flow sampling unit is used for carrying out random sampling operation on the separated cross-prone flow; the small flow recording unit is used for carrying out recording and fine-grained clustering operation on the separated small flows; the off-chip cross-flow-prone recording unit is used for receiving a result of the random sampling operation so as to form input of a learning model; the off-chip learning unit is used for dividing a training data set into a plurality of flow clusters according to the input of the learning model and performing machine learning on each flow cluster; and the per-stream size estimation unit is used for responding to the query request. Through the measurement system provided by the invention, the size estimation of the flow which is easily influenced by the Hash noise can be corrected in time, and the measurement precision of medium and small-scale flows is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of digital signal processing, and more particularly to a high-precision flow size measurement system based on machine learning. Background Art

[0002] With the rapid development of Internet and mobile communication technologies, the number of users accessing the Internet and the scale of devices have shown exponential growth, and the scale of network traffic has also expanded rapidly. At the same time, the transmission rate of network links has also increased rapidly, evolving from the initial Gigabit Ethernet to the current 400G or even 800G Ethernet, and the processing time of a single data packet has been shortened to the nanosecond level. In the face of such a massive and high-speed data flow scenario, real-time statistics of the number of elements in each flow (i.e., flow size measurement) has become a key task in network management, which can provide important data support and decision-making basis for traffic scheduling, anomaly detection, load balancing, etc.

[0003] However, the storage space for supporting high-speed network traffic measurement is very limited. For example, although the on-chip high-speed cache of network switching devices can achieve fine-grained per-flow statistics on the premise of matching the network flow rate, its storage space needs to be shared by functions such as flow meter billing, network security, and routing table configuration, resulting in less than 1MB of space available for flow size measurement. Obviously, such a serious shortage of storage resources greatly restricts the realization of high-precision traffic measurement and is difficult to meet the high-precision measurement requirements while ensuring low storage overhead. To address the problem of scarce high-speed processing resources, many studies have designed flow size measurement algorithms using compact data structures (Sketch) to record all network flow information through storage space sharing. However, this method inevitably introduces overestimation errors caused by hash noise, reducing the accuracy of traffic measurement, especially having a greater impact on the measurement of small and medium-sized flows. To mitigate the impact of hash noise on flow size estimation, some studies use sampling techniques to download part of the traffic information to off-chip memory with relatively sufficient storage resources. However, the processing speed of off-chip memory is often 2 to 3 orders of magnitude slower than that of on-chip high-speed cache. Therefore, the sampling frequency needs to be set very low, resulting in the inability to efficiently record traffic information and being difficult to meet the real-time requirements. Summary of the Invention

[0004] An object of an embodiment of the present invention is to provide a high-precision flow size measurement system based on machine learning, which can significantly improve the measurement accuracy of small and medium-sized flows.

[0005] To achieve the above object, an embodiment of the present invention provides a high-precision flow size measurement system based on machine learning, including:

[0006] A large and small flow separation and recording unit for performing large and small flow separation operations on the incoming data stream;

[0007] An error-prone flow sampling unit for performing a random sampling operation on the separated error-prone flow;

[0008] A small flow recording unit for recording and performing fine-grained clustering operations on the separated small flows;

[0009] An off-chip error-prone flow recording unit for receiving the result of the random sampling operation to form the input of the learning model;

[0010] An off-chip learning unit for dividing the training data set into multiple flow clusters according to the input of the learning model and performing machine learning on each flow cluster respectively;

[0011] A per-flow size estimation unit for responding to query requests.

[0012] Optionally, the size flow separation and recording unit is used for:

[0013] Constructing an array of buckets with a preset length, where each bucket includes a flow fingerprint field, a forward counter, a reverse counter, and an overflow flag bit;

[0014] Obtaining each data packet arriving at the on-chip cache;

[0015] Extracting a flow label field from the header of the data packet;

[0016] Calculating the subscript of the data packet in the bucket according to formula (1):

[0017] i = H(f) = h(f) % M, (1)

[0018] where i is the calculated subscript, H(f) is the label-bucket array mapping function, h(f) is an independent and uniform hash function, and M is the length of the bucket array;

[0019] Calculating the fingerprint of the flow to which the data packet belongs using formula (2):

[0020] fp = FP(f) = h fp (f) % X, (2)

[0021] where fp is the fingerprint, FP(f) is the label-fingerprint mapping function, h fp (f) is an independent and uniform hash function, and X represents the maximum value of the hash function;

[0022] Determining whether the fingerprint of the data packet matches the flow fingerprint field in the bucket it is mapped to;

[0023] When it is determined that the fingerprint of the data packet matches the flow fingerprint field in the bucket it is mapped to, increment the corresponding forward counter by one;

[0024] In the case where it is determined that the fingerprint of the data packet does not match the flow fingerprint field in the mapped storage bucket, determine whether the flow fingerprint field is empty;

[0025] In the case where it is determined that the flow fingerprint field is empty, update the flow fingerprint field to the fingerprint of the data packet and update the corresponding forward counter to one;

[0026] In the case where it is determined that the flow fingerprint field is not empty, decrement the corresponding reverse counter by one;

[0027] Calculate the ratio of the reverse counter and the forward counter of the storage bucket;

[0028] Determine whether the ratio is greater than or equal to a preset ratio threshold;

[0029] In the case where it is determined that the ratio is greater than or equal to the ratio threshold, form a binary tuple with the corresponding stored flow fingerprint and the forward counter, and clear the corresponding storage bucket.

[0030] Optionally, the error-prone flow sampling unit is used for:

[0031] Calculate a pseudo-random number associated with each binary tuple using formula (3):

[0032] r = Rand(fp) = rand(fp) % X, (3)

[0033] where r is the pseudo-random number, Rand(fp) is a preset fingerprint-random number mapping function, rand(fp) is an independent and uniform hash function, and X represents the maximum value of the hash function;

[0034] Determine whether the pseudo-random number is less than or equal to a preset random number threshold;

[0035] In the case where it is determined that the pseudo-random number is less than or equal to the preset random number threshold, input the binary tuple into the off-chip error-prone flow record unit and the small flow record unit.

[0036] Optionally, the small flow record unit is used for:

[0037] Calculate the subscript of each binary tuple in the small flow record unit using multiple formulas (4):

[0038] j (i) = H (i) (fp) = h (i) (fp) % w, (4)

[0039] where j (i) is the calculated subscript, H (i)(fp) is the fingerprint-column subscript mapping function of the small flow record unit, h (i) (fp) is an independent and uniform hash function, and w is the number of columns of the first two-dimensional counter array in the small flow record unit;

[0040] Successively obtain the counters mapped by the binary group in the first two-dimensional counter array.

[0041] Optionally, successively obtaining the counters mapped by the binary group in the first two-dimensional counter array includes:

[0042] For any mapped counter, increase the value of the positive counter of the binary group in the first preset number of bit fields before it;

[0043] Divide the value of the bit field after increasing the value of the positive counter by the water level threshold;

[0044] Update the calculation result of dividing by the water level threshold to the remaining bit fields of the counter.

[0045] Optionally, the off-chip error-prone flow record unit is used for:

[0046] Use multiple formulas (5) to calculate the subscript of each binary group in the off-chip error-prone flow record unit:

[0047] j (i),′ =G (i) (fp)=g (i) (fp) % w′, (5)

[0048] where j (i),′ is the subscript of the binary group in the off-chip error-prone flow record unit, G (i) (fp) is the fingerprint-column subscript function of the off-chip error-prone flow record unit, and g (i) (fp) is an independent and uniform hash function, and w′ is the number of columns of the second two-dimensional array of the off-chip error-prone flow record unit;

[0049] Successively obtain the counters mapped by the binary group in the second two-dimensional counter array;

[0050] Add the data stream corresponding to the binary group to the hash table, and increase the counter field value corresponding to the hash table by the count field included in the binary group to record the corresponding flow size.

[0051] Optionally, successively obtaining the counters mapped by the binary group in the second two-dimensional counter array includes:

[0052] For any mapped counter, increase the value of the positive counter of the binary group in the first two preset number of bit fields before it;

[0053] Divide the value of the bit field after increasing the value of the positive counter by the watermark threshold;

[0054] Update the calculation result of dividing by the watermark threshold to the remaining bit fields of the counter.

[0055] Optionally, the off-chip learning unit is configured to:

[0056] Retrieve the second two-dimensional counter array according to the flow label of the hash table of the off-chip error-prone flow recording unit to obtain the value of the counter of each flow in the second two-dimensional counter array;

[0057] Calculate the difference between the minimum value and the second minimum value of the value of the counter;

[0058] Determine whether the difference between the minimum value and the second minimum value is less than the error-prone flow threshold;

[0059] In the case where it is determined that the difference is greater than or equal to the error-prone flow threshold, determine that the flow corresponding to the flow label is an error-prone flow;

[0060] Classify the error-prone flow into multiple flow clusters according to the remaining bit fields of the error-prone flow;

[0061] Use multiple linear regression learners to learn the numerical relationship between the counter of the error-prone flow in each flow cluster and the actual flow size.

[0062] Optionally, the per-flow size estimation unit is configured to:

[0063] After the measurement time slice ends, store the output of the large-small flow separation recording unit, the output of the small flow recording unit, and the output of the off-chip learning unit;

[0064] Obtain the flow label to be queried;

[0065] Calculate the fingerprint of the flow label using formula (2):

[0066] fp = FP(f) = h fp (f) % X, (2)

[0067] where fp is the fingerprint, FP(f) is the label-fingerprint mapping function, h fp (f) is an independent and uniform hash function, and X represents the maximum value of the hash function;

[0068] Determine whether the fingerprint is in the bucket of the output of the size separation recording unit;

[0069] In the case where it is determined that the fingerprint is in the bucket of the output of the size separation recording unit, use the corresponding value of the positive counter as the returned estimated value;

[0070] In the case of determining that the fingerprint is not in the bucket output by the size separation recording unit, calculate the subscript of the flow label in the small flow recording unit according to formula (4):

[0071] j (i) = H (i) (fp)= h (i) (fp) % w, (4)

[0072] where j (i) is the calculated subscript, H (i) (fp) is the fingerprint-column subscript mapping function of the small flow recording unit, h (i) (fp) is an independent and uniform hash function, and w is the number of columns of the first two-dimensional counter array in the small flow recording unit;

[0073] Successively obtain the counters mapped by the binary tuple in the first two-dimensional counter array;

[0074] Calculate the difference between the minimum value and the second minimum value of the obtained counter values;

[0075] Determine whether the difference is less than a preset difference threshold;

[0076] In the case of determining that the difference is less than the difference threshold, use the minimum value as the returned estimated value;

[0077] In the case of determining that the difference is greater than or equal to the difference threshold, determine the flow cluster of the flow label according to the remaining bit fields corresponding to the flow label;

[0078] Determine the returned estimated value according to the linear regression learner corresponding to the determined flow cluster.

[0079] Through the above technical solutions, the embodiments of the present invention provide a high-precision flow size measurement system based on machine learning. The measurement system introduces machine learning technology into the on-chip memory, performs fine-grained partitioning and screening on the sampled traffic characteristics, and dynamically fits the distribution law of hash noise through a training model, thereby improving the fitting ability for hash noise. Through the measurement system provided by the present invention, it is possible to timely correct the size estimation of flows vulnerable to hash noise and significantly improve the measurement accuracy for small and medium-sized flows.

[0080] Other features and advantages of the embodiments of the present invention will be described in detail in the subsequent specific implementation part. Description of the Drawings

[0081] The accompanying drawings are used to provide a further understanding of the embodiments of the present invention, and constitute a part of the specification. Together with the following specific implementation manners, they are used to explain the embodiments of the present invention, but do not constitute a limitation to the embodiments of the present invention. In the accompanying drawings:

[0082] Figure 1 is a structural block diagram of a high-precision flow size measurement system based on machine learning according to an embodiment of the present invention;

[0083] Figure 2 is a flowchart of a method for performing size flow separation operation by a size flow separation recording unit according to an embodiment of the present invention;

[0084] Figure 3 is a flowchart of a method for performing random sampling operation by an error-prone flow sampling unit according to an embodiment of the present invention;

[0085] Figure 4 is a flowchart of a method for performing recording and fine-grained clustering operations by a small flow recording unit according to an embodiment of the present invention;

[0086] Figure 5 is a flowchart of a method for an error-prone flow recording unit to receive the result of a random sampling operation and form the input of a learning model according to an embodiment of the present invention;

[0087] Figure 6 is a flowchart of a method for performing machine learning by an on-chip learning unit according to an embodiment of the present invention;

[0088] Figure 7 is a flowchart of a method for a per-flow size estimation unit to respond to a query request according to an embodiment of the present invention. Specific Implementation Manner

[0089] The following details the specific implementation manners of the embodiments of the present invention in conjunction with the accompanying drawings. It should be understood that the specific implementation manners described herein are only used to illustrate and explain the embodiments of the present invention, and are not used to limit the embodiments of the present invention.

[0090] In the embodiments of the present application, some industry-existing solutions such as certain software, components, models, etc. may be mentioned. They should be regarded as exemplary, and their purpose is only to illustrate the feasibility in the implementation of the technical solutions of the present application, but it does not mean that the applicant has already or necessarily used this solution.

[0091] As Figure 1 shown is a structural block diagram of a high-precision flow size measurement system based on machine learning according to an embodiment of the present invention. In this Figure 1In this case, the measurement system may include a large / small flow separation and recording unit 1, an error-prone flow sampling unit 2, a small flow recording unit 3, an on-chip error-prone flow recording unit 4, an on-chip learning unit 5, and a per-flow size estimation unit 6. Among them, the large / small flow separation and recording unit 1 may be used to perform large / small flow separation operations on the incoming data stream. The error-prone flow sampling unit 2 may be used to perform random sampling operations on the separated error-prone flows. The small flow recording unit 3 may be used to record and perform fine-grained clustering operations on the separated small flows. The on-chip error-prone flow recording unit 4 may be used to receive the results of the random sampling operations to form the input of the learning model. The on-chip learning unit 5 may be used to divide the training data set into multiple flow clusters according to the input of the learning model and perform machine learning on each flow cluster separately. The per-flow size estimation unit 6 may be used to respond to query requests.

[0092] In this embodiment, the large / small flow separation and recording unit 1 may be used to perform large / small flow separation operations on the incoming data stream. Specifically, the large / small separation recording unit 1 may be composed of a bucket array B with a length of M, and each bucket may contain a flow fingerprint field B fp , a forward counter B + , a reverse counter B - and an overflow flag bit B flag in four parts. More specifically, when performing large / small flow separation operations, the large / small flow separation and recording unit 1 may execute the steps shown in Figure 2 . Specifically, in this Figure 2 , the method for performing large / small flow separation operations may include the following steps:

[0093] In step S10, construct a bucket array with a preset length. Among them, each bucket array includes a flow fingerprint field, a forward counter, a reverse counter, and an overflow flag bit;

[0094] In step S11, obtain each data packet arriving at the on-chip cache;

[0095] In step S12, extract the flow label field from the header of the data packet. Among them, the flow label field may be used to represent the network flow to which the data packet belongs. In addition, data packets with the same flow label belong to the same flow.

[0096] In step S13, calculate the subscript of the data packet in the bucket array according to formula (1):

[0097] i = H(f) = h(f) % M, (1)

[0098] where i is the calculated subscript, H(f) is the label-bucket array mapping function, h(f) is an independent and uniform hash function, and M is the length of the bucket array;

[0099] In step S14, the fingerprint of the flow to which the data packet belongs is calculated using formula (2):

[0100] fp = FP(f) = h fp (f) % X, (2)

[0101] where fp is the calculated fingerprint, FP(f) is the label - fingerprint mapping function, h fp (f) is an independent and uniform hash function, and X represents the maximum value of the hash function;

[0102] In step S15, it is determined whether the fingerprint of the data packet matches the flow fingerprint field of the bucket it is mapped to;

[0103] In step S16, when it is determined that the fingerprint of the data packet matches the flow fingerprint field in the bucket it is mapped to, the corresponding forward counter is incremented by one;

[0104] In step S17, when it is determined that the fingerprint of the data packet does not match the flow fingerprint field in the bucket it is mapped to, it is determined whether the flow fingerprint field is empty;

[0105] In step S18, when it is determined that the flow fingerprint field is empty, the flow fingerprint field is updated to the fingerprint of the data packet, and the corresponding forward counter is updated to one;

[0106] In step S19, when it is determined that the flow fingerprint field is not empty, the corresponding reverse counter is decremented by one;

[0107] In step S20, the ratio of the reverse counter and the forward counter of the bucket is calculated;

[0108] In step S21, it is determined whether the ratio is greater than or equal to a preset ratio threshold;

[0109] In step S22, when it is determined that the ratio is greater than or equal to the ratio threshold, the corresponding bucket is emptied, and the stored flow fingerprint and the forward counter are combined into a binary tuple, and the corresponding bucket is emptied.

[0110] In this embodiment, the error - prone flow sampling unit 2 can be used to perform a random sampling operation on the separated error - prone flows. Specifically, the error - prone flow sampling unit 2 can perform the random sampling operation by the method shown in Figure 3 As shown. In this Figure 3 The method for performing the random sampling operation may include the following steps:

[0111] In step S30, a pseudo - random number associated with each binary tuple is calculated using formula (3):

[0112] r = Rand(fp) = rand(fp) % X, (3)

[0113] Where r is a pseudo - random number, Rand(fp) is a preset fingerprint - random number mapping function, rand(fp) is an independent and uniform hash function, and X represents the maximum value of the hash function;

[0114] In step S31, it is judged whether the pseudo - random number is less than or equal to a preset random number threshold;

[0115] In step S32, when it is judged that the pseudo - random number is less than or equal to the preset random number threshold, the binary tuple is input into the on - chip error - prone flow recording unit and the small flow recording unit;

[0116] In step S33, when it is judged that the pseudo - random number is less than or equal to the preset random number threshold, sampling is not performed.

[0117] The small flow recording unit 3 can be used to record and perform fine - grained clustering operations on the separated small flows. In an example of the present invention, the small flow recording unit 3 may be composed of a first two - dimensional counter array S with d rows and w columns, each counter is composed of b bits, and is equipped with d independent hash functions H (i) (·). Where the first x bits of each counter are used to record the small flow size, and the subsequent b - x bits are used to record the coarse - grained flow size scale. Before the measurement starts, each counter in the counter array is initialized to 0. Specifically, the small flow recording unit 3 can be used to perform recording and fine - grained clustering operations through the method shown in Figure 4 . In this Figure 3 , the small flow recording unit 3 can be used to perform the following steps:

[0118] In step S40, multiple formulas (4) are used to calculate the subscript of each binary tuple in the small flow recording unit:

[0119] j (i) = H (i) (fp)= h (i) (fp) % w, (4)

[0120] Where j (i) is the calculated subscript, H (i) (fp) is the fingerprint - column subscript mapping function of the small flow recording unit, h (i) (fp) is an independent and uniform hash function, and w is the number of columns of the first two - dimensional counter array in the small flow recording unit;

[0121] In step S41, the counters mapped by the binary tuples in the first two-dimensional counter array are sequentially obtained. More specifically, in this example, the method for updating the counter can be that for any mapped counter, first increase the first preset value x number of bit fields in front of it by the forward counter B[i] of the binary tuple + value, then divide the value of the bit field after increasing the forward counter value by the water level threshold T, and finally update the calculation result of dividing by the water level threshold T to the remaining (b - x) bit fields of the counter.

[0122] The off-chip error-prone flow recording unit 4 can be used to receive the results of the random sampling operation to form the input of the learning model. In an example of the present invention, the off-chip error-prone flow recording unit 4 may be composed of a second two-dimensional counter array S' with d rows and w' columns and a lossless hash table HT. Each counter is composed of b' bits and is equipped with d independent hash functions G (i) (·). Among them, the first x' bits of each counter are used to record the small flow size, and the subsequent b' - x' bits are used to record the coarse-grained flow size scale. Where x' = x·p and b' = b·p. Before the measurement starts, each counter in the counter array is initialized to 0. Specifically, the off-chip error-prone flow recording unit 4 can be Figure 5 by the method shown in to receive the results of the random sampling operation to form the input of the learning model. In this Figure 5 , the off-chip error-prone flow recording unit 4 can be used to perform the following steps:

[0123] In step S50, multiple formulas (5) are used to calculate the subscript of each binary tuple in the off-chip error-prone flow recording unit:

[0124] j (i),′ =G (i) (fp)=g (i) (fp)%w′, (5)

[0125] Among them, j (i),′ is the subscript of the binary tuple in the off-chip error-prone flow recording unit, G (i) (fp) is the fingerprint-column subscript function of the off-chip error-prone flow recording unit, g (i) (fp) is an independent and uniform hash function, and w' is the number of columns of the second two-dimensional array of the off-chip error-prone flow recording unit;

[0126] In step S51, the counters mapped by the binary tuples in the second two-dimensional counter array are sequentially obtained. More specifically, in an example of the present invention, the method for updating the counter can be that for any mapped counter, first increase the first preset value x' number of bit fields in front of it by the forward counter B[i] of the binary tuple +· The value of p, then divide the value of the bit field (b′-x′) after increasing the value of the positive counter by the watermark threshold, and finally update the calculation result of dividing by the watermark threshold T to the remaining bit fields of the counter.

[0127] In step S52, add the data stream corresponding to the binary tuple to the hash table (lossless hash table HT), and increase the value of the corresponding counter field in the hash table by the count field (B[i] + · p) included in the binary tuple to record the corresponding stream size.

[0128] The on-chip learning unit 5 can be used to divide the training data set into multiple stream clusters according to the input of the learning model, and perform machine learning on each stream cluster separately. In an example of the present invention, the on-chip learning unit 5 may include multiple linear regression learners Among them, each linear regression learner is used to learn the numerical relationship between the values of d counters of error-prone streams in a certain stream cluster and the true stream size. Among them, n is the number of stream clusters, and lr i is the linear regression learner of the i-th stream cluster. Specifically, the on-chip learning unit 5 can perform machine learning by the method as Figure 6 shown. In this Figure 6 the on-chip learning unit 5 can be used to perform the following steps:

[0129] In step S60, retrieve the second two-dimensional counter array according to the stream label of the hash table of the on-chip error-prone stream recording unit 4 to obtain the values of the counters of each stream in the second two-dimensional counter array;

[0130] In step S61, calculate the difference between the minimum value and the second minimum value of the counter values;

[0131] In step S62, determine whether the difference between the minimum value and the second minimum value is less than the error-prone stream threshold;

[0132] In step S63, when it is determined that the difference is greater than or equal to the error-prone stream threshold, determine that the stream corresponding to the stream label is an error-prone stream;

[0133] In step S64, classify the error-prone streams into multiple stream clusters according to the remaining bit fields of the error-prone streams;

[0134] In step S65, use multiple linear regression learners to learn the numerical relationship between the counters of the error-prone streams in each stream cluster and the true stream size;

[0135] In step 66, when it is determined that the difference is less than the error-prone stream threshold, determine that the stream corresponding to the stream label is not an error-prone stream.

[0136] The per-flow size estimation unit 6 can be used to respond to a query request. Specifically, the per-flow size estimation unit 6 can respond to the query request by the method as Figure 7 shown. In this Figure 7 , the per-flow size estimation unit 6 can be used to perform the following steps:

[0137] In step S70, after the measurement time slice ends, store the outputs of the size-flow separation record unit 1, the small-flow record unit 3, and the off-chip learning unit 5;

[0138] In step S71, obtain the flow label to be queried;

[0139] In step S72, calculate the fingerprint of the flow label using formula (2):

[0140] fp = FP(f) = h fp (f) % X, (2)

[0141] where fp is the fingerprint, FP(f) is the label-fingerprint mapping function, h fp (f) is an independent and uniform hash function, and H represents the maximum value of the hash function;

[0142] In step S73, determine whether the fingerprint is in the bucket of the output of the size separation record unit 1;

[0143] In step S74, when it is determined that the fingerprint is in the bucket of the output of the size separation record unit 1, use the value of the corresponding forward counter as the returned estimated value;

[0144] In step S75, when it is determined that the fingerprint is not in the bucket of the output of the size separation record unit 1, calculate the subscript of the flow label in the small-flow record unit 3 according to formula (4):

[0145] J (i) = H (i) (fp) = h (i) (fp) % w, (4)

[0146] where j (i) is the calculated subscript, H (i) (fp) is the fingerprint-column subscript mapping function of the small-flow record unit 3, h (i) (fp) is an independent and uniform hash function, and w is the number of columns of the first two-dimensional counter array in the small-flow record unit 3;

[0147] In step S76, sequentially obtain the counters mapped by the binary tuples in the first two-dimensional counter array;

[0148] In step S77, calculate the difference between the minimum value and the second minimum value of the calculated counter value;

[0149] In step S78, determine whether the difference is less than a preset difference threshold;

[0150] In step S79, when it is determined that the difference is less than the difference threshold, use the minimum value as the returned estimated value.

[0151] In step S80, when it is determined that the difference is greater than or equal to the difference threshold, determine the flow cluster of the flow label according to the remaining bit fields corresponding to the flow label;

[0152] In step S81, determine the returned estimated value according to the linear regression learner corresponding to the determined flow cluster.

[0153] Through the above technical solution, the embodiment of the present invention provides a high-precision flow size measurement system based on machine learning. The measurement system introduces machine learning technology into the on-chip memory, performs fine-grained division and screening on the sampled traffic characteristics, and dynamically fits the distribution law of hash noise through a training model, thereby improving the fitting ability for hash noise. Through the measurement system provided by the present invention, it is possible to timely correct the size estimation of flows vulnerable to hash noise, and significantly improve the measurement accuracy for small and medium-sized flows.

[0154] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0155] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0156] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including an instruction means that implements the function specified in one or more of the blocks and / or processes. Figure 1 one or more of the processes and / or blocks Figure 1 specified in one or more of the blocks and / or processes.

[0157] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing steps for implementing the function specified in one or more of the blocks and / or processes. Figure 1 one or more of the processes and / or blocks Figure 1 specified in one or more of the blocks and / or processes.

[0158] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0159] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read only memory (ROM) or flash memory. Memory is an example of computer-readable media.

[0160] Computer-readable media includes both permanent and non-permanent, removable and non-removable media implemented by any method or technology for storage of information such as computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory or other memory technology, compact disc read only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0161] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus. Without further limitation, an element limited by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising the element.

[0162] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various modifications and variations can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A high-precision flow size measurement system based on machine learning, characterized in that, The measurement system includes: A large-small flow separation and recording unit for performing large-small flow separation operations on the incoming data stream; An error-prone flow sampling unit for performing random sampling operations on the separated error-prone flows; A small flow recording unit for recording and performing fine-grained clustering operations on the separated small flows; An on-chip error-prone flow recording unit for receiving the results of the random sampling operations to form the input of the learning model; An on-chip learning unit for dividing the training data set into multiple flow clusters according to the input of the learning model and performing machine learning on each flow cluster respectively; A per-flow size estimation unit for responding to query requests.

2. The measurement system according to claim 1, wherein, The large-small flow separation and recording unit is used for: Constructing an array of buckets with a preset length, where each bucket includes a flow fingerprint field, a forward counter, a reverse counter, and an overflow flag bit; Obtaining each data packet arriving at the on-chip cache; Extracting the flow label field from the header of the data packet; Calculating the subscript of the data packet in the bucket according to formula (1): i = H(f) = h(f) % M, (1) where i is the calculated subscript, H(f) is the label-bucket array mapping function, h(f) is an independent and uniform hash function, and M is the length of the bucket array; Calculating the fingerprint of the flow to which the data packet belongs using formula (2); fp = FP(f) = h fp (f) % X, (2) wherein, fp is the fingerprint, FP(f) is the tag-fingerprint mapping function, h fp (f) is an independent and uniform hash function, and X represents the maximum value of the hash function; Determining whether the fingerprint of the data packet matches the flow fingerprint field in the mapped bucket; When it is determined that the fingerprint of the data packet matches the flow fingerprint field in the mapped bucket, incrementing the corresponding forward counter by one; When it is determined that the fingerprint of the data packet does not match the flow fingerprint field in the mapped bucket, determining whether the flow fingerprint field is empty; When it is determined that the flow fingerprint field is empty, updating the flow fingerprint field to the fingerprint of the data packet and updating the corresponding forward counter to one; When it is determined that the flow fingerprint field is not empty, decrementing the corresponding reverse counter by one; Calculating the ratio of the reverse counter and the forward counter of the bucket; Determining whether the ratio is greater than or equal to a preset ratio threshold; When it is determined that the ratio is greater than or equal to the ratio threshold, forming a binary tuple of the corresponding stored flow fingerprint and the forward counter and clearing the corresponding bucket.

3. The measurement system according to claim 1, wherein The error-prone flow sampling unit is used for: Calculating a pseudo-random number associated with each binary tuple using formula (3): r = Rand(fp) = rand(fp) % X, (3) where r is the pseudo-random number, Rand(fp) is a preset fingerprint-random number mapping function, rand(fp) is an independent and uniform hash function, and X represents the maximum value of the hash function; Determining whether the pseudo-random number is less than or equal to a preset random number threshold; When it is determined that the pseudo-random number is less than or equal to the preset random number threshold, inputting the binary tuple into the on-chip error-prone flow recording unit and the small flow recording unit.

4. The measurement system according to claim 1, characterized in that, The small flow recording unit is used for: Calculating the subscript of each binary tuple in the small flow recording unit using multiple formula (4): j (i) = H (i) (fp) = h (i) (fp) %w, (4) where j (i) is the calculated subscript, H (i) (fp) is the fingerprint-column subscript mapping function of the small flow record unit, h (i) (fp) is an independent and uniform hash function, and w is the number of columns of the first two-dimensional counter array in the small flow record unit; Successively obtain the counters mapped by the binary tuples in the first two-dimensional counter array.

5. The measurement system according to claim 4, wherein Successively obtaining the counters mapped by the binary tuples in the first two-dimensional counter array includes: For any mapped counter, increase the value of the forward counter of the binary tuple by the value of the first preset number of bit fields before it; Divide the value of the bit field after increasing the value of the forward counter by the water level threshold; Update the calculation result of dividing by the water level threshold to the remaining bit fields of the counter.

6. The measurement system according to claim 1, wherein The off-chip error-prone flow recording unit is used for: Calculate the subscript of each binary tuple in the off-chip error-prone flow recording unit using multiple formulas (5); j (i),′ = G (i) (fp) = g (i) (fp)% w′, (5) where j (i),′ is the subscript of the binary tuple in the off-chip error-prone flow record unit, G (i) (fp) is the fingerprint-column subscript function of the off-chip error-prone flow record unit, g (i) (fp) is an independent and uniform hash function, and w′ is the number of columns of the second two-dimensional array of the off-chip error-prone flow record unit; Successively obtain the counters mapped by the binary tuples in the second two-dimensional counter array; Add the data stream corresponding to the binary tuple to the hash table, and increase the counter field value corresponding to the hash table by the count field included in the binary tuple to record the corresponding flow size.

7. The measurement system according to claim 6, characterized in that, Successively obtaining the counters mapped by the binary tuples in the second two-dimensional counter array includes: For any mapped counter, increase the value of the forward counter of the binary tuple by the value of the second preset number of bit fields before it; Divide the value of the bit field after increasing the value of the forward counter by the water level threshold; Update the calculation result of dividing by the water level threshold to the remaining bit fields of the counter.

8. The measurement system according to claim 1, characterized in that, The off-chip learning unit is used for: Retrieve the second two-dimensional counter array according to the flow label of the hash table of the off-chip error-prone flow recording unit to obtain the values of the counters of each flow in the second two-dimensional counter array; Calculate the difference between the minimum value and the second minimum value of the values of the counters; Judge whether the difference between the minimum value and the second minimum value is less than the error-prone flow threshold; In the case where it is judged that the difference is greater than or equal to the error-prone flow threshold, determine the flow corresponding to the flow label as an error-prone flow; Classify the error-prone flows into multiple flow clusters according to the remaining bit fields of the error-prone flows; Use multiple linear regression learners to learn the numerical relationship between the counters of the error-prone flows in each flow cluster and the true flow size.

9. The measurement system according to claim 1, characterized in that The per-flow size estimation unit is used for: After the measurement time slice ends, store the output of the large-small flow separation recording unit, the output of the small flow recording unit, and the output of the off-chip learning unit; Obtain the flow label to be queried; Calculate the fingerprint of the flow label using formula (2); fp = FP(f) = h fp (f) % X, (2) wherein, fp is the fingerprint, FP(f) is the label-fingerprint mapping function, h fp (f) is an independent and uniform hash function, and X represents the maximum value of the hash function; Judge whether the fingerprint is in the bucket of the output of the large-small separation recording unit; In the case where it is judged that the fingerprint is in the bucket of the output of the large-small separation recording unit, use the value of the corresponding forward counter as the returned estimated value; In the case where it is judged that the fingerprint is not in the bucket of the output of the large-small separation recording unit, calculate the subscript of the flow label in the small flow recording unit according to formula (4); J (i) = H (i) (fp) = h (i) (fp)%w, (4) where j (i) is the calculated subscript, H (i) (fp) is the fingerprint-column subscript mapping function of the small flow record unit, h (i) (fp) is an independent and uniform hash function, and w is the number of columns of the first two-dimensional counter array in the small flow record unit; Successively obtain the counters mapped by the binary tuples in the first two-dimensional counter array; The difference between the minimum value and the second minimum value of the calculated counter values; Judge whether the difference is less than the preset difference threshold; In the case where it is judged that the difference is less than the difference threshold, use the minimum value as the returned estimated value; In the case of determining that the difference is greater than or equal to the difference threshold, determine the flow cluster of the flow label according to the remaining bit fields corresponding to the flow label; Determine the returned estimated value according to the linear regression learner corresponding to the determined flow cluster.