A method of fine-grained network traffic measurement that filters both large and small flows
Patent Information
- Application Number
- CN202311641303.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-01
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-12-01
AI Technical Summary
而小流的数据包很少,所以容易发生测漏或者对数据包的多测,而这些情况都会引起很大的测量误差
[0038]本发明过滤大流和小流的精细化网络流量测量方法,首先使用级联哈希表结构来进行大小流分离,并对大流进行测量,然后采用布隆过滤器组来对极小流进行精确测量,最后采用sketch结构来对小流进行估计,从而克服小流的漏测或者多测问题,实现精细化网络流量测量,提高流量测量的准确性。
Smart Images

Figure CN117692369B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of network management technology, and more specifically, relates to a refined network traffic measurement method for filtering large and small flows. Background Technology
[0002] Network measurement is a method for evaluating and quantifying network performance, availability, security, and quality by collecting and analyzing network data. It plays a crucial role in assessing network performance, troubleshooting, security assessment, planning and optimization, and business decision-making. Through network measurement, network administrators and users can understand the efficiency and reliability of the network, promptly identify and resolve issues to improve network performance and availability. Furthermore, network measurement can help protect the network from malicious activities and provide valuable data support for business development and strategy formulation. In summary, network measurement can meet the needs of network performance and security assessment and provide effective guidance and support for network planning and decision-making.
[0003] Topology measurement, performance measurement, and traffic measurement are the three main aspects of network measurement, providing a comprehensive assessment of the network's capabilities based on their different measurement targets. Topology measurement focuses primarily on the network's structure and connections, revealing its topology by collecting and analyzing connection information between network devices. This measurement helps understand the network's physical layout, the relationships between nodes, and packet path selection. Topology measurement is crucial for network planning and troubleshooting, helping administrators identify potential single points of failure and bottlenecks, and providing optimization suggestions. Performance measurement is key to evaluating network performance, primarily focusing on metrics such as transmission rate, latency, throughput, and packet loss rate. By collecting and analyzing these metrics, the network's efficiency, reliability, and responsiveness can be assessed. Performance measurement is essential for network optimization and troubleshooting. It helps administrators identify network bottlenecks and performance issues, allowing them to take appropriate measures to improve network performance and user experience. Traffic measurement focuses on data traffic within the network. By monitoring and analyzing packet flow, it reveals the actual data traffic, traffic distribution, and traffic characteristics within the network. Traffic measurement is crucial for optimizing network resource allocation, identifying abnormal traffic, and network attacks. It can help administrators understand the network load, adjust bandwidth allocation in a timely manner, ensure the transmission of important data, and identify and isolate potential security threats.
[0004] Among the various aspects of network measurement, traffic measurement is a crucial task. By monitoring and analyzing data traffic within a network, we can gain in-depth insights into the actual data transmission, traffic distribution, and traffic characteristics of the network.
[0005] Among existing network measurement technologies, commonly used traffic measurement methods include CM-Sketch, CU-Sketch, and HashPipe. These methods have certain advantages and limitations in estimating the frequency or unique count of different elements in a data stream.
[0006] CM-Sketch is a probabilistic data structure that uses a hash function and an array of counters to estimate the frequency of different elements. However, the accuracy of CM-Sketch is limited by hash collisions and estimation errors of low-frequency elements.
[0007] CU-Sketch is an improved flow measurement method that focuses on estimating the unique count of distinct elements in a data stream. Compared to CM-Sketch, CU-Sketch reduces storage requirements and provides a more accurate unique count estimate. However, it is still subject to hash collisions and estimation errors.
[0008] To address the hash collisions and estimation errors inherent in traffic measurement methods such as CM-Sketch and CU-Sketch, researchers have proposed a method that separates large and small flows. This method processes large and small flows separately, placing large flows in an independent storage structure for measurement, while placing small flows in other storage structures. This effectively reduces hash collisions and counter overflows and balances storage resource utilization. For example, Hash Pipe uses a hash pipeline approach to separate large and small flows, effectively reducing hash collisions and counter overflows; however, it still has some errors when processing small flows. Eastic-Sketch separates large and small flows through traffic sampling, adapting to high-speed and large-scale data flows with lower memory requirements and estimation errors; however, it may increase estimation errors when processing high-frequency elements. Furthermore, both Hash Pipe and Eastic-Sketch exhibit significant measurement errors for extremely small flows, i.e., flows containing only 1, 2, 3, etc., with a very small number of packets.
[0009] Analysis of traffic traces captured from the research network revealed a large number of small flows and a small number of large flows. Small flows contain very few packets, making them prone to missed detections or over-measurement, both of which introduce significant measurement errors. Further analysis of existing measurement errors revealed numerous extremely small flows with only one or two packets per packet. When estimating these flows using sketching techniques, these flows often result in measurements with errors multiplied by a factor of two. Summary of the Invention
[0010] The purpose of this invention is to overcome the shortcomings of the prior art and provide a refined network traffic measurement method that filters large and small flows to overcome the problems of missed or excessive small flow measurement and achieve refined network traffic measurement.
[0011] To achieve the above-mentioned objectives, the present invention provides a refined network traffic measurement method for filtering large and small flows, characterized by comprising the following steps:
[0012] (1) Insertion mechanism
[0013] 1.1) Cascaded hash table insertion mechanism
[0014] The design includes two cascaded hash tables: a first-level hash table and a second-level hash table. Data packets arrive at the cascaded hash table first, leaving the large flow to be measured in the cascaded hash table, while the small flow is driven to the subsequent Bloom filter.
[0015] 1.1.1) For a first-level hash table, when a data packet arrives at the first-level hash table, its bucket position corresponding to its insertion into the first-level hash table is calculated using the hash function h1 based on its flow ID:
[0016] If the hash bucket at this location is empty, insert the data packet directly into the hash bucket and set the packet count of the flow to 1;
[0017] If the hash bucket at that location is not empty, the flow ID of the data packet is matched with the flow ID recorded in the hash bucket. If they are the same, it indicates that no hash collision has occurred, and the packet count of the corresponding flow is directly incremented by 1. If the flow ID of the data packet is different from the flow ID recorded in the hash bucket, it indicates that a hash collision has occurred. Based on the packet count of the corresponding flow recorded in the hash bucket, two sub-cases are considered: if the packet count of the corresponding flow recorded in the hash bucket is greater than or equal to 2, it indicates that the flow corresponding to the hash bucket is more likely to be a large flow, and the arriving data packet is expelled to the secondary hash table. Conversely, if the packet count of the corresponding flow recorded in the hash bucket is equal to 1, the arriving data packet is inserted into the hash bucket at that location, and the data packet in the hash bucket is expelled to the secondary hash table.
[0018] 1.1.2) For a two-level hash table, based on the flow ID of the packet evicted to the two-level hash table, the bucket position corresponding to the packet's insertion into the two-level hash table is calculated using the hash function h2:
[0019] If the hash bucket at this location is empty, insert the data packet directly into the hash bucket and set the packet count of the flow to 1;
[0020] If the hash bucket at that location is not empty, the flow ID of the arriving data packet is matched with the flow ID recorded in the hash bucket. If they are the same, it indicates that no hash collision has occurred, and the packet count of the corresponding flow is directly incremented by 1. If the flow ID of the data packet and the flow ID recorded in the hash bucket are different, it indicates that a hash collision has occurred. Based on the packet count of the corresponding flow recorded in the hash bucket, two sub-cases are considered: if the packet count of the corresponding flow recorded in the hash bucket is greater than or equal to 2, it indicates that the flow corresponding to the hash bucket is more likely to be a large flow, and the arriving data packet is expelled to the secondary hash table. Conversely, if the packet count of the corresponding flow recorded in the hash bucket is equal to 1, the arriving data packet is inserted into the hash bucket at that location, and the data packet in the hash bucket is expelled to the Bloom filter group.
[0021] 1.2) Bloom filter bank insertion mechanism
[0022] Design a Bloom filter group that includes three Bloom filters. When a data packet arrives at the Bloom filter group, first check whether the first Bloom filter contains the flow ID of the data packet. If it does not contain it, calculate the hash value of the flow ID of the data packet using the hash function h3 and set the corresponding counter value to 1. If it contains it, check whether the second Bloom filter contains the flow ID of the data packet.
[0023] If the second Bloom filter does not contain the flow ID of the packet, then the flow ID of the packet is hashed using the hash function h4, and the corresponding counter value is set to 1. If it is contained, then it is determined whether the third Bloom filter contains the flow ID of the packet.
[0024] If the third Bloom filter does not contain the flow ID of the packet, then the flow ID of the packet is hashed using the hash function h5, and the corresponding counter value is set to 1. If it is contained, the packet is expelled to the sketch for measurement.
[0025] 1.3) Sketch structure insertion mechanism
[0026] Packets expelled by the Bloom filter are used for traffic estimation using a sketch structure;
[0027] (2) Query mechanism
[0028] 2.1) Bloom filter group query mechanism
[0029] For stream Fn, the stream ID is hashed using hash function h3. The system checks if all the counter values corresponding to the hash values in the first Bloom filter are 1. If not all are 1, the first Bloom filter does not contain the stream, indicating that the stream has not been inserted into the Bloom filter group, and its reading is 0. If all are 1, the first Bloom filter contains the stream.
[0030] The stream ID is hashed using the hash function h4. The system then checks if all the counter values corresponding to the hash values in the second Bloom filter are 1. If not all are 1, the second Bloom filter does not contain the stream, meaning the stream has not been inserted into the second Bloom filter, and its reading is 1. If all are 1, the second Bloom filter contains the stream.
[0031] The stream ID is hashed using the hash function h5. The counter values corresponding to the hash values in the third Bloom filter are then checked to see if they are all 1. If they are not all 1, the third Bloom filter does not contain the stream, indicating that the stream has not been inserted into the third Bloom filter, and the stream reading is 2. If they are all 1, the third Bloom filter contains the stream, and the stream reading is 3.
[0032] 2.2) Query mechanism for flow measurement methods based on filtering large and small flows
[0033] For a stream, firstly, the bucket position corresponding to the stream in the first and second level hash tables is calculated based on the stream ID using hash functions h1 and h2, and then compared with the stream ID in the corresponding hash bucket. If they are the same, the number of packets in the stream is used as the cascade hash table reading of the stream. If they are different in both the first and second level hash tables, the cascade hash table reading of the stream is 0.
[0034] Then, based on the stream readings in the Bloom filter bank, the queries are divided into two sub-cases:
[0035] If the flow reading in the Bloom filter group is less than the number of Bloom filters, it means that the flow has not been expelled into the sketch structure by the Bloom filter group. The flow rate is: cascade hash table reading + Bloom filter group reading.
[0036] If the flow reading in the Bloom filter group is greater than or equal to the number of Bloom filters, it means that the flow may be expelled into the sketch structure by the Bloom filter group. Query the sketch structure according to the flow ID to get the sketch reading of the flow. The flow volume is: cascade hash table reading + Bloom filter group reading + sketch reading.
[0037] The objective of this invention is achieved as follows:
[0038] This invention provides a refined network traffic measurement method that filters large and small flows. First, a cascaded hash table structure is used to separate large and small flows and measure the large flows. Then, a Bloom filter array is used to accurately measure the small flows. Finally, a sketch structure is used to estimate the small flows, thereby overcoming the problem of missed or over-measured small flows, achieving refined network traffic measurement, and improving the accuracy of traffic measurement. Attached Figure Description
[0039] Figure 1 This is a flowchart of the refined network traffic measurement method for filtering large and small flows according to the present invention;
[0040] Figure 2 This is a schematic diagram illustrating the measurement principle of the refined network flow measurement method for filtering large and small flows according to the present invention.
[0041] Figure 3 This is a schematic diagram of data packet insertion into a Bloom filter;
[0042] Figure 4 This is a diagram illustrating a Bloom filter query.
[0043] Figure 5 This is a comparison chart of average relative measurement error curves;
[0044] Figure 6 This is a comparison chart of weighted average relative measurement error curves;
[0045] Figure 7 This is a comparison chart of entropy estimation relative to measurement error curves;
[0046] Figure 8 This is a comparison chart of high-volume detection curves;
[0047] Figure 9 This is a comparison chart of the relative errors of the base estimation;
[0048] Figure 10 This is a comparison chart of throughput estimations. Detailed Implementation
[0049] The specific embodiments of the present invention will now be described with reference to the accompanying drawings to enable those skilled in the art to better understand the invention. It should be particularly noted that in the following description, detailed descriptions of known functions and designs that might obscure the main content of the invention will be omitted here.
[0050] Analysis of existing measurement errors reveals numerous extremely small flows with only one or two packets per packet in the network. When estimating these flows using the sketch method, measurement errors are often multiplied. Therefore, this invention proposes a refined network traffic measurement method employing a cascaded hash table structure, a Bloom filter array, and a sketch structure. First, a cascaded hash table structure is used to separate large and small flows, and the large flows are measured. Then, a Bloom filter array is used to accurately measure the extremely small flows. Finally, a sketch structure is used to estimate the small flows, thereby overcoming the problems of missed or over-measured small flows, achieving refined network traffic measurement, and improving the accuracy of traffic measurement.
[0051] Compared with traditional network traffic measurement methods, this invention has the following characteristics:
[0052] (1) A cascaded hash table structure is used to separate large and small flows, and the large flow is measured;
[0053] (2) Propose a Bloom filter array to accurately count the minimum flow;
[0054] (3) Use the sketch structure to estimate the small flow.
[0055] In addition, the present invention also includes a cascaded hash table structure design for separating small and large flows, and a Bloom filter group design for recording minimal flows.
[0056] Figure 1 , 2 These are, respectively, a flowchart and a schematic diagram of the refined network traffic measurement method for filtering large and small flows according to the present invention.
[0057] In this embodiment, as Figure 1 , 2 As shown, the refined network traffic measurement method of the present invention, which filters large and small flows, includes the following steps:
[0058] Step S1: Insertion Mechanism
[0059] Step S1.1: Cascaded Hash Table Insertion Mechanism
[0060] The design includes two cascaded hash tables: a first-level hash table and a second-level hash table. Data packets first arrive at the cascaded hash table, where large flows are measured and small flows are driven to the subsequent Bloom filter.
[0061] In this embodiment, the flow ID is set as the key value of the hash table for hash calculation. The flow ID is the flow's quintuple, source and destination IPs, etc. In specific implementation, the number of packets in the flow can be other measurement characteristics.
[0062] Step S1.1.1: First-level hash table insertion mechanism
[0063] For a first-level hash table, when a data packet arrives at the first-level hash table, its bucket position for insertion into the first-level hash table is calculated based on its flow ID using the hash function h1.
[0064] If the hash bucket at this location is empty, insert the data packet directly into the hash bucket and set the packet count of the flow to 1;
[0065] If the hash bucket at that location is not empty, the flow ID of the data packet is matched with the flow ID recorded in the hash bucket. If they are the same, it indicates that no hash collision has occurred, and the packet count of the corresponding flow is directly incremented by 1. If the flow ID of the data packet and the flow ID recorded in the hash bucket are different, it indicates that a hash collision has occurred. Based on the packet count of the corresponding flow recorded in the hash bucket, two sub-cases are considered: if the packet count of the corresponding flow recorded in the hash bucket is greater than or equal to 2, it indicates that the flow corresponding to the hash bucket is more likely to be a large flow, and the arriving data packet is expelled to the secondary hash table. Conversely, if the packet count of the corresponding flow recorded in the hash bucket is equal to 1, the arriving data packet is inserted into the hash bucket at that location, and the data packet in the hash bucket is expelled to the secondary hash table.
[0066] Step S1.1.2: Second-level hash table insertion mechanism
[0067] For a two-level hash table, the bucket position corresponding to the packet to be inserted into the two-level hash table is calculated using the hash function h2, based on the flow ID of the packet evicted to the two-level hash table.
[0068] If the hash bucket at this location is empty, insert the data packet directly into the hash bucket and set the packet count of the flow to 1;
[0069] If the hash bucket at that location is not empty, the flow ID of the arriving data packet is matched with the flow ID recorded in the hash bucket. If they are the same, it indicates that no hash collision has occurred, and the packet count of the corresponding flow is directly incremented by 1. If the flow ID of the data packet and the flow ID recorded in the hash bucket are different, it indicates that a hash collision has occurred. Based on the packet count of the corresponding flow recorded in the hash bucket, two sub-cases are considered: if the packet count of the corresponding flow recorded in the hash bucket is greater than or equal to 2, it indicates that the flow corresponding to the hash bucket is more likely to be a large flow, and the arriving data packet is expelled to the secondary hash table. Conversely, if the packet count of the corresponding flow recorded in the hash bucket is equal to 1, the arriving data packet is inserted into the hash bucket at that location, and the data packet in the hash bucket is expelled to the Bloom filter group.
[0070] Step S1.2: Bloom filter bank insertion mechanism
[0071] Design a Bloom filter group that includes three Bloom filters. When a data packet arrives at the Bloom filter group, first check whether the first Bloom filter contains the flow ID of the data packet. If it does not contain it, calculate the hash value of the flow ID of the data packet using the hash function h3 and set the corresponding counter value to 1. If it contains it, check whether the second Bloom filter contains the flow ID of the data packet.
[0072] If the second Bloom filter does not contain the flow ID of the packet, then the flow ID of the packet is hashed using the hash function h4, and the corresponding counter value is set to 1. If it is contained, then it is determined whether the third Bloom filter contains the flow ID of the packet.
[0073] If the third Bloom filter does not contain the flow ID of the packet, then the flow ID of the packet is hashed using the hash function h5, and the corresponding counter value is set to 1. If it is contained, the packet is expelled to the sketch for measurement.
[0074] Bloom filter banks are primarily used for precise measurement of extremely small flows with only one or two packets. The principle is that a single Bloom filter can determine the existence of a flow. If only one Bloom filter contains the flow, it indicates that the flow has one packet; if two consecutive Bloom filters contain the flow, it indicates that the flow has two packets, and so on. A Bloom filter bank containing three Bloom filters can accurately measure flows with both one and two packets.
[0075] Figure 3 This is a diagram illustrating packet insertion into a Bloom filter.
[0076] like Figure 3 As shown in (a), the number of packets in flow F1.1 arriving at the Bloom filter group is 1. Therefore, firstly, it is determined whether the first Bloom filter contains flow F1. If not, the flow ID is hashed and the corresponding counter value is set to 1; if flow F1 is contained, then... Figure 3 As shown in (b), it is determined whether the second Bloom filter contains the flow F1. If it does not contain it, the flow F1.1 is inserted into the second Bloom filter.
[0077] If so Figure 3 As shown in (c), the number of packets in flow F2.2 arriving at the Bloom filter group is 2. First, it is determined whether the first Bloom filter contains flow F2. If not, flow F2.2 is inserted. Since the number of packets is 2, flow F2.2 is then inserted into the second Bloom filter.
[0078] If so Figure 3 As shown in (d), flow F1.3 arrives at the Bloom filter group with a packet count of 3. Since the first and second Bloom filters already contain flow F1, and the third Bloom filter does not contain flow F1, flow F1.3 is inserted into the third Bloom filter, and the remaining two packets (flow F1.2) that are not inserted into the filter are expelled into the sketch structure for measurement.
[0079] Step S1.3: Sketch structure insertion mechanism
[0080] Packets expelled by the Bloom filter indicate that the flow is not a large flow, and its number of packets is greater than 3, nor is it a very small flow. Therefore, the sketch structure is used for flow estimation.
[0081] Step S2: Query Mechanism
[0082] Step S2.1: Bloom filter group query mechanism
[0083] Figure 4 This is a diagram illustrating a Bloom filter query.
[0084] In this embodiment, as Figure 4 As shown, for stream Fn, the stream ID is hashed using hash function h3. The system checks if all the counter values corresponding to the hash values in the first Bloom filter BF1 are 1. If not all are 1 (false), then the first Bloom filter BF1 does not contain the stream, indicating that the stream has not been inserted into the Bloom filter group, and the stream's count (packet count) is 0. If all are 1 (true), then the first Bloom filter BF1 contains the stream.
[0085] The stream ID is hashed using the hash function h4. The system then checks if all the counter values corresponding to the hash values in the second Bloom filter BF2 are 1. If not all are 1 (false), the second Bloom filter BF2 does not contain the stream, meaning the stream has not been inserted into BF2, and its count is 1. If all are 1 (true), the second Bloom filter BF2 contains the stream.
[0086] The stream ID is hashed using the hash function h5. The system then checks if all the counter values corresponding to the hash values in the third Bloom filter BF3 are 1. If not all are 1 (false), the third Bloom filter BF3 does not contain the stream, meaning the stream has not been inserted into the third Bloom filter BF3, and the stream's reading is 2. If all are 1 (true), the third Bloom filter BF3 contains the stream, and the stream's reading is 3.
[0087] Step S2.2: Query mechanism for flow measurement methods based on filtering large and small flows
[0088] For a stream, firstly, the bucket position corresponding to the stream in the first and second level hash tables is calculated based on the stream ID using hash functions h1 and h2, and then compared with the stream ID in the corresponding hash bucket. If they are the same, the number of packets in the stream is used as the cascade hash table reading of the stream. If they are different in both the first and second level hash tables, the cascade hash table reading of the stream is 0.
[0089] Then, based on the stream readings in the Bloom filter bank, the queries are divided into two sub-cases:
[0090] If the flow reading in the Bloom filter group is less than the number of Bloom filters, it means that the flow has not been expelled into the sketch structure by the Bloom filter group. The flow rate is: cascade hash table reading + Bloom filter group reading.
[0091] If the flow reading in the Bloom filter group is greater than or equal to the number of Bloom filters, it means that the flow may be expelled into the sketch structure by the Bloom filter group. Query the sketch structure according to the flow ID to get the sketch reading of the flow. The flow volume is: cascade hash table reading + Bloom filter group reading + sketch reading.
[0092] Simulation results
[0093] This invention's refined network traffic measurement method for filtering large and small flows is evaluated using average relative error, weighted average relative error, entropy estimation based on flow size distribution, heavy hitter estimation, cardinality estimation, and throughput estimation. Simulation experiments measured 1 million packets, approximately 100,000 flows. In the Elastic Sketch method, the heavy part allocates 20% of the total memory. In this invention, the first and second-level hash tables also allocate 20% of the total memory; all sketches are allocated in 3 rows.
[0094] 1. Mean Relative Error (ARE)
[0095] The mean relative measurement error, or ARE, is calculated using the following formula:
[0096]
[0097] Where f i For the true value of the i-th stream, Let be the measurement value of the i-th flow. There are a total of n flows.
[0098] like Figure 5 As shown, this invention compares with existing methods such as Elastic Sketch, CM Sketch, CU Sketch, and CountSketch. Simulation results show that, under the same memory allocation, this invention (OUR) has higher accuracy than the other four methods.
[0099] 2. Weighted Average Relative Error (WMRE)
[0100] The weighted average relative error, or WMRE, is calculated using the following formula:
[0101]
[0102] Where n i For a stream of size i, Let z be the measured value of the flow, where z represents the maximum number of packets in the flow.
[0103] like Figure 6As shown, the measurement results indicate that, under the same memory size, the WMRE of this invention (OUR) is superior to other methods; and the larger the allocated memory, the higher the measurement accuracy.
[0104] 3. Entropy estimation based on flow size distribution
[0105] Entropy, an entropy estimation based on the flow size distribution, is calculated using the following formula:
[0106]
[0107] Where n i Let i be the number of flows of size i, and m represent the total number of flows measured.
[0108] The relative error between the actual flow entropy value and the measured flow entropy value was calculated. The measurement results show that, for example... Figure 7 As shown, the present invention (OUR) achieves better measurement accuracy than the other four methods when allocating the same amount of memory.
[0109] 4. Heavy Hitter Detection
[0110] Large flow detection identifies flows whose size exceeds a certain threshold.
[0111] Experimental results show that, Figure 8 As shown, when the memory is greater than 0.6M, the detection results of the present invention (OUR) and other measurement methods are all very good. When the memory is less than 0.4M, the measurement results of the present invention (OUR) are better than those of CM Sketch and Count Sketch.
[0112] 5. Baseline estimation
[0113] Cardinality estimation, i.e., estimating the number of flows. For the Sketch part, we use the Linear Count method to calculate the cardinality; for the heavy part of Elastic, its cardinality is the number of non-empty buckets in the hash table; for the Bloom filter in our measurement method, we propose the following formula for cardinality calculation:
[0114] Suppose the Bloom filter has m bits, b bits are marked as 1, and k hash functions, and that it contains e different source IPs.
[0115] After inserting e distinct source IPs, the probability that a certain bit is set to 1 is:
[0116]
[0117] Given that after inserting e distinct source IPs, the probability that a certain bit in the Bloom filter will be set to 1 is: Therefore, the number of different source IPs, e, can be deduced as:
[0118]
[0119] The relative error of the base value is now calculated for this invention and four other methods, such as... Figure 9 As shown in the experimental results, except for Count Sketch, which has a relatively large measurement error, the other four methods have relatively small measurement errors.
[0120] 6. Throughput estimation
[0121] This section describes the number of data packets a switch can process per unit of time using different measurement methods. For example... Figure 10 As shown in the experimental results, Elastic Sketch has the highest throughput, followed by this invention.
[0122] Although the illustrative specific embodiments of the present invention have been described above to enable those skilled in the art to understand the invention, it should be understood that the invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the invention as defined and determined by the appended claims, and all inventions utilizing the concept of the present invention are protected.
Claims
1. A refined network traffic measurement method that filters large and small flows, characterized by comprising the following steps: (1) Insertion mechanism; 1.1) Cascaded hash table insertion mechanism; The design includes two cascaded hash tables: a first-level hash table and a second-level hash table. Data packets arrive at the cascaded hash table first, leaving the large flow to be measured in the cascaded hash table, while the small flow is driven to the subsequent Bloom filter. 1.1.1) For a first-level hash table, when a data packet arrives at the first-level hash table, its bucket position corresponding to its insertion into the first-level hash table is calculated using the hash function h1 based on its flow ID: If the hash bucket at that location is empty, insert the data packet directly into that hash bucket and set the packet count of the corresponding stream to 1; If the hash bucket at that location is not empty, the flow ID of the data packet is matched with the flow ID recorded in the hash bucket. If they are the same, it indicates that no hash collision has occurred, and the packet count of the corresponding flow is directly incremented by 1. If the flow ID of the data packet is different from the flow ID recorded in the hash bucket, it indicates that a hash collision has occurred. Based on the packet count of the corresponding flow recorded in the hash bucket, two sub-cases are considered: if the packet count of the corresponding flow recorded in the hash bucket is greater than or equal to 2, it indicates that the flow corresponding to the hash bucket is more likely to be a large flow, and the arriving data packet is expelled to the secondary hash table. Conversely, if the packet count of the corresponding flow recorded in the hash bucket is equal to 1, the arriving data packet is inserted into the hash bucket at that location, and the data packet in the hash bucket is expelled to the secondary hash table. 1.1.2) For a two-level hash table, based on the flow ID of the packet evicted to the two-level hash table, the bucket position corresponding to the packet's insertion into the two-level hash table is calculated using the hash function h2: If the hash bucket at this location is empty, insert the data packet directly into the hash bucket and set the packet count of the flow to 1; If the hash bucket at that location is not empty, the flow ID of the arriving data packet is matched with the flow ID recorded in the hash bucket. If they are the same, it indicates that no hash collision has occurred, and the packet count of the corresponding flow is directly incremented by 1. If the flow ID of the data packet and the flow ID recorded in the hash bucket are different, it indicates that a hash collision has occurred. Based on the packet count of the corresponding flow recorded in the hash bucket, two sub-cases are considered: if the packet count of the corresponding flow recorded in the hash bucket is greater than or equal to 2, it indicates that the flow corresponding to the hash bucket is more likely to be a large flow, and the arriving data packet is expelled to the secondary hash table. Conversely, if the packet count of the corresponding flow recorded in the hash bucket is equal to 1, the arriving data packet is inserted into the hash bucket at that location, and the data packet in the hash bucket is expelled to the Bloom filter group. 1.2) Bloom filter bank insertion mechanism; Design a Bloom filter group that includes three Bloom filters. When a data packet arrives at the Bloom filter group, first check whether the first Bloom filter contains the flow ID of the data packet. If it does not contain it, calculate the hash value of the flow ID of the data packet using the hash function h3 and set the corresponding counter value to 1. If it contains it, check whether the second Bloom filter contains the flow ID of the data packet. If the second Bloom filter does not contain the flow ID of the packet, then the flow ID of the packet is hashed using the hash function h4, and the corresponding counter value is set to 1. If it is contained, then it is determined whether the third Bloom filter contains the flow ID of the packet. If the third Bloom filter does not contain the flow ID of the packet, the flow ID of the packet is hashed using the hash function h5, and the corresponding counter value is set to 1. If it is contained, the packet is expelled to the sketch for measurement. 1.3) Sketch structure insertion mechanism; Packets expelled by the Bloom filter are used for traffic estimation using a sketch structure; (2) Query mechanism; 2.1) Bloom filter group query mechanism; For stream Fn, the stream ID is hashed using hash function h3. The system checks if all the counter values corresponding to the hash values in the first Bloom filter are 1. If not all are 1, the first Bloom filter does not contain the stream, indicating that the stream has not been inserted into the Bloom filter group, and its reading is 0. If all are 1, the first Bloom filter contains the stream. The stream ID is hashed using the hash function h4. The system then checks if all the counter values corresponding to the hash values in the second Bloom filter are 1. If not all are 1, the second Bloom filter does not contain the stream, meaning the stream has not been inserted into the second Bloom filter, and its reading is 1. If all are 1, the second Bloom filter contains the stream. The stream ID is hashed using the hash function h5. The counter values corresponding to the hash values in the third Bloom filter are then checked to see if they are all 1. If they are not all 1, the third Bloom filter does not contain the stream, indicating that the stream has not been inserted into the third Bloom filter, and the stream reading is 2. If they are all 1, the third Bloom filter contains the stream, and the stream reading is 3. 2.2) Query mechanism based on flow measurement methods that filter large and small flows; For a stream, firstly, the bucket position corresponding to the stream in the first and second level hash tables is calculated based on the stream ID using hash functions h1 and h2 respectively, and then compared with the stream ID in the corresponding hash bucket. If they are the same, the number of packets in the stream is used as the cascade hash table reading of the stream. If they are different in both the first and second level hash tables, the cascade hash table reading of the stream is 0. Then, based on the stream readings in the Bloom filter bank, the queries are divided into two sub-cases: If the flow reading in the Bloom filter group is less than the number of Bloom filters, it means that the flow has not been expelled into the sketch structure by the Bloom filter group. The flow rate is: cascade hash table reading + Bloom filter group reading. If the flow reading in the Bloom filter group is greater than or equal to the number of Bloom filters, it means that the flow may be expelled into the sketch structure by the Bloom filter group. Query the sketch structure according to the flow ID to get the sketch reading of the flow. The flow volume is: cascade hash table reading + Bloom filter group reading + sketch reading.