A Hierarchical Large Flow Detection Method Compatible with Programmable Hardware

By adopting a structure based on storage buckets and storage slots in layered large traffic detection, the problem of low storage space utilization and high difficulty in compatible programmable hardware in the prior art is solved, and efficient traffic detection is achieved.

CN119544624BActive Publication Date: 2025-06-13SUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510068242.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-06-13
Estimated Expiration
2045-01-16

AI Technical Summary

Technical Problem

The existing hierarchical large-traffic detection methods have problems such as low storage space utilization, low throughput and high difficulty in compatible with programmable hardware.

Method used

Using a structure based on a bucket and a storage slot, the storage slot is used to store the identification field of the traffic to be detected. The identification field is generated based on the number of traffic layers and the flow label. The conditional count value is calculated by traversing the storage slot list to determine the target type traffic.

Benefits of technology

Improves space utilization, enhances detection efficiency, and is compatible with deployment in programmable switches.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119544624B_ABST
    Figure CN119544624B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of traffic detection, and specifically provides a hierarchical large flow detection method compatible with programmable hardware, including: placing storage slots into corresponding storage slot lists according to the traffic layers of the storage slots of a storage bucket in a data architecture, where the storage bucket includes multiple storage slots, and the storage slots are used to store identification fields of received traffic to be detected, and the identification fields are generated according to the traffic layers and flow labels of the traffic to be detected, and the number of storage slot lists is the same as the total number of traffic layers; traversing the storage slots in the storage slot lists to obtain the attribute fields of the currently traversed storage slots; calculating the conditional count value of the traffic to be detected according to the attribute fields; and determining the traffic to be detected corresponding to the conditional count value as the target type traffic when the conditional count value exceeds a preset threshold. To solve the problems in the related art of the hierarchical large traffic detection method, such as low utilization rate of storage space, resulting in low traffic detection efficiency, and great difficulty in being compatible with programmable hardware.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of traffic detection, and particularly to a hierarchical large flow detection method compatible with programmable hardware. Background Art

[0002] Existing hierarchical large flow detection methods can generally be divided into two categories: tree-based methods (Trie-based) and methods based on compact data structures (sketch-based).

[0003] Tree-based methods usually use a tree storage structure to store information of each network flow, dynamically track possible hierarchical large flows, and eliminate small flows to reduce space overhead. Such methods need to frequently create and delete tree nodes during the measurement period, so the throughput is very low; moreover, the dynamic storage structure cannot be deployed in programmable switches that only support static memory allocation.

[0004] Methods based on compact data structures such as SpaceSaving and MVPipe allocate a separate compact data structure (sketch) for network flows of each layer. Since the number of hierarchical large flows existing in each layer is not uniform, such methods will cause a large amount of storage space waste.

[0005] Based on the low space utilization rate, when the space size of the traffic forwarding device remains the same, the number of traffic detections performed at the corresponding same time will decrease, that is, the traffic detection efficiency is reduced.

[0006] In addition, there are also many limitations when deploying in programmable switches. Programmable switches generally do not support dynamic memory allocation and only contain a limited number of processing stages in each packet processing pipeline, and only one read / write operation is allowed in each stage, which poses higher requirements on the deployed traffic measurement algorithm, and thus makes it more difficult for existing traffic detection methods to be compatible with programmable hardware.

[0007] Regarding the problems of low storage space utilization rate, low throughput, and high difficulty in being compatible with programmable hardware in the hierarchical large flow detection method in the related art, no effective solution has been proposed yet. Summary of the Invention

[0008] A hierarchical large flow detection method compatible with programmable hardware provided by an embodiment of the present invention can at least solve the problems of low storage space utilization rate, low throughput, and high difficulty in being compatible with programmable hardware in the hierarchical large flow detection method in the related art.

[0009] The embodiment of the present invention provides a hierarchical large flow detection method compatible with programmable hardware, including: putting storage slots into corresponding storage slot lists according to the flow levels of the storage slots of the storage buckets in the data architecture, where the storage bucket includes multiple storage slots, and the storage slots are used to store the identification fields of the received traffic to be detected, and the identification fields are generated according to the flow levels and the flow tags of the traffic to be detected, and the number of the storage slot lists is the same as the total number of flow levels; traversing the storage slots in the storage slot lists to obtain the attribute fields of the currently traversed storage slot, where the attribute fields include a key field and a count field; calculating the conditional count value of the traffic to be detected according to the attribute fields; and when the conditional count value exceeds a preset threshold, determining the traffic to be detected corresponding to the conditional count value as the target type traffic, where the target type traffic is the traffic exceeding the preset traffic threshold.

[0010] As an optional embodiment, calculating the conditional count value of the traffic to be detected according to the attribute fields includes: taking the product of the value of the count field of the storage slot and the total number of flow levels of the storage slot list where the storage slot is located as the count estimate value; and obtaining the conditional count value by subtracting the estimated value of the descendant key in the output set from the count estimate value, where the descendant key is the descendant key of the key field of the storage slot, the output set is used to record the target type traffic, and the estimated value of the descendant key is recorded in the attribute parameters of the output set.

[0011] As an optional embodiment, after determining the traffic to be detected corresponding to the conditional count value as the target type traffic when the conditional count value exceeds the preset threshold, the method further includes: adding the target type traffic to the output set and updating the estimated value of the descendant key of the output set; determining whether all the storage slots of the storage bucket have been traversed; outputting the output set when all the storage slots of the storage bucket have been traversed; and when not all the storage slots of the storage bucket have been traversed, performing the step of traversing the storage slots in the storage slot lists to obtain the attribute fields of the currently traversed storage slot.

[0012] As an optional embodiment, before putting the storage slots of the storage bucket in the data architecture into the corresponding storage slot lists, the method further includes: receiving the traffic to be detected and obtaining the flow tag of the data packet of the traffic to be detected; determining the identification field of the traffic to be detected according to the flow tag and the randomly generated flow level; determining the storage bucket of the traffic to be detected according to the hash result of the identification field; storing the identification field in the storage slot in the storage bucket, and updating the attribute fields of the storage slot.

[0013] As an alternative embodiment, determining the identification field of the traffic to be detected according to the flow label and the randomly generated number of traffic layers includes: determining the value range of the number of traffic layers according to the traffic stratification method; based on the value range, randomly selecting a random number within the value range as the number of traffic layers; selecting a prefix field with a corresponding number of bits in the flow label according to the number of traffic layers as the identification field; determining the bucket of the traffic to be detected according to the hash result of the identification field, including: performing a hash operation on the identification field through a hash function to obtain the hash result; using the hash result as the bucket identifier to find the corresponding candidate bucket, where the candidate bucket is a virtual bucket provided by the data architecture for the traffic to be detected in advance; using the candidate bucket as the bucket of the traffic to be detected.

[0014] As an alternative embodiment, storing the identification field into a storage slot in the bucket and updating the attribute field of the storage slot includes: determining whether there is a storage slot in the candidate bucket that has already stored the identification field according to the identification field; in the case where there is a storage slot in the candidate bucket that has already stored the identification field, incrementing the count field of the storage slot by one; in the case where there is no storage slot in the candidate bucket that has already stored the identification field, determining whether there is an empty storage slot in the candidate bucket; in the case where there is an empty storage slot in the candidate bucket, setting the key field of the empty storage slot to the identification field, setting the layer number field to the number of traffic layers, and setting the count field to one; in the case where there is no empty storage slot in the candidate bucket, finding the storage slot with the smallest count field in the candidate bucket and replacing the key field of the storage slot with the smallest count field with the identification field with a preset probability, where the preset probability is the reciprocal of the minimum value of the count field.

[0015] As an alternative embodiment, a right child count field is further set in each storage slot for recording the number of right children of the key field in the storage slot; after determining the identification field of the traffic to be detected according to the flow label and the randomly generated number of traffic layers, the method further includes: determining an odd parameter according to the number of traffic layers, where in the case where the number of traffic layers is even, adding one to the number of traffic layers to obtain the corresponding odd parameter, and in the case where the number of traffic layers is odd, directly using the number of traffic layers as the odd parameter; selecting a prefix field with a corresponding number of bits in the flow label according to the odd parameter as the odd field; replacing the step of determining the bucket of the traffic to be detected according to the hash result of the identification field with: determining the bucket of the traffic to be detected according to the odd hash result of the odd field; storing the odd field into a storage slot in the bucket according to the identification field and updating the attribute field of the storage slot.

[0016] As an alternative embodiment, determining the bucket for the traffic to be detected according to the odd hash result of the odd field includes: performing a hash operation on the odd field through a hash function to obtain the odd hash result; using the odd hash result as a bucket identifier to find a corresponding candidate bucket, where the candidate bucket is a virtual bucket provided by the data architecture in advance for the traffic to be detected; and using the candidate bucket as the bucket for the traffic to be detected.

[0017] As an alternative embodiment, storing the odd field into a storage slot in the bucket according to the identification field and updating the attribute field of the storage slot includes: determining whether there is a storage slot in the candidate bucket that has already stored the odd field according to the odd field; in the case where there is a storage slot in the candidate bucket that has already stored the odd field, incrementing the count field of the storage slot by one for the odd field and determining whether the identification field is the right child of the odd field; in the case where there is a storage slot in the candidate bucket that has already stored the odd field and the identification field is the right child of the odd field, incrementing the right child count field of the storage slot by one; in the case where there is a storage slot in the candidate bucket that has already stored the odd field but the identification field is not the right child of the odd field, ending the operation of storing the identification field into the storage slot in the bucket and maintaining the attribute field of the storage slot; and in the case where there is no storage slot in the candidate bucket that has already stored the identification field, storing the odd field into the candidate bucket.

[0018] As an alternative embodiment, storing the odd field into the candidate bucket includes: determining whether there is an empty storage slot in the candidate bucket; in the case where there is an empty storage slot in the candidate bucket, setting the key field of the empty storage slot to the odd field, setting the layer number field to the odd parameter, and setting the count field to one; and in the case where there is no empty storage slot in the candidate bucket, finding the storage slot with the smallest count field in the candidate bucket and replacing the key field of the storage slot with the smallest count field with the odd field with a preset probability, where the preset probability is the reciprocal of the minimum value of the count field.

[0019] The traffic detection method provided by the embodiment of the present invention uses a structure based on buckets and storage slots. The storage slots are used to store the identification fields of the traffic to be detected received. The identification fields are generated according to the traffic layer number and the flow label of the traffic to be detected. The traffic can be stored according to the layer number, so that the traffic of different layers can share the storage space and improve the space utilization rate. Based on this, only one update operation is required for each data packet, and each update operation only needs to use a small number of processing stages, enabling it to be compatible with deployment in a programmable switch.

[0020] During detection, according to the traffic layer number of the storage slots in the storage bucket of the data architecture, the storage slots are placed into the corresponding storage slot list, the storage slots in the storage slot list are traversed, the attribute fields of the currently traversed storage slot are obtained, the conditional count value of the traffic to be detected is calculated, and then when the conditional count value reaches the preset threshold, the corresponding traffic to be detected is determined as the target type traffic. It realizes improving the number of detected traffic in the same time on the basis of high space utilization. It solves the problems in the hierarchical large traffic detection method in the related technology, such as low storage space utilization rate, resulting in low detection efficiency, and great difficulty in being compatible with programmable hardware, and achieves the technical effects of reducing the space occupied by traffic, improving the space utilization rate, further improving the detection efficiency, and being compatible with programmable hardware. Brief Description of the Drawings

[0021] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other embodiments according to these drawings without creative efforts.

[0022] Figure 1 It is a flowchart of a hierarchical large traffic detection method compatible with programmable hardware according to an embodiment of the present invention.

[0023] Figure 2 It is a schematic diagram of the traffic update process according to an embodiment of the present invention.

[0024] Figure 3 It is a schematic diagram of the traffic update process using a compression algorithm according to an embodiment of the present invention.

[0025] Figure 4 It is a schematic diagram of the traffic detection query process according to an embodiment of the present invention.

[0026] Figure 5 It is a schematic diagram of a hierarchical large traffic detection device compatible with programmable hardware according to an embodiment of the present invention.

[0027] Figure 6 It is a schematic diagram of the structure of an electronic device according to the present invention. Detailed Embodiments

[0028] Embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to more thoroughly and completely understand the present invention. It should be understood that the drawings and embodiments of the present invention are only for exemplary purposes and are not used to limit the protection scope of the present invention.

[0029] Network traffic measurement provides crucial data support for various network applications such as attack detection, load balancing, and traffic engineering. Packets in a network can be abstracted into different network flows, and each network flow contains a flow label, which can be defined according to different application requirements, usually the source IP address or the destination IP address.

[0030] In recent years, Heavy Hitter Detection has attracted extensive attention from researchers. The goal of the Heavy Hitter Detection task is to find a small number of network flows that contain a large proportion of the traffic. This task can help achieve load balancing, traffic engineering, etc. However, only performing Heavy Hitter Detection cannot complete anomaly detection, especially for tasks such as Distributed Denial of Service (DDoS) attacks.

[0031] For example, an attacker can control a large number of zombie machines, and each zombie machine only sends a small number of packets to the attack target, but the total number of packets they send can exhaust the resources of the attack target and make it unable to work properly. In this scenario, more complex traffic measurement is needed to identify potential DDoS attacks.

[0032] IP addresses have a hierarchical feature. Each IP address can be hierarchically divided by bytes (5 layers) or by bits (33 layers). Based on this feature, researchers have proposed the problem of Hierarchical Heavy Hitter detection. Compared with the ordinary Heavy Hitter detection problem, Hierarchical Heavy Hitter detection not only has to find the heavy hitters whose flow size exceeds the threshold, but also has to identify some sets of network flows identified by IP address prefixes and whose total flow size exceeds the threshold. By performing Hierarchical Heavy Hitter detection, even if each machine in the botnet only sends a small number of packets, the traffic anomaly situation of the entire botnet can be identified.

[0033] Similar to other network traffic measurement tasks, hierarchical large flow detection faces many challenges and difficulties. First, network traffic measurement usually requires the use of special network processing chips, and the storage space of such chips is very limited (less than 15 MiB). In a typical backbone network, the amount of data transmitted per second can reach dozens of GB, so it is impossible to store the information of each network flow on-chip.

[0034] In the hierarchical large flow detection methods in the related art, the storage space utilization rate is low, resulting in the problem of low flow detection efficiency, and no effective solution has been proposed yet.

[0035] To solve the above technical problems, an embodiment of the present invention provides a hierarchical large flow detection method compatible with programmable hardware. Figure 1 It is a flowchart of a hierarchical large flow detection method compatible with programmable hardware according to an embodiment of the present invention. As Figure 1 shown, the hierarchical large flow detection method compatible with programmable hardware includes the following steps:

[0036] Step S101, according to the flow layer number of the storage slots of the storage buckets in the data architecture, put the storage slots into the corresponding storage slot lists, where the storage bucket includes a plurality of storage slots, and the storage slots are used to store the identification fields of the received traffic to be detected. The identification fields are generated according to the flow layer number and the flow label of the traffic to be detected, and the total flow layer numbers of the storage slot lists are the same;

[0037] Step S102, traverse the storage slots in the storage slot list, and obtain the attribute fields of the currently traversed storage slot, where the attribute fields include a key field and a count field;

[0038] Step S103, calculate the conditional count value of the traffic to be detected according to the attribute fields;

[0039] Step S104, when the conditional count value exceeds a preset threshold, determine that the traffic to be detected corresponding to the conditional count value is the target type traffic, where the target type traffic is the traffic that exceeds the preset traffic threshold.

[0040] The hierarchical large flow detection method compatible with programmable hardware provided by the embodiment of the present invention adopts a structure based on storage buckets and storage slots. The storage slots are used to store the identification fields of the received traffic to be detected. The identification fields are generated according to the flow layer number and the flow label of the traffic to be detected. The traffic can be stored according to the layer number, so that the traffic of different layers can share the storage space and improve the space utilization rate. Based on this, only one update operation is required for each data packet, and each update operation only needs to use a small number of processing stages, enabling it to be compatible with deployment in a programmable switch.

[0041] During detection, according to the traffic layer number of the storage slots in the storage buckets of the data architecture, the storage slots are placed into the corresponding storage slot lists. The storage slots in the storage slot lists are traversed to obtain the attribute fields of the currently traversed storage slot, and the conditional count value of the traffic to be detected is calculated. Then, when the conditional count value reaches the preset threshold, the corresponding traffic to be detected is determined as the target type traffic. This realizes improving the number of traffic detected in the same time on the basis of a high space utilization rate. It solves the problems in the hierarchical large traffic detection method in the related art, such as low storage space utilization rate leading to low detection efficiency and great difficulty in being compatible with programmable hardware, and achieves the technical effects of reducing the space occupied by traffic, improving the space utilization rate, further improving the detection efficiency, and being compatible with programmable hardware.

[0042] The execution entity of the above steps can be a communication device such as a router or a gateway for traffic transmission and forwarding. The above execution entity can include a memory and a controller. The memory is used to store at least the program data running according to the above steps, and the controller can act according to the above steps.

[0043] The above data architecture is also the data architecture of the memory of the above execution entity, and data storage and reading are performed according to this data architecture. The above execution entity can receive and store the traffic passing through the execution entity according to the detection frequency and detection period, and detect the received and stored traffic at the end of the detection period, and query out the traffic data detected as the target type traffic. Therefore, the above data architecture will also affect the detection of the stored data.

[0044] The above target type traffic is the traffic exceeding the preset traffic threshold, and can also be called large traffic. The data volume of large traffic is relatively large, and more resources are required in the forwarding and recording processes, and different communication devices often need to cooperate. Therefore, the identification of large traffic is a key technology in traffic detection.

[0045] The above data architecture can be a compact data structure sketch. The compact data structure sketch includes multiple storage buckets, and each storage bucket includes multiple storage slots. The storage slots are used to store the identification fields of the traffic to be detected received, and the identification fields are generated according to the traffic layer number and the flow label of the traffic to be detected. The above data architecture can enable traffic at different layers to share the storage space and improve the space utilization rate.

[0046] When storing the received traffic into the storage buckets and storage slots, the corresponding storage buckets and storage slots are found according to the identification fields, and the traffic is stored. Storing the traffic data is also to update the stored traffic, and the specific steps will be described later.

[0047] Based on the above data architecture, when detecting the stored traffic, multiple storage slot lists can be generated first. The number of storage slot lists is the same as the total number of traffic layers. It should be noted that the value range of the traffic layer is different under different layering conditions. For example, the IP address of the traffic can be layered by byte, divided into 5 layers; or layered by bit, divided into 33 layers. The total number of traffic layers is the total number of traffic layers.

[0048] According to the traffic layer of the storage slot of the storage bucket in the data architecture, the storage slot is placed into the corresponding storage slot list. That is, the data content corresponding to the storage slot is placed into the storage slot list. In this way, each storage slot stores the traffic data of the same layer. When detecting, only according to the attribute fields in the storage slot, the conditional count value of the traffic to be detected can be estimated through the estimation algorithm.

[0049] The conditional count value can represent the estimated level of the data volume of the traffic to be detected. The higher the conditional count value, the larger the data volume of the traffic to be detected. On the contrary, the lower the conditional count value, the smaller the data volume of the traffic to be detected. Its specific calculation method will be described later.

[0050] The calculation of the conditional count value mainly depends on the attribute fields of the storage slot, including the key field and the count field. The identification field of the data packet in the traffic, such as the prefix part of the IP address, can be stored and counted. Therefore, the size of the same traffic with the same IP address can be reflected through calculation.

[0051] When the conditional count value exceeds the preset threshold, it can be determined that the traffic to be detected corresponding to the conditional count value is the target type traffic, that is, the large traffic exceeding the preset traffic threshold mentioned above.

[0052] As an optional embodiment, calculating the conditional count value of the traffic to be detected according to the attribute fields includes: using the product of the value of the count field of the storage slot and the total number of traffic layers of the storage slot list where the storage slot is located as the count estimation value; obtaining the conditional count value by subtracting the estimated value of the descendant key in the output set from the count estimation value, where the descendant key is the descendant key of the key field of the storage slot, the output set is used to record the target type traffic, and the estimated value of the descendant key is recorded in the attribute parameters of the output set.

[0053] When calculating the conditional count value, first use the product of the value of the count field of the storage slot and the total number of traffic layers of the storage slot list where the storage slot is located as the count estimation value. Since the traffic of different traffic layers corresponds to different fields of the IP address of the traffic data packet, this count estimation value can be understood as the maximum cluster of the data packets of this traffic in theory, that is, the maximum value of the traffic.

[0054] Then, subtract the estimated value of the descendant key in the output set from the count estimate value to obtain the conditional count value. The descendant key is the descendant key of the key field corresponding to the above-mentioned count field of the storage slot, which can represent the amount of duplicate calculation in the data packets of the traffic. Furthermore, the size of the traffic can be estimated by taking the difference, that is, the above-mentioned conditional count value.

[0055] The output set is used to record the target type traffic that has been found, and the estimated value of the descendant key is recorded in the attribute parameters of the output set. Or it can be obtained by querying through the output set. It can represent the data packets that have been counted in other traffic.

[0056] Through the above method, the conditional count value can be estimated simply and quickly to be used as the basis for traffic detection, improving the efficiency of traffic detection.

[0057] As an optional embodiment, when the conditional count value exceeds the preset threshold and it is determined that the traffic to be detected corresponding to the conditional count value is the target type traffic, the method further includes: adding the target type traffic to the output set and updating the estimated value of the descendant key of the output set; determining whether all the storage slots of the bucket have been traversed; when all the storage slots of the bucket have been traversed, outputting the output set; when not all the storage slots of the bucket have been traversed, performing the step of traversing the storage slots in the storage slot list to obtain the attribute fields of the currently traversed storage slot.

[0058] It should be noted that when the conditional count value does not exceed the preset threshold, it is determined that the traffic to be detected corresponding to the conditional count value is not the target type traffic.

[0059] Since there may be more than one large traffic received, that is, there may be more than one target type traffic. After determining whether the traffic to be detected is the target type according to the conditional count value, it is necessary to continue traversing the storage slot list until all traffic has been detected.

[0060] That is, when all the storage slots of the bucket have been traversed, outputting the output set; when not all the storage slots of the bucket have been traversed, performing the step of traversing the storage slots in the storage slot list to obtain the attribute fields of the currently traversed storage slot and continuing the traversal.

[0061] However, when the estimated value of the descendant key is recorded by the output set, after adding the target type traffic to the output set, it is necessary to update the estimated value of the descendant key of the output set for subsequent calculations.

[0062] As an alternative embodiment, before putting the storage slots of the buckets in the data architecture into the corresponding storage slot list, the method further includes: receiving the traffic to be detected, and obtaining the flow label of the data packet of the traffic to be detected; determining the identification field of the traffic to be detected according to the flow label and the randomly generated number of traffic layers; determining the bucket of the traffic to be detected according to the hash result of the identification field; storing the identification field into the storage slot in the bucket, and updating the attribute field of the storage slot.

[0063] Before putting the storage slots of the buckets in the data architecture into the corresponding storage slot list, the stored data can be updated according to the received traffic to be detected.

[0064] The flow label of the data packet of the traffic to be detected described above is f, and the flow label can be defined according to different task requirements. In this embodiment, the source IP address is used as the flow label.

[0065] Determine the identification field of the traffic to be detected according to the flow label and the randomly generated number of traffic layers. The prefix F(f,r) of the first r bits of the flow label f can be used as the identification field.

[0066] Taking F(f,r) as the input of the hash function to obtain a hash result h(F(f,r)), and using it as the identification of the allocated candidate bucket as the basis for looking up the bucket. That is, determine the bucket of the traffic to be detected according to the hash result of the identification field.

[0067] The compact data structure provides a candidate bucket for each traffic. In this embodiment, multiple candidate storage locations are allocated for each traffic, and an attempt will be made to find an empty storage location for storage during insertion. In some embodiments, the identification of the bucket can be a subscript, which is generated when the data structure generates candidate buckets, and the subscript has a one-to-one mapping relationship with the bucket.

[0068] Through the above steps, each data packet only needs to be updated once, which improves the throughput while reducing the number of processing stages required by the algorithm and has better hardware compatibility.

[0069] As an alternative embodiment, determining the identification field of the traffic to be detected according to the flow label and the randomly generated number of traffic layers includes: determining the value range of the number of traffic layers according to the traffic stratification method; based on the value range, randomly selecting a random number within the value range as the number of traffic layers; selecting the corresponding number of bits of the prefix field in the flow label according to the number of traffic layers as the identification field.

[0070] The flow label of the data packet of the traffic to be detected described above is f, and the flow label can be defined according to different task requirements. In this embodiment, the source IP address is used as the flow label.

[0071] Then a random number r is generated. The value range of the traffic layer number r varies according to different stratification methods. When stratified by bytes, the value range of r is from 0 to 32; when stratified by bits, the value range of r is from 0 to 4.

[0072] Select a prefix field with a corresponding number of bits in the flow label according to the traffic layer number as the identification field. It can be the first r bits of the flow label f, i.e., F(f, r), as the identification field.

[0073] As an alternative embodiment, determine the storage bucket of the traffic to be detected according to the hash result of the identification field, including: performing a hash operation on the identification field through a hash function to obtain a hash result; using the hash result as the storage bucket identifier to find the corresponding candidate storage bucket, where the candidate storage bucket is a virtual bucket provided by the data architecture in advance for the traffic to be detected; using the candidate storage bucket as the storage bucket of the traffic to be detected.

[0074] Taking the identification field F(f, r) as the input of the hash function to obtain a hash result h(F(f, r)), and this hash result h(F(f, r)) is the same as the subscript of an existing storage bucket or has a one-to-one mapping relationship. That is, performing a hash operation on the identification field through the hash function to obtain a hash result.

[0075] Using it as the identifier of the allocated candidate storage bucket as the basis for finding the storage bucket. That is, using the hash result as the storage bucket identifier to find the corresponding candidate storage bucket.

[0076] Then store the identification field in the storage slot of the storage bucket and update the attribute field of the storage slot. Specifically as follows.

[0077] As an alternative embodiment, storing the identification field in the storage slot of the storage bucket and updating the attribute field of the storage slot includes: determining whether there is a storage slot in the candidate storage bucket that has already stored the identification field according to the identification field; in the case where there is a storage slot in the candidate storage bucket that has already stored the identification field, incrementing the count field of the storage slot by one; in the case where there is no storage slot in the candidate storage bucket that has already stored the identification field, determining whether there is an empty storage slot in the candidate storage bucket; in the case where there is an empty storage slot in the candidate storage bucket, setting the key field of the empty storage slot to the identification field, setting the layer number field to the traffic layer number, and setting the count field to one; in the case where there is no empty storage slot in the candidate storage bucket, finding the storage slot with the smallest count field in the candidate storage bucket and replacing the key field of the storage slot with the smallest count field with the identification field with a preset probability, where the preset probability is the reciprocal of the minimum value of the count field.

[0078] After determining the storage bucket, traverse multiple storage slots in the storage bucket. During the traversal, if a storage slot that has already stored the F(f,r) information is found, directly increment the count field in the storage slot by 1. That is, when there is a storage slot in the candidate storage bucket that has already stored the identification field, increment the count field of the storage slot by one.

[0079] If an empty storage slot is found, record the index of the empty storage slot. After the traversal, if no storage slot that has already stored the F(f,r) information is found but the index of the empty storage slot is recorded, store F(f,r) in the empty storage slot, that is, set the key field of the storage slot to F(f,r), set the layer number field to r, and set the count field to 1.

[0080] That is, when there is no storage slot in the candidate storage bucket that has already stored the identification field, determine whether there is an empty storage slot in the candidate storage bucket; when there is an empty storage slot in the candidate storage bucket, set the key field of the empty storage slot to the identification field, set the layer number field to the traffic layer number, and set the count field to one.

[0081] If neither a storage slot that has already stored the F(f,r) information is found nor the index of the empty storage slot is recorded, select the storage slot with the smallest count field value among all storage slots for replacement, and the replacement probability is the reciprocal of the smallest count field value.

[0082] That is, when there is no empty storage slot in the candidate storage bucket, find the storage slot with the smallest count field in the candidate storage bucket and replace the key field of the storage slot with the smallest count field with the identification field with a preset probability. The preset probability is the reciprocal of the minimum value of the count field.

[0083] Due to replacement according to probability, the smaller the count field, the greater the replacement probability, and the larger the count field, the smaller the replacement probability. In this way, the probability of larger traffic being replaced is minimized from the probability perspective. If the replacement is successful, set the key field of the storage slot to F(f,r), set the layer number field to r, and keep the count field value unchanged.

[0084] That is, when replacing the key field of the storage slot with the smallest count field with the identification field, replace the key field of the storage slot with the smallest count field with the identification field.

[0085] By the above strategy, the selection of the storage space depends only on the identification field related to the flow label and the layer number, sharing the storage space for the data packets of different layer traffic, and improving the utilization rate of the storage space.

[0086] As an alternative embodiment, a right child count field is also set in each storage slot for recording the number of right children of the key field in the storage slot; after determining the identification field of the traffic to be detected according to the flow label and the randomly generated number of traffic layers, the method further includes: determining an odd parameter according to the number of traffic layers, wherein, when the number of traffic layers is even, adding one to the number of traffic layers to obtain the corresponding odd parameter, and when the number of traffic layers is odd, directly using the number of traffic layers as the odd parameter; selecting a prefix field with a corresponding number of bits in the flow label according to the odd parameter as the odd field; replacing the step of determining the storage bucket of the traffic to be detected according to the hash result of the identification field with: determining the storage bucket of the traffic to be detected according to the odd hash result of the odd field; storing the odd field into the storage slot in the storage bucket according to the identification field, and updating the attribute field of the storage slot.

[0087] The above right child can be the right child Right - Child field. Specifically, the right child count field is used to record the right child count of the key in the storage slot. For example, if the key field in the storage slot is 10.10.0.0 / 24, its right child is 10.10.0.1 / 25. It should be noted that in some other embodiments, it can also be the left child Left – Child. These are all concepts in the tree - shaped data structure.

[0088] In the case where there is a right child field, the steps and processes for updating the stored data according to the received traffic to be detected will also change. Specifically as follows.

[0089] After determining the identification field of the traffic to be detected according to the flow label and the randomly generated number of traffic layers, an odd parameter will also be determined according to the number of traffic layers. Specifically, when the number of traffic layers is even, adding one to the number of traffic layers to obtain the corresponding odd parameter, and when the number of traffic layers is odd, directly using the number of traffic layers as the odd parameter.

[0090] That is, the value of the odd parameter m is determined according to the number of traffic layers r. If the number of traffic layers r is even, then the odd parameter m = r + 1; otherwise, the odd parameter m = r.

[0091] Select a prefix field with a corresponding number of bits in the flow label according to the odd parameter as the odd field. It can be the prefix F(f,m) of the first m bits of the flow label f as the odd field.

[0092] Determine the storage bucket of the traffic to be detected according to the odd hash result of the odd field. Use the odd field F(f,m) as the input of the hash function to obtain an odd hash result h(F(f,m)). Determine the position of the storage bucket based on the odd hash result h(F(f,m)).

[0093] As an alternative embodiment, determining the storage bucket for the traffic to be detected according to the odd hash result of the odd field includes: performing a hash operation on the odd field through a hash function to obtain an odd hash result; using the odd hash result as the storage bucket identifier to find the corresponding candidate storage bucket, where the candidate storage bucket is a virtual bucket provided by the data architecture for the traffic to be detected in advance; and using the candidate storage bucket as the storage bucket for the traffic to be detected.

[0094] Taking F(f,m) as the input of the hash function to obtain an odd hash result h(F(f,m)), and this odd hash result h(F(f,m)) is the same as the subscript of an existing storage bucket or has a one-to-one mapping relationship. That is, performing a hash operation on the identification field through the hash function to obtain a hash result.

[0095] Using it as the basis for finding the storage bucket by taking the identifier of the allocated candidate storage bucket, and finding the corresponding storage bucket. That is, using the hash result as the storage bucket identifier to find the corresponding candidate storage bucket.

[0096] Then, according to the identification field, store the odd field into the storage slot in the storage bucket and update the attribute field of the storage slot. Specifically as follows.

[0097] As an alternative embodiment, storing the odd field into the storage slot in the storage bucket and updating the attribute field of the storage slot according to the identification field includes: determining whether there is a storage slot in the candidate storage bucket that has already stored the odd field according to the odd field; in the case where there is a storage slot in the candidate storage bucket that has already stored the odd field, incrementing the count field of the odd field storage slot and determining whether the identification field is the right child of the odd field; in the case where there is a storage slot in the candidate storage bucket that has already stored the odd field and the identification field is the right child of the odd field, incrementing the right child count field of the storage slot; in the case where there is a storage slot in the candidate storage bucket that has already stored the odd field but the identification field is not the right child of the odd field, ending the operation of storing the identification field into the storage slot in the storage bucket and keeping the attribute field of the storage slot; and in the case where there is no storage slot in the candidate storage bucket that has already stored the identification field, storing the odd field into the candidate storage bucket.

[0098] After calculating the index of the storage bucket, that is, after the odd field F(f,m), determine whether there is a storage slot in the corresponding candidate storage bucket that has already stored the odd field F(f,m). If there is a storage slot in the storage bucket that has already stored the odd field F(f,m), directly increment the count value of the storage slot by 1. And determine whether the identification field F(f,r) is the right child of the odd field F(f,m).

[0099] That is, in the case where there is a storage slot in the candidate storage bucket that has already stored the odd field, increment the count field of the odd field storage slot and determine whether the identification field is the right child of the odd field.

[0100] If the identification field F(f,r) is the right child of the odd field F(f,m), increment the right child count of the storage slot by 1. If the identification field F(f,r) is not the right child of the odd field F(f,m), end the update operation.

[0101] That is, when there is a storage slot in the candidate bucket that already stores an odd field and the identification field is the right child of the odd field, increment the right child count field of the storage slot by one. When there is a storage slot in the candidate bucket that already stores an odd field but the identification field is not the right child of the odd field, end the operation of storing the identification field into the storage slot of the bucket and maintain the attribute fields of the storage slot.

[0102] When there is no storage slot in the candidate bucket that already stores the identification field, store the odd field into the candidate bucket. Specifically as follows.

[0103] As an alternative embodiment, storing the odd field into the candidate bucket includes: determining whether there is an empty storage slot in the candidate bucket; when there is an empty storage slot in the candidate bucket, set the key field of the empty storage slot to the odd field, set the layer number field to the odd parameter, and set the count field to one; when there is no empty storage slot in the candidate bucket, find the storage slot with the smallest count field in the candidate bucket and replace the key field of the storage slot with the smallest count field with the odd field with a preset probability, where the preset probability is the reciprocal of the minimum value of the count field.

[0104] When there is no storage slot in the current bucket that already stores the odd field F(f,m), determine whether there is an empty storage slot in the current bucket. If there is an empty storage slot, directly set the key of the storage slot to the odd field F(f,m), set the layer number to m, and set the count value to 1.

[0105] That is, determine whether there is an empty storage slot in the candidate bucket; when there is an empty storage slot in the candidate bucket, set the key field of the empty storage slot to the odd field, set the layer number field to the odd parameter, and set the count field to one.

[0106] There is neither a storage slot in the bucket that already stores F(f,r) nor an empty storage slot. Find the storage slot with the smallest count value in the bucket and replace the key of the storage slot with F(f,m) with a probability that is the reciprocal of this count value.

[0107] That is, when there is no empty storage slot in the candidate bucket, find the storage slot with the smallest count field in the candidate bucket and replace the key field of the storage slot with the smallest count field with the odd field with a preset probability.

[0108] During the process of storing traffic in storage slots according to the identification field in this embodiment, by referring to the count of the right child, the identification fields of multiple traffic are stored in one bucket, achieving a certain degree of data compression and further improving the memory usage efficiency.

[0109] It should be noted that this embodiment also provides an alternative implementation manner, which will be described in detail below. This implementation manner provides a hierarchical large flow detection method that can achieve high memory efficiency and hardware compatibility, aiming to use a single compact data structure to store network flow information in different layers and improve the space utilization rate. At the same time, the present invention involves a sampling-based update method, which only requires one update operation for each data packet. Therefore, each update operation only needs to use a small number of processing stages and can be deployed in a programmable switch.

[0110] This implementation manner designs a hierarchical large flow detection method that can achieve high memory efficiency and hardware compatibility, using a single compact data structure to store network flow information in different layers, achieving higher space utilization rate and better hardware compatibility.

[0111] The key points of this implementation manner are as follows: (1) By designing a single compact data structure based on the bucket structure, the network flows in different layers share the storage space, improving the space utilization rate; (2) Design a sampling-based update method, which only needs to be updated once for each data packet, reducing the number of processing stages required by the algorithm while improving the throughput and having better hardware compatibility; (3) Design a compression method, by storing the information of multiple network flows or prefixes in one bucket, further improving the memory usage efficiency.

[0112] The data structure of this implementation manner mainly includes an initialization module, an update module, and a query module. The initialization module is responsible for initializing the compact data structure according to user requirements. The update module performs an insertion operation on the arriving data packets and inserts the flow elements contained in the data packets into the compact data structure; the query module is responsible for providing the query function of the hierarchical large flow after the measurement period ends, and users can obtain the flow labels and corresponding flow sizes of the detected hierarchical large flows through this module. The detailed technical solution is as follows:

[0113] This implementation manner provides an initialization module that can initialize the compact data structure according to user requirements.

[0114] (1) Initialize the number of buckets of the compact data structure: The compact data structure has w buckets.

[0115] (2) Initialize each bucket: Each bucket contains k storage slots. The compact data structure provides a candidate bucket for each network flow or prefix. Therefore, in this embodiment, multiple candidate storage locations are allocated for each network flow or prefix, and an attempt will be made to find an empty storage location for storage during insertion.

[0116] (3) Initialize each storage slot: Each storage slot contains three fields: key, level, and count. During initialization, the key, level, and count fields are all set to 0.

[0117] This embodiment provides an update module. The update module is responsible for performing update operations on the compact data structure for each incoming data packet. Its specific implementation is as follows:

[0118] (1) Extract the flow label from the data packet.

[0119] Each data packet contains a flow label f, and the flow label can be defined according to different task requirements. In the present invention, the source IP address is used as the flow label.

[0120] (2) Determine the bucket in which the flow is stored based on the hash result of the flow label.

[0121] When determining the bucket, first generate a random number r, and then select the prefix F(f,r) of the flow label f as the label to be updated. The value range of r varies according to different layering methods. When layering by byte, the value range of r is from 0 to 32; when layering by bit, the value range of r is from 0 to 4.

[0122] After determining the label to be updated, use it as the input of the hash function to obtain a hash result h(F(f,r)), and use it as the subscript of the bucket. By randomly selecting the prefix for update, the present invention achieves high update throughput and better hardware compatibility.

[0123] (3) Search for an available storage slot in the bucket for storage or replace a non-empty storage slot;

[0124] After determining the bucket, traverse the k storage slots in the bucket. During the traversal, if a storage slot that has already stored the information of F(f,r) is found, directly increment the count field in the storage slot by 1; if an empty storage slot is found, record the subscript of the empty storage slot.

[0125] After the traversal is completed, if a storage slot that has already stored the information of F(f,r) is not found, but the subscript of the empty storage slot is recorded, then store F(f,r) in the empty storage slot, that is, set the key field of the storage slot to F(f,r), set the level field to r, and set the count field to 1.

[0126] If neither a storage slot storing the F(f,r) information is found nor the subscript of an empty storage slot is recorded, then select the storage slot with the smallest count field value among all storage slots for replacement, and the replacement probability is the reciprocal of the smallest count field value.

[0127] If the replacement is successful, set the key field of the storage slot to F(f,r), set the layer number field to r, and keep the count field value unchanged.

[0128] This embodiment also provides an update module using a compression algorithm. Using the compression algorithm can improve the space utilization rate and detection accuracy of the present invention in the hierarchical large flow detection by bit - level stratification. The update module using the compression algorithm includes the following changes:

[0129] (1) In each storage slot, add a Right - Child field to record the right - child count of the key in the storage slot. For example, if the key in the storage slot is 10.10.0.0 / 24, then its right - child is 10.10.0.1 / 25.

[0130] (2) When determining the storage bucket, first generate a random number r, and then select the prefix F(f,m) of the flow label f as the label to be updated. The value of m is determined according to r. If r is even, then m = r + 1; otherwise, m = r. Then use F(f,m) as the input of the hash function to determine the position of the storage bucket.

[0131] (3) After determining the storage bucket, the process of finding the storage slot is the same as that without using the compression algorithm. However, when updating the value of the count field after determining the storage slot, it is necessary to additionally determine whether F(f,r) is the right - child of F(f,m). If so, it is necessary to additionally update the right - child count field of the storage slot; otherwise, no additional operation is performed.

[0132] This embodiment provides a query module to detect and output large flows:

[0133] (1) First, construct a list A of storage slots with the same number as the total number of layers 0 ,A 1 ,…,A H-1 , and initialize these storage slot lists to be empty. Then traverse each storage bucket in the compact data structure, and add each storage slot in each storage bucket to the corresponding storage slot list according to the layer number field therein.

[0134] (2) Initialize an empty output set.

[0135] (3) Starting from the list of storage slots with index 0, traverse each list of storage slots in ascending order of indices. When traversing each list of storage slots, first take out the key field in the storage slot, and then take out the count field in the storage slot. Multiply the count field by the total number of layers to obtain an estimated count value of the key.

[0136] Since the criterion for determining whether a key is a hierarchical heavy hitter is whether its conditional count value exceeds the threshold, it is necessary to additionally calculate a conditional count value of a key.

[0137] The method for calculating the conditional count value of a key is to subtract the estimated count values of all keys that are descendants of this key in the output set from the estimated count value. If the conditional count value of a key exceeds the given threshold, add it to the output set.

[0138] (4) If an update module with a compression algorithm is used, the query module also needs to be modified. When initializing the list of storage slots, each storage slot will be transformed into three storage slots. In addition to the original storage slot, assuming that the key in the original storage slot is key, the number of layers is r, the count value is c, and the right child count is rc, the keys of the additional two storage slots are the right child and the left child of key respectively, the number of layers is r - 1, and the count values are rc and c - rc respectively. The remaining steps of the update module remain unchanged.

[0139] Specifically, the embodiment of the present invention proposes a space - efficient and programmable - hardware - compatible hierarchical heavy - hitter detection method. To describe this method more clearly, the following will introduce it in more detail with reference to the accompanying drawings.

[0140] Figure 2 is a schematic diagram of the traffic update process of the implementation manner of the present invention. As Figure 2 shown, the specific implementation steps of the update operation in this embodiment are as follows:

[0141] S11: Obtain the data packet f to be inserted.

[0142] S12: Randomly select a prefix of f or itself F(f, r), where r is a random number, and F(f, r) represents the prefix of f at the r - th layer.

[0143] S13: Calculate the storage bucket index using the hash function: h(F(f, r));

[0144] S14: After calculating the index of the storage bucket, determine whether there is a storage slot in the corresponding candidate storage bucket that has already stored F(f, r). If so, execute S15; otherwise, execute S16.

[0145] S15: There is a storage slot in the storage bucket that has already stored F(f, r), directly increment the count value of this storage slot by 1.

[0146] S16: If there is no storage slot in the current bucket that has stored F(f, r), determine whether there is an empty storage slot in the current bucket. If so, execute S17; otherwise, execute S18.

[0147] S17: There is an empty storage slot. Directly set the key of the storage slot to F(f, r), the layer number to r, and the count value to 1.

[0148] S18: There is neither a storage slot in the bucket that has stored F(f, r) nor an empty storage slot. Find the storage slot with the smallest count value in the bucket, and replace the key of the storage slot with F(f, r) with the probability of the reciprocal of this count value.

[0149] Figure 3 It is a schematic diagram of the traffic update process using the compression algorithm in the implementation manner of the present invention. As Figure 3 shown, the specific implementation steps of the update operation using the compression algorithm in this implementation manner are as follows:

[0150] S201: Obtain the data packet f to be inserted.

[0151] S202: Randomly select a prefix of f or f itself F(f, r), where r is a random number, and F(f, r) represents the prefix of f at the r-th layer.

[0152] S203: Calculate the bucket subscript using the hash function: h(F(f, m)). If r is even, m = r + 1; otherwise, m = r.

[0153] S204: After calculating the index of the bucket, determine whether there is a storage slot in the corresponding candidate bucket that has stored F(f, m). If so, execute S25; otherwise, execute S28.

[0154] S205: There is a storage slot in the bucket that has stored F(f, m). Directly increment the count value of this storage slot by 1 and execute S26.

[0155] S206: Determine whether F(f, r) is the right child of F(f, m). If so, execute S27; otherwise, end the update operation.

[0156] S207: If F(f, r) is the right child of F(f, m), then increment the right child count of the storage slot by 1.

[0157] S208: If there is no storage slot in the current bucket that has stored F(f, m), determine whether there is an empty storage slot in the current bucket. If so, execute S209; otherwise, execute S210.

[0158] S209: There is an empty storage slot. Directly set the key of the storage slot to F(f, m), the layer number to m, and the count value to 1.

[0159] S210: There is neither a storage slot in the storage bucket that has stored F(f, r) nor an empty storage slot. Find the storage slot with the smallest count value in the storage bucket, and replace the key of the storage slot with F(f, m) with the probability that is the reciprocal of this count value.

[0160] Figure 4 It is a schematic diagram of the flow detection query process of the implementation manner of the present invention. As Figure 4 shown, the specific implementation steps of the query operation in this implementation manner are as follows:

[0161] S301: Initialize H empty storage slot lists, where H is the total number of layers.

[0162] S302: Traverse each storage slot in all buckets, and put it into the corresponding storage slot list according to the layer number of the storage slot. If a compression algorithm is used, each storage slot should be first converted into three new storage slots according to the conversion method described in the invention content and then put into the storage slot lists respectively.

[0163] S303: Initialize an empty output set.

[0164] S304: Start traversing from the storage slot list with index 0.

[0165] S305: For each traversed storage slot, first take out the key in it, and then calculate its count estimate value. The calculation method is to multiply the count value by the total number of layers.

[0166] S306: Calculate the conditional count value. The calculation method is to subtract the count estimate values of all keys that are descendants of this key in the output set from the count estimate value.

[0167] S307: Judge whether the calculated conditional count value is greater than the threshold. If so, execute S308; otherwise, execute 309.

[0168] S308: If it is judged that the calculated conditional count value is greater than the threshold, add the key to the output set and execute S309.

[0169] S309: Judge whether all storage slots have been traversed. If so, execute S310; otherwise, execute traversing the next storage slot and execute S305.

[0170] S310: After traversing all storage slots, return the output set.

[0171] The algorithm proposed in this embodiment no longer uses multiple compact data structures in traditional hierarchical large flow detection solutions. While improving space utilization, it also shows high friendliness to programmable hardware (such as FPGAs and P4 chips). Because all network flow and prefix information is stored in a compact data structure, and each update uses a sampling method to perform only one memory access, greatly reducing the processing stages required for deployment on programmable hardware.

[0172] In addition, the embodiment of the present invention also designs a compression algorithm based on the right child count field, which can store the identification fields of multiple storage slots in one storage slot, further improving the memory usage efficiency and the hierarchical large flow detection accuracy.

[0173] Figure 5 is a schematic diagram of a hierarchical large flow detection device compatible with programmable hardware according to an embodiment of the present invention. As Figure 5 shown, based on the above-mentioned hierarchical large flow detection method compatible with programmable hardware provided by the embodiment of the present invention, the embodiment of the present invention also provides a hierarchical large flow detection device compatible with programmable hardware, which is applied to hierarchical large flow detection. The device includes: a storage slot extraction module 501, an attribute field acquisition module 502, a count value calculation module 503, and a target flow determination module 504. The device will be described in detail below.

[0174] The storage slot extraction module 501 is used to place the storage slots into the corresponding storage slot lists according to the flow levels of the storage slots of the buckets in the data architecture. Among them, the bucket includes multiple storage slots, and the storage slots are used to store the identification fields of the received traffic to be detected. The identification fields are generated according to the flow levels and the flow labels of the traffic to be detected. The number of storage slot lists is the same as the total number of flow levels.

[0175] The attribute field acquisition module 502 is connected to the above-mentioned storage slot extraction module 501 and is used to traverse the storage slots in the storage slot lists to obtain the attribute fields of the currently traversed storage slot. Among them, the attribute fields include a key field and a count field.

[0176] The count value calculation module 503 is connected to the above-mentioned attribute field acquisition module 502 and is used to calculate the conditional count value of the traffic to be detected according to the attribute fields.

[0177] The target flow determination module 504 is connected to the above-mentioned count value calculation module 503 and is used to determine the traffic to be detected corresponding to the conditional count value as the target type traffic when the conditional count value exceeds a preset threshold. Among them, the target type traffic is the traffic whose flow exceeds the preset flow threshold.

[0178] The hierarchical large flow detection device compatible with programmable hardware provided by the embodiments of the present invention adopts a structure based on buckets and storage slots. The storage slots are used to store the identification fields of the received traffic to be detected. The identification fields are generated according to the traffic layer number and the flow label of the traffic to be detected, so that the traffic can be stored according to the layer number, and then the traffic of different layers can share the storage space, improving the space utilization rate. Based on this, only one update operation is required for each data packet, and each update operation only needs to use a small number of processing stages, enabling it to be compatible with deployment in a programmable switch.

[0179] During detection, according to the traffic layer number of the storage slots of the buckets in the data architecture, the storage slots are placed in the corresponding storage slot list, the storage slots in the storage slot list are traversed, the attribute fields of the currently traversed storage slot are obtained, the conditional count value of the traffic to be detected is calculated, and then when the conditional count value reaches the preset threshold, the corresponding traffic to be detected is determined as the target type traffic. It realizes improving the number of detected traffic in the same time on the basis of a high space utilization rate. It solves the problems in the hierarchical large flow detection method in the related technology, such as low storage space utilization rate, resulting in low detection efficiency, and great difficulty in being compatible with programmable hardware, achieving the technical effects of reducing the space occupied by traffic, improving the space utilization rate, further improving the detection efficiency, and being compatible with programmable hardware.

[0180] The embodiments of the present invention also provide a non-transitory machine-readable medium storing a computer program, wherein the computer program is used to cause the computer to execute the method of the embodiments of the present invention when executed by a processor of the computer.

[0181] The embodiments of the present invention also provide a computer program product, including a computer program, wherein the computer program is used to cause a computer to execute the method of the embodiments of the present invention when executed by a processor of the computer.

[0182] The embodiments of the present invention also provide an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program capable of being executed by the at least one processor, and the computer program is used to cause the electronic device to execute the method of the embodiments of the present invention when executed by the at least one processor.

[0183] Reference Figure 6, a block diagram of an electronic device of a server or a client that can be an embodiment of the present invention will now be described. It is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device is intended to represent various forms of digital electronic computer devices, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0184] As Figure 6 shown, the electronic device includes a computing unit 601, which can execute various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0185] Multiple components in the electronic device are connected to the I / O interface 605, including: an input unit 606, an output unit 607, a storage unit 608, and a communication unit 609. The input unit 606 can be any type of device that can input information into the electronic device. The input unit 606 can receive input digital or character information, and generate key signal inputs related to the user settings and / or function controls of the electronic device. The output unit 607 can be any type of device that can present information, and can include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 608 can include, but is not limited to, magnetic disks, optical disks. The communication unit 609 allows the electronic device to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks, and can include, but is not limited to, a modem, a network card, an infrared communication device, and / or a wireless communication transceiver, such as a Bluetooth device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.

[0186] The computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a CPU, a graphics processing unit (GPU), various special artificial intelligence (AI) computing units, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 executes the various methods and processes described above. For example, in some embodiments, the method embodiments of the present invention can be implemented as a computer program, which is tangibly contained in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device via the ROM 602 and / or the communication unit 609. In some embodiments, the computing unit 601 can be configured to execute the above-described method in any other suitable manner (e.g., by means of firmware).

[0187] The computer program for implementing the method of the embodiments of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer program can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0188] In the context of the embodiments of the present invention, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable signal medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, or infrared system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0189] It should be noted that the term "including" and its variants used in the embodiments of the present invention are open-ended, that is, "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The modifications of "one" and "a plurality" mentioned in the embodiments of the present invention are illustrative rather than restrictive. Those skilled in the art should understand that, unless clearly specified otherwise in the context, it should be understood as "one or more".

[0190] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the embodiments of the present invention are all information and data that have been authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for the user to choose to authorize or reject.

[0191] The various steps described in the method embodiments provided by the embodiments of the present invention can be executed in different orders and / or executed in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The protection scope of the present invention is not limited in this regard.

[0192] The term "embodiment" in this specification means that the specific features, structures, or characteristics described in connection with the embodiment may be included in at least one embodiment of the present invention. The phrase appears in various positions in the specification does not necessarily mean the same embodiment, nor does it mean being independent or alternative to other embodiments and mutually exclusive. The various embodiments in this specification are all described in a related manner, and the same or similar parts between the embodiments are referred to each other. In particular, for the device, equipment, and system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts refer to the partial description of the method embodiments.

[0193] The above-described embodiments only represent several implementation manners of the present invention, and the description is relatively specific and detailed, but it should not be construed as a limitation on the protection scope. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the appended claims.

Claims

1. A hierarchical large flow detection method compatible with programmable hardware, characterized in that: include: According to the number of traffic layers of the storage slots of the storage bucket of the data architecture, the storage slots are placed in a corresponding storage slot list, wherein the storage bucket includes a plurality of storage slots, the storage slots are used to store an identification field of the received traffic to be detected, the identification field is generated according to the number of traffic layers and a flow label of the traffic to be detected, and the number of the storage slot list is the same as the total number of traffic layers; Traversing the storage slots in the storage slot list to obtain the attribute fields of the currently traversed storage slot, wherein the attribute fields include a key field and a count field; Calculate the conditional count value of the flow to be detected according to the attribute field; In the case where the condition count value exceeds a preset threshold, it is determined that the to-be-detected traffic corresponding to the condition count value is a target type traffic, wherein the target type traffic is traffic that exceeds a preset traffic threshold.

2. The method according to claim 1, characterized in that Calculating the conditional count value of the flow to be detected according to the attribute field includes: The product of the value of the count field of the storage slot and the total number of flow layers of the storage slot list where the storage slot is located is used as the count estimate value; The conditional count value is obtained by subtracting the estimated value of the descendant key in the output set from the count estimate, wherein the descendant key is the descendant key of the key field of the storage slot, the output set is used to record the target type traffic, and the estimated value of the descendant key is recorded in the attribute parameters of the output set.

3. The method according to claim 2, characterized in that When the condition count value exceeds the preset threshold, after determining that the to-be-detected traffic corresponding to the condition count value is the target type traffic, the method further includes: Add the target type traffic to the output set, and update the estimated value of the descendant key of the output set; Determine whether the storage slots of the storage bucket have been traversed; When the storage slot traversal of the storage bucket is completed, outputting the output set; In the case that the storage slots of the storage bucket have not been traversed completely, a step of traversing the storage slots in the storage slot list to obtain the attribute fields of the currently traversed storage slot is performed.

4. The method according to any one of claims 1 to 3, characterized in that Before placing the storage slots of the storage buckets of the data architecture into the corresponding storage slot list, the method further includes: Receiving the traffic to be detected, and obtaining the flow label of the data packet of the traffic to be detected; Determining the identification field of the to-be-detected traffic according to the flow label and the randomly generated traffic layer number; Determine the storage bucket of the traffic to be detected according to the hash result of the identification field; The identification field is stored in a storage slot in the storage bucket, and the attribute field of the storage slot is updated.

5. The method according to claim 4, characterized in that Determining, according to the flow label and the randomly generated flow layer number, an identification field of the flow to be detected, including: Determine the value range of the number of traffic layers according to the traffic stratification method; Based on the value range, randomly selecting a random number within the value range as the traffic layer number; According to the traffic layer number, a prefix field with a corresponding number of bits is selected in the flow label as the identification field; Determining the storage bucket of the to-be-detected traffic according to the hash result of the identification field includes: Performing a hash operation on the identification field by using a hash function to obtain the hash result; Using the hash result as a bucket identifier, searching for a corresponding candidate bucket, wherein the candidate bucket is a virtual bucket pre-provided by the data architecture for the traffic to be detected; The candidate storage bucket is used as the storage bucket of the traffic to be detected.

6. The method according to claim 4, characterized in that Storing the identification field in a storage slot in the storage bucket and updating the attribute field of the storage slot includes: Determine, according to the identification field, whether there is a storage slot in the candidate storage bucket that has stored the identification field; If there is a storage slot in the candidate storage bucket that has stored the identification field, increment the count field of the storage slot by one; If no storage slot in the candidate storage bucket has stored the identification field, determining whether there is an empty storage slot in the candidate bucket; In the case where there is an empty storage slot in the candidate storage bucket, the key field of the empty storage slot is set to the identification field, the layer number field is set to the traffic layer number, and the count field is set to one; When there is no empty storage slot in the candidate storage bucket, find the storage slot with the smallest count field in the candidate storage bucket, and replace the identification field with the key field of the storage slot with the smallest count field with a preset probability, wherein the preset probability is the inverse of the minimum value of the count field.

7. The method according to claim 4, characterized in that Each storage slot is also provided with a right child count field for recording the number of right children of the key field in the storage slot; After determining the identification field of the to-be-detected traffic according to the flow label and the randomly generated traffic layer number, the method further includes: Determine an odd parameter according to the number of traffic layers, wherein, when the number of traffic layers is an even number, add one to the number of traffic layers to obtain a corresponding odd parameter, and when the number of traffic layers is an odd number, use the number of traffic layers directly as the odd parameter; According to the odd parameter, a prefix field of a corresponding number of bits is selected in the flow label as an odd field; The step of determining the storage bucket of the traffic to be detected according to the hash result of the identification field is replaced by: determining the storage bucket of the traffic to be detected according to the odd hash result of the odd field; According to the identification field, the odd field is stored in a storage slot in the storage bucket, and the attribute field of the storage slot is updated.

8. The method according to claim 7, characterized in that Determining a storage bucket of the traffic to be detected according to the odd hash result of the odd field includes: Performing a hash operation on the odd field using a hash function to obtain the odd hash result; Using the odd hash result as a bucket identifier, searching for a corresponding candidate bucket, wherein the candidate bucket is a virtual bucket pre-provided by the data architecture for the traffic to be detected; The candidate storage bucket is used as the storage bucket of the traffic to be detected.

9. The method according to claim 7, characterized in that: According to the identification field, storing the odd field into a storage slot in the storage bucket, and updating the attribute field of the storage slot, including: Determine, according to the odd field, whether there is a storage slot in the candidate storage bucket that has stored the odd field; In the case that there is a storage slot in the candidate storage bucket that has stored the odd field, incrementing the count field of the storage slot of the odd field by one, and determining whether the identification field is the right child of the odd field; If there is a storage slot in the candidate storage bucket that has stored the odd field, and the identification field is the right child of the odd field, increment the right child count field of the storage slot by one; If there is a storage slot in the candidate storage bucket that has stored the odd field, but the identification field is not the right child of the odd field, the operation of storing the identification field into the storage slot in the storage bucket is terminated, and the attribute field of the storage slot is retained; In the case that no storage slot in the candidate storage bucket has stored the identification field, the odd field is stored in the candidate storage bucket.

10. The method according to claim 9, characterized in that Storing the odd field in the candidate storage bucket includes: Determining whether there is an empty storage slot in the candidate storage bucket; In the case where there is an empty storage slot in the candidate storage bucket, setting the key field of the empty storage slot to the odd field, setting the layer number field to the odd parameter, and setting the count field to one; When there is no empty storage slot in the candidate storage bucket, find the storage slot with the smallest count field in the candidate storage bucket, and replace the odd field with the key field of the storage slot with the smallest count field with a preset probability, wherein the preset probability is the inverse of the minimum value of the count field.

Citation Information

Patent Citations

  • A method for scalable distributed network traffic analytics in telco

    CN105917632A

  • A data processing method and apparatus

    CN109460406A