Super host traceable identification method based on improved bloom filter

By introducing improved Bloom filters and three-dimensional cardinality analysis data structures in network security exception detection, the efficiency and accuracy of super host identification in high-speed networks are solved, and efficient traceability and multi-dimensional network security analysis of super hosts are realized.

CN120128346APending Publication Date: 2025-06-10TIANJIN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311667020.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-07
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently and accurately identify super hosts in high-speed networks and realize overall and multi-dimensional analysis of network security in heterogeneous network scenarios.

Method used

A super host traceability recognition method based on an improved Bloom filter is designed, using a probabilistic data structure and a three-dimensional cardinality analysis data structure, and mapping the packet address into the partition through a hash function, which realizes rapid estimation of source cardinality and destination cardinality and identification of multiple super hosts.

Benefits of technology

It realizes efficient identification of super hosts in high-speed networks, improves traceability efficiency, and can simultaneously estimate source cardinality and destination cardinality, ensures security of the network space, and achieves high accuracy with low computing and storage overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120128346A_ABST
    Figure CN120128346A_ABST
Patent Text Reader

Abstract

The invention discloses a super host traceability identification method based on an improved bloom filter, and designs an identification model capable of rapidly tracing and simultaneously identifying various super hosts, and the model comprises a probabilistic data structure and a three-dimensional cardinal number analysis data structure. The probabilistic data structure is used for filtering repeated data packets and ensuring the uniqueness of the data packets recorded by the three-dimensional cardinal number analysis data structure so as to improve the accuracy of host cardinal number estimation. The three-dimensional cardinality analysis data structure can simultaneously estimate a source cardinality and a destination cardinality to identify various super hosts, and can directly return its address and host cardinality at the end of a measurement period. The super host traceable identification method supports high-speed updating of each data packet, can realize accurate and rapid host cardinal number estimation and super host identification under relatively low calculation and storage overhead, and remarkably improves the performance of network security anomaly detection under the condition of mass data packets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of network security anomaly detection, and particularly relates to a super host traceable identification method based on an improved Bloom filter. Background Art

[0002] In a high-speed network, traffic anomalies can easily lead to security threats to the Internet, such as network congestion, worm viruses, distributed denial-of-service attacks, etc. Anomaly detection plays a crucial role in network security, and its purpose is to find patterns in network traffic that do not conform to expected normal behavior. How to design an effective network security anomaly detection mechanism is an important research issue.

[0003] In network security anomaly detection, super host identification is an important task. A super host refers to a host that sends a large amount of traffic or communicates with a large number of other hosts. Therefore, how to achieve a rapid and accurate estimation of the host cardinality and an efficient traceable identification of super hosts is the key issue of this task.

[0004] Sketch is a hash-based probabilistic data structure that performs fast host cardinality estimation operations by using probability instead of exact algorithms and data structures, and it has been widely used in super host identification. However, the irreversibility of the hash algorithm and the storage of a large amount of network traffic information pose challenges to the efficient traceability of super hosts. In addition, the current main research results are mostly two-dimensional data structures, which can only estimate the target cardinality to identify super spreaders and cannot be used for the overall and multi-dimensional analysis of network security in heterogeneous network scenarios. Therefore, designing an efficient traceable multi-super host identification model to improve network security has become an urgent problem to be solved in this field. Summary of the Invention

[0005] Aiming at the problems existing in the prior art, the present invention designs a super host traceable identification method based on an improved Bloom filter. This method can improve the traceability efficiency on the basis of accurately returning all super host addresses and their cardinality. Its three-dimensional structure can simultaneously estimate the source cardinality and the destination cardinality to identify different types of super hosts, effectively controlling the security risks in the network space.

[0006] To solve the above technical problems, a super host traceable identification model based on an improved Bloom filter proposed by the present invention includes a probabilistic data structure and a three-dimensional cardinality analysis data structure;

[0007] The probabilistic data structure includes two improved Bloom filters. Among them, the Bloom filter for filtering the source address S of the data packet IP is denoted as LBF, and the Bloom filter for filtering the destination address D of the data packet IPThe Bloom filter is represented as an RBF. Each Bloom filter corresponds to d hash functions, which are respectively mapped to d partitions. The size of each partition in the LBF is r, and the size of each partition in the RBF is c. By querying the source address S of the data packet IP and the destination address D IP to determine whether the values in the d partitions of the two Bloom filters are all 1, so as to judge whether the current data packet (S IP , D IP ) is a newly arrived data packet;

[0008] The three-dimensional radix analysis data structure is a three-dimensional bit array of size r×c×d. Where, i∈{1, 2,..., r}, j∈{1, 2,..., c}, k∈{1, 2,..., d}. Denote the value of each bucket in the three-dimensional bit array as B(i, j, k), and B(i, j, k)∈{0, 1}. The value of the bucket is conducive to merging the radix information of each node to provide a network-wide radix measurement view. In addition, each row and each column of the three-dimensional bit array contains an additional information group: the information group corresponding to the row is denoted as (cs, ha, is, flag), where cs represents the total destination radix of the source address mapped to this row; ha records the source address mapped to this row, which is used for quickly tracing the super spreader; is is an indication information, which helps to determine the source address recorded in ha and is used to check and update ha of this row; the role of flag is to track the position of S IP in the subsequent layer, that is, flag = h k+1 (S IP ). In addition, the row information group of the d-th layer does not include flag. When there are multiple abnormal addresses in each layer, by matching the modulus of S IP recorded in flag layer by layer, it can be determined whether it is abnormal. The information group corresponding to the column is denoted as (cs, ha, id, flag), where cs represents the total source radix of the destination address mapped to this column; ha records the destination address mapped to this column, which is used for quickly tracing the super receiver; id is an indication information, which helps to determine the destination address recorded in ha and is also used to check and update ha of this column; flag is used to record the position of D IP in the next layer, that is, flag = h k+1 (D IP ). In addition, the column information group of the d-th layer still does not include flag.

[0009] Meanwhile, the present invention also proposes a traceable identification method for super hosts based on an improved Bloom filter, which can identify various super hosts such as super spreaders and super receivers by simultaneously estimating the source cardinality and the destination cardinality to ensure the security of the entire cyberspace; at the end of a measurement cycle, the three-dimensional cardinality analysis data structure can directly return all super host addresses and their cardinalities to promptly respond to any possible ongoing attacks; the information group is used to store packet information, and at the same time, the index results of the hash function are reused to support the rapid estimation of the host cardinality under the condition of less memory resource occupation, which is beneficial to comprehensively detecting network security anomalies.

[0010] The traceable identification method for super hosts based on the improved Bloom filter includes: Network traffic screening stage: The initially arriving network packets are screened through the update operation of the probabilistic data structure to filter out the repeatedly appearing packets, ensuring high accuracy with low computational and storage overheads; Host cardinality estimation stage, using the three-dimensional cardinality analysis data structure to record the information of all initially arriving packets, simultaneously estimating the source cardinality and the destination cardinality to identify various super hosts, and performing security analysis on the entire network; Super host identification stage: Since the three-dimensional cardinality analysis data structure can quickly return the super host addresses and their corresponding cardinalities without complex reversible calculation operations, ensuring fast update and identification speeds, it is applicable to processing network traffic from high-speed network environments.

[0011] Furthermore, constructing the traceable identification method for super hosts based on the improved Bloom filter includes the following steps:

[0012] Step 1. Update the packet to the probabilistic data structure: Before anomaly detection, the values of the probabilistic data structures LBF and RBF are both 0. When the packet (S IP , D IP ) arrives, for LBF, d hash functions map the source address S IP to d partitions respectively and update the corresponding positions of each partition to 1; for RBF, the same d hash functions map the destination address D IP to d partitions respectively and update the corresponding positions of each partition to 1;

[0013] Step 2. Packet screening: Query the probabilistic data structure. If the values of the 2*d positions mapped by a certain packet are all 1, it means that the packet appears repeatedly, and at this time, it cannot pass the probabilistic data structure; if at least one of the 2*d positions mapped has a value of 0, it can be confirmed that the packet is initially arriving, and at this time, it can enter the three-dimensional cardinality analysis data structure;

[0014] Step 3. Update the three-dimensional cardinality analysis data structure: Each bucket of the three-dimensional bit array and the additional information group are initialized to 0. When the packet (SIP , D IP ) When reaching the three-dimensional cardinality analysis data structure, it can directly index to the positions of d buckets by reusing the hash calculation result in Step 1, i.e., i = h k (S IP ), j = h k (D IP ), and at the same time update the value of bucket B(i, j, k) to 1;

[0015] Step Four, Information Group Update: Add 1 to the cs values in the row information group and column information group where the bucket B(i, j, k) indexed by the data packet (S IP , D IP ) is located; and update ha according to the majority voting algorithm (MJRTY): Check whether S IP is the same as ha in the row information group. If ha is the same, add 1 to is, otherwise subtract 1. When the value of is is less than 0, replace the source address in ha with S IP , and reset the value of is to its absolute value; the update operation of ha in the column information group is the same as that in the row information group. In addition, record the flag in the row information group as h k+1 (S IP ), and the flag in the column information group as h k+1 (D IP );

[0016] Step Five, Host Cardinality Estimation: Taking the estimation of the destination cardinality as an example, first index the position of the data packet (S IP , D IP ) in the d layer through the hash function to obtain d row information groups recording the destination cardinality of this data packet. Subsequently, compare the source address S IP with ha in the row information group. If the source address S IP is the same as ha, then DCk(S IP ) = (cs + is) / 2; otherwise DCk(S IP ) = (cs - is) / 2, where k ∈ [1, d]. Finally, estimate the destination cardinality of the data packet (S IP , D IP ) as the minimum value of the estimation results in the d layer, i.e.: DC(S IP ) = min{DC 1 (S IP ), DC 2 (S IP ),..., DC d (S IP )}. The estimation process of the source cardinality is the same;

[0017] Step 6. Super host identification: Super hosts include super spreaders, super receivers, and super changers. To identify super spreaders, first locate the abnormal rows in layer d where cs is greater than the predefined threshold. The source address S stored in ha in the abnormal row information group IP is more likely to be a super spreader. At this time, match the modulo coefficient of the source address S IP in flag. If the cs of all rows in layer d where S IP is indexed is greater than the threshold, the destination base number of the source address S can be further estimated IP . If the base number exceeds the threshold, determine this host as a super spreader.

[0018] The identification process of super receivers is similar to that of super spreaders. The difference is that it is necessary to locate the abnormal columns where cs is greater than the predefined threshold. If the source base number of the destination address exceeds the predefined threshold, it is regarded as a super receiver.

[0019] The identification process of super changers is after two consecutive measurement cycles. First, locate the rows where is in the row information group is greater than 0 in the current measurement cycle; then, estimate the host base number stored in the ha address in the row information group and calculate the difference in the base number of the address between the current measurement cycle and the previous measurement cycle; finally, if the base number difference exceeds the predefined threshold, report it as a super changer.

[0020] Compared with the prior art, the present invention has the following advantages:

[0021] (1) Traceability: After determining the super hosts, the three-dimensional base analysis data structure can directly return all the super host addresses and their base numbers at the end of a measurement cycle without using additional algorithms to calculate and reconstruct abnormal addresses.

[0022] (2) Versatility: The three-dimensional base analysis data structure can estimate both the source base number and the destination base number simultaneously to achieve the identification of multiple super hosts such as super spreaders and super receivers.

[0023] (3) The present invention mainly uses the information group to store packet information and does not need to use all the buckets in the three-dimensional base analysis data structure. Compared with the existing solutions, the present invention occupies less memory resources, meets the storage consumption requirements for the monitoring of a large number of packets on high-speed links, and can significantly improve the network security anomaly detection performance of a large number of packets.

[0024] (4) When the present invention locates the packet index to the three-dimensional base analysis data structure, it reuses the hash value of the probabilistic data structure, reduces the number of hash calculations, and realizes the high-speed update of packets and the efficient identification of super hosts.

[0025] (5) The present invention improves the Bloom filter, which can effectively reduce its false positive probability and ensure the uniqueness of data packets to the greatest extent during the network traffic screening stage. In addition, during the host cardinality estimation stage, by applying the majority voting algorithm (MJRTY), an accurate and rapid estimation of the host cardinality is achieved with relatively low computational and storage overheads, ensuring performance under extreme network conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 FIG. is a flowchart of a method for traceable identification of super hosts based on an improved Bloom filter provided by an embodiment of the present invention;

[0027] Figure 2 FIG. is a schematic structural diagram of a method for traceable identification of super hosts based on an improved Bloom filter provided by an embodiment of the present invention;

[0028] Figure 3 FIG. is a structural design diagram of a method for traceable identification of super hosts based on an improved Bloom filter provided by an embodiment of the present invention;

[0029] Figure 4 FIG. is an update method diagram for traceable identification of super hosts provided by an embodiment of the present invention;

[0030] Figure 5 FIG. is an estimation method diagram for traceable identification of super hosts provided by an embodiment of the present invention.

[0031] ( Figure 1 is the drawing for the abstract) DETAILED DESCRIPTION OF THE EMBODIMENTS

[0032] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, but the following embodiments are by no means limiting to the present invention.

[0033] As Figure 1 shown, the method for traceable identification of super hosts based on an improved Bloom filter provided by an embodiment of the present invention includes the following processes:

[0034] (1) During the network traffic screening stage, the initially arriving network data packets are screened through the update operation of the probabilistic data structure to filter out the repeatedly occurring data packets, ensuring high accuracy with relatively low computational and storage overheads;

[0035] (2) After the screening, the data packet information entering the host cardinality estimation stage is updated. The three-dimensional cardinality analysis data structure is used to record all the initially arriving data packets, and at the same time, their source cardinality and destination cardinality are estimated to identify various super hosts, and the security of the entire network is analyzed;

[0036] (3) In the host cardinality estimation stage, an estimated value of the packet address cardinality is obtained by applying the estimation operation of the majority voting algorithm (MJRTY) in the three-dimensional cardinality analysis data structure;

[0037] (4) In the superhost identification stage, the three-dimensional cardinality analysis data structure can quickly trace the superhost address and its cardinality without performing complex reversible calculation operations, ensuring fast update and identification speeds, and thus can be applied to handle network traffic from high-speed network environments.

[0038] As Figure 2 shown, the superhost traceable identification method based on the improved Bloom filter provided by the embodiment of the present invention includes:

[0039] Network traffic screening stage 1, screening the initially arriving network packets through the update operation of the probabilistic data structure to filter out the repeatedly occurring packets, ensuring high accuracy with low computational and storage overheads;

[0040] Host cardinality estimation stage 2, using the three-dimensional cardinality analysis data structure to record all the information of the initially arriving packets, and simultaneously estimating the source cardinality and the destination cardinality to identify multiple superhosts and perform security analysis on the entire network;

[0041] Superhost identification stage 3, since the three-dimensional cardinality analysis data structure can quickly return the superhost address and its corresponding cardinality without performing complex reversible calculation operations, ensuring fast update and identification speeds, and thus can be applied to handle network traffic from high-speed network environments.

[0042] As Figure 3 shown, the superhost traceable identification model proposed by the embodiment of the present invention includes a probabilistic data structure and a three-dimensional cardinality analysis data structure. The probabilistic data structure includes two improved Bloom filters. Among them, the Bloom filter for filtering the source address S IP of the packet is denoted as LBF, and the Bloom filter for filtering the destination address D IP of the packet is denoted as RBF. Each Bloom filter corresponds to d hash functions, which are respectively mapped to d partitions. The size of each partition in LBF is r, and the size of each partition in RBF is c. By querying whether the values in the d partitions of the two Bloom filters for the source address S IP and the destination address D IP are all 1 to determine whether the current packet (S IP , D IPWhether it is a packet arriving for the first time; the three-dimensional cardinality analysis data structure is a three-dimensional bit array of size r×c×d. Where i∈{1, 2,..., r}, j∈{1, 2,..., c}, k∈{1, 2,..., d}. Denote the value of each bucket in this three-dimensional bit array as B(i, j, k), and B(i, j, k)∈{0, 1}. The value of the bucket is conducive to merging the cardinality information of each node to provide a network-wide cardinality measurement view. In addition, each row and each column of this three-dimensional bit array contains an additional information group: the information group corresponding to the row is denoted as (cs, ha, is, flag), where cs represents the total destination cardinality of the source address mapped to this row; ha records the source address mapped to this row, which is used for quickly tracing the super spreader; is is an indication information, which helps to determine the source address recorded in ha and is used to check and update ha of this row; the role of flag is to track the position of S IP in the subsequent layer, that is, flag = h k+1 (S IP ). In addition, the row information group of the d-th layer does not include flag. When there are multiple abnormal addresses in each layer, by matching the modulus of S IP recorded in flag layer by layer, it can be determined whether it is abnormal. The information group corresponding to the column is denoted as (cs, ha, id, flag), where cs represents the total source cardinality of the destination address mapped to this column; ha records the destination address mapped to this column, which is used for quickly tracing the super receiver; id is an indication information, which helps to determine the destination address recorded in ha and is also used to check and update ha of this column; flag is used to record the position of D IP in the next layer, that is, flag = h k+1 (D IP ). In addition, the column information group of the d-th layer still does not include flag.

[0043] The method for traceable identification of a super host based on an improved Bloom filter according to an embodiment of the present invention includes the following steps:

[0044] Step 1. Update the packet to the probabilistic data structure: Before anomaly detection, the values of the probabilistic data structures LBF and RBF are both 0. When the packet (S IP , D IP ) arrives, for LBF, d hash functions map the source address S IP to d partitions respectively and update the corresponding positions of each partition to 1; for RBF, the same d hash functions map the destination address D IP to d partitions respectively and update the corresponding positions of each partition to 1;

[0045] Step 2, Packet Screening: Query the probabilistic data structure. If the values of the 2*d positions mapped by a certain packet are all 1, it indicates that the packet appears repeatedly, and at this time, it cannot pass the probabilistic data structure; if at least one of the values of the 2*d positions mapped is 0, it can be confirmed that the packet arrives for the first time, and at this time, it can enter the three-dimensional radix analysis data structure;

[0046] Step 3, Update of the Three-Dimensional Radix Analysis Data Structure: Each bucket of the three-dimensional bit array and the additional information group are initialized to 0. As Figure 4 shown, when the packet (S IP , D IP ) arrives at the three-dimensional radix analysis data structure, it can directly index to the positions of d buckets by reusing the hash calculation result in Step 1, that is, i = h k (S IP ), j = h k (D IP ), and at the same time, update the value of the bucket B(i, j, k) to 1;

[0047] Step 4, Update of the Information Group: Add 1 to the cs values in the row information group and column information group where the bucket B(i, j, k) indexed by the packet (S IP , D IP ) is located; and update ha according to the majority voting algorithm (MJRTY): Check whether S IP is the same as ha in the row information group. If ha is the same, add 1 to is, otherwise subtract 1. When the value of is is less than 0, replace the source address in ha with S IP , and reset the value of is to its absolute value; the update operation of ha in the column information group is the same as that in the row information group. In addition, record the flag in the row information group as h k+1 (S IP ), and the flag in the column information group as h k+1 (D IP );

[0048] Step 5, Host Cardinality Estimation: As Figure 5 shown, taking the estimation of the destination cardinality as an example, first index the row position of the packet (S IP , D IP ) in the d layer through the hash function to obtain d row information groups recording the destination cardinality of this packet. Subsequently, compare the source address S IP with ha in the row information group. If the source address S IP is the same as ha, then DC k (S IP ) = (cs + is) / 2; otherwise DCk(S IP ) = (cs - is) / 2, where k ∈ [1, d]. Finally, estimate the packet (SIP , D IP ) has a destination base number that is the minimum of the d-layer estimation results, i.e.: DC(S IP ) = min{DC 1 (S IP ), DC 2 (S IP ),..., DC d (S IP )}. The estimation process of the source base number is the same;

[0049] Step Six, Super Host Identification: Super hosts include super spreaders, super receivers, and super changers. To identify super spreaders, first locate the abnormal rows in the d-layer where cs is greater than a predefined threshold. The source address S stored in ha in the abnormal row information group IP is more likely to be a super spreader. At this time, match the modulo coefficient of the source address S IP in flag. If the cs of all rows indexed by S IP in the d-layer is greater than the threshold, the destination base number of the source address S IP can be further estimated. If the base number exceeds the threshold, then determine this host as a super spreader.

[0050] The identification process of super receivers is similar to that of super spreaders. The difference is that it is necessary to locate the abnormal columns where cs is greater than a predefined threshold. If the source base number of the target address exceeds the predefined threshold, then regard it as a super receiver.

[0051] The identification process of super changers is after two consecutive measurement periods. First, locate the rows in the current measurement period where is in the row information group is greater than 0; then, estimate the host base number of the address stored in ha in the row information group, and calculate the difference in the base number of the address between the current measurement period and the previous measurement period; finally, if the base number difference exceeds the predefined threshold, then report it as a super changer.

[0052] Although the present invention has been described above in conjunction with the accompanying drawings, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many deformations without departing from the purpose of the present invention, and these all fall within the protection scope of the present invention.

Claims

1. A super host traceable identification model based on an improved Bloom filter, comprising a probabilistic data structure and a three-dimensional cardinality analysis data structure; Characterized in that: The probabilistic data structure includes two improved Bloom filters. Among them, the Bloom filter for filtering the source address S of the data packet IP is denoted as LBF, and the Bloom filter for filtering the destination address D of the data packet IP is denoted as RBF. Each Bloom filter corresponds to d hash functions, which are respectively mapped to d partitions. The size of each partition in LBF is r, and the size of each partition in RBF is c. By querying whether the values in the d partitions of the two Bloom filters for the source address S IP and the destination address D IP are all 1, it is determined whether the current data packet (S IP , D IP ) is a newly arrived data packet; The three-dimensional radix analysis data structure is a three-dimensional bit array of size r×c×d. Where, i ∈ {1, 2,..., r}, j ∈ {1, 2,..., c}, k ∈ {1, 2,..., d}. Denote the value of each bucket in this three-dimensional bit array as B(i, j, k), and B(i, j, k) ∈ {0, 1}. The value of the bucket is conducive to merging the radix information of each node to provide a network-wide radix measurement view. In addition, each row and each column of this three-dimensional bit array contains an additional information group: the information group corresponding to the row is represented as (cs, ha, is, flag), where cs represents the sum of the destination radixes of the source addresses mapped to this row; ha records the source addresses mapped to this row, which is used for quickly tracing the super spreaders; is is an indication information, which helps to determine the source addresses recorded in ha and is used to check and update the ha of this row; the role of flag is to track the position of S IP in the subsequent level, that is, flag = h k+1 (S IP ). In addition, the row information group of the d-th layer does not include flag. When there are multiple abnormal addresses in each layer, the modulus of S IP recorded in flag can be matched layer by layer to determine whether it is abnormal. The information group corresponding to the column is represented as (cs, ha, id, flag), where cs represents the sum of the source radixes of the destination addresses mapped to this column; ha records the destination addresses mapped to this column, which is used for quickly tracing the super receivers; id is an indication information, which helps to determine the destination addresses recorded in ha and is also used to check and update the ha of this column; flag is used to record the position of D IP in the next layer, that is, flag = h k+1 (D IP ). In addition, the column information group of the d-th layer still does not include flag.

2. A super host traceable identification method based on an improved Bloom filter, Characterized in that, For the super host traceable identification model based on the improved Bloom filter as described in claim 1, the super host traceable identification method can identify various super hosts such as super spreaders and super receivers by simultaneously estimating the source cardinality and the destination cardinality, ensuring the security of the entire cyberspace; at the end of a measurement cycle, the three-dimensional cardinality analysis data structure can directly return all super host addresses and their cardinalities to promptly respond to any possible ongoing attacks; Use information groups to store packet information, and at the same time reuse the index results of hash functions to support quickly estimating the host cardinality under the condition of less occupied memory resources, which is beneficial to comprehensively detecting network security anomalies. The super host traceable identification method based on the improved Bloom filter includes: Network traffic screening stage: Screen the initially arriving network packets through the update operation of the probabilistic data structure to filter out the repeatedly appearing packets, ensuring high accuracy under low computational and storage overheads; Host cardinality estimation stage, use the three-dimensional cardinality analysis data structure to record the information of all initially arriving packets, and simultaneously estimate the source cardinality and the destination cardinality to identify various super hosts and perform security analysis on the entire network; Super host identification stage: Since the three-dimensional cardinality analysis data structure can quickly return the super host addresses and their corresponding cardinalities without complex reversible calculation operations, ensuring fast update and identification speeds, it is applicable to processing network traffic from high-speed network environments.

3. For the super host traceable identification method based on the improved Bloom filter as described in claim 2, Characterized in that, The construction of the super host traceable identification method based on the improved Bloom filter includes the following steps: Step 1. Update the data packet to the probabilistic data structure: Before anomaly detection, the values of the probabilistic data structures LBF and RBF are both 0. When the data packet (S IP , D IP ) arrives, for LBF, d hash functions map the source address S IP to d partitions respectively and update the corresponding positions of each partition to 1; for RBF, the same d hash functions map the destination address D IP to d partitions respectively and update the corresponding positions of each partition to 1; Step 2: Packet screening: Query the probabilistic data structure. If the values of the 2*d positions mapped to by a certain packet are all 1, it means that the packet appears repeatedly, and at this time it cannot pass through the probabilistic data structure; if at least one of the values of the 2*d positions mapped to is 0, it can be confirmed that the packet is initially arriving, and at this time it can enter the three-dimensional cardinality analysis data structure; Step 3. Update of the 3D radix analysis data structure: Each bucket of the 3D bit array and the additional information group are initialized to 0. When a data packet (S IP , D IP ) arrives at the 3D radix analysis data structure, it can directly index to the positions of d buckets by reusing the hash calculation result in Step 1, i.e., i = h k (S IP ), j = h k (D IP ), and at the same time update the value of bucket B(i, j, k) to 1; Step 4. Information group update: Increment the cs values in the row information group and column information group where the bucket B(i, j, k) indexed by the data packet (S IP , D IP ) is located; and update ha according to the majority voting algorithm (MJRTY): Check whether S IP is the same as ha in the row information group. If ha is the same, increment is by 1; otherwise, decrement it by 1. When the value of is is less than 0, replace the source address in ha with S IP , and reset the value of is to its absolute value; the update operation for ha in the column information group is the same as that in the row information group. In addition, record the flag in the row information group as h k+1 (S IP ), and the flag in the column information group as h k+1 (D IP ); Step 5. Host Cardinality Estimation: Taking the estimated destination cardinality as an example, first, index the data packet (S IP , D IP ) in the d-layer through a hash function to obtain d row information groups that record the destination cardinality of this data packet. Subsequently, compare the source address S IP with ha in the row information group. If the source address S IP is the same as ha, then DC k (S IP ) = (cs + is) / 2; otherwise, DC k (S IP ) = (cs - is) / 2, where k ∈ [1, d]. Finally, estimate the destination cardinality of the data packet (S IP , D IP ) as the minimum value of the d-layer estimation results, that is: DC(S IP ) = min{DC 1 (S IP ), DC 2 (S IP ),..., DC d (S IP )}. The estimation process of the source cardinality is the same; Step 6. Super host identification: Super hosts include super spreaders, super receivers, and super changers. To identify super spreaders, first locate the abnormal lines in layer d where cs is greater than the predefined threshold. The source address S stored in ha in the abnormal line information group IP is more likely to be a super spreader. At this time, match the modular coefficient of the source address S IP in flag. If the cs of all lines indexed in layer d is greater than the threshold, the destination base number of the source address S IP can be further estimated. If the base number exceeds the threshold, determine that this host is a super spreader. IP ​ The identification process of super receivers is similar to that of super spreaders, except that it is necessary to locate the abnormal columns where cs is greater than the predefined threshold. If the source cardinality of the target address exceeds the predefined threshold, it is regarded as a super receiver. The identification process of super changers is after two consecutive measurement cycles. First, locate the rows in the row information group where is is greater than 0 in the current measurement cycle; then, estimate the host cardinality of the ha address stored in the row information group and calculate the difference in cardinality of the address between the current measurement cycle and the previous measurement cycle; finally, if the cardinality difference exceeds the predefined threshold, report it as a super changer.