A method for real-time finding of persistent and low-frequency elements, an APT attack detection method, and an FRP detection method
Through the PISketch data structure, combined with the Bloom filter and hash table, the weight activation rules are used to identify continuous and low-frequency elements, solving the problems of large memory consumption and slow processing speed in the prior art, and achieving efficient and accurate APT attacks and FRP detection.
Patent Information
- Application Number
- CN202111479997.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-06
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2041-12-06
AI Technical Summary
When searching for APT attacks and FRP detection, it is difficult to efficiently and accurately identify continuous and low-frequency elements, resulting in large memory consumption, slow processing speed and complex parameter configuration.
Using PISketch compact data structure, the weight activation rules of elements are defined through the combination of Bloom filter and hash table, rewards appear for the first time and repetitions are punished, and elements with total weights exceeding the threshold are extracted as continuous and low-frequency elements.
It realizes efficient and accurate search of continuous and low-frequency elements under small memory usage, supports real-time APT attacks and FRP detection, and improves processing speed and detection efficiency.
Smart Images

Figure CN115529149B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of compact data structures and network measurement, and specifically to a method for using Sketch to real-time find persistent and low-frequency elements, an APT attack detection method, and an FRP detection method, which can be used for screening against APT attacks and FRP. Background Art
[0002] With the rapid development of network technology and the popularization of Internet applications, network security has become increasingly important. In recent years, frequent Advanced Persistent Threats (abbreviated as APT attacks) have received extensive attention due to their huge harmfulness. The complex and hidden attack means of APT attacks require processing and analyzing a large amount of data when detecting APT attacks. At the same time, with the rapid expansion of the network scale, the amount of network information data is also rising rapidly, and the huge amount of data brings great challenges to APT attack detection technology. Since network measurement technology is the main means to obtain the data required for detecting APT attacks, if APT attack detection can be achieved by using network measurement technology, the detection scope can be reduced, thereby improving the efficiency of the detection work and reducing the detection cost. It should be noted that the typical feature of APT attacks is that they prefer to persistently and covertly invade the target flow database to avoid detection. Therefore, finding persistent and low-frequency elements is important for detecting such attacks. However, existing work has not focused on finding persistent and low-frequency elements. The vast majority of the most advanced technologies are about finding frequent elements or finding persistent elements.
[0003] FRP (Fast Reverse Proxy) is an intranet penetration tool that will conduct short-term communications from time to time to confirm the connection when it is turned on. Why is the intranet penetration function needed? Because accessing one's own private device from the public network is often difficult and troublesome. For a company, this may involve confidentiality, so such an operation is definitely not allowed (or at least very cautious), so it is necessary to detect FRP. However, the existing solutions are often offline and cannot achieve real-time detection.
[0004] Currently, a direct solution for finding persistent and low-frequency elements is to separately find persistent elements and low-frequency elements and then take their intersection. The state-of-the-art algorithms for finding the set of persistent elements are On-Off and PIE, while the state-of-the-art algorithms for estimating the frequency of elements are algorithms based on compact data structures (Sketch), such as Count-Min (CM) Sketch. However, they are usually used to find frequently occurring elements. Since data in the real world usually follows a Zipf distribution, low-frequency elements are always very large in number but are not as important as frequent elements. To our knowledge, there is no work focusing on finding low-frequency elements. Taking the intersection of the above two types of algorithms is obviously not an ideal method for finding persistent and low-frequency elements. In fact, because the set of persistent elements and the low-frequency set may be very large, storing these two sets will result in huge memory consumption. The large memory consumption will directly lead to a slow processing speed and a high time complexity because such algorithms need to run on fast on-chip SRAM to achieve high-speed processing, and the memory size of on-chip memory is very limited. Additionally, this method may require the user to configure a large number of parameters. In summary, an ideal algorithm for finding persistent and low-frequency elements in a data stream needs to meet the following three requirements: high accuracy, small memory occupancy, and high processing speed. Summary of the Invention
[0005] To overcome the problems of slow processing speed, high time complexity, large memory space occupancy of data structures, and many configuration parameters of existing methods, the present invention provides a method using a compact data structure. This method realizes, for the first time, accurately and efficiently finding persistent and low-frequency elements, only requires a very small memory occupancy of the data structure while having a high processing speed, and can achieve APT attack detection.
[0006] The key idea of the present invention is to define a weight and its corresponding activation rule for each element that appears in any window, so as to cleverly combine and balance the information about persistence and low frequency. This activation rule is a reward and punishment strategy that only "welcomes" the first arrival of an element in each window: when an element first appears in a given window, add a large initial value L (for example, take L = 5) to its weight W in the current window as a reward. However, if it appears more than once in the same window (i.e., all occurrences except the first one), subtract 1 from its weight W i as a punishment. In this way, the relationship between persistence and low frequency is well balanced. Next, extract the elements whose total weight is greater than a threshold T set by the user as persistent and low-frequency (Persistent and Infrequent, PI) elements for output, where the total weight W is the sum of the weights of the element in all its windows: i The greater the total weight of an element, the more likely it is to be a PI element. To retain as many PI elements as possible, the present invention stores elements in a designed compact data structure, and when the data structure is full, a replacement strategy is used to make room for other potential PI elements. To achieve the above functions, the present invention proposes a compact data structure called PISketch to meet this requirement.
[0007] A method for real-time searching of persistent and low-frequency elements in the present invention includes the following steps:
[0008] Establish a Sketch-based compact data structure, namely PISketch, which includes two parts. The first part of the data structure is a Bloom filter, and the second part of the data structure is a hash table with several buckets;
[0009] The first part of the data structure queries whether an element first appears in the current time window, and the second part of the data structure calculates the weight of each element according to the query result of the first part of the data structure: when the element first appears in the current time window, increase the weight of the element in the current window as a reward; when the element appears more than once in the same time window, reduce the weight of the element as a penalty;
[0010] Extract the elements whose total weight is greater than the threshold set by the user as persistent and low-frequency elements for output.
[0011] Furthermore, PISketch consists of two parts: the first part reports whether an element first appears in the current time window. The second part calculates the weight of each element and finds out the elements with the most potential to obtain a higher weight. Before that, the present invention first divides the data stream into V non-overlapping parts on average, so there are a total of V time windows.
[0012] Furthermore, the data structure of the first part is a Bloom filter. The Bloom filter is a compact representation of the elements that have arrived in the current time window. Whenever a new time window starts, the present invention clears the Bloom filter. Each time an element comes, the present invention queries whether it comes for the first time (i.e., it has not been inserted yet). If so, it inserts the element into the Bloom filter. Then, the result is passed to the next part for weight calculation. The Bloom filter is used to remove duplicate elements from the incoming elements. Removing duplicate elements is necessary: because in the previous definition, the operations of first arrival and subsequent arrivals are completely different. If the Bloom filter reports true, it means that the element has appeared in this time window.
[0013] Furthermore, the data structure of the second part is a hash table with U buckets: B1, B2, …, B U。Based on the query results of the first part, these buckets calculate the total weight for each element and retain, as much as possible, the elements whose weights may exceed the threshold T. Each bucket has p cells, and each cell has a 3-tuple: the element ID (key), the total weight W of the element, and the number of windows N in which the element appears. The hash function h(.) randomly maps the elements to one of the buckets.
[0014] Furthermore, for the insertion operation of an element, the following steps are included:
[0015] First, query the Bloom filter to check whether the element appears in the current window. If the Bloom filter reports false, indicating that the element does not appear in the current window, then insert the element into the Bloom filter and perform two operations on this window: initialize the weight W in this window i = L and increment the window count N = N + 1. Otherwise, it means that the element has already appeared in the current window: decrement the weight W of the window corresponding to this element i by 1, i.e., W i = W i - 1.
[0016] Then, attempt to store the information of this element into the bucket B h to which it is mapped. According to the content in B h , it can be divided into three different cases: 1) If there is a cell containing this element, directly update the fields of this cell: add W i to the total weight W, i.e., W ← W + W i , and update the stored window count N to the current N. 2) If the element cannot be found in B h and B h is not full, then store this element in any empty cell. Set W to W i , i.e., W ← W + W i . At this time: W = L, N = 1. 3) If none of the cells in B h contains this element and B h has no empty cells, then attempt to make room for this element in B h . To retain as many potential PI elements as possible in the data structure, select an element m whose weight is the smallest among all the elements in B h . Although the weight of m is the smallest, it is not yet certain whether m is a PI element. Therefore, use the replacement strategy to evict the element m with the smallest weight in B h : whenever a replacement occurs, decrement its weight W' by 1. If W' is less than 0, it means the replacement is successful. Then evict m, and this element occupies the position of m. Set W to W i + 1 and N to 1. If the replacement is not successful, then this element leaves.
[0017] Finally, the Bloom filter is cleared by setting all bits to 0 at the end of each time window. For the reporting of elements: To obtain these PI elements, PISketch only needs to traverse the buckets.
[0018] The present invention also provides an APT attack detection method, which takes the original data stream to be screened as input, and uses the above method of the present invention to find persistent and low-frequency elements, which are suspected APT attack streams.
[0019] The present invention also provides an FRP detection method, which takes the original data stream to be screened as input, and uses the above method of the present invention to find persistent and low-frequency elements, which are suspected FRPs.
[0020] The beneficial effects of the present invention are: By using the compact data structure PISketch, users only need a given threshold T (according to their own needs and application scenarios) to quickly, efficiently and accurately obtain the desired PI elements. Application: By obtaining PI elements, preliminary screening of suspected APT attack streams and FRPs can be carried out. Description of the Drawings
[0021] Figure 1 is a schematic diagram of the data structure of PISketch of the present invention.
[0022] Figure 2 is the overall flowchart of the algorithm for finding persistent and low-frequency elements based on PISketch of the present invention.
[0023] Figure 3 is an example diagram of the insertion process of PISketch of the present invention. Detailed Embodiments
[0024] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the following further details the present invention in conjunction with the examples in the accompanying drawings.
[0025] Figure 1 is the data structure of PISketch of the present invention. As Figure 1 shown, PISketch consists of two parts. The data structure of the first part is the Bloom filter. This Bloom filter consists of M bits and k hash functions h1(.), h2(.), …, h k(.) composition. Whenever a new time window starts, the Bloom filter is emptied. Each time an element (such as element e in the figure) arrives, it is queried whether it comes for the first time. If so, it is inserted into the Bloom filter. Then, the result is passed to the second part for weight calculation. The Bloom filter is used to remove duplicate elements from the incoming elements. Removing duplicate elements is necessary because, as defined earlier in the present invention, the operations for the first arrival and subsequent arrivals are completely different. If the Bloom filter reports true, it means that the element has appeared in this time window. The data structure of the second part is a hash table with U buckets: B1, B2, …, B U . According to the query result of the first part, these buckets calculate the total weight for each element and retain, as much as possible, the elements whose weights may exceed the threshold T. Each bucket has p cells, and each cell has a 3-tuple: element ID (key), the total weight W of the element, and the number of windows N in which the element appears. The corresponding hash function h(.) of this hash table randomly maps the element to one of the buckets (such as bucket 3 in the figure). The specific insertion operation is described below.
[0026] Figure 2 is the overall process of the method for finding persistent and low-frequency elements based on PISketch of the present invention. As Figure 2 shown, it includes the following steps:
[0027] First, query the Bloom filter to check whether the element (i.e., a certain element arriving at the i-th time window) appears in the current window. If the Bloom filter reports false, it means that the element does not appear in the current window. Then, insert the element into the Bloom filter and perform two operations on this window: initialize the weight W in this window i = L and increment the window count N = N + 1. Otherwise, it means that the element has already appeared in the current window: decrement the weight W of the window corresponding to the element i , i.e., W i = W i - 1.
[0028] Then, attempt to store the information of the element into the bucket B h to which it is mapped. According to the content in B h , it can be divided into three different cases:
[0029] 1) If there is a cell containing the element, directly update the fields of this cell: add W i to the total weight W, i.e., W ← W + W i , and update the stored window count N to the current N.
[0030] 2) If the element cannot be found in B h and B hIf it is not full, then store the element in any empty cell. Set W to W i , that is, W←W+W i . At this time: W=L, N=1.
[0031] 3) If B h No cell in B contains the element and h If there are no empty cells, try h Make room for this element. To keep as many potential PI elements as possible in the data structure, select an element m whose weight is in B h The smallest among all elements of . Although m has the smallest weight, it is not certain whether m is a PI element. Therefore, the replacement strategy is used to evict B. h The element m with the smallest weight in the set is: Whenever a replacement occurs, its weight W' is reduced by 1. If W' is less than 0, it means the replacement is successful. Then m is expelled and the element takes the position of m. W is set to W i +1, N is set to 1. If the replacement is unsuccessful, the element is gone.
[0032] Finally, the Bloom filter is cleared by setting all bits to 0 at the end of each time window. For the report of elements: To obtain these PI elements, PISketch only needs to traverse the bucket.
[0033] Figure 3 This is an example of the insertion process of PISketch of the present invention. Figure 3 As shown, set L = 5, U = 3, p = 2. For convenience, we assume that elements e, r, and g enter this window for the first time (i.e., the Bloom filter report in the first part of PISketch is false). When element e arrives, PISketch maps it to bucket 1 through the hash function h(e). Since bucket 1 already has a cell storing e, PISketch directly updates the three fields of the cell: W e Add 5, N e Add 1. When element r arrives, PISketch maps it to bucket 2 through the hash function h(r). At this time, since r is not stored in bucket 2 and there is an empty cell, PISketch stores r in the empty cell and sets ID, W r 、N rThey are set to: the ID of r, 5, and 1 respectively. When it comes to element g, PISketch first calculates the hash function h(g) to map it to bucket 3. Since all cells in bucket 3 do not store g, PISketch attempts to evict the element stored in bucket 3 to make room for g. At this time, there are two elements in bucket 3: d and c, and the weight of d is significantly greater than that of c. c is the element with the smallest weight. Therefore, PISketch subtracts 1 from the weight W c of c, and W c becomes -1. Since the weight of c is less than 0 at this time, c is evicted, and g occupies the cell that originally stored c. Finally, PISketch stores the information of g in this cell: set the ID, W g , N g to: the ID of g, 6, and 1 respectively.
[0034] Experimental data:
[0035] Experimental settings: The method of the present invention is implemented in C++. The hash function used is Bob Hash. The experiment is carried out on a server with two CPUs (Intel Xeon E5-2620V3@2.4GHZ) and 62G DRAM.
[0036] Experimental data sets: Two real data sets (CAIDA, MAWI) and one APT simulation data set are used to evaluate the performance of PISketch. Each real data set contains approximately 10 million elements. APT simulation data set: The synthetic method is to mix real APT attack flows / elements into normal flows / elements.
[0037] CAIDA data set link: https: / / www.caida.org / catalog / datasets / overview /
[0038] MAWI data set link: http: / / mawi.wide.ad.jp / mawi /
[0039] Real APT attack data set:
[0040] https: / / www.mediafire.com / folder / c2az029ch6cke / TRAFFIC_PATTERNS_COLLECTION#734479hwy1b97
[0041] Performance evaluation metrics: (1) Recall Rate (RR), which is used to evaluate accuracy. (2) Average Relative Error (ARE), which is used to evaluate precision. (3) Throughput, which is used to evaluate speed. Its unit is Million of insertions per second (Mips).
[0042] Experimental results of three datasets are shown as follows:
[0043] 1. Experimental results on the CAIDA dataset are shown in Table 1 as follows:
[0044] Table 1
[0045] Memory (KB) RR ARE Throughput 100 0.95843 0.01168 8.4546 Mips 150 0.98687 0.00464 8.49846 Mips 200 0.99125 0.00371 8.45829 Mips 250 0.99344 0.00251 8.39263 Mips
[0046] 2. Experimental results on the MAWI dataset are shown in Table 2 as follows:
[0047] Table 2
[0048] Memory (KB) RR ARE Throughput 100 0.92333 0.01637 8.74034 Mips 150 0.96333 0.00723 8.68319 Mips 200 0.97333 0.00452 8.65948 Mips 250 0.98667 0.00329 8.61539 Mips
[0049] 3. Experimental results on the APT simulation dataset are shown in Table 3 as follows:
[0050] Table 3
[0051] Memory (KB) RR ARE Throughput 100 0.73667 0.06173 11.2604 Mips 150 0.84833 0.05381 11.00776 Mips 200 0.92333 0.04799 10.89731 Mips 250 0.95167 0.04358 10.54761 Mips
[0052] In other embodiments of the present invention, the replacement strategy can be converted into a probabilistic form of replacement. Specifically, with probability P, the weight W' of the minimum-weight element m is decreased by 1. Where P is user-defined and only needs to be ensured to be between (0, 1). For example, here is a reference provided: P = C -W′ + 0.01, where C is a certain user-defined constant.
[0053] Another embodiment of the present invention provides an APT attack detection method. When detecting APT attacks, directly adopt the above method of the present invention, take the original data stream to be screened as the input, input it into the PISketch of the present invention, and the elements with persistence and low frequency output at last are the suspected APT attack streams.
[0054] Another embodiment of the present invention provides an FRP detection method. When detecting FRP, directly adopt the above method of the present invention, take the original data stream to be screened as the input, input it into the PISketch of the present invention, and the elements with persistence and low frequency output at last are the suspected FRPs.
[0055] The method of the present invention can also be applied to other scenarios, such as detecting crawler scripts with regular polling, etc.
[0056] Another embodiment of the present invention provides an electronic device (such as a computer, a server, a smart phone, etc.), which includes a memory and a processor. The memory stores a computer program, and the computer program is configured to be executed by the processor. The computer program includes instructions for performing the steps in the method of the present invention.
[0057] Another embodiment of the present invention provides a computer-readable storage medium (such as ROM / RAM, a disk, an optical disc). When the computer program stored in the computer-readable storage medium is executed by a computer, the various steps of the method of the present invention are implemented.
[0058] The specific embodiments of the present invention disclosed above are intended to help understand the content of the present invention and implement it accordingly. Those of ordinary skill in the art can understand that various substitutions, changes, and modifications are possible without departing from the spirit and scope of the present invention. The present invention should not be limited to the content disclosed in the embodiments of this specification, and the protection scope of the present invention is subject to the scope defined by the claims.
Claims
1. A method for detecting APT attacks, characterized in that, Taking the original data stream to be screened as the input, the method of continuously and low-frequencyly searching in real time is adopted to find the continuous and low-frequency elements, which are the suspected APT attack streams; the method of continuously and low-frequencyly searching in real time includes the following steps: Establish a compact data structure based on Sketch, which consists of two parts. The first part of the data structure is a Bloom filter, and the second part of the data structure is a hash table with several buckets; The first part of the data structure queries whether an element first appears in the current time window, and the second part of the data structure calculates the weight of each element according to the query result of the first part of the data structure: when the element first appears in the current time window, the weight of the element in the current window is increased as a reward; when the element appears more than the second time in the same time window, the weight of the element is decreased as a punishment; the number of times exceeding the second time refers to all the appearance times except the first appearance; Extract the elements whose total weight is greater than the threshold set by the user as the continuous and low-frequency elements for output.
2. A FRP detection method, characterized in that, Taking the original data stream to be screened as the input, the method of continuously and low-frequencyly searching in real time is adopted to find the continuous and low-frequency elements, which are the suspected FRP; the method of continuously and low-frequencyly searching in real time includes the following steps: Establish a compact data structure based on Sketch, which consists of two parts. The first part of the data structure is a Bloom filter, and the second part of the data structure is a hash table with several buckets; The first part of the data structure queries whether an element first appears in the current time window, and the second part of the data structure calculates the weight of each element according to the query result of the first part of the data structure: when the element first appears in the current time window, the weight of the element in the current window is increased as a reward; when the element appears more than the second time in the same time window, the weight of the element is decreased as a punishment; the number of times exceeding the second time refers to all the appearance times except the first appearance; Extract the elements whose total weight is greater than the threshold set by the user as the continuous and low-frequency elements for output.
3. The method according to claim 1 or 2, characterized in that, The Bloom filter consists of M bits and k hash functions. Whenever a new time window starts, the Bloom filter is cleared. Each time an element comes, it is queried whether it comes for the first time. If so, it is inserted into the Bloom filter, and then the result is passed to the second part of the data structure for weight calculation; the hash table is a hash table with U buckets. According to the query result of the first part of the data structure, these buckets calculate the total weight for each element and retain the elements whose weight exceeds the threshold T set by the user; each bucket has p cells, and each cell has a triple: element ID, the total weight W of the element, and the number of windows N in which the element appears; The hash function h(.) corresponding to the hash table randomly maps the element to one of the buckets.
4. The method according to claim 3, wherein The second part of the data structure calculates the weight of each element according to the query result of the first part of the data structure, including: when an element first appears in a given window, a large initial value L is added to the weight W of the element in the current window as a reward; when the element appears more than once in the same window, its weight W is decreased by 1 as a penalty. i The Bloom filter checks whether the element appears in the current window. If the Bloom filter reports false, indicating that the element does not appear in the current window, the element is inserted into the Bloom filter, and two operations are performed on the window: initializing the weight W in the window to L and incrementing the window count N = N + 1. i Otherwise, it means that the element has already appeared in the current window, and the weight W of the window corresponding to the element is decreased by 1, i.e., W i = W i - 1. i = W i - 1.
5. The method according to claim 4, wherein The following steps are used to map an element to bucket B of the hash table of the second part data structure h : If there is a cell containing this element, directly update the field of this cell: add W i to the total weight W, that is, W ← W + W i , and update the stored number of windows N to the current N; If in B h The element is not found in B h If it is not full, then store the element in any empty cell and set W to W. i , that is, W←W+W i , at this time: W=L, N=1; If B h has no cell that contains the element and B h has no empty cell, then make room for the element from B h 6. The method according to claim 5, characterized in that, Said from B h To make room for the element, an element m is selected whose weight is the smallest among all elements in B h and then the replacement strategy is used to evict the element m with the smallest weight in B h The replacement strategy includes: whenever a replacement occurs, its weight W’ is decremented by 1. If W’ is less than 0, it means the replacement is successful, then m is evicted, the element occupies the position of m, W is set to W i +1, and N is set to 1; if the replacement is not successful, the element leaves.
7. The method according to claim 6, wherein The replacement strategy is a probabilistic form of replacement, that is, the weight W' of the element m with the minimum weight is decreased by 1 with probability P, where P is user-defined and guaranteed to be between (0,1).
8. An electronic device, characterized in that, It includes a memory and a processor. The memory stores a computer program, and the computer program is configured to be executed by the processor. The computer program includes instructions for executing the method according to any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a computer, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Data storage and query method based on matrix hash
CN108287840A
Data storage method and query method of stream sliding window
CN110532307A