Efficient approximate counting method for data stream

By designing an efficient approximate counting method for data flow in the Sketch algorithm, using the exchange probability data structure and the approximate sliding window model, data operations are performed in parallel, and memory updates are performed using single-instruction multi-data flow technology, which solves the problem of sliding window model increasing memory consumption and blocking insertion and query operations in the existing technology, and realizes efficient and low-latency data processing.

CN120030017AActive Publication Date: 2025-05-23DALIAN MARITIME UNIVERSITY

Patent Information

Application Number
CN202510122316.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-05-23
Estimated Expiration
2045-01-26

AI Technical Summary

Technical Problem

In the prior art, the Sketch algorithm solves the problem of updating expired data by introducing a sliding window model, but this increases memory consumption and operation time, and may block insertion and query operations, affecting system efficiency and response speed.

Method used

An efficient approximate counting method for data flow is designed, using exchange probability data structure and approximate sliding window model, and data insertion, update and query operations are performed in parallel, and memory update is performed using single-instruction multi-data flow technology to avoid blocked insertion and query operations.

Benefits of technology

Reduces memory consumption and operation time, improves processing speed, avoids blocked insertion and query operations, thereby improving system efficiency and response speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030017A_ABST
    Figure CN120030017A_ABST
Patent Text Reader

Abstract

The invention discloses a data stream-oriented efficient approximate counting method. The method comprises the following steps: S1, constructing an exchange probability data structure; s2, acquiring a data stream; s3: based on the exchange probability data structure, executing insertion, updating and query operations of data in the data stream in parallel; the data updating comprises the following steps: starting a data updating thread to perform data updating on each array; the data updating thread is used for executing memory updating operation and executing data exchange operation between a memory and an external memory by adopting a single-instruction multi-data-stream technology; through the exchange probability data structure and the approximate sliding window model constructed on the basis of the exchange probability data structure, data insertion, query and updating are performed in parallel, so that the throughput is not influenced; meanwhile, the exchange probability data structure ensures that when data stored in the approximate sliding window model reaches a set threshold value, a data updating thread is started to perform data updating on each bucket, and a single-instruction multi-data-stream technology is adopted to perform memory updating on the single-instruction multi-data-stream compatible array according to a set memory updating rule; according to the memory updating process, memory consumption is reduced, the processing speed is increased, meanwhile, blocking insertion and query operation is avoided, and therefore the query speed is increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of big data analysis, and in particular to an efficient approximate counting method for data streams. Background Art

[0002] In the field of network traffic monitoring, with the increase of data traffic and the acceleration of flow rate, real-time data processing has become particularly critical. Sketch algorithm is a type of probabilistic data structure widely used in the field of network measurement. It uses probabilistic methods such as hash functions to map elements to continuous memory space, and sacrifices a certain degree of accuracy to achieve small space consumption and extremely fast constant-level processing time. Such characteristics make Sketch algorithm well applied to the estimation of large-volume data flows such as networks and databases. At the same time, by combining the sliding window model, Sketch can only retain data from the most recent period of time, and continuously update data frequency information, focusing on the latest changes in data flows, so as to better realize real-time network data monitoring. For example, the Sketch algorithm can use the IP address of a network data packet as the unique identifier of the data packet, calculate its bucket index through a hash function, and insert it into the array of the corresponding bucket according to the bucket index to count the access frequency of this IP.

[0003] In practice, as time goes by, the importance of historical data will decrease relatively, but it takes up a lot of memory. In the existing technology, the Sketch algorithm solves the problem of updating expired data by introducing an approximate sliding window model. However, the introduction of the sliding window model will increase memory consumption and operation time. At the same time, the data update process usually blocks insertion and query operations, thereby affecting the efficiency and response speed of the system. Summary of the invention

[0004] The present invention provides an efficient approximate counting method for data streams to overcome the technical problem that the Sketch algorithm in the prior art solves the problem of deleting expired data by introducing a sliding window model, but the introduction of the sliding window model increases memory consumption and operation time, and usually blocks insertion and query operations, thereby affecting the efficiency and response speed of the system.

[0005] In order to achieve the above object, the technical solution of the present invention is:

[0006] An efficient approximate counting method for data streams, comprising:

[0007] S1: constructing an exchange probability data structure and an approximate sliding window model, and initializing the exchange probability data structure, wherein the exchange probability data structure includes a bucket array, the bucket array includes a plurality of arrays, each of the arrays includes n data items, and the n data items are used to store n different data occurrence frequencies;

[0008] S2: Get data stream;

[0009] S3: executing insert, update and query operations of data in the data stream in parallel based on the exchange probability data structure;

[0010] Inserting data includes: inserting the data occurrence frequency into the data item of the corresponding array according to the unique identifier of each data in the data stream;

[0011] The data updating includes: determining whether the frequency of occurrence of the data stored in the approximate sliding window model reaches a set threshold, if so, starting a data updating thread to update the data of each array, otherwise, not updating the data;

[0012] The data update thread includes using single instruction multiple data flow technology to perform memory update operations and execute data exchange operations between the internal memory and the external memory;

[0013] The use of single instruction multiple data flow technology to perform the memory update operation includes:

[0014] Create SIMD compatible arrays;

[0015] storing the data in a plurality of the arrays into the SIMD compatible array;

[0016] A memory update rule is set, and a single instruction multiple data stream technology is used to perform a memory update on the single instruction multiple data stream compatible array according to the set memory update rule;

[0017] The query of data includes: querying the data occurrence frequency of the data to be queried according to the data stored in the exchange probability data structure, and the data occurrence frequency is the approximate counting result.

[0018] Furthermore, each of the data items is a tuple, and the tuple is represented as <h f (e i .id),E,S,C>;

[0019] Among them, h f (e i .id) represents the compressed identifier; E is the exit counter, which is used to save the frequency of data that is about to be discarded in the approximate sliding window model; S is the sum counter, which is used to save the sum of the data occurrence frequencies in the first m′-1 subwindows in the approximate sliding window model; C is the cur counter, which is used to save the data occurrence frequency in the m′th subwindow, i.e., the current subwindow, in the approximate sliding window model, and m′ is greater than 0.

[0020] Furthermore, in the updating of the data, the SIMD compatible array established includes: an Aid array, an Aexit array, an Asum array, an A cur Arrays and memory variables X array;

[0021] Among them, the A id The array is used to store n node identifiers, i.e., the values ​​of node.id; exit The array is used to store the values ​​of n exit counter records; the A sum The array is used to store the values ​​of n sum counter records; the A cur The array is used to store the values ​​recorded by n cur counters; the memory variable X array is used to temporarily store A cur The values ​​of the array record.

[0022] Furthermore, the memory update rule includes:

[0023] Using A cur The values ​​stored in the array correspond to the values ​​in the updated memory variable X array;

[0024] Update A sum The value A stored at the i-th position in the array sum [i] is:

[0025] A sum [i]=A sum [i]'-A exit [i]+A cur [i]

[0026] In the formula, A sum [i] is the updated A sum The value recorded at position i in the array, A sum [i]' is A before update sum The value recorded at position i in the array, A exit [i] is A exit The value recorded at position i in the array, A cur [i] is A cur The value recorded at the i-th position in the array, i∈[0,n-1], and n is greater than 1;

[0027] A cur The values ​​stored in the array are set to 0.

[0028] Furthermore, in the updating of the data, a data exchange operation between the internal memory and the external memory is performed, including:

[0029] Slide each sub-window stored in the external memory to discard the frequency of the data in the first sub-window, and replace the value of the previous sub-window with the value of the next sub-window in turn to empty the last sub-window;

[0030] Update the value of the last empty subwindow with the value stored in the memory variable X array;

[0031] Update the exit counter with the value of the first subwindow.

[0032] Furthermore, in inserting the data, the data occurrence frequency is inserted into the data item of the corresponding array according to the unique identifier of each data in the data stream, including:

[0033] Assume that the data stream S consists of a series of data arriving in real time, S = {e 1 ,e 2 ,...,e i ...}, each data e i Each has a unique identifier;

[0034] Use hash function h to calculate data e i The hash value h(e i ), the hash value h(e i ) According to the hash value h(e i ) Determine the data e i The frequency of occurrence should be inserted into the nth array of the bucket array;

[0035] Using hash function h f For data i The identifier is compressed to obtain the compressed identifier h f (e i .id);

[0036] Traverse the data i The frequency of occurrence of the array to be inserted, comparing the compressed identifier h f (e i .id) and the data i The frequency of occurrence of each node in the array to be inserted is determined by the stored identifier node.id. If there is a match, node.id = h f (e i .id), then the cur counter of the node is incremented by 1, indicating that data e i appears again in the current sliding window; if there is no match, continue searching in the array to be inserted, and when the first empty node is found, define the identifier of the empty node as h f (e i.id), and initialize the cur counter of the node to 1, indicating that the data appears for the first time in the current sub-window.

[0037] Furthermore, in the query of the data, querying the data occurrence frequency of the data to be queried according to the data stored in the exchange probability data structure includes:

[0038] The hash function h is used to calculate the hash value of the data to be queried, and the array where the data to be queried is located is determined according to the hash value of the data to be queried;

[0039] Using hash function h f Compressing the identifier of the data to be queried to obtain a compressed identifier of the data to be queried;

[0040] Traverse the array where the data to be queried is located, compare the compressed identifier of the data to be queried with the identifier node.id stored in each node in the array where the data to be queried is located to see if there is a match. If there is a match, calculate the sum of the values ​​recorded by the sum counter and the cur counter, and return it as the query result, indicating the frequency of occurrence of the data in the current sliding window; if there is no match, the query result returns 0.

[0041] Furthermore, the constructed approximate sliding window model is a time-based approximate sliding window model, and constructing the time-based approximate sliding window model includes:

[0042] Set the total time length of the sliding window to W;

[0043] Divide the sliding window into m equal sub-windows, each sub-window records The frequency of data occurrence per unit time, and the amplitude of each slide is the length of a sub-window;

[0044] In the time-based approximate sliding window model, the data occurrence frequencies in the [1, m-1]th sub-window are all stored in the external memory, the data occurrence frequencies in the mth sub-window are all stored in the internal memory, and m is greater than 1.

[0045] Furthermore, the constructed approximate sliding window model is a count-based approximate sliding window model, and constructing the count-based approximate sliding window model includes:

[0046] Set the total count length of the sliding window to N;

[0047] Divide the sliding window into p equal sub-windows, each sub-window records The frequency of occurrence of a specific data in a data item, and the amplitude of each slide is the length of a sub-window;

[0048] In the count-based approximate sliding window model, the data occurrence frequencies in the [1, p-1]th sub-window are all stored in the external memory, the data occurrence frequencies in the pth sub-window are all stored in the internal memory, and p is greater than 1.

[0049] Beneficial effects: The present invention designs an exchange probability data structure with low memory consumption based on the existing Sketch algorithm, and constructs an approximate sliding window model based on the exchange probability data structure, and the window sliding process is parallel to data insertion and query, so it does not affect the throughput; at the same time, the exchange probability data structure can ensure that when the data stored in the approximate sliding window model reaches a set threshold, the data update thread is started to update the data for each bucket, wherein the single instruction multiple data stream technology is used to perform the memory update operation, including: establishing a single instruction multiple data stream compatible array, and storing data in several of the arrays into the single instruction multiple data stream compatible array; using the single instruction multiple data stream technology and updating the memory of the single instruction multiple data stream compatible array according to the set memory update rules, the memory update process reduces memory consumption and improves the processing speed, while avoiding blocking insertion and query operations, thereby improving the query speed. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0051] Figure 1 A flowchart of an efficient approximate counting method for data streams in the present invention;

[0052] Figure 2 A schematic diagram of an exchange probability data structure in an embodiment of the present invention;

[0053] Figure 3 A schematic diagram of a data structure of exchange probability based on single instruction multiple data stream technology operation in an embodiment of the present invention;

[0054] Figure 4 It is an ARE indicator diagram under memory changes based on the CAIDA data set in an embodiment of the present invention;

[0055] Figure 5 This is a SPEED indicator diagram under memory changes based on the CAIDA data set in an embodiment of the present invention;

[0056] Figure 6It is an ARE index diagram under the change of the number of sliding windows based on the CAIDA data set in an embodiment of the present invention;

[0057] Figure 7 This is a SPEED indicator diagram based on the change of the number of sliding windows of the CAIDA data set in an embodiment of the present invention;

[0058] Figure 8 This is a diagram showing the acceleration effect of using the single instruction multiple data technology in an embodiment of the present invention. DETAILED DESCRIPTION

[0059] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0060] This embodiment provides an efficient approximate counting method for data streams, such as Figure 1 As shown, the specific steps include:

[0061] S1: constructing an exchange probability data structure and an approximate sliding window model, and initializing the exchange probability data structure, wherein the exchange probability data structure includes a bucket array, wherein the bucket array includes a plurality of arrays, i.e., a plurality of buckets, each of the arrays includes n data items, and the n data items are used to store n different data occurrence frequencies;

[0062] Specifically, in this embodiment, the swap probability data, namely Swap Sketch, is obtained by improving the existing Sketch algorithm. When constructing the swap probability data structure, firstly, the number b of bucket arrays and the length n of each array are calculated according to the memory size limit; then, n data items, namely nodes, are created for each array;

[0063] In a specific embodiment, Figure 2 As shown, each of the data items is a tuple, and the tuple is represented as <h f (e i .id),E,S,C>

[0064] Among them, hf(e i.id) represents the compressed identifier; E is the exit counter, which is used to save the frequency of data to be discarded in the approximate sliding window model, that is, the value in the earliest sub-window in the approximate sliding window model; S is the sum counter, which is used to save the sum of the data occurrence frequencies in the first m′-1 sub-windows in the approximate sliding window model; C is the cur counter, which is used to save the data occurrence frequency in the m′th sub-window in the approximate sliding window model, that is, the current sub-window (Currentsub-window), and m′ is greater than 0.

[0065] Specifically, the Swap Sketch structure is initialized, that is, all nodes are initially assigned values, and the exit counter, sum counter, and cur counter in each node are all set to 0.

[0066] Specifically, in this embodiment, the constructed approximate sliding window model is a time-based approximate sliding window model or a count-based approximate sliding window model, including:

[0067] Constructing the time-based approximate sliding window model includes:

[0068] Set the total time length of the sliding window to W;

[0069] Divide the sliding window into m equal sub-windows, each sub-window records The frequency of data occurrence per unit time, and the amplitude of each sliding is the length of a sub-window; each sub-window records the frequency of data occurrence of the same data in different time periods;

[0070] In the time-based approximate sliding window model, the data occurrence frequencies in the [1, m-1]th sub-window are all stored in the external memory, the data occurrence frequencies in the mth sub-window are all stored in the internal memory, and m is greater than 1.

[0071] Specifically, for example, in the time-based approximate sliding window model, the sliding window only stores the data of the most recent 5 minutes (the total length is 5), and the sliding window is divided into 5 sub-windows, so each sub-window records the frequency of occurrence of 1 minute of data; when the sliding window reaches the threshold, such as inserting to the 6th minute, the frequency of occurrence of the data in the 1st minute is about to expire, and the data update operation will be triggered.

[0072] Constructing the count-based approximate sliding window model includes:

[0073] Set the total count length of the sliding window to N;

[0074] Divide the sliding window into p equal sub-windows, each sub-window records The frequency of occurrence of a specific data in a data item, and the amplitude of each slide is the length of a sub-window; The value of is rounded down to an integer.

[0075] In the count-based approximate sliding window model, the data occurrence frequencies in the [1, p-1]th sub-window are all stored in the external memory, the data occurrence frequencies in the pth sub-window are all stored in the internal memory, and p is greater than 1.

[0076] S2: Get data stream;

[0077] S3: executing insert, update and query operations of data in the data stream in parallel based on the exchange probability data structure;

[0078] Inserting data includes: inserting the data occurrence frequency into the data item of the corresponding array according to the unique identifier of each data in the data stream;

[0079] The data updating includes: determining whether the frequency of occurrence of the data stored in the approximate sliding window model reaches a set threshold, if so, starting a data updating thread to update the data of each array, otherwise, not updating the data;

[0080] The data update thread includes using single instruction multiple data flow technology to perform memory update operations and execute data exchange operations between the internal memory and the external memory;

[0081] The use of single instruction multiple data flow technology to perform the memory update operation includes:

[0082] Create SIMD compatible arrays;

[0083] storing the data in a plurality of the arrays into the SIMD compatible array;

[0084] The set memory update rules adopt single instruction multiple data flow technology and perform memory update on the single instruction multiple data flow compatible array according to the set memory update rules, and store the memory update results accordingly in each array in the bucket array to obtain updated results.

[0085] Specifically, if the data stored in the time-based approximate sliding window model reaches the maximum total time length W, sliding begins, thereby starting the data update thread to update the data for each bucket; if the data stored in the count-based approximate sliding window model reaches the maximum total count length N, sliding begins, thereby starting the data update thread to update the data for each bucket; specifically, any approximate sliding window model can be selected in practice.

[0086] The query of data includes: querying the data occurrence frequency of the data to be queried according to the data stored in the exchange probability data structure, and the data occurrence frequency is the approximate counting result.

[0087] Specifically, data query refers to the query of a given data item e i For example, when you need to query its frequency of occurrence in the current sliding window, you can calculate its number of occurrences in the most recent N data items (count-based approximate sliding window model) or the most recent data items that arrived within W time (time-based approximate sliding window model).

[0088] Specifically, in this embodiment, data insertion, updating and querying do not affect each other, and query operations can be performed at any time to return approximate counts of query data. At the same time, in order to further improve the efficiency of update operations, single instruction multiple data flow technology is used to accelerate this process, thereby reducing time consumption.

[0089] In a specific embodiment, in inserting the data, inserting the data occurrence frequency into the data item of the corresponding array according to the unique identifier of each data in the data stream includes:

[0090] Assume that the data stream S consists of a series of real-time arriving data, S = {e 1 ,e 2 ,...,e i ...}, each data e i Each has a unique identifier;

[0091] Use hash function h to calculate data e i The hash value h(e i ), the hash value h(e i ) is the bucket index, according to the hash value h(e i ) Determine the data e i The frequency of occurrence should be inserted into the nth array of the bucket array;

[0092] Using hash function h f For data i The identifier is compressed to obtain the compressed identifier h f (e i .id);

[0093] Traverse the data i The frequency of occurrence of the array to be inserted, comparing the compressed identifier h f (e i .id) and the data i The frequency of occurrence of each node in the array to be inserted is determined by the stored identifier node.id of each node, thereby determining whether the data e iWhether it already exists; if there is a match, i.e., node.id = h f (e i .id), then increment the cur counter of this node by 1, indicating that the data e i appears again in the current sliding window; if there is no match, continue to search sequentially in the array to be inserted. When the first empty node is found, define the identifier of this empty node as h f (e i .id), and initialize the cur counter of this node to 1, indicating that this data appears for the first time in the current sub-window.

[0094] In a specific embodiment, in the update of the data, the established single instruction multiple data stream compatible array, i.e., the SIMD compatible array, includes: an Aid array, an Aexit array, an Asum array, an Acur array, and a memory variable X array arranged sequentially along the vertical direction;

[0095] Among them, the Aid array is used to store the values of n node identifiers, i.e., node.id; the Aexit array is used to store the values recorded by n exit counters; the Asum array is used to store the values recorded by n sum counters; the Acur array is used to store the values recorded by n cur counters; the memory variable X array is used to temporarily store the values recorded by the Acur array.

[0096] In a specific embodiment, the memory update rule includes:

[0097] Use the values stored in the Acur array to correspondingly update the values in the memory variable X array;

[0098] Update the value Asum[i] stored in the i-th position of the Asum array to:

[0099] Asum[i] = Asum[i]' - Aexit[i] + Acur[i]

[0100] In the formula, Asum[i] is the value recorded in the i-th position of the updated Asum array, Asum[i]' is the value recorded in the i-th position of the Asum array before update, Aexit[i] is the value recorded in the i-th position of the Aexit array, Acur[i] is the value recorded in the i-th position of the Acur array, and i ∈ [0, n - 1];

[0101] Set the values stored in the Acur array to 0.

[0102] Specifically, Asum[i] is the value recorded in the i-th position of the updated Asum array, that is, the value recorded by the sum counter of the i-th node, and its update formula is:

[0103] node[i].S=node[i].S'-node[i].E+node[i].C

[0104] Wherein, node[i].S is the value recorded by the sum counter of the i-th node after the update, node[i].S' is the value recorded by the sum counter of the i-th node before the update, node[i].E is the value recorded by the exit counter of the i-th node, and node[i].C is the value recorded by the cur counter of the i-th node;

[0105] Specifically, since the value recorded in node[i].S before the update includes the value of the expired sub-window, when discarding data, it is necessary to subtract the discarded value of node[i].E and add the value of the current sub-window to update the value of node[i].S and store it in the sum counter of the i-th node.

[0106] Specifically, in this embodiment, single instruction multiple data technology, namely SIMD instructions, is used in executing the memory update operation. In order to effectively utilize SIMD instructions and thus accelerate the memory update process, this embodiment establishes a SIMD compatible array based on the bucket array to better align with SIMD instructions and support parallel processing, that is, convert each array in the bucket array into a SIMD compatible array. Figure 3 As shown, the SIMD compatible array includes an Aid array, an Aexit array, an Asum array, an Acur array and a memory variable X array (not shown in the figure) arranged in sequence along the vertical direction; the vertical grouping structure of the SIMD compatible array reflects the structure of the node node of each array in the bucket array, and the data corresponding to the same subscript index in the four arrays of the SIMD compatible array logically corresponds to a node node of each array in the bucket array, and the data is effectively aligned according to the attributes of each node, so that the value of the node node in each array in the bucket array is correspondingly stored in the corresponding position in the Aid array, the Aexit array, the Asum array and the Acur array, such as: the current node data e 2In the second position of the bucket array, its values ​​should be stored in the corresponding second position of the Aid array, Aexit array, Asum array, and Acur array. This reorganization method avoids the original process of processing each node in each array in the bucket array separately and one by one, but enables Swap Sketch to be better aligned with SIMD instructions to optimize the processing of SIMD instructions, so that the data of all nodes can be processed simultaneously through SIMD instructions, without updating the memory of each node in turn, reducing the calculation time and improving the data update efficiency. Through this adjustment, Swap Sketch can make full use of SIMD instructions, thereby reducing the time consumed by memory update operations.

[0107] For example, consider Figure 3 Node e in 2 , after the data is reorganized into a SIMD-compatible array, node e 2 The values ​​corresponding to each counter in are as follows:

[0108] Aid[1]=hf(e 2 ); Aexit[1]=14; Asum[1]=38; Acur[1]=0

[0109] By utilizing SIMD instructions, the entire memory update process becomes very efficient, and all node data can be processed in parallel with minimal overhead and a small number of instructions. This method significantly reduces the time complexity from O(n) (where n represents the length of array A) to O(1). Unlike the traditional method of processing each element one by one, the SIMD-based method allows calculations to be performed on multiple nodes simultaneously, thereby significantly accelerating the update operation. The parallelism provided by SIMD ensures that all updates can be completed with only a few CPU cycles, which brings significant improvements in efficiency compared to the traditional element-by-element processing method.

[0110] In a specific embodiment, in the updating of the data, performing a data exchange operation between the internal memory and the external memory includes:

[0111] Slide each sub-window stored in the external memory to discard the frequency of the data in the first sub-window, and replace the value of the previous sub-window with the value of the next sub-window in turn to empty the last sub-window;

[0112] Update the value of the last empty subwindow with the value stored in the memory variable X array;

[0113] Update the exit counter with the value of the first subwindow.

[0114] Specifically, in the data exchange operation between the internal memory and the external memory, the previous data is mainly replaced by the latter data to simulate the "sliding" process of the standard sliding window model, thereby discarding the value in the first sub-window. After sliding, the position of the last sub-window is vacated to save the value temporarily stored in the memory variable X array, that is, the value recorded by the cur counter, and the data of the current first sub-window is stored in the exit counter.

[0115] Specifically, the update process is as follows: Figure 2 As shown, at this time the data e 2 Stored in the second position of array A, data e 2 The frequency in the sliding window is: node[1].S+node[1].C=48. At this time, according to the time-based sliding window setting, the earliest inserted data is about to expire, and the update operation is triggered. The update process is: first, the memory update operation is performed. The first step is to store the value 5 of node[1].C into the memory variable X. The second step is to update the value of node[1].S to node[1].S'-node[1].E+node[1].C. The third step is to set the value of node[1].C to 0. Finally, the values ​​of the counters in the node node are<node[1].E:10,node[1].S:43,node[1].C:5> becomes<node[1].E:10,node[1].S:38,node[1].C:0> ; Secondly, execute the internal and external memory exchange process, slide each sub-window stored in the external memory, replace the value of the previous sub-window with the value of the next sub-window, discard the value 10 of the first sub-window, and vacate the position of the last sub-window to store the value in the memory variable X. Finally, use the value of the first sub-window in the external memory to update the value of node[1].E to 14. At this time, the data stored in the external memory changes from {10,14,6,9,4} to {14,6,9,4,5}, and the value of node in the memory changes from<node[1].E10,node[1].S:38,node[1].C:0> Became<node[1].E:14,node[1].S:38,node[1].C:0> At this time, the internal and external memory data exchange process is completed.

[0116] In a specific embodiment, in the query of the data, querying the data occurrence frequency of the data to be queried according to the data stored in the exchange probability data structure includes:

[0117] The hash function h is used to calculate the hash value of the data to be queried, and the array where the data to be queried is located is determined according to the hash value of the data to be queried;

[0118] Using hash function h fCompressing the identifier of the data to be queried to obtain a compressed identifier of the data to be queried;

[0119] Traverse the array where the data to be queried is located, compare the compressed identifier of the data to be queried with the identifier node.id stored in each node in the array where the data to be queried is located to see if there is a match. If there is a match, calculate the sum of the values ​​recorded by the sum counter and the cur counter, and return it as the query result, indicating the frequency of occurrence of the data in the current sliding window; if there is no match, the query result returns 0, indicating that the data has not been observed recently.

[0120] Specifically, Figure 2 The following table shows examples of insert and query operations. Data insertion operations usually include the first insertion and subsequent insertion based on existing insertions. 1 is an example of the first case, where SwapSketch first calculates h(e 1 )mod b to locate the data in the array, then traverse array A to find the corresponding node node. If the node is empty, set node.id = h f (e 1 ) and node[0].C=1, thus completing the insertion operation, where b is the length of an array. Since h(e 1 ) The calculated value may exceed this length, thus causing the array to be accessed out of bounds. Therefore, this embodiment needs to perform a mod b operation, i.e., a remainder b operation, to ensure that the array index is valid and will not cause an out-of-bounds access. Insert e 3 is an example of the second case, where Swap Sketch first calculates h(e 3 )mod b to locate the data in the array, and then traverse array A. During the traversal, the node.id of the first two nodes and hf(e 3 ) does not match, but when traversing to the third node, it is found that the id matches, and node[2].C'=node[2].C+1 is set to complete the insertion operation. For the query operation, Query(e 3 ) is 10+2=12, which is obtained by traversing the nodes to find the matching id and then recording the value of the counter. If no matching id is found after traversing the entire array A, the query result returns 0.

[0121] Specifically, since there is almost no competition between the timing of Swap Sketch's insertion and query operations accessing the memory space and the timing of update operations accessing the memory space, the two threads do not affect each other. This mechanism allows Swap Sketch to actively process expired data without blocking Swap Sketch's insertion and query operations, greatly improving the system's throughput.

[0122] Beneficial effects:

[0123] (1) Efficient memory structure design: In order to minimize memory consumption when processing large-scale data, this embodiment designs a streamlined memory structure, namely, the Swap Sketch structure, which only stores the data required for query processing in the memory. Specifically, the Swap Sketch structure only stores the sum of the data occurrence frequencies in these sub-windows in the memory, and stores the data occurrence frequencies in the remaining sub-windows except the current sub-window in the external memory (such as a hard disk or distributed storage). This structure optimizes memory usage and avoids unnecessary data occupation, thereby reducing memory consumption while ensuring system performance. This method is particularly effective for scenarios with limited memory and can improve the overall efficiency of the system when processing massive data.

[0124] (2) Independent update operation thread design: In this embodiment, an independently running dedicated update thread is designed, which is responsible for managing memory data updates and data exchange between the memory and the external storage. When the storage of a sub-window in the memory reaches a preset threshold, the update thread will be automatically triggered to perform tasks such as data cleanup and data exchange between the memory and the external storage. This operation is non-blocking, which means that it will not affect the insertion operation of the main thread, thereby ensuring high throughput and continuous inflow of data. In this way, the system can maintain efficient management of internal data while continuously processing new data. The design of the data update thread makes full use of the parallelism of the system, improves processing capabilities, and can run stably under high load conditions.

[0125] (3) Single instruction multiple data technology is applied to update operations: In this embodiment, SIMD instructions are applied to the memory update step in the data update of SwapSketch to improve the parallelism of memory updates. In traditional single-threaded updates, each operation is usually processed for a single data item, but after using SIMD instructions, multiple data can be processed simultaneously in one clock cycle, which greatly improves the data update efficiency. SIMD instructions can remove expired data in parallel, reducing computing time and processing delays. Due to the efficient parallel computing capabilities of SIMD instructions, update operations are no longer a bottleneck, thereby further improving the throughput and processing capabilities of the system.

[0126] In summary, Swap Sketch adds an approximate sliding window model that simulates a sliding window on the basis of the hash-table structured sketch algorithm, enabling it to discard expired data. While ensuring high accuracy, it significantly reduces memory consumption. Compared with other traditional Sketch algorithms, Swap Sketch reduces the storage of redundant information and intermediate data through a streamlined data storage structure, thereby effectively reducing memory requirements. In addition, the update operation of Swap Sketch is executed by an independent thread, and an active update strategy is adopted to ensure that the update process does not block insertion and query operations, thereby maintaining a high throughput of the system.

[0127] Specifically, in order to verify the effectiveness of this embodiment, the CAIDA dataset is used for experimental verification. The CAIDA dataset is a large network traffic dataset that contains CAIDA anonymized Internet traffic data from 2018. Each data packet in the dataset consists of a five-tuple, including the source IP address, destination IP address, source port, destination port, and protocol type. There are a total of about 27 million data packets, covering 1.3 million unique entries, which is suitable for network traffic analysis and performance evaluation research.

[0128] Based on the CAIDA dataset, the performance of five Sketch methods, namely Mic-CM, Mic-CU, Sl-CM, Sl-CU and Swap Sketch, are evaluated. Different memory sizes are tested. The experimental indicators are ARE (average relative error) and throughput. The calculation formula of ARE is:

[0129]

[0130] Where n' is the number of related terms, f 2 represents the actual frequency, Represents the estimated frequency.

[0131] like Figure 4 As shown in the figure, Swap Sketch maintains a low and smooth ARE value under different memory sizes, indicating that it has good robustness. In contrast, the ARE values ​​of the other four Sketches are higher and fluctuate greatly with the change of memory size, indicating that they perform poorly under low memory constraints.

[0132] like Figure 5 As shown in the figure, the throughput of all Sketches remains relatively stable as the memory size increases, indicating that the memory size has little impact on the throughput. It is worth noting that Swap Sketch always maintains the highest throughput, which is several times faster than Sl-Sketches and Mic-Sketches, or even more than ten times faster.

[0133] like Figure 6 As shown in the figure, Swap Sketch maintains a low ARE value under different memory sizes and gradually decreases with the increase of the number of sliding windows. In contrast, the changes of Mic-CM and Mic-CU are relatively stable, but their ARE values ​​are still higher than SwapSketch, while the ARE values ​​of Sl-CM and Sl-CU gradually increase with the increase of the number of sliding windows.

[0134] like Figure 7 As shown in the figure, the throughput of all Sketches remains relatively stable as the number of sliding windows increases, indicating that the number of sliding windows has little impact on the throughput. It is worth noting that Swap Sketch always maintains the highest throughput, which is always higher than Sl-Sketches and Mic-Sketches.

[0135] In summary, the reason why Swap Sketch can maintain a low error (ARE) is due to its unique structural design, which allows it to operate normally by storing only the necessary data in memory, thus ensuring high accuracy under low memory constraints. The reason why Swap Sketch can maintain a high throughput is that its unique structural design allows two threads to perform data insertion and update operations in parallel, thus achieving a high throughput.

[0136] like Figure 8 As shown in the figure, by utilizing SIMD instructions, the memory update operations of Swap Sketch can be executed in parallel, making full use of the CPU's vector unit, reducing the memory access frequency, improving resource utilization, significantly reducing the operation cycle, and improving the performance by about 2 times.

[0137] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An efficient approximate counting method for data streams, characterized in that: include: S1: constructing an exchange probability data structure and an approximate sliding window model, and initializing the exchange probability data structure, wherein the exchange probability data structure includes a bucket array, the bucket array includes a plurality of arrays, each of the arrays includes n data items, and the n data items are used to store n different data occurrence frequencies; S2: Get data stream; S3: executing insert, update and query operations of data in the data stream in parallel based on the exchange probability data structure; Inserting data includes: inserting the data occurrence frequency into the data item of the corresponding array according to the unique identifier of each data in the data stream; The data updating includes: determining whether the frequency of occurrence of the data stored in the approximate sliding window model reaches a set threshold, if so, starting a data updating thread to update the data of each array, otherwise, not updating the data; The data update thread includes using single instruction multiple data flow technology to perform memory update operations and execute data exchange operations between the internal memory and the external memory; The use of single instruction multiple data flow technology to perform the memory update operation includes: Create SIMD compatible arrays; storing the data in a plurality of the arrays into the SIMD compatible array; A memory update rule is set, and a single instruction multiple data stream technology is used to perform a memory update on the single instruction multiple data stream compatible array according to the set memory update rule; The query of data includes: querying the data occurrence frequency of the data to be queried according to the data stored in the exchange probability data structure, and the data occurrence frequency is the approximate counting result.

2. The efficient approximate counting method for data streams according to claim 1, characterized in that: Each of the data items is a tuple, and the tuple is represented as <h f (e i .id),E,S,C>; Among them, h f (e i .id) represents the compressed identifier; E is the exit counter, which is used to save the frequency of data that is about to be discarded in the approximate sliding window model; S is the sum counter, which is used to save the sum of the data occurrence frequencies in the first m′-1 subwindows in the approximate sliding window model; C is the cur counter, which is used to save the data occurrence frequency in the m′th subwindow, i.e., the current subwindow, in the approximate sliding window model, and m′ is greater than 0.

3. The efficient approximate counting method for data stream according to claim 2, characterized in that: In the updating of the data, the established SIMD compatible array includes: an Aid array, an Aexit array, an Asum array, an A cur Arrays and memory variables X array; Among them, the A id The array is used to store n node identifiers, i.e., the values ​​of node.id; exit The array is used to store the values ​​of n exit counter records; the A sum The array is used to store the values ​​of n sum counter records; the A cur The array is used to store the values ​​recorded by n cur counters; the memory variable X array is used to temporarily store A cur The values ​​of the array record.

4. The efficient approximate counting method for data streams according to claim 3 is characterized in that: The memory update rules include: Using A cur The values ​​stored in the array correspond to the values ​​in the updated memory variable X array; Update A sum The value A stored at the i-th position in the array sum [i] is: A sum [i]=A sum [i]'-A exit [i]+A cur [i] In the formula, A sum [i] is the updated A sum The value recorded at position i in the array, A sum [i]' is A before update sum The value recorded at position i in the array, A exit [i] is A exit The value recorded at position i in the array, A cur [i] is A cur The value recorded at the i-th position in the array, i∈[0,n-1], and n is greater than 1; A cur The values ​​stored in the array are set to 0.

5. The efficient approximate counting method for data streams according to claim 4 is characterized in that: In the data update, a data exchange operation between the internal memory and the external memory is performed, including: Slide each sub-window stored in the external memory to discard the frequency of the data in the first sub-window, and replace the value of the previous sub-window with the value of the next sub-window in turn to empty the last sub-window; Update the value of the last empty subwindow with the value stored in the memory variable X array; Update the exit counter with the value of the first subwindow.

6. The efficient approximate counting method for data streams according to claim 5, characterized in that: In inserting the data, inserting the data occurrence frequency into the data item of the corresponding array according to the unique identifier of each data in the data stream includes: Assume that the data stream S consists of a series of real-time arriving data, S = {e1, e2, ..., e i ...}, each data e i Each has a unique identifier; Use hash function h to calculate data e i The hash value h(e i ), the hash value h(e i ) According to the hash value h(e i ) Determine the data e i The frequency of occurrence should be inserted into the nth array of the bucket array; Using hash function h f For data i The identifier is compressed to obtain the compressed identifier h f (e i .id); Traverse data e i The frequency of occurrence of the array to be inserted, comparing the compressed identifier h f (e i .id) and the data i The frequency of occurrence of each node in the array to be inserted is determined by the stored identifier node.id. If there is a match, node.id = h f (e i .id), then the cur counter of the node is incremented by 1, indicating that data e i appears again in the current sliding window; if there is no match, continue searching in the array to be inserted, and when the first empty node is found, define the identifier of the empty node as h f (e i .id), and initialize the cur counter of the node to 1, indicating that the data appears for the first time in the current sub-window.

7. The efficient approximate counting method for data streams according to claim 6, characterized in that: In the data query, querying the data occurrence frequency of the data to be queried according to the data stored in the exchange probability data structure includes: The hash function h is used to calculate the hash value of the data to be queried, and the array where the data to be queried is located is determined according to the hash value of the data to be queried; Using hash function h f Compressing the identifier of the data to be queried to obtain a compressed identifier of the data to be queried; Traverse the array where the data to be queried is located, compare the compressed identifier of the data to be queried with the identifier node.id stored in each node in the array where the data to be queried is located to see if there is a match. If there is a match, calculate the sum of the values ​​recorded by the sum counter and the cur counter, and return it as the query result, indicating the frequency of occurrence of the data in the current sliding window; if there is no match, the query result returns 0.

8. The efficient approximate counting method for data streams according to claim 1, characterized in that: The constructed approximate sliding window model is a time-based approximate sliding window model, and constructing the time-based approximate sliding window model includes: Set the total time length of the sliding window to W; Divide the sliding window into m equal sub-windows, each sub-window records The frequency of data occurrence per unit time, and the amplitude of each slide is the length of a sub-window; In the time-based approximate sliding window model, the data occurrence frequencies in the [1, m-1]th sub-window are all stored in the external memory, the data occurrence frequencies in the mth sub-window are all stored in the internal memory, and m is greater than 1.

9. The efficient approximate counting method for data streams according to claim 1, characterized in that: The constructed approximate sliding window model is a count-based approximate sliding window model, and constructing the count-based approximate sliding window model includes: Set the total count length of the sliding window to N; Divide the sliding window into p equal sub-windows, each sub-window records The frequency of occurrence of a specific data in a data item, and the amplitude of each slide is the length of a sub-window; In the count-based approximate sliding window model, the data occurrence frequencies in the [1, p-1]th sub-window are all stored in the external memory, the data occurrence frequencies in the pth sub-window are all stored in the internal memory, and p is greater than 1.

Citation Information

Patent Citations

  • Multiple data stream processing method based on MIC co-processor

    CN105204822A

  • Network flow measurement method and system based on approximate zero error probability measurement data structure Sketch

    CN110830322A

  • Element counting method and device, readable medium and equipment

    CN112597201A

  • Methods for optimized variable-size deduplication using two stage content-defined chunking and devices thereof

    US20200081868A1

  • Storing objects in data structures

    US20200301594A1

Cited By

  • Distribution measurement method and device based on similarity dynamic compression, equipment and medium

    CN120455306A