A data flow cardinality estimation method supporting element-level deletion operations
By adopting the virtual bucket structure and design of guard units and conventional units in the data flow cardinality estimation method, the deletion operation of data elements is supported, and the problem of not being able to effectively manage dynamic data in the prior art is solved, and efficient and accurate data flow cardinality estimation is achieved.
Patent Information
- Application Number
- CN202510038648.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-01-10
AI Technical Summary
Existing data flow cardinality estimation methods cannot effectively support the deletion of data elements, resulting in the inability to achieve fast and accurate data management in a dynamic environment.
Using a virtual bucket structure, the minimum sampling value and fingerprint are recorded through the guard unit and the conventional unit respectively, supporting element-level deletion operations, and calculating the global data cardinality of the data stream through the bucket-level cardinality.
It realizes efficient data element management, supports O(1) time complexity insertion and deletion operations, improves memory utilization and calculation speed, and is suitable for cardinality estimation in dynamic data environments.
Smart Images

Figure CN119484336B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a data stream cardinality estimation method supporting element-level deletion operations. Background Art
[0002] The data stream cardinality refers to the number of different data elements in a data stream. In database management, network measurement, and security systems, data stream cardinality estimation is a fundamental task that can provide critical support for data processing and resource management, and data stream cardinality estimation is widely used in real-time data analysis. Exemplarily, a database system can monitor the number of different queries, transactions, and column values to dynamically adjust the indexing strategy and resource allocation, thereby effectively coping with high-concurrency dynamic loads. In network traffic monitoring, cardinality estimation is used to track the uniqueness of IP addresses and connections to identify abnormal behaviors. In the field of network security, cardinality estimation techniques are used to analyze packets in an intrusion detection system (IDS for short) in real time to detect abnormal activities in a timely manner and ensure system stability.
[0003] Traditional data stream cardinality estimation methods include PCSA, LogLog, and HyperLogLog, etc. Although these traditional data stream cardinality estimation methods have shown great advantages in terms of insertion operations and memory efficiency, due to their inability to handle the deletion requirements of data elements, the adaptability of these methods in dynamic application scenarios is limited. There is also a "right-to-be-forgotten" data stream model (RFDS for short) based on a hash table that can support manual deletion operations on specific data elements, but the processing speed of RFDS shows obvious deficiencies in terms of complexity and memory utilization. Summary of the Invention
[0004] A data stream cardinality estimation method supporting element-level deletion operations provided by an embodiment of the present invention can at least solve the problem that existing data stream cardinality estimation methods cannot achieve fast and accurate data management in a dynamic environment. Based on virtual buckets, it dynamically updates target data and calculates the global data cardinality of the data stream, with high memory utilization and fast and accurate calculation.
[0005] The present invention provides a data stream cardinality estimation method supporting element-level deletion operations, including updating the bucket data sequence stored in a virtual bucket in response to an update operation; wherein, a plurality of virtual buckets are provided, and each virtual bucket includes a guard unit and a plurality of regular units, the guard unit is used to record the minimum sampling value, and the regular unit is used to store fingerprints; obtaining the number of non-empty units, and calculating the bucket-level cardinality of each virtual bucket according to the number of non-empty units and the updated bucket data sequence; wherein, the number of non-empty units is the number of regular units storing the fingerprints in the virtual bucket; accumulating the bucket-level cardinalities of each virtual bucket to calculate the global data cardinality of the data stream.
[0006] In an embodiment of the present invention, the update operation includes an insertion operation; in response to the insertion operation, updating the bucket data sequence stored in the virtual bucket includes: processing the data element to be inserted according to a bucket index hash function to obtain a first bucket index; wherein, the data element to be inserted is the data element to be inserted by the insertion operation; confirming a first target bucket according to the first bucket index; wherein, the first target bucket is the virtual bucket corresponding to the first bucket index; processing the data element to be inserted according to a fingerprint hash function to obtain a first fingerprint; judging whether the first fingerprint meets the sampling condition; wherein, when the sampling condition is met, the first fingerprint is not stored in the first regular unit and the first fingerprint is less than the first target value; the first regular unit belongs to the first target bucket, and the first target value is the current minimum sampling value of the first target bucket; if the first fingerprint meets the sampling condition, obtaining a second target value and judging whether the first fingerprint is less than the second target value; wherein, the second target value is the current maximum fingerprint value of the first target bucket; if the first fingerprint is not less than the second target value, updating the first target value to the second target value and ending the current insertion operation; if the first fingerprint does not meet the sampling condition, ending the current insertion operation.
[0007] In an embodiment of the present invention, when the first fingerprint meets the sampling condition: if the first fingerprint is less than the second target value, it is determined whether the target regular unit is a null value; wherein, the target regular unit is the regular unit corresponding to the second target value, and the null value is the maximum value of the data type. When the target regular unit is the null value, the fingerprint is not stored in the target regular unit; if the target regular unit is the null value, the first fingerprint is inserted into the target regular unit, and the current insertion operation is ended; if the target regular unit is not the null value, the first target value is updated to the second target value; after updating the first target value, the second target value is updated to the first fingerprint, and the current insertion operation is ended.
[0008] In an embodiment of the present invention, before determining whether the first fingerprint meets the sampling condition, it further includes preprocessing the first target bucket; preprocessing the first target bucket includes: obtaining the additional attribute to be inserted; the additional attribute to be inserted is the additional attribute of the data element to be inserted; verifying the second regular unit according to a preset policy function to confirm whether the second regular unit stores an expired fingerprint; wherein, the second regular unit belongs to the first target bucket; the expired fingerprint is a fingerprint that fails to pass the verification of the preset policy function; if the second regular unit stores the expired fingerprint, the second regular unit is set to the null value, the expired additional attribute is updated to the additional attribute to be inserted, and the current preprocessing is ended; wherein, when the second regular unit is the null value, the fingerprint is not stored in the second regular unit; the expired additional attribute is the additional attribute of the second regular unit; if the second regular unit does not store the expired fingerprint, the current preprocessing is ended.
[0009] In an embodiment of the present invention, before verifying the first target bucket according to the preset policy function, it further includes: determining whether the additional attribute to be inserted is null; if the additional attribute to be inserted is null, it proceeds to determine whether the first fingerprint meets the sampling condition; if the additional attribute to be inserted is not null, it proceeds to verify the first target bucket according to the preset policy function.
[0010] In an embodiment of the present invention, the additional attribute to be inserted includes a timestamp; the preset policy function is set as a time policy verification function, and the time policy verification function is expressed as:
[0011] ,
[0012] wherein, the value of Indicates that the fingerprint of the current second regular unit is verified. The value of Indicates that the fingerprint of the current second regular unit fails to be verified. Is the current timestamp; Is the fingerprint stored in the current second regular unit, Is a historical timestamp, and the historical timestamp is the timestamp of the current second regular unit; Is a preset time window.
[0013] In an embodiment of the present invention, the additional attribute to be inserted includes a timestamp and a spatial region; the preset policy function is set as a spatio-temporal policy verification function, and the spatio-temporal policy verification function Is expressed as:
[0014] ,
[0015] Wherein, The value of Indicates that the fingerprint of the current second regular unit is verified. The value of Indicates that the fingerprint of the current second regular unit fails to be verified. Is the current timestamp; Is the fingerprint stored in the current second regular unit, Is a historical timestamp, and the historical timestamp is the timestamp of the current second regular unit; Is a preset time window; Is a historical spatial region, and the historical spatial region is the spatial region of the current second regular unit; Is a preset spatial region.
[0016] In one embodiment of the present invention, the update operation includes a passive deletion operation; in response to the passive deletion operation, updating the bucket data sequence stored in the virtual bucket includes: processing the data element to be deleted according to the bucket index hash function to obtain a second bucket index; wherein, the data element to be deleted is the data element to be deleted by the passive deletion operation; confirming a second target bucket according to the second bucket index; wherein, the second target bucket is the virtual bucket corresponding to the second bucket index; processing the data element to be deleted according to the fingerprint hash function to obtain a second fingerprint; determining whether the second fingerprint meets the deletion condition; wherein, the deletion condition is that the second fingerprint is stored in a third regular unit, and the third regular unit belongs to the second target bucket; if the deletion condition is met, setting the third regular unit to a null value and ending the current passive deletion operation; in the case where the third regular unit is the null value, the third regular unit does not store the fingerprint; if the deletion condition is not met, ending the current passive deletion operation.
[0017] In one embodiment of the present invention, according to the number of non-empty units and the updated bucket data sequence, calculate the bucket-level cardinality of each virtual bucket , expressed as: ,
[0018] wherein, is the bit width of the fingerprint, is the minimum sampling value of the current virtual bucket, is the number of non-empty units of the current virtual bucket, is the number of regular units of the virtual bucket, , , is the number of virtual buckets;
[0019] Accumulate the bucket-level cardinalities of each virtual bucket to calculate the global data cardinality of the data stream , expressed as: .
[0020] In one embodiment of the present invention, before responding to the update operation, it further includes: setting the number of virtual buckets and the number of regular units of each virtual bucket; wherein, the number of regular units of each virtual bucket is equal; setting each regular unit to a null value; wherein, the null value is the maximum value of the data type; setting the value of each guard unit to the initial sampling value; wherein, the initial sampling value is the maximum value of the data type; defining a bucket index hash function and a fingerprint hash function; wherein, the bucket index hash function is used to obtain the bucket index corresponding to the data element, and the fingerprint hash function is used to obtain the fingerprint corresponding to the data element.
[0021] The above technical solution of the present invention has the following beneficial effects compared with the prior art:
[0022] The method for estimating the data stream cardinality that supports element-level deletion operations according to the present invention calculates the bucket-level cardinality of each virtual bucket based on the bucket data sequence and the number of non-empty cells after virtual bucket update, and thus calculates the global data cardinality of the data stream according to the bucket-level cardinality. The overall memory utilization rate is high, and the calculation is fast and accurate. It can achieve efficient data element management under limited memory conditions. In addition, it can support data element insertion and deletion operations with an O(1) time complexity, effectively meeting the data management requirements in a dynamic environment and ensuring real-time performance and specific data management requirements of users. The method for estimating the data stream cardinality that supports element-level deletion operations according to the present invention has good scalability and adaptability, provides an efficient and reliable solution for cardinality estimation in a dynamic data environment, and is particularly suitable for various dynamic data processing scenarios such as real-time stream processing, high-concurrency database management, and security monitoring. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other embodiments can be obtained based on these drawings. In the drawings:
[0024] Figure 1 is one of the flow diagrams of the method for estimating the data stream cardinality that supports element-level deletion operations in the preferred embodiment of the present invention.
[0025] Figure 2 is the second flow diagram of the method for estimating the data stream cardinality that supports element-level deletion operations in the preferred embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] The embodiments of the present invention will be described in more detail below with reference to the drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present invention. It should be understood that the drawings and embodiments of the present invention are only for exemplary purposes and are not used to limit the protection scope of the present invention.
[0027] It should be noted that traditional data stream cardinality estimation methods, such as PCSA, LogLog, and HyperLogLog, etc., can only perform insertion operations and cannot handle the deletion requirements of data elements. Therefore, the adaptability of these methods in dynamic application scenarios is limited, especially in systems that require continuous updates, such as real-time streaming media monitoring and user behavior tracking. The fundamental limitation is that traditional data stream cardinality estimation methods map data elements to statistical indicators in the storage pool. Once inserted, the data element no longer has an independent identity in the shared memory pool, and this design is difficult to distinguish and delete individual data elements.
[0028] While RFDS can support the manual deletion operation of specific data elements, enabling users to request the deletion of their historical data from the model to meet privacy protection requirements. However, RFDS uses a hash table to store sampled data elements, resulting in instability in its performance under a large amount of data, especially in programmable hardware (such as FPGA) or systems with high concurrency requirements. The operation complexity O(q) of RFDS is related to the size q of the hash table. Therefore, the processing speed of RFDS shows obvious deficiencies in terms of complexity and memory utilization.
[0029] Existing cardinality estimation methods lack the ability to efficiently delete data with an O(1) time complexity, which also becomes a bottleneck for achieving fast and accurate data management in a dynamic environment.
[0030] To solve the above problems, as shown in Figure 1 the present invention provides a data stream cardinality estimation method that supports element-level deletion operations, including:
[0031] Responding to an update operation, updating the bucket data sequence stored in the virtual bucket.
[0032] Obtaining the number of non-empty cells, and calculating the bucket-level cardinality of each virtual bucket according to the number of non-empty cells and the updated bucket data sequence.
[0033] Accumulating the bucket-level cardinalities of each virtual bucket to calculate the global data cardinality of the data stream.
[0034] It should be noted that in the data stream cardinality estimation method that supports element-level deletion operations of the present invention, the virtual bucket array structure is used to store data elements when performing data stream cardinality estimation, which can support users' insertion and deletion operations, has sublinear memory consumption, and supports accelerated deployment through programmable hardware, ensuring high adaptability in various application scenarios.
[0035] The virtual bucket array structure includes A number of virtual buckets, each virtual bucket includes a guard unit and a regular unit. Preferably, in each virtual bucket, there is only one guard unit and regular units are provided. This design allows each virtual bucket to independently perform adaptive sampling and deletion operations, enabling efficient and flexible data management, with high space utilization, meeting the requirements of data stream cardinality estimation. To facilitate distinguishing each virtual bucket, the virtual buckets are denoted as , .
[0036] The guard unit is used to store a fingerprint, which corresponds to the minimum sampling value of the virtual bucket corresponding to this guard unit for adjusting the sampling rate. To facilitate distinguishing the minimum sampling value of each virtual bucket , the minimum sampling value of the virtual bucket is denoted as . During the process of cardinality estimation, the minimum sampling value is dynamically updated, and the fingerprints corresponding to new data elements smaller than the current sampling value are preferentially recorded. Therefore, during the process of processing each updated operation data element, the memory utilization is high.
[0037] Each regular unit is used to store a fingerprint and has an optional field for additional attributes to facilitate policy-driven deletion operations. The additional attributes are such as timestamps, spatial regions, etc. To facilitate distinguishing, the fingerprint stored in the corresponding regular unit is denoted as , . During the process of cardinality estimation, each regular unit is responsible for tracking the status of the corresponding data element to achieve accurate cardinality estimation. The number of non-empty units in each virtual bucket is the number of regular units storing fingerprints in this virtual bucket.
[0038] After the cardinality estimation system receives the updated operation data sequence, it sequentially processes each updated operation data element of the updated operation data sequence, hashes it into the virtual bucket, generates the corresponding fingerprint, thereby updating the bucket data sequence stored in the virtual bucket.
[0039] Among them, the updated operation data sequence includes updated operation data elements, the types of processing operations required by the updated operation data elements, and additional types. Among them, the types of processing operations include insert operations and passive deletion operations. The additional attributes are such as timestamps, spatial regions, etc., and the additional attributes are optional. The bucket data sequence stored in the virtual bucket includes the minimum sampling value stored in the guard unit, the fingerprints stored in the regular units, and additional attributes.
[0040] After the updated operation data sequence is executed, or when the user has a query requirement, according to the current number of non-empty units and the minimum sampling values of each virtual bucket in the bucket data sequence, calculate the bucket-level cardinality of the corresponding virtual bucket.
[0041] The data stream cardinality estimation method supporting element-level deletion operations according to the present invention calculates the bucket-level cardinality of each virtual bucket based on the bucket data sequence and the number of non-empty cells after virtual bucket update, thereby calculating the global data cardinality of the data stream. It has a high overall memory utilization rate, fast and accurate calculation. It can achieve efficient data element management under limited memory conditions. In addition, it can support data element insertion and deletion operations with a time complexity of O(1), effectively meeting the data management requirements in a dynamic environment and ensuring real-time performance and specific data management requirements of users. The data stream cardinality estimation method supporting element-level deletion operations according to the present invention has good scalability and adaptability, providing an efficient and reliable solution for cardinality estimation in a dynamic data environment, and is particularly suitable for various dynamic data processing scenarios such as real-time stream processing, high-concurrency database management, and security monitoring.
[0042] Referring to Figure 2 As shown, in some embodiments, before responding to an update operation, the data stream cardinality estimation method supporting element-level deletion operations according to the present invention further includes an initialization operation to provide initial conditions for subsequent update operations. Specifically, the initialization operation includes:
[0043] First, set the number of virtual buckets. Through this step, a basis for space division can be provided for the insertion and deletion of subsequent data elements.
[0044] Second, set the number of regular cells of each virtual bucket . It should be noted that the number of regular cells of each virtual bucket is equal, all being . Through this step, it is convenient for regular cells to store inserted data elements, ensuring the dispersion and uniformity of data. After setting the number, set each regular cell to a null value. Specifically, set the
[0045] initial value of each regular cell to , that is, the maximum value of the data type, to ensure the default empty state of the regular cell before data insertion. Similarly, set the value of each guard cell to the initial sampling value. Specifically, set the
[0046] initial value of each regular cell to , that is, the maximum value of the data type, to ensure that new data elements default to meet the sampling conditions, facilitating the system to effectively capture new data in the initial stage of operation. During the subsequent cardinality estimation process, the value of the guard cell is updated dynamically. wherein
[0047] is the bit width of the fingerprint, i.e., the bit length of the fingerprint, which is also the number of bits in a cell. Considering that there is no concept of infinity in a computer, infinity is here equivalent to the maximum value of the current data type. Exemplarily, for an unsigned 8-bit data type, infinity is equivalent to 11111111 in binary, which is -1. The minimum sampling value is related to the adaptive sampling rate. Specifically, the adaptive sampling rate is expressed as , and when is set to , the adaptive sampling rate is 1. As the minimum sampling value of the guard cell decreases, the adaptive sampling rate also decreases.
[0048] After the settings are completed, define the bucket index hash function and the fingerprint hash function to process data elements. Among them, the bucket index hash function is used to obtain the bucket index corresponding to the data element, and the fingerprint hash function is used to obtain the fingerprint corresponding to the data element.
[0049] After initialization is completed, the update operation data elements can be processed. The processing operation types of the update operation include insert operation and passive deletion operation. Considering that different update operation data elements correspond to different processing operation types and the processing operations performed on them are also different, preferably, before processing the update operation data elements, first determine whether to perform a passive deletion operation or an insert operation according to its processing operation type to meet the specific data management needs of the user.
[0050] In some embodiments of the data stream cardinality estimation method of the present invention, in response to an insert operation, update the bucket data sequence stored in the virtual bucket, including:
[0051] Process the data element to be inserted according to the bucket index hash function to obtain the first bucket index. Among them, the data element to be inserted is the data element to be inserted in the insert operation. After it is processed by the bucket index hash function, the first bucket index can be obtained.
[0052] After obtaining the first bucket index, the first target bucket into which the data element to be inserted is to be inserted can be determined according to the first bucket index. Among them, the first target bucket is the virtual bucket corresponding to the first bucket index.
[0053] At the same time, obtain the additional attribute to be inserted. The additional attribute to be inserted is the additional attribute of the data element to be inserted. The additional attribute to be inserted includes a timestamp, a spatial region, etc. In some cases, the additional attribute to be inserted may also be empty.
[0054] After obtaining the additional attribute to be inserted, it is determined whether the additional attribute to be inserted is empty or non-empty, and then transferred to the corresponding steps to process the data element to be inserted. Specifically, if the additional attribute to be inserted is empty, it is transferred to the adaptive sampling insertion sub-operation. If the additional attribute to be inserted is non-empty, it is first transferred to the policy-driven deletion sub-operation. After preprocessing the first target bucket, the adaptive sampling insertion sub-operation is performed to achieve efficient processing of the data stream.
[0055] Among them, the policy-driven deletion sub-operation is to realize automatic data cleaning based on a preset policy function, and automatically remove the data when the data meets the cleaning conditions to achieve efficient processing of the data stream.
[0056] Specifically, performing the policy-driven deletion sub-operation and preprocessing the first target bucket includes:
[0057] Verify the second regular unit according to the preset policy function to confirm whether the second regular unit stores an expired fingerprint. Among them, the second regular unit belongs to the first target bucket, that is, the second regular unit is a regular unit of the first target bucket. Preferably, the second regular unit is a non-empty regular unit, that is, during verification, each non-empty regular unit of the first target bucket is verified.
[0058] The expired fingerprint is a fingerprint that fails to pass the verification of the preset policy function. The data element corresponding to the expired fingerprint is invalid and needs to be deleted; correspondingly, the data element corresponding to the unexpired fingerprint is valid.
[0059] If the second regular unit stores an expired fingerprint, set the second regular unit to a null value. Specifically, set the fingerprint of the second regular unit to , that is, the maximum value of the data type, to delete the expired fingerprint, release the corresponding storage space. After completion, the second regular unit does not store a fingerprint.
[0060] Preferably, while deleting the expired fingerprint, update the expired additional attribute to the additional attribute to be inserted to facilitate the insertion of the subsequent data element to be inserted. Among them, the expired additional attribute is the additional attribute of the second regular unit. After completing the operation on the second regular unit, end the current preprocessing of the first target bucket and perform the adaptive sampling insertion sub-operation on the data element to be inserted.
[0061] If the second regular unit does not store an expired fingerprint, there is no need to delete it, and directly end the current preprocessing of the first target bucket and perform the adaptive sampling insertion sub-operation on the data element to be inserted.
[0062] In some embodiments of the data stream cardinality estimation method supporting element-level deletion operations described in the present invention, the additional attribute to be inserted includes a time stamp. Correspondingly, the preset policy function is set as a time policy verification function, and the time policy verification function Expressed as:
[0063] ,
[0064] Wherein, The value of Indicates that the fingerprint of the current second regular unit has passed the verification, and deletion is not required at this time. The value of Indicates that the fingerprint of the current second regular unit has not passed the verification and needs to be deleted. Indicates if. Is the current timestamp. How to obtain the current timestamp will not be elaborated here. Is the fingerprint stored in the current second regular unit, Is the historical timestamp, that is, the timestamp of the current second regular unit. Is the preset time window. When the difference between the two timestamps is not less than the preset time window, it indicates that the current fingerprint has expired and needs to be deleted.
[0065] In some embodiments of the data flow cardinality estimation method supporting element-level deletion operations according to the present invention, the additional attributes to be inserted include a timestamp and a spatial region. The preset policy function is set as a spatio-temporal policy verification function, and the spatio-temporal policy verification function Expressed as:
[0066] ,
[0067] Wherein, The value of Indicates that the fingerprint of the current second regular unit has passed the verification, and deletion is not required at this time. The value of Indicates that the fingerprint of the current second regular unit has not passed the verification and needs to be deleted. Is the historical spatial region, that is, the spatial region of the current second regular unit. Is the preset spatial region. Only when the difference between the two timestamps is less than the preset time window and the historical spatial region is within the preset spatial region, the current fingerprint is valid; otherwise, it has expired and is invalid and needs to be deleted.
[0068] In the case of completing the policy-driven deletion sub-operation, preprocessing the first target bucket, or the additional attributes to be inserted being empty, an adaptive sampling insertion sub-operation is performed. Specifically, it includes:
[0069] Process the data element to be inserted according to the fingerprint hash function to obtain the first fingerprint. The first fingerprint corresponds to the data element to be inserted.
[0070] After obtaining the first fingerprint, it is determined whether the first fingerprint meets the sampling conditions, so as to achieve efficient processing of the data stream according to the determination result. When the sampling conditions are met, the first fingerprint is not stored in the first regular unit, and the first fingerprint is less than the first target value.
[0071] The first regular unit belongs to the first target bucket, and the first target value is the current minimum sampling value of the first target bucket. Since the initialization operation is performed, the minimum sampling value of the first target bucket is the maximum value at the beginning, so that new data elements default to meet the sampling conditions, thus facilitating the system to effectively capture new data in the initial stage of operation.
[0072] If the first fingerprint does not meet the sampling conditions, the current insertion operation is ended.
[0073] If the first fingerprint meets the sampling conditions, the second target value is obtained, and it is determined whether the first fingerprint is less than the second target value. According to the comparison result of the first fingerprint and the second target value, it is determined how to update the minimum sampling value of the first target bucket. Among them, the second target value is the current maximum fingerprint value of the first target bucket.
[0074] When the first fingerprint meets the sampling conditions, if the first fingerprint is less than the second target value, it is determined whether the target regular unit is a null value. Among them, the target regular unit is the regular unit corresponding to the second target value, and the null value is the maximum value of the data type. When the target regular unit is a null value, the target regular unit does not store a fingerprint. If the target regular unit is a null value and it does not store a fingerprint, the first fingerprint is inserted into the target regular unit, and the current insertion operation is ended.
[0075] If the target regular unit is not a null value, the current first target value needs to be updated, and the first target value is updated to the second target value. After updating the first target value, the second target value is updated to the first fingerprint, and the current insertion operation is ended.
[0076] If the first fingerprint is not less than the second target value, the current first target value needs to be updated, and the first target value is updated to the second target value, and the current insertion operation is ended.
[0077] In some embodiments of the data stream cardinality estimation method supporting element-level deletion operations according to the present invention, in response to a passive deletion operation, the bucket data sequence stored in the virtual bucket is updated, including:
[0078] The second bucket index is obtained by processing the data element to be deleted according to the bucket index hash function. At the same time, the second fingerprint is obtained by processing the data element to be deleted according to the fingerprint hash function. Among them, the data element to be deleted is the data element to be deleted by the passive deletion operation. After being processed by the bucket index hash function, the second bucket index can be obtained, and after being processed by the fingerprint hash function, the second fingerprint can be obtained.
[0079] After obtaining the second bucket index, the second target bucket where the data element to be deleted is located can be determined according to the second bucket index. Among them, the second target bucket is the virtual bucket corresponding to the second bucket index.
[0080] Subsequently, it is determined whether the second fingerprint meets the deletion condition. Among them, the deletion condition is that the second fingerprint is stored in the third regular unit, and the third regular unit belongs to the second target bucket.
[0081] If the deletion condition is met, the third regular unit is set to a null value. Specifically, the fingerprint of the third regular unit is set to , that is, the maximum value of the data type, to realize the deletion of the fingerprint and release the corresponding storage space, meeting the user's specific data management requirements. After completion, the third regular unit does not store a fingerprint, and the current passive deletion operation ends. Preferably, if the third regular unit stores additional attributes, they are also cleared when the fingerprint is deleted.
[0082] If the deletion condition is not met, the deletion fails, and the current passive deletion operation ends.
[0083] In some embodiments of the data stream cardinality estimation method supporting element-level deletion operations according to the present invention, after completing the passive deletion or insertion operation, the data stream cardinality is estimated in real time. By combining the virtual bucket array structure and the dynamic update of the guard unit value, it is possible to quickly and accurately calculate the cardinality of the data stream without directly storing all data elements, meeting the requirements of limited memory conditions and high memory utilization, and ensuring efficient data analysis and real-time monitoring capabilities.
[0084] Specifically, according to the number of non-empty units and the updated bucket data sequence, the bucket-level cardinality of each virtual bucket is calculated , expressed as:
[0085] ,
[0086] Among them, is the bit width of the fingerprint, is the minimum sampling value of the current virtual bucket, is the number of non-empty units of the current virtual bucket, is the number of regular units of the virtual bucket, , , is the number of virtual buckets.
[0087] After obtaining the bucket-level cardinality of each virtual bucket , they are accumulated to calculate the global data cardinality of the data stream , expressed as:
[0088] .
[0089] It should be noted that the term "including" and its variants used in the embodiments of the present invention are open-ended, that is, "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The modifications of "one" and "a plurality" mentioned in the embodiments of the present invention are illustrative rather than restrictive. Those skilled in the art should understand that unless clearly stated otherwise in the context, it should be understood as "one or more".
[0090] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the embodiments of the present invention are all information and data authorized by the user or fully authorized by all parties, and corresponding operation entrances are provided for the user to select authorization or rejection.
[0091] The various steps described in the method embodiments provided by the embodiments of the present invention can be executed in different orders and / or executed in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The protection scope of the present invention is not limited in this regard.
[0092] The term "embodiment" in this specification means that the specific features, structures or characteristics described in connection with the embodiment may be included in at least one embodiment of the present invention. The phrase appears in various positions in the specification does not necessarily mean the same embodiment, nor does it mean being independent or alternative to other embodiments and mutually exclusive. The various embodiments in this specification are described in a related manner, and the same or similar parts between the various embodiments are referred to each other. In particular, for the device, equipment, and system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts refer to the partial description of the method embodiments.
[0093] The above-described embodiments only represent several implementation manners of the present invention, and the description is relatively specific and detailed, but it should not be construed as a limitation of the protection scope. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the appended claims.
Claims
1. A data stream cardinality estimation method supporting element-level deletion operations, characterized in that: include: In response to the update operation, the bucket data sequence stored in the virtual bucket is updated; wherein the virtual bucket is provided with a plurality of virtual buckets, each of the virtual buckets includes a guard unit and a plurality of regular units, the guard unit is used to record the minimum sampling value, and the regular unit is used to store fingerprints; Obtain the number of non-empty cells, and calculate the bucket-level cardinality of each virtual bucket based on the number of non-empty cells and the updated bucket data sequence. , expressed as: ; wherein the number of non-empty cells is the number of regular cells in the virtual bucket storing the fingerprint, is the bit width of the fingerprint, is the minimum sampling value of the current virtual bucket, is the number of non-empty units in the current virtual bucket, is the number of the regular units of the virtual bucket, , , is the number of the virtual buckets; The bucket-level cardinalities of the virtual buckets are accumulated to calculate the global data cardinality of the data flow.
2. The data stream cardinality estimation method supporting element-level deletion operation according to claim 1, characterized in that: The update operation includes an insert operation; In response to the insert operation, updating the bucket data sequence stored in the virtual bucket includes: Processing the data element to be inserted according to the bucket index hash function to obtain a first bucket index; wherein the data element to be inserted is the data element to be inserted by the insertion operation; According to the first bucket index, a first target bucket is determined; wherein the first target bucket is the virtual bucket corresponding to the first bucket index; Processing the data element to be inserted according to the fingerprint hash function to obtain a first fingerprint; Determine whether the first fingerprint satisfies a sampling condition; wherein, if the sampling condition is met, the first fingerprint is not stored in a first regular unit, and the first fingerprint is less than a first target value; the first regular unit belongs to the first target bucket, and the first target value is the current minimum sampling value of the first target bucket; If the first fingerprint meets the sampling condition, a second target value is obtained, and it is determined whether the first fingerprint is less than the second target value; wherein the second target value is the current maximum fingerprint value of the first target bucket; If the first fingerprint is not less than the second target value, the first target value is updated to the second target value, and the current insertion operation is terminated; If the first fingerprint does not meet the sampling condition, the current insertion operation is terminated.
3. The data stream cardinality estimation method supporting element-level deletion operation according to claim 2, characterized in that: When the first fingerprint satisfies the sampling condition: If the first fingerprint is less than the second target value, determining whether the target regular unit is a null value; wherein the target regular unit is the regular unit corresponding to the second target value, the null value is the maximum value of the data type, and when the target regular unit is the null value, the target regular unit does not store the fingerprint; If the target regular unit is the null value, inserting the first fingerprint into the target regular unit and ending the current inserting operation; If the target regular unit is not the null value, the first target value is updated to the second target value; after the first target value is updated, the second target value is updated to the first fingerprint, and the current insertion operation is terminated.
4. The data stream cardinality estimation method supporting element-level deletion operation according to claim 2 or 3, characterized in that: Before determining whether the first fingerprint meets the sampling condition, the method further includes preprocessing the first target bucket; Preprocessing the first target bucket includes: Acquire the additional attribute to be inserted; the additional attribute to be inserted is the additional attribute of the data element to be inserted; Verify the second regular unit according to the preset policy function to confirm whether the second regular unit stores an expired fingerprint; wherein the second regular unit belongs to the first target bucket; the expired fingerprint is a fingerprint that has not been verified by the preset policy function; If the second regular unit stores the expired fingerprint, the second regular unit is set to a null value, the expired additional attribute is updated to the additional attribute to be inserted, and the current preprocessing is terminated; wherein, when the second regular unit is the null value, the second regular unit does not store the fingerprint; the expired additional attribute is the additional attribute of the second regular unit; If the second regular unit does not store the expired fingerprint, the current preprocessing is terminated.
5. The data stream cardinality estimation method supporting element-level deletion operation according to claim 4, characterized in that: Before verifying the first target bucket according to the preset strategy function, the method further includes: Determine whether the additional attribute to be inserted is empty; If the additional attribute to be inserted is empty, proceed to determining whether the first fingerprint meets the sampling condition; If the additional attribute to be inserted is not empty, the process proceeds to verifying the first target bucket according to a preset strategy function.
6. The data stream cardinality estimation method supporting element-level deletion operation according to claim 4, characterized in that: The additional attribute to be inserted includes a timestamp; The preset strategy function is set as a time strategy verification function, and the time strategy verification function It is expressed as: , in, The value of Indicates that the fingerprint of the current second conventional unit has passed verification, The value of Indicates that the fingerprint of the current second conventional unit has not passed verification; is the current timestamp; the fingerprint stored for the current second conventional unit, is a historical timestamp, and the historical timestamp is the timestamp of the current second regular unit; The preset time window.
7. The data stream cardinality estimation method supporting element-level deletion operation according to claim 4, characterized in that: The additional attributes to be inserted include a timestamp and a spatial region; The preset strategy function is set as a spatiotemporal strategy verification function. It is expressed as: , in, The value of Indicates that the fingerprint of the current second conventional unit has passed verification, The value of Indicates that the fingerprint of the current second conventional unit has not passed verification; is the current timestamp; the fingerprint stored for the current second conventional unit, is a historical timestamp, and the historical timestamp is the timestamp of the current second regular unit; is the preset time window; is a historical space area, and the historical space area is the space area of the current second conventional unit; It is the preset space area.
8. The data stream cardinality estimation method supporting element-level deletion operation according to claim 1, characterized in that: The update operation includes a passive deletion operation; In response to the passive deletion operation, updating the bucket data sequence stored in the virtual bucket includes: Processing the data element to be deleted according to the bucket index hash function to obtain a second bucket index; wherein the data element to be deleted is the data element to be deleted by the passive deletion operation; According to the second bucket index, confirm a second target bucket; wherein the second target bucket is the virtual bucket corresponding to the second bucket index; Process the data element to be deleted according to the fingerprint hash function to obtain a second fingerprint; Determine whether the second fingerprint satisfies a deletion condition; wherein the deletion condition is that the second fingerprint is stored in a third regular unit, and the third regular unit belongs to the second target bucket; If the deletion condition is met, the third regular unit is set to a null value, and the current passive deletion operation is terminated; when the third regular unit is the null value, the third regular unit does not store the fingerprint; If the deletion condition is not met, the current passive deletion operation is terminated.
9. The data stream cardinality estimation method supporting element-level deletion operation according to claim 1, characterized in that: The bucket-level cardinality of each virtual bucket is accumulated to calculate the global data cardinality of the data flow , expressed as: 。 10. The data stream cardinality estimation method supporting element-level deletion operation according to claim 1, characterized in that: Before responding to the update operation, it also includes: Setting the number of the virtual buckets and the number of the regular units of each of the virtual buckets; wherein the number of the regular units of each of the virtual buckets is equal; Setting each of the regular cells to a null value; wherein the null value is the maximum value of the data type; Setting the value of each guard unit to an initial sampling value; wherein the initial sampling value is the maximum value of the data type; A bucket index hash function and a fingerprint hash function are defined; wherein the bucket index hash function is used to obtain the bucket index corresponding to the data element, and the fingerprint hash function is used to obtain the fingerprint corresponding to the data element.
Citation Information
Patent Citations
Data stream processing method and device, computer equipment and readable storage medium
CN119211055A
Method for estimating cardinal number in dynamic data stream
CN119248860A