Data processing method of Bloom filter based on chameleon hash function, medium and equipment

By introducing the chameleon hash function and multi-dimensional Bronze filter structure into the Bronze filter, the problems of long query time and large storage overhead of the Bronze filter on CPS devices are solved, and the false alarm rate and memory overhead are achieved.

CN120179675APending Publication Date: 2025-06-20NORTHWESTERN POLYTECHNICAL UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510238386.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing Bloom filters have a long query time on resource-constrained network physical system (CPS) devices, and the problem of a sharp increase in false alarm rate and memory overhead on resource-constrained network physical system (CPS) devices.

Method used

A multi-dimensional Bloom filter based on the chameleon hash function is used to hash messages and random seeds through a lightweight hash function, and the hash results are mapped into a two-dimensional Bloom filter, realizing data insertion and query while reducing storage and time overhead.

Benefits of technology

It significantly reduces the false positive rate and memory overhead, improves query efficiency, and solves the performance bottleneck of traditional Bloom filters on CPS devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179675A_ABST
    Figure CN120179675A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method of a Bloom filter based on a chameleon hash function, a medium and equipment, the Bloom filter is a multi-dimensional Bloom filter which dispersedly maps v hash values to v two-dimensional Bloom filters, the dimension of each two-dimensional Bloom filter is a and b, and the size of each DF unit in the two-dimensional Bloom filter is z-bit; when data are inserted into the multi-dimensional Bloom filter, Hash operation is executed for v times, v positions are set as 1, and the execution mode is as follows: firstly, the Hash value of the inserted data is calculated, and the coordinates of the inserted data in the two-dimensional Bloom filter are obtained by combining the dimensions a and b of the two-dimensional Bloom filter; and then, based on the hash values of the inserted data and the random number seed0, calculating the position which should be set to be 1 in the z bit of the element corresponding to the coordinate. The multidimensional bloom filter designed by the invention can realize lower false alarm rate and memory overhead for network physical system equipment with limited resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data processing, and particularly relates to a data processing method, medium and device of a Bloom filter based on a chameleon hash function. Background Art

[0002] A Bloom Filter (BF) is a widely used probabilistic data structure for efficiently testing whether an element belongs to a set. An empty Bloom filter is a fixed-length bit array, and all bits are initially set to 0. It defines a set of hash functions, each of which maps or hashes an element to a certain position in the array, resulting in a uniform random distribution. The number of hash functions usually depends on the required false positive rate. The probabilistic nature of the Bloom filter may lead to false positive matches, but no false negatives (the query result is either "possibly in the set" or "definitely not in the set"). In a Bloom filter with a fixed length, as the number of added elements increases, the probability of false positives increases.

[0003] The current state-of-the-art Bloom filters improve query efficiency and reduce the false positive rate by leveraging two-dimensional characteristics or modifying the Murmur hash function. Fan et al. proposed a new data structure for approximate set membership queries, the CuckooFilter, where the hash value of each element is mapped to two possible buckets (storage locations). If the first bucket is full, one of the existing elements is kicked out via Cuckoo Hashing and attempted to be relocated to another location, thus supporting dynamic adjustment and relocation. Its advantage is that it only stores the hash fingerprint of the element, thus significantly saving storage space. In the case of a low false positive rate (≤3%), the storage efficiency of the Cuckoo Filter is usually higher than that of the traditional BF, its query time is close to that of the traditional BF, but it additionally supports the delete operation. The disadvantage is that when the load factor is high, insertions may fail and the filter needs to be rebuilt or expanded, and it is less efficient than the traditional BF in scenarios with extremely low false positive rates. Patgiri et al. proposed a data structure called r-dimensional Bloom Filter (rDBF) to solve the membership query problem for large-scale sets. By using multiple independent hash functions, the input element is mapped to an r-dimensional bit vector (instead of the one-dimensional bit array of the traditional BF), and each dimension independently stores relevant information, thus enhancing space utilization. The r-dimensional structure reduces the probability of hash collisions in the traditional BF, thus significantly reducing the false positive rate, which decreases exponentially with the increase in the number of dimensions r and is suitable for large-scale set queries; the r-dimensional structure allows dynamic adjustment of the size of each dimension and the number of hash functions to adapt to different data scales. The disadvantages are increased storage overhead, more computational resources required, and design skills needed. Nayak et al. proposed an improved Bloom filter called RobustBF (Robust Bloom Filter), which effectively reduces hash collisions and the false positive rate by expanding the traditional one-dimensional bit array into a two-dimensional matrix, where each element is mapped to two dimensions of the matrix by two independent hash functions; by optimizing the hash function and matrix layout, bit collisions and redundant storage are further reduced, supporting complex queries based on multi-dimensional information and enhancing adaptability in two-dimensional data scenarios. Its false positive rate is inversely proportional to the size of the two-dimensional matrix and proportional to the number of hash functions in each dimension. The false positive rate of RobustBF is 20%-60% lower than that of the traditional BF, and the storage requirement is reduced by about 30%-50%, making it more suitable for processing two-dimensional or multi-dimensional data. The disadvantage is that the two-dimensional structure design and parameter tuning are more complex than those of the traditional BF, and the performance will degrade in some low-performance hardware. Gebretsadik et al. proposed an improved Bloom filter called eBF (Enhanced Bloom Filter) to enhance the intrusion detection ability in the Internet of Things (IoT) environment.Support the deletion and dynamic update of elements by introducing additional structures (such as counting arrays), design the filter as a multi-level structure, store different types or priorities of data in each level, and use an improved hash function to reduce hash collisions between different elements, further reducing the false positive rate; adopt compression technology and segmented storage strategy to save memory while ensuring a low false positive rate. The false positive rate is reduced by 30%-60% compared with traditional BF, the memory usage is saved by about 30%-40%, and the query speed is 10%-20% faster. The disadvantages are complex implementation and large latency on low-performance IoT devices. For resource-constrained cyber-physical system (CPS) devices, the false positive rate and memory overhead of the above Bloom filter still need to be further reduced to meet the efficient continuous message authentication requirements of CPS devices. Summary of the Invention

[0004] The present invention aims at the deficiencies in the prior art and provides a data processing method, medium and device of a Bloom filter based on a chameleon hash function.

[0005] To achieve the above object, the present invention adopts the following technical solutions:

[0006] In a first aspect, the present invention provides a data processing method of a Bloom filter based on a chameleon hash function, where the Bloom filter is a multi-dimensional Bloom filter that distributes v hash values into v two-dimensional Bloom filters, the dimensions of the two-dimensional Bloom filter are a and b, and the size of each DF unit in the two-dimensional Bloom filter is z bits;

[0007] When inserting data into the multi-dimensional Bloom filter, perform v hash operations and set v positions to 1. The execution method is as follows: First, calculate the hash value of the inserted data, and combine the dimensions a and b of the two-dimensional Bloom filter to obtain the coordinates of the inserted data in the two-dimensional Bloom filter; then calculate the position in the z bits of the element corresponding to the coordinates that should be set to 1, and set it to 1.

[0008] Optionally, inserting data into the multi-dimensional Bloom filter specifically includes the following steps:

[0009] S1: Input the inserted data data, the number of verification points v of the two-dimensional Bloom filter, the random number seed0, and the dimensions a and b of the two-dimensional Bloom filter;

[0010] S2: Perform a hash operation on the inserted data data and the random number seed0 through the lightweight hash function Lhash(·) to obtain the position parameter lo0 = Lhash(data, seed0);

[0011] S3: Perform modulo operations on dimensions a and b of the two-dimensional Bloom filter to obtain the coordinates of the inserted data in the two-dimensional Bloom filter: a_lo0 = lo0 mod a, b_lo0 = lo0 mod b;

[0012] S4: Based on the coordinates a_lo0 and b_lo0, calculate the positions in the z bits corresponding to the elements at these coordinates and set them to 1;

[0013] S5: For the i-th verification point next, where i < v, calculate the position parameter lo i = Lhash(data, lo i-1 ), calculate the coordinates of the inserted data in the two-dimensional Bloom filter: a_lo i = lo i mod a, b_lo i = lo i mod b, and based on a_lo i and b_lo i , calculate the positions where the inserted data is mapped in the z bits and set them to 1; until the insertion for all v verification points is completed.

[0014] Optionally, in S2, the lightweight hash function Lhash(·) uses Murmur Hash.

[0015] Optionally, in S4, the calculation of the positions where the inserted data is mapped in the z bits and setting them to 1 is specifically:

[0016]

[0017] where 2DBF0 represents the first two-dimensional Bloom filter, << represents the left shift operation, represents the largest prime number not exceeding z.

[0018] Optionally, in S5, the calculation of the positions where the inserted data is mapped in the z bits and setting them to 1 is specifically:

[0019]

[0020] where 2DBF i represents the (i + 1)-th two-dimensional Bloom filter, << represents the left shift operation, represents the largest prime number not exceeding z.

[0021] Optionally, the probability of an error in querying an element not inserted into the multi-dimensional Bloom filter is:

[0022]

[0023] Where ε represents the probability of error, TM represents the theoretical memory requirement of the multi-dimensional Bloom filter, and N represents the number of inserted elements.

[0024] Optionally, the dimensions a and b of the two-dimensional Bloom filter are calculated by the following formula:

[0025]

[0026] Where pri max (x) represents the largest prime number less than x, pri min (x) represents the smallest prime number greater than x, TM represents the theoretical memory requirement of the multi-dimensional Bloom filter, and K is the number of hash functions.

[0027] Optionally, the actual memory requirement of the multi-dimensional Bloom filter is:

[0028] AM = a * b * z * K

[0029] Where AM represents the actual memory requirement.

[0030] In a second aspect, the present invention provides a computer-readable storage medium storing a computer program, and the computer program causes a computer to execute the data processing method of the Bloom filter based on the chameleon hash function as described in the first aspect.

[0031] In a third aspect, the present invention provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the data processing method of the Bloom filter based on the chameleon hash function as described in the first aspect is implemented.

[0032] The beneficial effects of the present invention are as follows: The present invention performs a hash operation on a message and a random seed through a lightweight hash function, and maps the hash result to a two-dimensional Bloom filter by performing a modulo operation on the dimensions a and b of the two-dimensional Bloom filter, thereby reducing the storage and time overhead while realizing data insertion and query, solving the problems of large query time, large required storage overhead, and sharp increase in storage overhead with the reduction of the false positive rate in the existing Bloom filter on CPS resource-constrained devices, and being able to achieve a lower false positive rate and memory overhead. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 is a schematic structural diagram of the multi-dimensional Bloom filter proposed by the present invention.

[0034] Figure 2 is a schematic diagram of the insertion process algorithm of the multi-dimensional Bloom filter. DETAILED DESCRIPTION

[0035] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application.

[0036] In one embodiment, the present invention proposes a data processing method for a Bloom filter based on a chameleon hash function. This Bloom filter is a multi-dimensional Bloom filter (mBF) that distributes v hash values into v two-dimensional Bloom filters (2DBF). Let a and b be the dimensions of the 2DBF, and the size of each BF unit is z bits. First, calculate the hash value of the inserted data, and define its position in the 2DBF as i = h(data) mod a, j = h(data) mod b. Then, calculate the specific position in z bits as where << represents the left shift operation, represents the largest prime number not exceeding z. The proposed mBF structure is as Figure 1 shown.

[0037] When inserting data, the data will undergo v hash operations in sequence, and each hash value will be mapped to the corresponding position in the 2DBF. This method significantly reduces the false positive rate compared to the traditional BF and further reduces the memory overhead. Its insertion algorithm is as Figure 2 shown in Algorithm 1.

[0038] In Algorithm 1, first, the data and the random number seed0 are hashed through the lightweight hash function Lhash(·) to obtain the position parameter lo0. To map lo0 into the 2DBF, through modulo operations on the dimensions a and b of the 2DBF, the specific coordinates a_lo0 and b_lo0 of the data in the 2DBF are obtained, and the corresponding position of the coordinate element is set to 1.

[0039] It should be noted that in the computing unit, each position of the BF can be divided into z bits. Therefore, to set a specific bit to 1, a bitwise OR operation needs to be performed on the target bit position. In addition, this algorithm can obtain an integer less than z by taking the modulo of to ensure that the bit position of the pointed element does not exceed the length of the element.

[0040] For the subsequent 2DBF, the same operation is performed, but the random number seed0 will be replaced by the previous position parameter lo0. Finally, a certain bit of the corresponding coordinate element of each 2DBF in v different 2DBF is set to 1, thus completing the insertion of the Bloom filter.

[0041] Here, Lhash(·) is a lightweight hash function, such as MurmurHash. To study the relationship between the false positive rate and memory usage of this filter, assume that the theoretical memory requirement of mBF is TM, and the number of inserted elements is N. When inserting data into mBF, v hash operations are performed, setting a certain position of the elements at v positions to 1 respectively. Therefore, the probability that any specific bit in BF is not set is And the probability that a certain bit is set to 1 is At this time, the probability of an error in querying an element not inserted into mBF is defined by formula (1):

[0042]

[0043] Generally, assume that the optimal number of hash functions is Substituting it into the formula, we get

[0044] Therefore, theoretically, the memory requirement TM depends on the inserted data volume N and the false positive rate ε. The length and width of 2DBF can be set according to TM.

[0045] However, the actual memory requirement of mBF depends not only on N and ε, but also on the sizes of a and b. Through TM and z, a and b can be calculated using formula (2) and formula (3):

[0046]

[0047]

[0048] Among them, pri max (x) represents the largest prime number less than x, and pri min (x) represents the smallest prime number greater than x. K is the number of hash functions. Therefore, a and b should be prime numbers, and a ≠ b. If a and b are not prime numbers, the false positive rate will increase. The actual memory requirement of mBF is denoted as AM, and its calculation formula is (4).

[0049] AM = a * b * z * K (4)

[0050] To verify the efficiency of the proposed mBF, on the ESP32-D0WDQ6 platform with 2 Xtensa 32-bit LX6 MCUs, 448KB ROM, 520KB SRAM, and a main frequency of 240MHz, the query time overhead and storage overhead of mBF and traditional BF, robustBF, and eBF were analyzed and compared, and the storage overhead and query time of mBF at different false positive rates were tested. The results are as follows:

[0051] (1) The single query time of mBF is 20% lower than that of eBF, 40% lower than that of robustBF, and 45% lower than that of traditional BF.

[0052] (2) For the storage overhead of 1 million verification points, mBF is 52.5% lower than eBF, 94.1% lower than robustBF, and 97.1% lower than traditional BF.

[0053] (3) When storing 10,000 verification points, the storage overhead of mBF is 0.574 KB when the false positive rate is 0.001, and 0.902 KB when the false positive rate is 0.000001; when storing 50,000 verification points, the storage overhead of mBF is 2.589 KB when the false positive rate is 0.001, and 5.121 KB when the false positive rate is 0.000001.

[0054] (4) When storing 10,000 verification points, the execution time overhead of mBF is approximately 95.89 milliseconds when the false positive rate is in the range of [0.000001, 0.001]; when storing 30,000 verification points, the execution time overhead of mBF is approximately 287.69 milliseconds when the false positive rate is in the range of [0.000001, 0.001], and its time overhead is hardly affected by the false positive rate.

[0055] The present invention performs a hashing operation on the message and the random seed through a lightweight hash function, and maps the hashing result to 2DBF by performing a modulo operation on the dimensions a and b of 2DBF, so as to reduce the storage and time overhead while realizing data insertion and query. Possible alternatives are to change the size of mBF or expand BF from two dimensions to multiple dimensions, but this will increase the computational complexity.

[0056] In another embodiment, the present invention proposes a computer-readable storage medium storing a computer program, and the computer program causes a computer to execute the data processing method of the Bloom filter based on the chameleon hash function in the foregoing embodiment.

[0057] In another embodiment, the present invention proposes an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the data processing method of the Bloom filter based on the chameleon hash function in the foregoing embodiment is realized.

[0058] In the embodiments disclosed in the present application, the computer storage medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The computer storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any suitable combination of the foregoing. More specific examples of the computer storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0059] Those of ordinary skill in the art will recognize that the units and algorithm steps of the examples described in connection with the embodiments disclosed in the present application can be implemented in electronic hardware or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled artisans may use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.

[0060] The above are only the preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the concept of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art in the technical field, several improvements and refinements made without departing from the principle of the present invention should be regarded as within the protection scope of the present invention.

Claims

1. A data processing method of a bloom filter based on a chameleon hash function, characterized in that: The Bloom filter is a multidimensional Bloom filter that distributes v hash values ​​into v two-dimensional Bloom filters, the dimensions of the two-dimensional Bloom filter are a and b, and the size of each DF unit in the two-dimensional Bloom filter is z bits; When inserting data into the multidimensional Bloom filter, v hash operations are performed and v positions are set to 1. The execution method is: first calculate the hash value of the inserted data, and combine the dimensions a and b of the two-dimensional Bloom filter to obtain the coordinates of the inserted data in the two-dimensional Bloom filter; then calculate the position in the z position of the element corresponding to the coordinate that should be set to 1, and set it to 1.

2. The data processing method of the bloom filter based on the chameleon hash function according to claim 1, characterized in that: Inserting data into the multidimensional Bloom filter specifically includes the following steps: S1: Input insertion data data, the number of verification points v of the two-dimensional Bloom filter, the random number seed0, and the dimensions a and b of the two-dimensional Bloom filter; S2: Perform a hash operation on the inserted data data and the random number seed0 through the lightweight hash function Lhash(·) to obtain the position parameter lo0=Lhash(data,seed0); S3: Perform a modulo operation on the dimensions a and b of the two-dimensional Bloom filter to obtain the coordinates of the inserted data in the two-dimensional Bloom filter a_lo0=lo0 mod a, b_lo0=lo0 mod b; S4: According to the coordinates a_lo0 and b_lo0, calculate the position in the z position of the element corresponding to the coordinate that should be set to 1, and set it to 1; S5: For the next i-th verification point, i<v, calculate the position parameter lo i = Lhash(data,lo i-1 ), calculate the coordinate a_lo of the inserted data in the two-dimensional Bloom filter i =lo i mod a, b_lo i =lo i mod b, according to a_lo i and B_LO i , calculate the position of the inserted data mapping in the z position and set it to 1; until a total of v verification points have completed the insertion.

3. The data processing method of the bloom filter based on the chameleon hash function as claimed in claim 2, characterized in that: In S2, the lightweight hash function Lhash(·) adopts Murmur Hash.

4. The data processing method of the bloom filter based on the chameleon hash function according to claim 2, characterized in that: In S4, the calculation inserts the position of the data mapping in the z position and sets it to 1, specifically: Where 2DBF0 represents the first two-dimensional Bloom filter, << represents the left shift operation, Represents the largest prime number not exceeding z.

5. The data processing method of the bloom filter based on the chameleon hash function according to claim 2, characterized in that: In S5, the calculation inserts the position of the data mapping in the z position and sets it to 1, specifically: Where, 2DBF i represents the i+1th two-dimensional Bloom filter, << represents a left shift operation, Represents the largest prime number not exceeding z.

6. The data processing method of the bloom filter based on the chameleon hash function according to claim 1, characterized in that: The probability of error in querying elements that are not inserted into the multidimensional Bloom filter is: Where ε represents the probability of error, TM represents the theoretical memory requirement of the multidimensional Bloom filter, and N represents the number of elements inserted.

7. The data processing method of the bloom filter based on the chameleon hash function according to claim 1, characterized in that: The dimensions a and b of the two-dimensional Bloom filter are calculated by the following formula: In the formula, pri max (x) represents the largest prime number less than x, pri min (x) represents the smallest prime number greater than x, TM represents the theoretical memory requirement of the multidimensional Bloom filter, and K is the number of hash functions.

8. The data processing method of the bloom filter based on the chameleon hash function according to claim 7, characterized in that: The actual memory requirement of the multidimensional Bloom filter is: AM=a*b*z*K Where AM represents the actual memory requirement.

9. A computer-readable storage medium storing a computer program, characterized in that: The computer program enables a computer to execute the data processing method of a Bloom filter based on a chameleon hash function as described in any one of claims 1 to 8.

10. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the data processing method of the bloom filter based on the chameleon hash function as described in any one of claims 1 to 8 is implemented.