Hash table data processing method and device
By constructing a two-level hash table, dynamically adjusting the number of hash buckets and subslots, and combining status flags and address pointer mapping, the problems of high hash table storage resource consumption and high hash conflict rate are solved, and efficient and reliable data processing is achieved.
Patent Information
- Application Number
- CN202510606528.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-10-14
AI Technical Summary
When processing ultra-large-scale data sets, existing hash tables consume high storage resources and have high hash collision rates, making it difficult to simultaneously achieve low storage resource consumption and high hash collision processing efficiency.
A hash table with a two-level table structure includes a first-level table and a second-level table. The first-level table contains multiple hash buckets and first subslots, and the second-level table contains multiple address columns and second subslots. Private keys and signature values are generated by combining keywords, and free subslots are dynamically selected for storage. Status flags and address pointers are used for mapping, reducing storage resource consumption and hash conflict risks.
It realizes flexible configuration of storage space, reduces storage resource consumption, improves hash table access efficiency and data processing reliability, adapts to data storage needs of different scales, reduces hash conflict risks, and improves data processing accuracy and efficiency.
Smart Images

Figure CN120780228A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a method and device for processing data of a hash table. Background Art
[0002] With the growing demand for efficient classification and processing of massive data packets in modern data networks, hash tables, with their fast search performance and flexible structural design, have become a core technology for packet classification. However, traditional hash tables, when dealing with extremely large data sets, need to store complete key-value pairs, which consumes a large amount of storage resources. Furthermore, traditional hash tables often experience hash collisions. When multiple keys are mapped to the same hash bucket through a hash function, a conflict resolution strategy is required.
[0003] To address the problem of high storage resource consumption, existing methods mainly use compressed hash technology to reduce storage requirements. Compressed hashing uses an additional bitmap to mark valid data in the key value, and uses valid data to replace the complete key value to achieve data compression, thereby reducing the storage occupancy of the hash table. In addition, its key-value separation strategy can also eliminate storage redundancy to a certain extent. However, compressed hashing is limited by the design of the bitmap marking rule, and the efficiency of hash construction, search, and insertion is low, which easily leads to a higher hash conflict rate problem. In addition, existing compressed hash algorithms usually adopt a strategy of one bitmap marking rule corresponding to one hash table, resulting in a large processing delay caused by insertion serialization.
[0004] To address hash collisions, existing methods include open addressing, chain addressing, public overflow, and Cuckoo Hash. Cuckoo Hash uses hash values to select hash buckets and subslots to store key-value pairs, and employs a culling mechanism to resolve hash collisions. While this approach offers high table utilization and high search performance, maintaining multiple hash tables results in high storage consumption, making it difficult to meet resource constraints in high-performance network equipment. Summary of the Invention
[0005] The main purpose of the present invention is to provide a hash table data processing method and device, aiming to solve the problem that existing hash tables are difficult to simultaneously take into account low storage resource consumption and hash conflict processing efficiency.
[0006] A first aspect of the present invention provides a data processing method for a hash table, which includes: constructing a hash table; wherein the hash table includes a first-level table and a second-level table, the first-level table includes multiple hash buckets, each of the hash buckets includes multiple first subslots, and the second-level table includes multiple address bars, the address bars correspond one-to-one to the first subslots, and each of the address bars includes multiple second subslots; obtaining a first key-value pair, the first key-value pair includes a first keyword and a first data item; combining data of preset bits in the first keyword to generate a first private key, and combining data of non-preset bits in the first keyword to generate a first signature value; selecting a target subslot from the idle first subslots; storing the first signature value in the target subslot, and storing the first private key and the first data item in the second subslot of the address bar corresponding to the target subslot.
[0007] Optionally, in a first implementation of the first aspect of the present invention, constructing a hash table includes: allocating the first-level table and the second-level table in the storage space based on a preset compression factor; setting a plurality of hash buckets in the storage space of the first-level table, and allocating a plurality of first subslots to each hash bucket, and configuring an alternative bucket for each first subslot from the plurality of hash buckets; setting a plurality of address bars in the storage space of the second-level table; wherein each address bar includes a plurality of second subslots, the number of the address bars is equal to the number of the first subslots, and the number of second subslots of each address bar is the preset compression factor; presetting different address pointers for different first subslots; wherein the address pointer corresponds one-to-one to the address of the address bar, so that the matching second subslot in the corresponding address bar can be accessed through the first subslot.
[0008] Optionally, in a second implementation of the first aspect of the present invention, a status flag is also preset in the first subslot, and the status flag is used to map the storage status of multiple second subslots in the address bar corresponding to the first subslot. After presetting different address pointers for different first subslots, it also includes: detecting the storage status of multiple second subslots in each address bar; and mapping the storage status of each second subslot to the corresponding status flag in the first subslot.
[0009] Optionally, in a third implementation of the first aspect of the present invention, before combining the data of the preset bits in the first keyword to generate a first private key, and combining the data of non-preset bits in the first keyword to generate a first signature value, it also includes: obtaining a bitmap marking rule list; wherein the bitmap marking rule list includes multiple preset bitmap marking rules, and different preset bitmap marking rules are used to mark different preset bits in the keyword; performing simulation tests on each of the preset bitmap marking rules in the bitmap marking rule list to screen out target bitmap marking rules that meet preset performance conditions; determining the preset bits in the first keyword based on the target bitmap marking rules; after selecting the target subslot from the idle first subslot, it also includes: storing the target bitmap marking rule to the second subslot of the address bar corresponding to the address pointer of the target subslot.
[0010] Optionally, in a fourth implementation of the first aspect of the present invention, the types of the idle first subslots include a first idle subslot, a second idle subslot, and a third idle subslot. Before selecting a target subslot from the idle first subslot, the method includes: calculating a first primary hash bucket address and a first secondary hash bucket address based on the first signature value to determine the primary hash bucket and the secondary hash bucket corresponding to the first signature value in the hash bucket of the first-level table; judging whether there is a first idle subslot in the first subslot in the primary hash bucket and the secondary hash bucket; wherein the status flag of the first idle subslot indicates that the second subslots corresponding to the address bar are all vacant, or the signature value of the first idle subslot matches the first signature value, and the status flag of the first idle subslot indicates that the corresponding address bar There is at least one vacant second subslot in the primary hash bucket; if the first free subslot does not exist in the first subslots in the primary hash bucket and the secondary hash bucket, then determine whether there is a second free subslot in the first subslot of each primary alternative bucket; wherein, the primary alternative bucket is the alternative bucket corresponding to each of the first subslots of the primary hash bucket, and the status flag of the second free subslot indicates that the second subslots corresponding to the address bar are all vacant; if the second free subslot does not exist in the first subslot of each of the primary alternative buckets, then search for the third free subslot in the first subslot of the secondary alternative bucket step by step; wherein, the secondary alternative bucket is the alternative bucket corresponding to each of the first subslots of the secondary hash bucket, and the status flag of the third free subslot indicates that the second subslots corresponding to the address bar are all vacant.
[0011] Optionally, in a fifth implementation of the first aspect of the present invention, the step-by-step search for the third free subslot in the first subslot of the secondary alternative bucket includes: determining whether the third free subslot exists in the first subslot of the secondary alternative bucket; if the third free subslot does not exist in the secondary alternative bucket, using the first subslot at the target position in the secondary hash bucket as a discarded subslot; cyclically shifting each of the first subslots in the secondary hash bucket; replacing the secondary hash bucket with the secondary alternative bucket corresponding to the discarded subslot, and replacing the secondary alternative bucket with the alternative bucket corresponding to each of the first subslots in the replaced secondary hash bucket, so as to determine whether the third free subslot exists in the first subslot of each of the replaced secondary alternative buckets, until it is confirmed that the third free subslot exists in the first subslot of the secondary alternative bucket.
[0012] Optionally, in the sixth implementation of the first aspect of the present invention, the data processing method further includes: recording the correspondence between the culled sub-slots and each of the sub-alternative buckets and the corresponding culled sub-slots in sequence, and counting the number of the culled sub-slots; if the number of the culled sub-slots is greater than a preset culling threshold, or the first sub-slot at the target position in the current sub-hash bucket has been recorded as the culled sub-slot, reporting insertion failure information.
[0013] Optionally, in a seventh implementation of the first aspect of the present invention, selecting a target subslot from the idle first subslot includes: determining the type of the idle first subslot; if the idle first subslot is the first idle subslot, selecting a first target hash bucket from the primary hash bucket and the secondary hash bucket based on a first priority, and selecting the first idle subslot closest to the tail in the first target hash bucket as the target subslot; if the idle first subslot is the second idle subslot, selecting a second target hash bucket from the primary alternative bucket based on a second priority, and moving the element of the second target hash bucket corresponding to the first subslot in the primary hash bucket to the second idle subslot closest to the tail in the second target hash bucket, and using the first subslot corresponding to the second target hash bucket in the primary hash bucket as the target subslot, and using the address pointer of the second idle subslot closest to the tail in the second target hash bucket before the move as the address of the target subslot. pointer; if the idle first subslot is the third idle subslot, a third target hash bucket is selected from the secondary candidate bucket based on the third priority, and the elements in the culled subslot in the secondary hash bucket corresponding to the third target hash bucket are moved to the third idle subslot closest to the tail of the third target hash bucket, and the secondary candidate bucket is replaced by the secondary hash bucket, and the hash bucket where the culled subslot corresponding to the replaced secondary candidate bucket is located is replaced by the hash bucket where the replaced secondary candidate bucket is located, so that the elements in the culled subslot in the replaced secondary hash bucket are moved to the first subslot that is vacant after the secondary candidate bucket is moved, until the elements in the culled subslot in the most basic secondary hash bucket are moved to the first subslot that is vacant after the move in the corresponding secondary candidate bucket, and the culled subslot in the most basic secondary hash bucket is used as the target subslot, and the address pointer of the third idle subslot closest to the tail of the third target hash bucket before the move is used as the address pointer of the target subslot.
[0014] Optionally, the data processing method in the eighth implementation manner of the first aspect of the present invention further includes: upon receiving a search instruction, obtaining a second keyword and a to-be-determined map marking rule corresponding to the search instruction, and determining the to-be-determined bit corresponding to the to-be-determined map marking rule; combining the data of the to-be-determined bit in the second keyword to generate a second private key, and combining the data of the non-to-be-determined bit in the second keyword to generate a second signature value; calculating the second primary hash bucket address and the second secondary hash bucket address based on the second signature value to determine the to-be-determined primary hash bucket and the to-be-determined secondary hash bucket corresponding to the second signature value; judging whether there is the first signature value matching the second signature value in each of the first subslots in the to-be-determined primary hash bucket and the to-be-determined secondary hash bucket; if there is a matching first signature value, then determining the address of the first subslot corresponding to the first signature value based on the address of the first subslot The address pointer locates the address bar in the corresponding second-level table, and compares the to-be-positioned bitmap marking rule with the target bitmap marking rule of each second subslot in the corresponding address bar, and compares the second private key with the first private key of each second subslot in the corresponding address bar; if the target bitmap marking rule of the second subslot is consistent with the to-be-positioned bitmap marking rule, and the first private key and the second private key of the same second subslot are consistent, then the first data item corresponding to the second subslot is output; when an update instruction is received, the first data item is updated; when a delete instruction is received, the data stored in the second subslot is deleted, and the status flag in the corresponding first subslot is updated. If all the second subslots mapped by the status flag are empty, the first signature value in the corresponding first subslot is cleared.
[0015] A second aspect of the present invention further provides a data processing device for a hash table, wherein the hash table is obtained based on the data processing method for a hash table as described above, and the data processing device for the hash table includes: a construction module for constructing a hash table; wherein the hash table includes a first-level table and a second-level table, the first-level table includes multiple hash buckets, each of the hash buckets includes multiple first subslots, and the second-level table includes multiple address bars, the address bars correspond one-to-one to the first subslots, and each of the address bars includes multiple second subslots; an acquisition module for acquiring a first key-value pair, wherein the first key-value pair includes a first keyword and a first data item; a compression processing module for combining data of preset bits in the first keyword to generate a first private key, and combining data of non-preset bits in the first keyword to generate a first signature value; a control module for selecting a target subslot from the idle first subslots, and storing the first signature value in the target subslot, and storing the first private key and the first data item in the second subslot corresponding to the address bar of the target subslot; and a storage module for storing the hash table.
[0016] The embodiment of the present invention provides a hash table data processing method and device, which first constructs a hash table of a specific structure, then obtains a first key-value pair, and reorganizes the first keyword in the first key-value pair to obtain a first private key and a first signature value, and finally selects a target subslot from an idle first subslot, stores the first signature value in the target subslot, and stores the first private key and the first data item in the second subslot of the address bar corresponding to the target subslot, thereby realizing the processing of hash table data. The hash table constructed by this method realizes flexible configuration of storage space, and can dynamically adjust the number of hash buckets and subslots according to actual needs to adapt to data storage requirements of different scales, which can effectively reduce storage resource consumption and improve storage utilization. The mapping setting of the two-level table also reduces the risk of hash conflicts and improves the access efficiency of the hash table. Data processing based on the hash table also realizes efficient storage and access to key-value pair data, which can effectively improve the reliability and accuracy of data processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 A flowchart of a method for processing data in a hash table according to an embodiment of the present invention;
[0018] Figure 2 for Figure 1 A schematic flow chart of an embodiment of step 101 in the embodiment;
[0019] Figure 3 for Figure 2 An example schematic diagram of an embodiment of a hash table constructed in the embodiment;
[0020] Figure 4 for Figure 2 An example schematic diagram of another embodiment of the hash table constructed in the embodiment;
[0021] Figure 5 for Figure 1 An exemplary schematic diagram of an embodiment of step 103 in the embodiment;
[0022] Figure 6 for Figure 1 A schematic diagram of a flow chart of an embodiment before step 104 in the embodiment;
[0023] Figure 7 for Figure 6 A flow chart of an embodiment of step 1034 in the embodiment;
[0024] Figure 8 for Figure 1 A flow chart of an embodiment of step 104 in the embodiment;
[0025] Figure 9 for Figure 1An exemplary schematic diagram of an embodiment of an embodiment;
[0026] Figure 10 1 is a flow chart of another embodiment of a method for processing data of a hash table according to an embodiment of the present invention;
[0027] Figure 11 Schematic diagram of functional modules of an embodiment of a data processing device for a hash table according to an embodiment of the present invention. DETAILED DESCRIPTION
[0028] The terms "first," "second," "third," "fourth," and the like in the specification and claims of the present invention and in the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to these processes, methods, products, or apparatus.
[0029] For ease of understanding, the specific process of the data processing method of the hash table in the embodiment of the present invention is described below. Figure 1 , Figure 1 1 is a flow chart of an embodiment of a method for processing data of a hash table according to an embodiment of the present invention. In this embodiment, the method for processing data of a hash table includes:
[0030] 101. Construct a hash table; wherein the hash table includes a first-level table and a second-level table, the first-level table includes multiple hash buckets, each hash bucket includes multiple first subslots, and the second-level table includes multiple address columns, the address columns correspond one-to-one to the first subslots, and each address column includes multiple second subslots;
[0031] In this embodiment, the data processing method is implemented by specifically relying on a hash table with a specific structure constructed in this embodiment. Specifically, the hash table has a two-level table structure, comprising a first-level table and a second-level table. The first-level table contains multiple hash buckets, each of which contains multiple first subslots. These first subslots are used to store key data information, such as signature values. The second-level table contains multiple address fields, each corresponding one-to-one with the first subslots in the first-level table. Each address field further contains multiple second subslots, which are used to store specific data content, such as private keys and data items. This two-level hash table structure effectively distinguishes and stores subsequently generated signature values and private keys, and performs insertion and access control on data items, achieving efficient data organization and storage. This facilitates more flexible and rapid data processing based on the correspondence between first and second subslots, as well as the association between private keys and keywords, thereby improving hash table utilization and reducing storage space usage, while also enhancing data processing efficiency and accuracy.
[0032] Optional, see Figure 2 In some embodiments, constructing a hash table includes:
[0033] 1011. Allocate the first-level table and the second-level table in the storage space based on a preset compression factor;
[0034] 1012. Set a plurality of hash buckets in the storage space of the first-level table, allocate a plurality of first subslots to each hash bucket, and configure a candidate bucket for each first subslot from the plurality of hash buckets;
[0035] 1013. Set a plurality of address columns in the storage space of the second-level table; each address column includes a plurality of second subslots, the number of address columns is equal to the number of first subslots, and the number of second subslots in each address column is a preset compression factor;
[0036] 1014. Preset different address pointers for different first sub-slots, wherein the address pointers correspond to the addresses in the address column one by one, so that the first sub-slot can be used to access the matching second sub-slot in the corresponding address column.
[0037] In this optional embodiment, the preset compression factor is a parameter used to allocate storage space between the first-level table and the second-level table of the hash table, and can establish a compression relationship between the first-level table and the second-level table. By decoupling the data into two parts for storage, the common part of some data stored in the first-level table is allocated, and the characteristic part of the data with the same common part is allocated in the second-level table. By setting a part of the data used to associate the two-level table structure, the data in the second-level table is compressed, and the data in the first-level table can be quickly located to the corresponding data in the second-level table, which facilitates the management of the entire hash table by operating the first-level table, thereby improving the efficiency and flexibility of subsequent data processing and reducing the storage space occupied. The preset compression factor can be set according to actual storage requirements and the characteristics of the storage data structure to achieve reasonable and efficient use of storage space.
[0038] In this optional embodiment, further optionally, in some embodiments, the preset compression factor is the ratio of the expected maximum number of entries in the second-level table to the maximum number of entries in the first-level table, that is, c=P / D, where P is the expected maximum number of entries in the second-level table, which is used to characterize the maximum number of key-value pairs that the constructed hash table needs to support. For example, in ultra-large-scale data classification scenarios, P can be set to millions, such as 1M, to meet the needs of massive data processing; D is the expected maximum number of entries in the first-level table, that is, the address depth of the hash bucket, which is determined by the storage resources and the hash conflict rate. For example, if D=128K, the first-level table manages 128K signatures through the hash bucket and subslot structure, and each hash bucket contains multiple subslots to improve the conflict handling capability; c is the preset compression factor, which is used to control the storage compression multiple of the second-level table. By adjusting the preset compression factor, a dynamic balance can be achieved between storage resources and performance, resulting in a first-level table with lower storage resource consumption for data processing. For example, when c=8, the capacity of the first-level table is 1 / 8 of that of the second-level table, thereby significantly reducing storage redundancy and resource consumption.
[0039] In this optional embodiment, the storage space can specifically be a hierarchical storage architecture within a storage medium such as static random-access memory (SRAM), dynamic random-access memory (DRAM), or flash memory. SRAM is preferably used in the present invention due to its high speed, low latency, high reliability, low power consumption, and support for parallel access. By allocating first-level tables and second-level tables within this storage space based on a preset compression factor, storage resources can be fully utilized within a limited storage space, enabling efficient hash table construction to meet the data storage and processing requirements in different application scenarios. For example, all data corresponding to 4 bits require 16 subslots to cover the corresponding 16 types of data in a traditional single-table system. Each table entry stores 4 bits of data, that is, 16 types of data 0000, 0001, 0010, 0011, 0100, 0101, 0110, 0111, 1000, 1001, 1010, 1011, 1100, 1101, 1110, and 1111 are stored respectively, requiring a total of 64 bits of storage space. In the traditional two-level table structure, the second-level table usually still uses the same storage space size to store data, and the data of the corresponding 16 subslots need to be managed separately through the first-level table, which occupies more storage space. This embodiment, however, employs a two-level table structure for compressed data storage. For example, when the compression factor is 8, in the optimal case, only two subslots are needed in the first-level table to manage the second-level table. These correspond to two address fields in the second-level table, each storing eight 3-bit data items. This means that the second-level table only requires 48 bits of storage space. This achieves compression in both the storage depth of the first-level table and the storage width of the second-level table, effectively reducing storage space usage and further minimizing storage redundancy and resource consumption.
[0040] In this optional embodiment, the hash bucket is the core storage unit of the first-level table, and the number of hash buckets set determines the balance between conflict handling capability and storage efficiency. Each hash bucket is composed of multiple first subslots, and the number of first subslots is dynamically configured according to system resources and performance requirements. For example, each hash bucket is allocated 4 first subslots. The first subslot is subsequently used to store signature values generated based on key-value pairs to support fast search and access of the hash table. On this basis, the number of table entries in the first-level table is the product of the number of hash buckets and the number of first subslots. For example, the first-level table includes 32k hash buckets, and each hash bucket includes 4 first subslots. The number of table entries in the first-level table is 128k, which can be used to store and manage 128k signature values. In addition, each first subslot is configured with a hash bucket in the first-level table as its corresponding alternative bucket, which is used to store data migrated when the hash bucket where the first subslot is located overflows or conflicts, ensuring the efficient operation of the hash table.
[0041] In this optional embodiment, the second-level table stores data in the second subslots corresponding to different first subslots by setting up multiple address columns. Each address column includes multiple second subslots, which are subsequently used to store private keys and corresponding data items with the same signature value, as well as target bitmap marking rules for splitting keywords to obtain signature values and private keys and recombining signature values and private keys to obtain keywords, to support hash table conflict resolution and data storage. The number of address columns is equal to the number of first subslots, and the number of second subslots in each address column is a preset compression factor. For example, based on the previous example, the second-level table includes 128k address columns, each address column includes 8 second subslots, and a total of 1M table entries. The number of table entries in this second-level table represents the maximum number of key-value pairs that the hash table can support, meeting the requirement of the expected maximum number of table entries.
[0042] In this optional embodiment, the first subslots of the first-level table and the address columns of the second-level table are mapped one-to-one via address pointers. The address pointers correspond one-to-one with the addresses in the address columns, allowing access to the matching second subslots in the corresponding address columns through the first subslot. This address pointer can be implemented through hardware-level address management and address remapping, such as base address binding and offset calculation. By presetting different address pointers for different first subslots, fast indexing from the first-level table to the second-level table is achieved. This structure not only improves hash table access efficiency but also effectively reduces the risk of hash conflicts. When a new key-value pair needs to be inserted, the first signature value generated by the key-value pair can be used to quickly locate the corresponding first subslot. The corresponding address column in the second-level table can then be accessed using the pre-set address pointer. A search is then performed to determine whether a matching private key exists in the multiple second subslots of the corresponding address column. This hash table data processing method based on two-level table mapping not only enables flexible configuration and efficient utilization of storage space, but also improves the data processing capability and adaptability of hash tables, making it widely applicable in various data storage and processing scenarios.
[0043] Optionally, in one embodiment, a status flag is also preset in the first subslot, and the status flag is used to map the storage status of multiple second subslots in the address bar corresponding to the first subslot. After presetting different address pointers for different first subslots, it also includes: detecting the storage status of multiple second subslots in each address bar; based on the storage status, mapping the storage status of each second subslot to the corresponding status flag in the first subslot.
[0044] In this optional embodiment, the status flag includes multiple bits that can be mapped one-to-one to the storage status of the second subslot in the corresponding address bar. For example, a 1 indicates that the corresponding second subslot stores data, and a 0 indicates that the corresponding second subslot is vacant. For example, if a status flag is 00000111, it means that the last three second subslots in the corresponding address bar store data, while the other second subslots are vacant. The status flag is mapped to the storage status of the second subslot in the corresponding address bar in real time. When data in the second subslot is updated or deleted, the status flag can sense it in real time and make corresponding changes. For example, when data in a second subslot is deleted, the corresponding status flag changes from 1 to 0, indicating that the second subslot is now vacant. The corresponding storage status is then fed back to the first subslot to adjust the status flag. This real-time mapping and status update mechanism enables the hash table to dynamically manage storage space. When processing data within the hash table, the status flag can be used to determine whether the corresponding address column is completely empty, partially empty, or full. This allows for more accurate search for free subslots and selection of target subslots for data processing, ensuring the hash table's storage efficiency and data access speed. This dynamic management mechanism based on status flags further enhances the hash table's flexibility and adaptability, enabling it to better meet the data storage and processing requirements of diverse application scenarios. By enabling real-time updates to the status flags, the hash table can accurately locate the corresponding data location when deleting or updating data, avoiding invalid data access and wasted storage space.
[0045] For details of this embodiment, please refer to Figure 3 , Figure 3This is an example diagram of an embodiment of a hash table constructed in this embodiment, wherein the left side is an example of a first-level table, which includes 32k hash buckets from 0 to 32k-1, and each hash bucket includes 4 first subslots, such as hash bucket 0 includes four first subslots slot00, slot01, slot02 and slot03. There are 128k first subslots in the first-level table, that is, 128k table entries. Each first subslot is pre-stored with an address pointer ptr, such as the first subslot slot00 stores a pointer ptr0, the first subslot slot01 stores a pointer ptr2, and each A subslot also includes a status flag bit st to map the storage status of the second subslot in the address column of the corresponding second-level table. In the figure, the second-level table includes 128k address columns from 0 to 128k-1, and each address column includes 8 second subslots. For example, the address column 0 corresponding to the first subslot slot00 of hash bucket 0 includes 5 vacant second subslots and 3 second subslots with data stored. Therefore, the corresponding status flag bit st0 in slot00 is 00000111. The data stored in the second subslot includes the private key Skey, the preset bitmap marking rule BM and the data item Value introduced later, as shown in FIG. Figure 2 The five second subslots shown in the figure include (Skey0, BM0, Value0), (Skey1, BM1, Value1), (Skey2, BM2, Value2), (Skey3, BM3, Value3), and (Skey4, BM4, Value4), respectively. The data stored in the corresponding first subslot also includes the signature value sig introduced later, such as sig0, sig1, and sig2, and the corresponding status flags are st0, st1, st2, etc. It should be noted that the contents of some first subslots and second subslots in this figure are not shown. It is understandable that each first subslot not shown in the figure includes a corresponding address pointer, which may include a status flag and a signature value. The second subslot not shown includes a private key and data item corresponding to the signature value in the corresponding first subslot, and may also include a corresponding preset bitmap marking rule, which will not be repeated here.
[0046] This embodiment can also refer to Figure 4 , Figure 4 This is a schematic diagram of another embodiment of the hash table constructed in this embodiment. Figure 4The mapping relationship of the example first level table and the second level table is shown. In the first level table, a certain hash bucket includes four first sub-slots, slot0 (st0, sig0, pt0), slot1 (st1, sig1, pt1), slot2 (st2, sig2, pt2), and slot3 (st3, sig3, pt3). In the second level table, the address column corresponding to the first sub-slot slot0 is included, the first sub-slot is pointed to the address column by the address pointer ptr0, and the second sub-slot in the address column corresponds to the state flag bit st0 of the first sub-slot slot0 one by one. There are 8 groups of data (Skey0, BM0, Value0), (Skey1, BM1, Value1), (Skey2, BM2, Value2), (Skey3, BM3, Value3), (Skey4, BM4, Value4), (Skey5, BM5, Value5), (Skey6, BM6, Value6), and (Skey7, BM7, Value7), which correspond to the 0th bit to the 7th bit of the state flag bit one by one.
[0047] In the embodiment, the above-mentioned way of constructing the hash table can construct a compressed cuckoo hash table data structure, which realizes the compression of storage resources and the balance of conflict processing efficiency through a two-level table architecture, key-value split storage, and a dynamic elimination mechanism. Specifically, the data structure is composed of a first level table and a second level table. The first level table is used to store signature values and address pointers based on key-value pairs. The second level table stores data items of key-value pairs through an address column and private keys and corresponding bitmap marking rules based on key-value pairs. Under the data structure, the compression multiple of the first level table relative to the second level table is controlled by a preset compression factor. For example, when the preset compression factor is 8, the second level table of 1M table items is compressed to the first level table of 128k table items. Through the recombination of bit positions, the storage space occupation of the two-level table is reduced, and efficient use of storage resources is realized. The first level table and the second level table realize fast data insertion, searching, deletion, and modification through the mapping of address pointers and state flag bits, which is beneficial to improving data processing efficiency and can be widely applied to various data storage and processing scenarios to meet the data storage and processing needs in different application scenarios.
[0048] 102. Obtain a first key-value pair, the first key-value pair including a first key and a first data item;
[0049] In the embodiment, the data insertion processing can be implemented based on the hash table constructed above. Specifically, before the data insertion, the target data to be inserted into the hash table, i.e., the first key-value pair, needs to be obtained. The first key-value pair is composed of a first keyword and a first data item. The first keyword can be a specific string or number, which can be converted into a binary number with multiple bits. The first data item is the data content associated with the first keyword.
[0050] 103. combining the data of the preset bits in the first keyword to generate a first private key, and combining the data of the non-preset bits in the first keyword to generate a first signature value;
[0051] In the embodiment, the first private key can be obtained by extracting and combining the data of the preset bits in the first keyword, and the first signature value can be obtained by extracting and combining the data of the non-preset bits in the first keyword. The first signature value is used to quickly locate the corresponding storage location in the first-level table. For example, the first keyword is 11001010, and the preset bits are 1, 3, and 6, i.e., 01001010. In this case, the combination method can obtain the first private key 111 and the first signature value 10000. For another example, the preset bits are 1, 3, and 7, i.e., 10001010. In this case, the first private key and the first signature value obtained by the combination method are the same as in the previous case, i.e., the first private key 111 and the first signature value 10000. Unlike the traditional mask method, in the traditional mask method, the signature value retains the data of the non-preset bits and clears the other bits. The signature values obtained by the two preset bits are 10000000 and 01000000, respectively, which occupy more bit widths, and the extracted data with the same signature value is less, resulting in lower table utilization. In the present solution, the combination method is selected to reduce the bit width and the storage resource consumption. In the combination method, the bit width of the signature value is significantly smaller than that in the traditional mask method, so that one signature value corresponds to more private keys. The signature values of the same private keys and data items can be stored in the address bar corresponding to one first sub-slot, which can effectively improve the utilization of the hash table. After the preset bits are determined, the first signature value and the first private key can be reorganized according to the corresponding preset bits to uniquely determine the corresponding first keyword, and the corresponding first data item can be quickly located in the hash table, ensuring the accurate and reliable data processing of the hash table.
[0052] In this embodiment, the preset bits can be fixed rules or rules that can be directly obtained from relevant storage or records during data processing, or can be flexibly selected, such as randomly selecting from all bits of each keyword, or selecting a rule for determining the preset bits from multiple preset rules. Among them, when data security requirements are low or for data of a specific key-value pair set, such as the complete 16 table entries 0000-1111 in the previous example, when effective data processing can be performed, fixed rules can be used or preset bits can be obtained externally to further save storage space and simplify the data search process. When data security requirements are high or the key-value pair set changes dynamically, a method of determining the preset bits based on preset rules can be adopted, and the corresponding preset rules or preset bits can be stored in the second subslot. In the search and match process, matching verification of the corresponding preset rules or comparison of the preset bits can be added to further improve the accuracy and security of data processing, adapt to dynamically changing key-value pair sets, and be applicable to a wider range of data processing scenarios.
[0053] In this embodiment, in the former case, since the preset bits are known, when the preset bits are fixed, although different keywords may obtain the same signature value, the private keys corresponding to different keywords with the same signature value will still be different, thereby achieving unique matching of keywords. Similarly, in the case where the preset bits are different but the rules for the corresponding preset bits can be directly obtained, the keywords and private keys can be reorganized according to the corresponding rules of the preset bits to obtain uniquely matching keywords. Therefore, in the former case, you only need to enter the keyword when searching. In the latter case, when the preset bits are different, the keyword and private key may be the same during the insertion process. For example, when the first keyword is 101011 and the preset bits are 101010, and when the first keyword is 011011 and the preset bits are 011001, the private keys obtained are both 111 and the keywords are both 001. Since the preset bits are not fixed and cannot be directly obtained, it is necessary to store the corresponding preset bits or the rules used to obtain the preset bits in the second-level table. In subsequent searches, it is necessary to input the keyword and the corresponding preset bits or the rules used to obtain the preset bits at the same time. When it is determined that the input preset bits are consistent with the stored preset bits and the signature value and private key match, the corresponding first keyword is obtained using the preset bits, signature value, and private key, thereby obtaining the corresponding first data item. In this way, the risk of hash collisions can be effectively reduced and the accuracy of the search process can be improved, ensuring the security and reliability of data.
[0054] 104. Select a target subslot from the idle first subslot;
[0055] 105. Store the first signature value in the target subslot, and store the first private key and the first data item in the second subslot of the address column corresponding to the target subslot.
[0056] In this embodiment, after obtaining the first signature value and the first private key, the first signature value needs to be stored in the first subslot of the target in the first-level table, and the first private key needs to be stored in the corresponding second subslot of the second-level table. In this solution, the first subslot of the target, i.e., the target subslot, is selected from an idle first subslot. The idle first subslot refers to a first subslot that currently does not store any signature value, or a first subslot that has an empty second subslot in the address column of the corresponding second-level table. The idle first subslot can be determined specifically by traversal or by a specific matching mechanism. During the traversal or matching process, the status flag can be used to more accurately determine whether the corresponding first subslot is idle, thereby achieving rapid positioning. As long as an idle first subslot capable of storing the first signature value can be finally selected, it will be sufficient.
[0057] In this embodiment, the target subslot is determined in the free subslot according to the specific hash conflict resolution strategy and data storage and processing requirements. For example, in some cases, the first free subslot found may be selected as the target subslot to ensure efficient processing of data insertion, while in other cases, the target subslot may be strategically selected based on the position, access frequency or depth of the subslot to ensure effective resolution of the hash conflict problem. In this hash table, in order to improve the utilization and search efficiency of each first subslot, it is usually necessary to select the target subslot from the free first subslot in the hash bucket corresponding to the signature value. If there is no free subslot in the corresponding hash bucket, the elements in the subslot can be moved to the alternative bucket with a free first subslot according to the relationship between each first subslot and the alternative bucket, freeing up the first subslot in the corresponding hash bucket and using the freed subslot as the target subslot, thereby realizing the reuse of the first subslot and avoiding the low storage and access efficiency caused by hash conflicts.
[0058] In this embodiment, after determining the target subslot, the first signature value is stored in the preset signature value field of the target subslot. The first private key and the first data item are then stored in the appropriate second subslot in the address bar according to the guidance of the address pointer. In this address bar, an empty slot search is typically performed from the end to the beginning to ensure the orderliness of subsequent insertions and queries, reducing potential hash conflicts and data movement. Furthermore, since the aforementioned determination of an empty first subslot takes into account the availability of the second subslot in the address bar corresponding to the first subslot based on the status flag, the insertion of data into the second subslot is generally prevented from being occupied, ensuring the accuracy and efficiency of the data insertion process in the hash table. After data insertion is completed, the hash table can check and update the storage status of each second subslot in real time based on the status flag, ensuring that the hash table's status flag accurately reflects the storage status of the second subslot in the corresponding address bar, thereby ensuring the effectiveness and reliability of data processing.
[0059] Optionally, in some embodiments, before combining the data of preset bits in the first keyword to generate a first private key, and combining the data of non-preset bits in the first keyword to generate a first signature value, it also includes: obtaining a bitmap marking rule list; wherein the bitmap marking rule list includes multiple preset bitmap marking rules, and different preset bitmap marking rules are used to mark different preset bits in the keyword; performing simulation tests on each preset bitmap marking rule in the bitmap marking rule list to screen out a target bitmap marking rule that meets the preset performance conditions; and determining the preset bit in the first keyword based on the target bitmap marking rule.
[0060] In this embodiment, a manner is provided for determining preset bit positions by using a preset bitmap marking rule to split and reorganize a first keyword. The bitmap marking rule list is a set containing a plurality of preset bitmap marking rules, which define how to extract specific bit positions from a keyword to generate a private key and a signature value. By simulating and testing these rules, their performance in actual application, such as collision rate, storage efficiency, etc., can be evaluated, so as to filter out the optimal rule that meets the preset performance condition, i.e., the target bitmap marking rule. The simulation test method can be an exhaustive method, i.e., trying all possible bitmap marking rule combinations, or a heuristic search algorithm, such as genetic algorithm, simulated annealing algorithm, etc., to find an approximate optimal solution within a reasonable calculation time. The preset performance condition can be that the collision rate of the hash table is lower than a certain threshold, or the storage efficiency is higher than a certain standard, or the data access speed meets a certain requirement, etc., which can be flexibly set according to actual needs in specific implementation. Finally, the optimal target bitmap marking rule can be selected from the preset bitmap marking rules that meet the condition, i.e., based on the rule, it can be determined which bit positions in the first keyword are preset bit positions for generating a private key and which bit positions are non-preset bit positions for generating a signature value, so as to achieve accurate split and reorganization of the first keyword, which is beneficial to subsequent fast construction of the hash table and efficient data processing.
[0061] The optional embodiment can refer to Figure 5 The bitmap marking rule list includes 64 preset bitmap marking rules, i.e., BM1, BM2, BM3 to BM64, and the selected bitmap marking rule, i.e., the target bitmap marking rule, is BM3, which corresponds to the preset marking positions [3, 6, 9], indicating that the 3rd, 6th and 9th bit positions in the keyword are marked and extracted as a private key. The first keyword includes 96 bit positions, i.e., 0-95, and the first data item includes 24 items, i.e., 0-23, so that by combination processing, the data in the first signature value is 0-2, 4-5, 7-8 and 10-95 bit positions of the first keyword, and the data in the first private key is the 3rd, 6th and 9th bit positions of the first keyword. The first keyword will be written into the signature value field of the target sub-slot, and the first private key and the first data item will be stored in the second sub-slot of the address column corresponding to the address pointer of the target sub-slot in the second-level table.
[0062] Optionally, in some embodiments, after selecting the target sub-slot from the idle first sub-slot, the method further includes: storing the target bitmap marking rule to the second sub-slot of the address column corresponding to the address pointer of the target sub-slot.
[0063] After determining the target bitmap marking rule through the above steps, this optional embodiment can effectively identify the combination rule of the first private key and the first data item by storing the target bitmap marking rule in the second subslot of the address bar corresponding to the address pointer of the target subslot. In the matching process of the second-level table, it can also verify whether the bitmap marking rule matches, avoiding hash conflicts caused by incorrect matching, and avoiding the situation where the keywords with the same signature value and private key obtained through different preset bits are judged as the same keyword, further improving the accuracy and security of data access.
[0064] Optional, see Figure 6 In some embodiments, the types of the idle first subslots include a first idle subslot, a second idle subslot, and a third idle subslot. Before selecting a target subslot from the idle first subslots, the process includes:
[0065] 1031. Calculate a first primary hash bucket address and a first secondary hash bucket address based on the first signature value to determine the primary hash bucket and the secondary hash bucket corresponding to the first signature value in the hash buckets of the first-level table.
[0066] 1032. Determine whether there is a first free subslot in the first subslots in the primary hash bucket and the secondary hash bucket; wherein the status flag of the first free subslot indicates that all second subslots in the corresponding address column are free, or the signature value of the first free subslot matches the first signature value, and the status flag of the first free subslot indicates that there is at least one free second subslot in the corresponding address column;
[0067] 1033. If there is no first free subslot in the first subslots of the primary hash bucket and the secondary hash bucket, determine whether there is a second free subslot in the first subslot of each primary candidate bucket; wherein the primary candidate bucket is the candidate bucket corresponding to each first subslot of the primary hash bucket, and the status flag of the second free subslot indicates that the second subslot of the corresponding address column is vacant;
[0068] 1034. If there is no second free subslot in the first subslot of each primary alternative bucket, the third free subslot is searched for in the first subslot of the secondary alternative bucket of the secondary hash bucket step by step; wherein, the secondary alternative bucket is the alternative bucket corresponding to each first subslot of the secondary hash bucket, and the status flag of the third free subslot indicates that the second subslot of the corresponding address bar is vacant.
[0069] In this embodiment, the hash bucket includes a main hash bucket and a secondary hash bucket. The main hash bucket and the secondary hash bucket have different hash bucket addresses. The address of the main hash bucket refers to the hash bucket directly calculated based on the first signature value, and the address of the secondary hash bucket is the hash bucket used to store overflow data when the first sub-slot in the main hash bucket cannot meet the storage requirements. The hash bucket address can be calculated by the main hash function and the secondary hash function. The hash function can be CRC (Cyclic Redundancy Check, cyclic redundancy check), Toeplitz Hash (Toeplitz Hash algorithm) or MD5 algorithm (Message Digest Algorithm 5, Information Digest Algorithm 5), etc. This application does not impose specific restrictions on this. At the same time, each hash bucket includes multiple first sub-slots, and each first sub-slot is preset with a status flag bit for identifying the storage status of the first sub-slot and the storage status of the second sub-slot of its corresponding address bar. In addition, each first sub-slot also corresponds to an alternative bucket, which is used to provide additional storage space when the first sub-slot of the main hash bucket or the secondary hash bucket cannot meet the storage requirements.
[0070] In this optional embodiment, before selecting the target sub-slot, the corresponding primary hash bucket address and secondary hash bucket address are first calculated based on the first signature value, thereby determining the primary hash bucket and the secondary hash bucket. Then, it is determined whether there is an idle sub-slot that meets the conditions in the first sub-slots in the primary hash bucket and the secondary hash bucket, that is, the first idle sub-slot. Among them, the status flag of the first idle sub-slot needs to indicate that the second sub-slots of the corresponding address column are all vacant, or the signature value of the first idle sub-slot matches the first signature value, and its status flag indicates that there is at least one vacant second sub-slot in the corresponding address column. For example, if the primary hash bucket or the secondary hash bucket contains the first subslots slot00, slot01, slot02, and slot03, where the status flag bits corresponding to slot00 are all 0, slot01 stores a signature value that matches the first signature value, and the corresponding status flag bits have bits that are 0, the status flag bits of slot02 are not all 0, and its signature value does not match the first signature value, and the status flag bits of slot03 are all 1, then slot00 and slot01 are both the first free subslots, while slot02 and slot03 are not the first free subslots. The existence of the first free subslot indicates that the data item corresponding to the first signature value has not yet been stored in the primary hash bucket or the secondary hash bucket and there is an empty first subslot, or that a data item with the same first signature value has been stored, but the second subslot of its corresponding address bar is still empty, so there is a condition to store the corresponding data in the primary hash bucket or the secondary hash bucket.
[0071] In this optional embodiment, if the first free subslot does not exist in both the main hash bucket and the secondary hash bucket, it means that the corresponding main hash bucket and secondary hash bucket cannot provide insertion for the current first key-value pair, that is, both store data corresponding to other signature values that do not match the first signature value, or store data that matches the first signature value but the corresponding second subslot is full. At this time, a search will be conducted to determine whether there is a free subslot, that is, a second free subslot, in the alternative bucket of the main hash bucket, that is, the main alternative bucket. Among them, the status flag of the second free subslot needs to indicate that the second subslot of its corresponding address bar is vacant to ensure that there is enough space to store the first private key and the first data item. During the search process for the second free subslot, the main alternative bucket corresponding to the main hash bucket will be traversed to determine whether there is a first subslot in each main alternative bucket that satisfies the requirement that the corresponding second subslot is vacant. For example, a primary candidate bucket contains first subslots slot 10 and slot 11. Slot 10's status flag is all 0, while slot 11's status flag is not all 0. In this case, slot 10 is the second free subslot and can be selected as a candidate for the target subslot. The primary candidate bucket is searched before the secondary candidate bucket because the primary candidate bucket typically takes priority, while the secondary candidate bucket is used only when the primary hash bucket, secondary hash bucket, and primary candidate bucket all fail to meet storage requirements. Therefore, prioritizing the primary candidate bucket allows for more efficient use of storage space and reduces the number of data transfers.
[0072] In this optional embodiment, if the second free sub-slot does not exist in the main alternative bucket, the alternative bucket of the secondary hash bucket is further searched, that is, whether there is a free sub-slot in the secondary alternative bucket, that is, the third free sub-slot. Similarly, the status flag of the third free sub-slot also needs to indicate that the second sub-slots of its corresponding address bar are all vacant. During the search process of the secondary alternative bucket, all the secondary alternative buckets corresponding to the first sub-slots in the secondary hash bucket will be traversed, and whether the first sub-slots in these secondary alternative buckets meet the conditions will be judged one by one. If the third free sub-slot that meets the conditions is finally found in a certain secondary alternative bucket, it can be selected as the alternative object of the target sub-slot. This step-by-step search method can ensure that the free sub-slots that meet the conditions are found in the smallest possible range, thereby improving the storage efficiency and access speed of the hash table.
[0073] In this optional embodiment, if no free subslot that meets the conditions can be found in the main hash bucket, the secondary hash bucket, the main alternative bucket, and the secondary alternative bucket, a free subslot will be further searched in the deeper secondary alternative bucket of the secondary alternative bucket. During the search, it will be determined whether there is a free first subslot in the alternative bucket corresponding to the specific first subslot in the secondary alternative bucket, wherein the specific first subslot refers to the first subslot at a specific position or in a specific order in the secondary alternative bucket, which can be the first subslot determined by the target position and obtained by cyclic shift transformation, or a single first subslot selected according to relevant records or shift rules. By traversing the alternative buckets corresponding to the specific first subslot rather than all the first subslots, the number of alternative buckets and subslots accessed and processed can be reduced, the consumption of computing resources can be reduced, and the search for the third free subslot can be made more orderly, reducing the possibility of hash conflicts and data movement, while ensuring the integrity of the hash table data, maximizing the storage and access efficiency. If the candidate bucket corresponding to the secondary candidate bucket also does not have an idle subslot, the idle subslot will be further searched in the deeper secondary candidate buckets step by step until an idle subslot that meets the conditions is found or the preset restriction condition is reached.
[0074] Optional, see Figure 7 In some embodiments, searching for a third free subslot in the first subslot of the secondary candidate bucket of the secondary hash bucket step by step includes:
[0075] 1061. Determine whether there is a third idle subslot in the first subslot of the secondary candidate bucket;
[0076] 1062. If there is no third free subslot in the secondary candidate bucket, the first subslot at the target position in the secondary hash bucket is used as the eliminated subslot;
[0077] 1063. Circularly shift each first subslot in the secondary hash bucket;
[0078] 1064. Replace the secondary hash bucket with the secondary candidate bucket corresponding to the eliminated subslot, and replace the secondary candidate bucket with the candidate bucket corresponding to each first subslot in the replaced secondary hash bucket, so as to determine whether there is a third free subslot in the first subslot of each replaced secondary candidate bucket, until it is confirmed that there is a third free subslot in the first subslot of the secondary candidate bucket.
[0079] In the optional embodiment, a step-by-step lookup method for the third idle sub-slot is provided. Specifically, it is determined whether the third idle sub-slot exists in the first sub-slot of the secondary candidate bucket. If the third idle sub-slot exists in the first sub-slot of the secondary candidate bucket, a rejection sub-slot is determined in the first sub-slot of the secondary hash bucket, which is used to move the rejection sub-slot to the idle sub-slot of the corresponding candidate bucket in the subsequent process, so as to free the first sub-slot for storing data. The rejection sub-slot is selected from a target position, such as the head, tail or middle fixed position in the first sub-slot. The target position is the same position in each hash bucket, so as to ensure that the step-by-step lookup process can be orderly and efficiently performed. The tail of the hash bucket is preferentially selected as the target position. Since the first sub-slot of the tail usually has a higher priority in the insertion process, the lookup and insertion of the empty slot are performed from the tail to the head. Therefore, selecting the tail of the hash bucket as the target position helps to reduce the number of data movements and improve the insertion efficiency of the hash table.
[0080] In the optional embodiment, after the rejection sub-slot is selected, a cyclic shift operation is performed on each first sub-slot in the secondary hash bucket, that is, the contents in the first sub-slot are sequentially moved, so that the contents originally in the rejection sub-slot position are moved to the next order of the hash bucket, and the rejection sub-slot position, that is, the target position, is replaced by other first sub-slots. In the subsequent lookup process, when the same hash bucket is reached again, the first sub-slots in the secondary hash bucket can be sequentially rejected and cyclically shifted based on the ordered cyclic shift, so as to ensure that each first sub-slot has the opportunity to be checked and updated. For example, the first sub-slots in the original secondary hash bucket from the tail to the head are slot00, slot01, slot02 and slot03, wherein slot00 is in the target position and is selected as the rejection sub-slot. After the cyclic shift, the first sub-slots from the tail to the head are slot01, slot02, slot03 and slot00, that is, the data in the rejection sub-slot is moved to the head of the secondary hash bucket, and the data of slot01 after the cyclic shift is in the target position, that is, slot00.
[0081] In this embodiment, after the circular shift, in order to be able to effectively access the deeper alternative bucket of the secondary hash alternative bucket, the secondary alternative bucket corresponding to the eliminated sub-slot will be used to replace the current secondary hash bucket, and the alternative buckets corresponding to the first sub-slots in the replaced secondary hash bucket will be used as new secondary alternative buckets for subsequent judgment of whether there is a third free sub-slot in the first sub-slot of these new secondary alternative buckets. This replacement operation can ensure that when the third free sub-slot is continued to be searched step by step in the secondary hash bucket and its secondary alternative buckets, the search can be performed based on the order of the first sub-slot that has been circularly shifted, thereby avoiding repeated searches or missed searches. At the same time, since the eliminated sub-slot has been moved to a new position, it can also be performed again in the same hash bucket based on the new eliminated sub-slot in the subsequent search process to further update the order of the first sub-slots in the secondary hash bucket, thereby improving the storage efficiency and access speed of the hash table. After replacing the secondary hash bucket and its secondary candidate buckets, the system checks again to see if a third free subslot is found in the replaced secondary candidate buckets. If not, the system repeats the above level-by-level search process until a matching third free subslot is found in a secondary candidate bucket or the preset limit is met. This level-by-level search and circular shifting ensures that matching free subslots are found within the smallest possible range, while maintaining hash table storage efficiency and data consistency.
[0082] This optional embodiment implements a mechanism for detecting empty slots through the above steps. By searching the primary hash bucket, secondary hash bucket, primary alternative bucket, secondary alternative bucket, and multiple levels of alternative buckets for the secondary alternative bucket and judging the status flag, the free subslots in the hash table are searched, effectively improving the storage efficiency and data access speed of the hash table. If a free subslot that meets the conditions is finally found successfully, the target subslot will be selected based on the corresponding first free subslot, second free subslot, or third free subslot. If a free subslot that meets the conditions cannot be found, the hash table expansion operation can be optionally set according to the expansion strategy of the hash table to provide more storage space. This application does not impose specific restrictions on the specific strategy of expansion. Through the above-mentioned step of searching for the free first subslot, the target subslot that meets the storage requirements can be effectively searched and determined in the primary hash bucket, secondary hash bucket, and their corresponding alternative buckets. This process makes full use of the information of the status flag. The logical bucket-by-bucket search method in this way not only ensures storage efficiency, but also ensures the accuracy and consistency of the data. At the same time, this optional embodiment also implements a principle of element removal priority, by preferentially removing the first subslot at the end, to ensure that the number of data moves can be reduced as much as possible when data is inserted, thereby improving the insertion efficiency of the hash table.
[0083] Optionally, in some embodiments, the data processing method of the hash table also includes: recording the correspondence between the eliminated sub-slots and each sub-alternative bucket and the corresponding eliminated sub-slot in sequence, and counting the number of eliminated sub-slots; if the number of eliminated sub-slots is greater than the preset elimination threshold, or the first sub-slot at the target position in the current sub-hash bucket has been recorded as an eliminated sub-slot, then reporting the insertion failure information.
[0084] In this optional embodiment, two methods are provided to terminate the search for free subslots. These methods can be used only during the search process for the third free subslot, ensuring that the corresponding elimination, circular shift, and search operations are carried out in an orderly and traceable manner. The eliminated subslots can be recorded using RAM (Random Access Memory) or other storage devices. The recording process uses a sequential recording method, i.e., the eliminated subslots are recorded in an orderly manner according to the selection order or elimination timestamp order of the eliminated subslots. This ensures that the redundant description and management of the eliminated subslots in the hash table are clearer and more accurate. Furthermore, the recording of the corresponding relationship between each secondary candidate bucket and the corresponding eliminated subslot also facilitates the accurate determination of the corresponding relationship between each level of secondary candidate buckets and the eliminated subslots in the corresponding secondary hash bucket during the subsequent move and insert process. This allows the elements in the corresponding eliminated subslot to be accurately and efficiently moved to the vacant first subslot of the corresponding secondary candidate bucket after confirming the eliminated subslot of the hash bucket of the previous level of the secondary candidate bucket.
[0085] In this optional embodiment, based on the record of eliminated subslots, it is also convenient to count the number of eliminated subslots. If the number of eliminated subslots exceeds the preset elimination threshold, it means that there may not be enough free space in the hash table to store new data items, or the storage structure of the hash table has become too complex, resulting in a significant reduction in search efficiency. The subsequent insertion of eliminated subslot elements also requires more complex operations, which is too complex and difficult to meet the requirements of efficient storage. If the first subslot at the target position in the current secondary hash bucket has been recorded as a eliminated subslot, that is, all the first subslots in the current hash bucket are the first subslots that previously existed at the target position, and all have been selected and recorded as eliminated subslots, and their corresponding alternative buckets have been traversed, it means that each first subslot in the current hash bucket has been correspondingly eliminated, and one cycle has been eliminated. Further elimination will only repeat the previous elimination process. At this time, the hash table may be highly saturated, or there may be problems such as a broken circular linked list or an error in the storage status flag. Continuing to insert new data items may cause the performance of the hash table to seriously degrade, and even lead to the risk of data conflict and data loss. Therefore, to avoid the above situation, when it is detected that the number of discarded subslots exceeds the threshold or that all the first subslots to be discarded are already recorded discarded subslots, an insertion failure message will be promptly reported to the user or related application to inform them that the current storage state of the hash table can no longer meet the new storage requirements. This prompts the user or administrator to take appropriate measures, such as checking the storage state of the hash table, expanding the hash table, or rebuilding the hash table, to ensure the normal operation of the hash table and data consistency. In this way, problems such as data storage failure or data loss caused by abnormal hash table storage structure can be effectively avoided, thereby improving the reliability and stability of the hash table.
[0086] Optional, see Figure 8 In some embodiments, after confirming the free subslots in the above manner, selecting a target subslot from the free first subslot includes:
[0087] 1041. Determine the type of the idle first subslot.
[0088] 1042. If the idle first subslot is the first idle subslot, select a first target hash bucket from the primary hash bucket and the secondary hash bucket based on the first priority, and select the first idle subslot closest to the tail in the first target hash bucket as the target subslot;
[0089] 1043. If the free first subslot is the second free subslot, select a second target hash bucket from the primary candidate bucket based on the second priority, move the element of the second target hash bucket corresponding to the first subslot in the primary hash bucket to the second free subslot closest to the end of the second target hash bucket, set the first subslot in the primary hash bucket corresponding to the second target hash bucket as the target subslot, and set the address pointer of the second free subslot closest to the end of the second target hash bucket before the move as the address pointer of the target subslot.
[0090] 1044. If the free first subslot is the third free subslot, a third target hash bucket is selected from the secondary candidate bucket based on the third priority, and the elements in the culled subslot corresponding to the third target hash bucket in the secondary hash bucket are moved to the third free subslot closest to the tail of the third target hash bucket, and the secondary candidate bucket is replaced with the secondary hash bucket, and the hash bucket where the culled subslot corresponding to the replaced secondary candidate bucket is located is replaced with the secondary hash bucket, so that the elements in the culled subslot in the replaced secondary hash bucket are moved to the first subslot that is vacant after the secondary candidate bucket is moved, until the elements in the culled subslot in the most basic secondary hash bucket are moved to the first subslot that is vacant after the corresponding secondary candidate bucket is moved, and the culled subslot in the most basic secondary hash bucket is used as the target subslot, and the address pointer of the third free subslot closest to the tail of the third target hash bucket before the move is used as the address pointer of the target subslot.
[0091] In this optional embodiment, the types of the idle first subslots are the aforementioned first idle subslots, second idle subslots, and third idle subslots, and different types of idle subslots represent different storage locations and storage states. During specific implementation, a judgment is made based on the type of the idle subslots, and different types of idle subslots correspond to different selection strategies. Based on information such as the position of the hash bucket in which these idle first subslots are located, the number of these idle first subslots in the hash bucket, and the number of these idle first subslots in the secondary candidate buckets corresponding to each first subslot of the hash bucket, the optimal target subslot suitable for different types of idle subslots can be selected to achieve efficient storage and access of data.
[0092] In this optional embodiment, the first, second, and third priorities are used to select a unique target hash bucket from a hash bucket with a corresponding first, second, or third free subslot. This allows the target subslot to be selected from the free subslots of the target hash bucket, or to store subslot elements moved from the previous hash bucket, thereby freeing up space for newly inserted data in the primary or secondary hash bucket. The first, second, and third priorities can be flexibly configured based on requirements such as the hash bucket's storage status, storage efficiency, and data access speed, so that the selected target hash bucket can better meet the storage and access requirements of the hash table. For example, a hash bucket with a relatively free storage status, high storage efficiency, and fast data access speed can be preferentially selected as the target hash bucket. In a specific implementation, a comprehensive evaluation can be performed based on information such as the number of free subslots in the hash bucket, the number of stored data items, and the data access frequency to determine the optimal target hash bucket. In this way, the storage and access requirements of the hash table can be fully considered when selecting the target subslot, thereby achieving efficient data storage and access.
[0093] In this optional embodiment, if the idle first subslot is the first idle subslot, that is, the first idle subslot is located in the main hash bucket and / or the secondary hash bucket, then the first target hash bucket will be selected from the main hash bucket and the secondary hash bucket based on the first priority, and the corresponding target subslot will be further determined. Optionally, in some embodiments, the first priority is used to preferentially select the hash bucket whose signature value of the existing first idle subslot matches the first signature value. If the signature values corresponding to the main hash bucket and the secondary hash bucket are both matched or both mismatched, the hash bucket with the largest total number of vacancies in the second subslots corresponding to each first idle subslot is preferentially selected. If the main hash bucket and the secondary hash bucket are both matched or both mismatched, and the total number of vacancies in the corresponding second subslots is the same, the main hash bucket is preferentially selected. The first priority will first consider the matching of the signature values, that is, determine whether the signature value of the hash bucket where the idle subslot is located matches the signature value of the data item to be stored. If there are free subslots with matching signature values in the primary hash bucket and the secondary hash bucket, these matching hash buckets will be selected first, because the matching signature values usually means that the storage location of the data item in the hash table is more closely related to its logic, which is conducive to improving the speed and efficiency of data access. If the signature values of the primary hash bucket and the secondary hash bucket are both matched or both mismatched, other factors need to be further considered to determine the target hash bucket. Specifically, the hash buckets with the largest total number of vacant second subslots will be selected first. The total number of vacant second subslots can be determined based on the status flag bits of each first subslot in the hash bucket. The more vacant second subslots there are, the higher the storage efficiency and flexibility of the hash table will be when more data items need to be stored later, and the better it can meet storage needs. If the total number of vacant second subslots of the primary hash bucket and the secondary hash bucket is the same, the primary hash bucket will be selected as the target hash bucket based on the primary and secondary relationship of the hash buckets.
[0094] In this optional embodiment, it should be noted that the target hash bucket is selected based on the free subslots. For example, if one of the primary hash bucket and the secondary hash bucket does not include the first free subslot, and the other includes the first free subslot, the hash bucket including the first free subslot can be directly selected. According to the first priority mentioned above, the first priority of the hash bucket including the first free subslot is also higher than that of the hash bucket not including the first free subslot, and the target subslot can be directly selected from the corresponding hash bucket. The selection of the target subslot follows a tail-first principle, and the first free subslot closest to the tail is preferentially selected as the target subslot, ensuring that when data is inserted, the number of data moves can be reduced as much as possible, thereby improving the insertion efficiency of the hash table. This method also causes the data items from the tail to the head in the hash bucket to be arranged from old to new, which is also conducive to improving the efficiency of data access, making the corresponding search, deletion, and modification operations more efficient and convenient.
[0095] In this optional embodiment, if the free first subslot is a second free subslot, that is, the free subslot is located in the primary candidate bucket, then a second target hash bucket will be selected from each primary candidate bucket corresponding to each first subslot of the primary hash bucket according to the second priority, and the corresponding target subslot and the address pointer of the target subslot will be further determined. Optionally, in some embodiments, the second priority is used to preferentially select the primary candidate bucket with the largest number of second free subslots. If there are multiple primary candidate buckets with the same maximum number of second free subslots, the primary candidate bucket whose first subslot in the corresponding primary hash bucket is closest to the end of the primary hash bucket will be preferentially selected. This second priority is primarily determined based on the number of second free subslots. A greater number of second free subslots means that the primary candidate bucket can provide more storage space, which is more conducive to meeting subsequent storage needs. Therefore, the primary candidate bucket with the largest number of second free subslots is preferentially selected. If there are multiple primary candidate buckets with the same number of second free subslots, other factors need to be further considered to determine the target hash bucket. Specifically, the primary candidate buckets whose first subslots in the corresponding primary hash buckets are closest to the tail of the primary hash bucket will be preferentially selected. For example, if two primary candidate buckets correspond to the first subslots slot00 and slot01 in the primary hash bucket, and slot00 is closer to the tail of the primary hash bucket, the second target hash bucket corresponding to slot00 will be selected. The free subslots in the second target hash bucket selected in this way can minimize the number of data movements when inserting new data items, thereby improving the insertion efficiency of the hash table.
[0096] In this optional embodiment, after determining the second target hash bucket, it is necessary to move the elements of the first subslot in the main hash bucket corresponding to the hash bucket to the second free subslot closest to the tail of the second target hash bucket, and use the corresponding first subslot in the main hash bucket as the target subslot for data storage. During the process of this move, the integrity and consistency of the data will be ensured to avoid data loss or conflict. At the same time, the move operation will also update the corresponding address pointer, and use the second free subslot closest to the tail of the second target hash bucket as the address pointer of the target subslot, so that the newly inserted element uses the new pointer, and the address pointer is moved together with the corresponding element during the move, so that the signature value and the address column of the second-level table corresponding to the address pointer always correspond, ensuring that in the subsequent data access process, the required data can be accurately located, thereby improving the speed and accuracy of data access.
[0097] In the optional embodiment, if the idle first sub-slot is a third idle first sub-slot, i.e., the idle first sub-slot is located in a secondary candidate bucket or a deeper level candidate bucket, at this time, a third target hash bucket is selected from each secondary candidate bucket corresponding to each first sub-slot of the secondary hash bucket based on a third priority, and the corresponding target sub-slot and the address pointer of the target sub-slot are further determined. Optionally, in some embodiments, the third priority is used to preferentially select the secondary candidate bucket with the largest number of third idle first sub-slots, and if there are multiple secondary candidate buckets with the same largest number of third idle first sub-slots, the secondary candidate bucket with the first sub-slot position closest to the tail of the corresponding secondary hash bucket is preferentially selected. The third priority is mainly determined based on the number of third idle first sub-slots. The larger the number of third idle first sub-slots, the larger the storage space that the secondary candidate bucket can provide, and the more conducive to meeting subsequent storage requirements, so the secondary candidate bucket with the largest number of third idle first sub-slots is preferentially selected. If there are multiple secondary candidate buckets with the same number of third idle first sub-slots, other factors need to be further considered to determine the target hash bucket. Specifically, the secondary candidate buckets with the first sub-slot position closest to the tail of the corresponding secondary hash bucket are preferentially selected to ensure that the number of data moves is reduced as much as possible when inserting new data items, thereby improving the insertion efficiency of the hash table.
[0098] In the optional embodiment, after determining the third target hash bucket, a series of complex moving operations need to be performed, i.e., moving the elements in the excluded sub-slot of the hash bucket corresponding to the secondary hash bucket to the third idle first sub-slot closest to the tail of the third target hash bucket, and updating the correspondence of the hash bucket, replacing the secondary hash bucket with the replaced secondary candidate bucket, and replacing the secondary hash bucket with the hash bucket corresponding to the excluded sub-slot of the replaced secondary candidate bucket, so as to continue moving the elements in the excluded sub-slot of the replaced secondary hash bucket to the first sub-slot emptied after moving in the secondary candidate bucket of the next level, until all the elements in the excluded sub-slot of the initial secondary hash bucket are moved to the first sub-slot emptied after moving in the corresponding secondary candidate bucket, so that the excluded sub-slot in the initial secondary hash bucket is emptied and used as a target sub-slot for data storage.
[0099] In this optional embodiment, the most basic sub-hash bucket refers to the sub-hash bucket determined by hashing the first signature value during the corresponding data insertion process. In the process of moving the eliminated elements, the most basic sub-hash bucket and the hash bucket where the elimination sub-slot of the upper level corresponding to the sub-hash bucket where the vacant first sub-slot is located are determined by the elimination sub-slots recorded in sequence previously and the correspondence between each sub-alternative bucket and the corresponding elimination sub-slot, thereby ensuring that the storage and movement of data in the hash table can be carried out in an orderly manner. The movement operation will also ensure the integrity and consistency of the data and avoid data loss or conflict problems. During the movement process, the corresponding address pointer will also be updated so that the newly inserted element adopts the new pointer, and the address pointer will be moved together with the corresponding element during the movement process, ensuring that the required data can be accurately located in the subsequent data access process, thereby improving the speed and accuracy of data access. In addition, the move operation also takes into account the hierarchical structure of the hash table. By moving and replacing layer by layer, it ensures that the hash table can maintain data integrity and consistency during data processing, avoiding data loss or access anomalies caused by changes in the hash table structure. At the same time, during the insertion process, insertion is performed from a shallower hash bucket. By giving priority to insertion and removal from the same target position such as the tail, as well as circular shift processing and removing sub-slots and moving them step by step to the first vacant sub-slot of the next level alternative bucket, the newly inserted data can always be in the main hash bucket or the secondary hash bucket, and the older data will be in the deeper alternative bucket, so that the newly inserted data can be accessed faster, the storage space of the hash table can be effectively utilized, and the data access efficiency can be improved.
[0100] In this optional embodiment, a new element priority mechanism is implemented through the above-mentioned target subslot selection method. When selecting the target subslot, not only the storage location and storage status of the free subslot are considered, but also the signature value matching of each subslot in the hash table, the correspondence between the first subslot of the hash bucket and the alternative bucket, the status flag information and the priority principle of element selection, such as the tail priority principle, are combined. This ensures that when data is inserted into the hash table, multiple factors can be comprehensively considered, and the hierarchical structure of the first-level table and the correspondence between the two-level tables are utilized to achieve efficient storage and access to hash bucket data, improve the table utilization of the hash table, maintain the integrity and consistency of the data in the hash table, and make the hash table more efficient and stable during data processing.
[0101] In this optional embodiment, the above process can be referred to Figure 9 , Figure 9An example of a data insertion method for processing data in the above-mentioned hash table is provided. In this illustration, the primary hash bucket is hash bucket 2, and the secondary hash bucket is hash bucket 0. These two hash buckets and hash bucket 3 are all full. For details, please refer to the first subslot number and the stored status flag, signature value, and address pointer shown in the figure. This application will not elaborate on this. Hash bucket 1 has a first subslot slot 10 with data stored, and three idle first subslots slot 11, slot 12, and slot 13. These idle first subslots are each preset with address pointers ptr15, ptr16, and ptr17, respectively. At this time, it is necessary to insert the signature value sig15 into the hash table. If the primary candidate bucket corresponding to the primary hash bucket 2 is a plurality of hash buckets including hash bucket 1, after determining that the aforementioned second free subslot exists in these primary candidate buckets, and hash bucket 1 is determined as the second target candidate bucket according to the second priority as mentioned above, if it is determined that the first subslot in the primary hash bucket corresponding to hash bucket 1 is slot 20, then the element in the first subslot slot 20 can be moved to the second free subslot closest to the tail of hash bucket 1 according to the mechanism of searching from the tail to the head. In the idle sub-slot slot11, the element in slot11 is the element in the original slot20, namely (st5, sig5, ptr5), and slot20, as the target sub-slot, will use the address pointer ptr15 of slot11 as its address pointer. After the signature value sig15 is inserted, the element in the target sub-slot is (st15, sig15, ptr15), and the data in other first sub-slots except slot20 and slot11 will not change.If the third free subslot does not exist in the candidate bucket corresponding to hash bucket 2, hash bucket 3 is one of the multiple secondary candidate buckets corresponding to secondary hash bucket 0, and there is no third free subslot in other secondary candidate buckets of secondary hash bucket 0, that is, there is no empty slot in the subslots of the main / secondary hash buckets and the main / secondary candidate buckets, then it will start from secondary hash bucket 0 to remove, select the tail slot00 for recording, and perform cyclic shift, and use slot03 as the removed subslot. If it is determined that the candidate bucket corresponding to slot03 is hash bucket 3, and the candidate bucket corresponding to the first subslot slot30 in hash bucket 3 is hash bucket 1, there is a free subslot, that is, the third free subslot, then it is determined that there is a third free subslot in hash bucket 3, and hash bucket 1 is determined as the third free subslot according to the third priority as mentioned above. After the target alternative bucket is selected, the aforementioned removal and move process will be performed, and the element (st9, sig9, ptr9) in the first subslot slot30 will be moved to the second idle subslot slot11 closest to the tail in hash bucket 1, and the element (st0, sig0, ptr0) in slot03 will be moved to the subslot slot30 that has been vacated after the move, until the secondary hash bucket is the most basic secondary hash bucket 0. The secondary alternative bucket includes the corresponding hash bucket 3, and the removed subslot 03 of hash bucket 0 will be used as the target subslot, using the address pointer ptr15 of slot11 as its address pointer to insert the signature value sig15. At this time, the element in the target subslot slot03 is (st15, sig15, ptr15). By determining the target subslot from the idle subslots and inserting it in the above-mentioned way, the storage space can be managed and utilized more effectively, and the optimal target subslot can be selected for data storage. This mechanism not only improves storage efficiency, but also ensures the accuracy and consistency of the data. At the same time, by moving and recording the eliminated sub-slots step by step, the storage structure is further optimized and the reliability and stability of the hash table are improved.
[0102] This optional embodiment provides the processing steps for data insertion in data processing. Similarly, effective data query, update and deletion can also be performed based on the hash table. Figure 10 , the data processing method further includes:
[0103] 1071. Upon receiving a search instruction, obtain a second keyword and a pending bitmap marking rule corresponding to the search instruction, and determine a pending bit corresponding to the pending bitmap marking rule;
[0104] 1072. Combine the data of the undetermined bits in the second keyword to generate a second private key, and combine the data of the non-undetermined bits in the second keyword to generate a second signature value;
[0105] 1073、based on the second signature value, calculate a second primary hash bucket address and a second secondary hash bucket address to determine the pending primary hash bucket and the pending secondary hash bucket corresponding to the second signature value;
[0106] 1074、determine whether there is a first signature value matching the second signature value in each first sub-slot in the pending primary hash bucket and the pending secondary hash bucket;
[0107] 1075、if there is a matching first signature value, based on the address pointer of the first sub-slot corresponding to the first signature value, locate the address column in the corresponding second-level table, and compare the pending bitmap marking rule with the target bitmap marking rule of each second sub-slot in the corresponding address column, and compare the second private key with the first private key of each second sub-slot in the corresponding address column;
[0108] 1076、if the target bitmap marking rule of the second sub-slot is consistent with the pending bitmap marking rule, and the first private key and the second private key of the same second sub-slot are consistent, output the first data item corresponding to the second sub-slot;
[0109] 1077、when receiving an update instruction, update the first data item;
[0110] 1078、when receiving a deletion instruction, delete the data stored in the second sub-slot, and update the state flag bit in the corresponding first sub-slot, if all the second sub-slots mapped by the state flag bit are empty, clear the first signature value in the corresponding first sub-slot.
[0111] In this optional embodiment, the search instruction, the update instruction and the deletion instruction can be issued based on user interaction, program internal triggering or other external system requests, etc. Since the corresponding data needs to be found first when updating the data and deleting the data, the search instruction can also be an instruction manually or automatically issued before the update instruction and the deletion instruction. The search instruction contains the second key to be queried and the pending bitmap marking rule used to locate the data. The pending bitmap marking rule defines which bit data in the second key will be used to generate the private key and which bit data will be used to generate the signature value, which helps to more accurately locate the sub-slot storing the target data in the subsequent search process. After receiving the search instruction, the second private key and the second signature value will be generated according to the second key and the pending bitmap marking rule. The second key refers to the key information of the key-value pair that needs to be queried, updated or deleted, which is used to uniquely identify the data item to be queried, updated or deleted. By processing and matching the bit data of the second key, the specific location of the data item in the hash table can be located, the search of the corresponding data item can be realized, and then the corresponding update or deletion operation can be performed on the found data item and the matching key and private key.
[0112] In this embodiment, the second keyword is first subjected to the same splitting and reassembling process as previously described for the pending bits, generating the corresponding second private key and second signature value. Then, based on the second signature value, the second primary hash bucket address and the second secondary hash bucket address are calculated, thereby determining the pending primary hash bucket and the pending secondary hash bucket corresponding to the second signature value in the hash buckets of the first-level table. The pending primary hash bucket and the pending secondary hash bucket are then searched for a first subslot matching the second signature value. If a matching first subslot is found, the address pointer of the first subslot is used to locate the corresponding address column in the second-level table. In the address column, the second private key is further compared with the first private key in the corresponding second subslot, as well as the pending bitmap marking rule and the target bitmap marking rule in each second subslot. If the first and second private keys are consistent, and the pending bitmap marking rule and the target bitmap marking rule are consistent, then it can be determined that the input second keyword is consistent with the first keyword used during the insertion process. The corresponding first data item stored in the second subslot can then be retrieved and updated or deleted accordingly. If an update instruction is subsequently received, the first data item will be updated; if a delete instruction is received, the data stored in the second subslot will be deleted, and the elements in the corresponding first subslot will also be updated, such as the update status flag. When there is no data in each second subslot of the corresponding address bar, the signature value of the corresponding first subslot can also be deleted, making the first subslot vacant and freeing up the corresponding storage space. The first subslot can then be reused to insert key-value pairs with new signature values. In this way, data can be quickly and accurately searched, and data update and deletion operations can be implemented, ensuring the accuracy and consistency of data in the hash table, which is conducive to flexible data management and efficient data processing.
[0113] The data processing method of the hash table is used to process the constructed compressed cuckoo hash table, and significant optimization can be achieved in table utilization, time complexity and storage consumption compared with the traditional compressed hash table and the cuckoo hash table. Specifically, in terms of table utilization, the compressed cuckoo hash table can reach a utilization rate of about 90%, which is close to the performance of the cuckoo hash table and significantly better than the traditional compressed hash table. In terms of time complexity, the data processing methods of searching, updating and deleting have the same complexity of O(1), and the double matching verification method is used compared with the traditional compressed hash table and the cuckoo hash table, which can effectively improve the safety and accuracy of the related processing operation. In terms of the data processing method of insertion, the complexity is determined by the preset compression factor, which is limited by the bit number of the preset compression factor corresponding to the bitmap marking rule. Generally, the complexity is lower than that of the traditional compressed hash table limited by the number of bitmap rules and the traditional cuckoo hash table limited by the load factor. In terms of storage consumption, since the hierarchical storage and the mechanism of moving and removing sub-slots are used, the application can intelligently manage and utilize the storage space while ensuring the data access speed, thereby reducing unnecessary storage overhead. When storing large-scale data, the storage consumption of the structure is about 20% of that of the traditional cuckoo hash table, and the redundant storage problem of the traditional compressed hash table under the same scale is avoided. The optimized compressed cuckoo hash table structure makes the hash table have higher efficiency and lower resource consumption when processing large-scale data, and is suitable for various high-performance computing and storage application scenarios. In addition, through the removal mechanism and the priority strategy, the application can balance high throughput and low delay in the scenario of ultra-large-scale, such as million-level data packet classification, and solve the bottleneck problem of balancing storage efficiency and performance in the prior art.
[0114] To implement the above method embodiments and corresponding steps in various possible embodiments, an implementation of a hash bucket data processing device is provided as follows. Figure 11 , Figure 11 The figure is a functional module schematic diagram of an embodiment of the hash bucket data processing device 200 in the embodiment of the application. In the embodiment, the hash table data processing device 200 includes:
[0115] The construction module 201 is configured to construct a hash table, wherein the hash table includes a first-level table and a second-level table, the first-level table includes a plurality of hash buckets, each hash bucket includes a plurality of first sub-slots, the second-level table includes a plurality of address columns, each address column corresponds to a first sub-slot, and each address column includes a plurality of second sub-slots;
[0116] The acquisition module 202 is configured to acquire a first key-value pair, the first key-value pair including a first key and a first data item;
[0117] The compression processing module 203 is configured to combine data of preset bit positions in the first key to generate a first private key, and combine data of non-preset bit positions in the first key to generate a first signature value.
[0118] The control module 204 is configured to select a target sub-slot from the idle first sub-slot, store the first signature value in the target sub-slot, and store the first private key and the first data item in a second sub-slot corresponding to the address bar of the target sub-slot.
[0119] The storage module 205 is configured to store the hash table.
[0120] Optionally, in some embodiments, the construction module 201 is specifically configured to: allocate the first-level table and the second-level table in the storage space based on a preset compression factor; set a plurality of hash buckets in the storage space of the first-level table, and allocate a plurality of first sub-slots for each hash bucket, and configure an alternative bucket for each first sub-slot from the plurality of hash buckets; and set a plurality of address bars in the storage space of the second-level table; each address bar includes a plurality of second sub-slots, the number of address bars is equal to the number of first sub-slots, and the number of second sub-slots of each address bar is the preset compression factor; different address pointers are preset for different first sub-slots; the address pointers correspond to the addresses of the address bars one by one, so as to access the matching second sub-slot in the corresponding address bar through the first sub-slot.
[0121] Optionally, in some embodiments, the construction module 201 is specifically configured to: allocate the first-level table and the second-level table in the storage space based on a preset compression factor; set a plurality of hash buckets in the storage space of the first-level table, and allocate a plurality of first sub-slots for each hash bucket, and configure an alternative bucket for each first sub-slot from the plurality of hash buckets; and set a plurality of address bars in the storage space of the second-level table; each address bar includes a plurality of second sub-slots, the number of address bars is equal to the number of first sub-slots, and the number of second sub-slots of each address bar is the preset compression factor; different address pointers are preset for different first sub-slots; the address pointers correspond to the addresses of the address bars one by one, so as to access the matching second sub-slot in the corresponding address bar through the first sub-slot.
[0122] Optionally, in some embodiments, the storage module 205 includes a first storage module and a second storage module, the first storage module is configured to store the first-level table, and the second storage module is configured to store the second-level table.
[0123] Optionally, in some embodiments, the data processing apparatus 200 of the hash table further includes a hash calculation module, which is connected with the control module 204; when the control module 204 calls the hash calculation module, the hash calculation module is configured to calculate the main hash bucket address and the auxiliary hash bucket address according to the received signature value.
[0124] Optionally, in some embodiments, the acquisition module 202 is further specifically used to: obtain a bitmap marking rule list; wherein the bitmap marking rule list includes multiple preset bitmap marking rules, and different preset bitmap marking rules are used to mark different preset bits in the keyword; simulate and test each preset bitmap marking rule in the bitmap marking rule list to screen out a target bitmap marking rule that meets the preset performance conditions; determine the preset bit in the first keyword based on the target bitmap marking rule; after selecting the target subslot from the idle first subslot, it also includes: storing the target bitmap marking rule in the second subslot of the address bar corresponding to the address pointer of the target subslot.
[0125] Optionally, in some embodiments, a status flag is preset in each of the first subslots, and the types of idle first subslots include first idle subslots, second idle subslots, and third idle subslots. The control module 204 is further specifically used to: calculate the first primary hash bucket address and the first secondary hash bucket address based on the first signature value to determine the primary hash bucket and the secondary hash bucket corresponding to the first signature value in the hash bucket of the first-level table; wherein the primary hash bucket and the secondary hash bucket each include multiple first subslots, and each first subslot corresponds to an alternative bucket; determine whether there is a first idle subslot in the first subslots in the primary hash bucket and the secondary hash bucket; wherein the status flag of the first idle subslot indicates that the second subslots of the corresponding address bar are all vacant, or the signature value of the first idle subslot is equal to the first signature value. Match, and the status flag of the first free subslot indicates that there is at least one vacant second subslot in the corresponding address bar; if there is no first free subslot in the first subslots in the main hash bucket and the secondary hash bucket, then determine whether there is a second free subslot in the first subslot of each main alternative bucket; wherein, the main alternative bucket is the alternative bucket corresponding to each first subslot of the main hash bucket, and the status flag of the second free subslot indicates that the second subslots of the corresponding address bar are all vacant; if there is no second free subslot in the first subslot of each main alternative bucket, then search for the third free subslot in the first subslot of the secondary alternative bucket step by step; wherein, the secondary alternative bucket is the alternative bucket corresponding to each first subslot of the secondary hash bucket, and the status flag of the third free subslot indicates that the second subslots of the corresponding address bar are all vacant.
[0126] Optionally, in some embodiments, the control module 204 is further specifically used to: determine whether there is a third free subslot in the first subslot of the secondary alternative bucket; if there is no third free subslot in the secondary alternative bucket, use the first subslot at the target position in the secondary hash bucket as the eliminated subslot; cyclically shift each first subslot in the secondary hash bucket; replace the secondary hash bucket with the secondary alternative bucket corresponding to the eliminated subslot, and replace the secondary alternative bucket with the alternative bucket corresponding to each first subslot in the replaced secondary hash bucket, so as to determine whether there is a third free subslot in the first subslot of each replaced secondary alternative bucket, until it is confirmed that there is a third free subslot in the first subslot of the secondary alternative bucket.
[0127] Optionally, in some embodiments, the control module 204 is further specifically used to: record the eliminated sub-slots and count the number of eliminated sub-slots; if the number of eliminated sub-slots is greater than a preset elimination threshold, or the first sub-slot to be eliminated in the hash bucket is a eliminated sub-slot, then report the insertion failure information.
[0128] Optionally, in some embodiments, the control module 204 is further configured to: determine the type of the idle first subslot; if the idle first subslot is the first idle subslot, select the first target hash bucket from the primary hash bucket and the secondary hash bucket based on the first priority, and select the first idle subslot closest to the tail in the first target hash bucket as the target subslot; if the idle first subslot is the second idle subslot, select the second target hash bucket from the primary candidate bucket based on the second priority, and move the element of the first subslot in the primary hash bucket corresponding to the second target hash bucket to the second idle subslot closest to the tail in the second target hash bucket, and use the first subslot in the primary hash bucket corresponding to the second target hash bucket as the target subslot, and use the address pointer of the second idle subslot closest to the tail in the second target hash bucket before the move as the address pointer of the target subslot; if the idle first subslot is the second idle subslot, select the second target hash bucket from the primary candidate bucket based on the second priority, and move the element of the first subslot in the primary hash bucket corresponding to the second target hash bucket to the second idle subslot closest to the tail in the second target hash bucket If the idle first subslot is the third idle subslot, the third target hash bucket is selected from the secondary candidate bucket based on the third priority, and the elements in the culled subslot in the secondary hash bucket corresponding to the third target hash bucket are moved to the third idle subslot closest to the tail in the third target hash bucket, and the secondary candidate bucket is replaced by the secondary hash bucket, and the hash bucket where the culled subslot corresponding to the replaced secondary candidate bucket is located is used to replace the secondary hash bucket, so that the elements in the culled subslot in the replaced secondary hash bucket are moved to the first subslot that is vacant after the secondary candidate bucket is moved, until the elements in the culled subslot in the most basic secondary hash bucket are moved to the first subslot that is vacant after the corresponding secondary candidate bucket is moved, and the culled subslot in the most basic secondary hash bucket is used as the target subslot, and the address pointer of the third idle subslot closest to the tail in the third target hash bucket before the move is used as the address pointer of the target subslot.
[0129] Optionally, in some embodiments, the control module 204 is further specifically used to: upon receiving a search instruction, obtain the second keyword and the to-be-determined graph marking rule corresponding to the search instruction, and determine the to-be-determined bit corresponding to the to-be-determined graph marking rule; combine the data of the to-be-determined bit in the second keyword to generate a second private key, and combine the data of the non-to-be-determined bit in the second keyword to generate a second signature value; calculate the second main hash bucket address and the second sub-hash bucket address based on the second signature value to determine the to-be-determined main hash bucket and the to-be-determined sub-hash bucket corresponding to the second signature value; determine whether there is a first signature value that matches the second signature value in each first sub-slot in the to-be-determined main hash bucket and the to-be-determined sub-hash bucket; if there is a matching first signature value, then determine the to-be-determined bit based on the second signature value corresponding to the first sub-slot. The address pointer of a subslot locates the address bar in the corresponding second-level table, and compares the to-be-positioned bitmap marking rule with the target bitmap marking rule of each second subslot in the corresponding address bar, and compares the second private key with the first private key of each second subslot in the corresponding address bar; if there is a second subslot whose target bitmap marking rule is consistent with the to-be-positioned bitmap marking rule, and the first private key and the second private key of the same second subslot are consistent, then the first data item corresponding to the second subslot is output; when an update instruction is received, the first data item is updated; when a delete instruction is received, the data stored in the second subslot is deleted, and the status flag in the corresponding first subslot is updated. If all the second subslots mapped by the status flag are empty, the first signature value in the corresponding first subslot is cleared.
[0130] Since the embodiments of the device part correspond to the embodiments of the above-mentioned hash table data processing method, please refer to the above-mentioned hash table data processing method embodiments for the introduction of the hash bucket data processing device provided by the embodiments of the present invention. The embodiments of the present invention will not be repeated here, and it has the same beneficial effects as the above-mentioned hash table data processing method.
[0131] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for processing data of a hash table, characterized in that: include: Constructing a hash table; wherein the hash table includes a first-level table and a second-level table, the first-level table includes a plurality of hash buckets, each of the hash buckets includes a plurality of first subslots, the second-level table includes a plurality of address columns, the address columns correspond one-to-one to the first subslots, and each of the address columns includes a plurality of second subslots; Obtaining a first key-value pair, the first key-value pair including a first keyword and a first data item; Combining data of preset bits in the first keyword to generate a first private key, and combining data of non-preset bits in the first keyword to generate a first signature value; Select a target subslot from the idle first subslots; The first signature value is stored in the target subslot, and the first private key and the first data item are stored in the second subslot of the target subslot corresponding to the address bar.
2. The method for processing hash table data according to claim 1, wherein: The hash table construction includes: allocating the first-level table and the second-level table in a storage space based on a preset compression factor; Setting a plurality of hash buckets in the storage space of the first-level table, allocating a plurality of the first subslots to each of the hash buckets, and configuring an alternative bucket for each of the first subslots from the plurality of hash buckets; A plurality of address columns are provided in the storage space of the second-level table; wherein each address column includes a plurality of second sub-slots, the number of the address columns is equal to the number of the first sub-slots, and the number of the second sub-slots of each address column is the preset compression factor; Different address pointers are preset for different first sub-slots; wherein the address pointers correspond one-to-one to the addresses of the address bar, so that the second sub-slot that matches the first sub-slot in the address bar can be accessed through the first sub-slot.
3. The method for processing hash table data according to claim 2, wherein: The first subslot is also preset with a status flag, and the status flag is used to map the storage status of the first subslot corresponding to the plurality of second subslots in the address column. After presetting different address pointers for different first subslots, the method further includes: detecting a storage status of a plurality of the second sub-slots of each of the address bars; Mapping the storage status of each second sub-slot to the corresponding status flag bit in the first sub-slot.
4. The method for processing hash table data according to claim 1, wherein: Before combining the data of the preset bits in the first keyword to generate the first private key, and combining the data of the non-preset bits in the first keyword to generate the first signature value, the method further includes: Obtaining a bitmap marking rule list; wherein the bitmap marking rule list includes a plurality of preset bitmap marking rules, and different preset bitmap marking rules are used to mark different preset bits in the keyword; Performing simulation tests on each of the preset bitmap marking rules in the bitmap marking rule list to screen out a target bitmap marking rule that meets preset performance conditions; Determining the preset bit in the first keyword based on the target bitmap marking rule; After selecting the target subslot from the idle first subslot, the method further includes: The target bitmap marking rule is stored in the second sub-slot of the address column corresponding to the address pointer of the target sub-slot.
5. The method for processing hash table data according to claim 3, wherein: Types of the idle first subslots include a first idle subslot, a second idle subslot, and a third idle subslot. Before selecting a target subslot from the idle first subslots, the method includes: Calculate a first primary hash bucket address and a first secondary hash bucket address based on the first signature value to determine the primary hash bucket and the secondary hash bucket corresponding to the first signature value in the hash buckets of the first-level table; Determining whether there is a first free subslot in the first subslots in the primary hash bucket and the secondary hash bucket; wherein the status flag of the first free subslot indicates that the second subslots corresponding to the address bar are all free, or the signature value of the first free subslot matches the first signature value, and the status flag of the first free subslot indicates that there is at least one free second subslot in the corresponding address bar; If the first free subslot does not exist in the first subslots of the primary hash bucket and the secondary hash bucket, determining whether there is a second free subslot in the first subslot of each primary candidate bucket; wherein the primary candidate bucket is the candidate bucket corresponding to each first subslot of the primary hash bucket, and the status flag of the second free subslot indicates that the second subslot corresponding to the address bar is vacant; If the second free subslot does not exist in the first subslot of each of the primary alternative buckets, the third free subslot is searched for in the first subslots of the secondary alternative buckets step by step; wherein, the secondary alternative bucket is the alternative bucket corresponding to each of the first subslots of the secondary hash bucket, and the status flag of the third free subslot indicates that the second subslots of the corresponding address bar are all vacant.
6. The method for processing hash table data according to claim 5, wherein: The step of searching for the third free subslot in the first subslots of the secondary candidate buckets step by step includes: Determining whether the third idle subslot exists in the first subslot of the secondary candidate bucket; If the third free sub-slot does not exist in the secondary candidate bucket, the first sub-slot at the target position in the secondary hash bucket is used as a sub-slot to be eliminated; cyclically shifting each of the first subslots in the secondary hash bucket; The secondary candidate bucket corresponding to the eliminated sub-slot is used to replace the secondary hash bucket, and the secondary candidate bucket is replaced with the candidate bucket corresponding to each of the first sub-slots in the replaced secondary hash bucket, so as to determine whether the third free sub-slot exists in the first sub-slot of each of the replaced secondary candidate buckets, until it is confirmed that the third free sub-slot exists in the first sub-slot of the secondary candidate bucket.
7. The method for processing hash table data according to claim 6, wherein: The data processing method further includes: Recording the corresponding relationship between the rejection sub-slots and each of the secondary candidate buckets and the corresponding rejection sub-slots in sequence, and counting the number of the rejection sub-slots; If the number of the eliminated sub-slots is greater than a preset elimination threshold, or the first sub-slot to be eliminated in the hash bucket is the eliminated sub-slot, insertion failure information is reported.
8. The method for processing hash table data according to claim 7, wherein: The selecting a target subslot from the idle first subslots includes: Determining the type of the idle first subslot; If the first free subslot is the first free subslot, selecting a first target hash bucket from the primary hash bucket and the secondary hash bucket based on the first priority, and selecting the first free subslot closest to the tail in the first target hash bucket as the target subslot; If the free first subslot is the second free subslot, a second target hash bucket is selected from the primary candidate bucket based on the second priority, and the element of the second target hash bucket corresponding to the first subslot in the primary hash bucket is moved to the second free subslot closest to the tail in the second target hash bucket, and the first subslot in the primary hash bucket corresponding to the second target hash bucket is used as the target subslot, and the address pointer of the second free subslot closest to the tail in the second target hash bucket before the move is used as the address pointer of the target subslot; If the free first subslot is the third free subslot, a third target hash bucket is selected from the secondary candidate bucket based on the third priority, and the elements in the culled subslot in the secondary hash bucket corresponding to the third target hash bucket are moved to the third free subslot closest to the tail of the third target hash bucket, and the secondary candidate bucket is replaced by the secondary hash bucket, and the hash bucket where the culled subslot corresponding to the replaced secondary candidate bucket is located is replaced by the hash bucket where the replaced secondary candidate bucket is located, so that the elements in the culled subslot in the replaced secondary hash bucket are moved to the first subslot that becomes vacant after the secondary candidate bucket is moved, until the elements in the culled subslot in the most basic secondary hash bucket are moved to the first subslot that becomes vacant after the corresponding secondary candidate bucket is moved, and the culled subslot in the most basic secondary hash bucket is used as the target subslot, and the address pointer of the third free subslot closest to the tail of the third target hash bucket before the move is used as the address pointer of the target subslot.
9. The method for processing hash table data according to any one of claims 5 to 8, characterized in that: The data processing method further includes: Upon receiving a search instruction, obtaining a second keyword and a pending bitmap marking rule corresponding to the search instruction, and determining a pending bit corresponding to the pending bitmap marking rule; Combining data of undetermined bits in the second keyword to generate a second private key, and combining data of non-undetermined bits in the second keyword to generate a second signature value; Calculate a second primary hash bucket address and a second secondary hash bucket address based on the second signature value to determine a pending primary hash bucket and a pending secondary hash bucket corresponding to the second signature value; Determine whether there is a first signature value matching the second signature value in each of the first subslots in the pending primary hash bucket and the pending secondary hash bucket; If a matching first signature value exists, locating the address column in the corresponding second-level table based on the address pointer of the first subslot corresponding to the first signature value, comparing the to-be-bitmap marking rule with the target bitmap marking rule corresponding to each second subslot in the address column, and comparing the second private key with the first private key corresponding to each second subslot in the address column; If the target bitmap marking rule of the second subslot is consistent with the to-be-marked bitmap marking rule, and the first private key and the second private key of the same second subslot are consistent, then output the first data item corresponding to the second subslot; When receiving an update instruction, updating the first data item; When a delete instruction is received, the data stored in the second subslot is deleted, and the status flag corresponding to the first subslot is updated. If the second subslots mapped by the status flag are all empty, the first signature value corresponding to the first subslot is cleared.
10. A data processing device for a hash table, characterized in that: include: a construction module configured to construct a hash table; wherein the hash table includes a first-level table and a second-level table, the first-level table including a plurality of hash buckets, each of the hash buckets including a plurality of first subslots, and the second-level table including a plurality of address columns, the address columns corresponding one-to-one to the first subslots, and each of the address columns including a plurality of second subslots; An acquisition module, configured to acquire a first key-value pair, where the first key-value pair includes a first keyword and a first data item; a compression processing module, configured to combine data of preset bits in the first keyword to generate a first private key, and to combine data of non-preset bits in the first keyword to generate a first signature value; a control module, configured to select a target subslot from the free first subslots, store the first signature value in the target subslot, and store the first private key and the first data item in the second subslot corresponding to the address column of the address pointer of the target subslot; A storage module is used to store the hash table.