Hash table data insertion method and system and electronic equipment

By deeply expanding and storing signature value on the cuckoo hash table, the problem of decreasing insert success rate caused by excessive load factor of hash table is solved, and the utilization rate of hash table and data insertion efficiency is improved.

CN119961264APending Publication Date: 2025-05-09DAPUSTOR CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411994794.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

After the load factor of the cuckoo hash table exceeds a certain value, the success rate of inserting new key-value pairs is significantly reduced, resulting in a decrease in the utilization rate of the hash table.

Method used

By deeply expanding the first-level table, keywords and data items are obtained, and the address and free sub-slot of the hash bucket are obtained based on the keywords, the signature value and address pointer are stored in the free sub-slot of the hash bucket, and the keywords and data items are stored in the second-level table.

Benefits of technology

Improve the utilization of hash tables, enhance the efficiency of data insertion, and avoid degradation of insertion performance caused by excessive load factor.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961264A_ABST
    Figure CN119961264A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to the technical field of data processing, and discloses a data insertion method and system for a hash table and electronic equipment, and the method comprises the steps: carrying out the deep expansion of a first-level table, obtaining an expanded first-level table, obtaining a keyword and a data item, obtaining the address of a hash bucket based on the keyword, and carrying out the data insertion of the hash table. Obtaining the idle sub-slot of one Hash bucket of the expanded first-level table and the signature value of the keyword, obtaining an address pointer, storing the address pointer and the signature value in the idle sub-slot of the Hash bucket, and storing the keyword and the data item in the sub-slot of the second-level table pointed by the address pointer, so that the utilization rate of the Hash table can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a method, system and electronic device for inserting data into a hash table. Background Art

[0002] In the fields of computer algorithms, database management and information processing technology, cuckoo hashing is a commonly used data structure for fast table lookup and insertion, which stores key-value pairs by selecting hash buckets and subslots through hash values.

[0003] When inserting a new key-value pair and a hash collision occurs, Cuckoo Hash can solve the hash collision problem by replacing the subslot. However, since the number of subslot replacements is limited by the configured threshold, the greater the number of replacements, the load factor of the hash table will gradually increase. When the load factor exceeds a certain value, for example, more than 80%, when new key-value pairs continue to be inserted, the number of free slots in the hash table decreases, and the number of subslot replacements increases significantly, resulting in an exponential decrease in the success rate of inserting new key-value pairs, reducing the utilization of the hash table. Summary of the invention

[0004] The embodiments of the present application provide a method, system and electronic device for inserting data into a hash table, which can improve the efficiency of adding data.

[0005] To solve the above technical problems, the present application provides the following technical solutions:

[0006] In a first aspect, an embodiment of the present application provides a method for inserting data into a hash table, wherein the hash table includes a first-level table and a second-level table, the first-level table includes multiple hash buckets, the hash bucket includes multiple sub-slots, and the second-level table includes multiple sub-slots, and the method includes:

[0007] Performing a deep expansion on the first-level table to obtain an expanded first-level table;

[0008] Obtaining a first key-value pair, wherein the key-value pair includes a first keyword and a first data item;

[0009] According to the first keyword, obtain a free subslot of a hash bucket in the expanded first-level table and a first signature value corresponding to the first keyword;

[0010] Obtaining a first address pointer, wherein the first address pointer is used to store an address of a subslot of the second-level table;

[0011] The first signature value and the first address pointer corresponding to the first keyword are stored in the free subslot of the hash bucket of the first-level table;

[0012] According to the first address pointer, the first keyword and the first data item are stored in the second level table.

[0013] In some embodiments, the first-level table is deeply expanded, including:

[0014] Get the load factor;

[0015] Get the maximum depth of the hash table;

[0016] Based on the load factor and the maximum depth of the hash table, the depth of the first-level table is expanded, wherein the depth of the first-level table = (1 / σ)*D, σ is the load factor, and D is the maximum depth of the hash table.

[0017] In some embodiments, the hash bucket includes a primary hash bucket and a secondary hash bucket, and obtaining, according to the first keyword, a free subslot of a hash bucket of the expanded first-level table includes:

[0018] According to the first keyword, an address of a primary hash bucket and an address of a secondary hash bucket are obtained, the primary hash bucket is obtained according to the address of the primary hash bucket, and the secondary hash bucket is obtained according to the address of the secondary hash bucket, wherein the primary hash bucket and the secondary hash bucket each include a plurality of subslots, each subslot corresponds to an alternative bucket, and the subslot is used to store the first signature value;

[0019] When there are no free subslots in the primary hash bucket, the secondary hash bucket, and the candidate bucket corresponding to the primary hash bucket, it is determined whether there is an empty slot in the candidate bucket corresponding to the secondary hash bucket;

[0020] If there is an idle subslot in the candidate bucket corresponding to the secondary candidate bucket, the first elimination subslot and the first candidate bucket are determined according to the priority principle of the newly added elements, wherein the first elimination subslot is one of the multiple subslots in the secondary candidate bucket, and the first candidate bucket is the candidate bucket corresponding to the first elimination subslot;

[0021] Add the elements of the first eliminated sub-slot to the first candidate bucket, and use the first eliminated sub-slot as the sub-slot for storing the first signature value;

[0022] If there is no free subslot in the candidate bucket corresponding to the secondary candidate bucket, the second elimination subslot and the second candidate bucket corresponding to the secondary candidate bucket are determined according to the elimination element priority principle, wherein the second elimination subslot is a subslot in the secondary candidate hash bucket, and the second candidate bucket is the candidate bucket corresponding to the second elimination subslot;

[0023] After all subslots in the secondary candidate hash bucket are shifted in sequence, they are overwritten with the candidate hash bucket, and the number of removals is increased by one;

[0024] The second candidate bucket is replaced with the secondary candidate bucket, and it is determined whether there is an empty slot in the secondary candidate bucket until there is an empty sub-slot in the candidate bucket corresponding to the secondary candidate bucket.

[0025] In some embodiments, the hash table includes a plurality of domain segments, each domain segment includes a plurality of hash buckets, and the priority principle of the newly added element includes the principle of the largest number of free subslots of the hash bucket, the principle of the largest number of free subslots of the domain segment, and the tail priority principle;

[0026] According to the priority principle of the newly added elements, the first elimination sub-slot and the first candidate bucket are determined, including:

[0027] According to the principle that the number of free sub-slots of the hash bucket is the largest, a third candidate bucket is determined, wherein the third candidate bucket is the candidate bucket with the largest number of free sub-slots among all candidate buckets corresponding to the secondary hash bucket;

[0028] If the number of the third candidate buckets is greater than the preset number, a fourth candidate bucket is determined according to the principle that the number of free subslots in the domain segment is the largest, wherein the fourth candidate bucket is a candidate bucket in the third candidate bucket;

[0029] If the number of the fourth candidate buckets is greater than the preset number, the fifth candidate bucket is determined according to the tail priority principle, and the fifth candidate bucket is used as the first candidate bucket, wherein the fifth candidate bucket is one of the multiple fourth candidate buckets;

[0030] The subslot in the secondary hash bucket corresponding to the first candidate bucket is used as the first elimination subslot.

[0031] In some embodiments, the element culling priority principle includes a tail-first culling principle or a head-first culling principle;

[0032] According to the priority principle of culling elements, the corresponding second culling sub-slot and the second candidate bucket in the secondary candidate bucket are determined, including:

[0033] If the element removal priority principle is the tail first removal principle, the first subslot at the tail of the secondary hash bucket is used as the second removal subslot;

[0034] If the element removal priority principle is the head first removal principle, the first subslot at the head of the secondary hash bucket is used as the second removal subslot;

[0035] The candidate bucket corresponding to the second rejection sub-slot is used as the second candidate bucket.

[0036] In some embodiments, all subslots in the secondary candidate hash bucket are sequentially shifted cyclically, including:

[0037] All subslots in the secondary candidate hash bucket are moved one position toward the tail of the secondary candidate hash bucket in sequence until all subslots in the secondary candidate hash bucket are shifted.

[0038] In some embodiments, the method further comprises:

[0039] If the number of the third candidate buckets is equal to the preset number, the third candidate bucket is used as the first candidate bucket;

[0040] If the number of the fourth candidate buckets is equal to the preset number, the fourth candidate bucket is used as the first candidate bucket.

[0041] In some embodiments, the method further comprises:

[0042] If there is an idle subslot in the primary hash bucket and there is no idle subslot in the secondary hash bucket, the first signature value and the first address pointer are stored in the idle subslot of the primary hash bucket;

[0043] If there is no free subslot in the primary hash bucket and there is a free subslot in the secondary hash bucket, the first signature value and the first address pointer are stored in the free subslot of the secondary hash bucket;

[0044] If both the primary hash bucket and the secondary hash bucket have free subslots, determine the hash bucket according to the priority principle of the newly added elements, and insert the first signature value and the first address pointer into the hash bucket, where the hash bucket is one of the primary hash bucket and the secondary hash bucket;

[0045] If there is no free subslot in the primary hash bucket and no free subslot in the secondary hash bucket, and there is a free subslot in the primary candidate bucket corresponding to the subslot of the primary hash bucket, insert the first signature value and the first address pointer into the free subslot of the primary candidate bucket corresponding to the subslot of the primary hash bucket;

[0046] If there is no free subslot in the main hash bucket, and there is no free subslot in the secondary hash bucket, there is no free subslot in the main candidate bucket corresponding to the subslot of the main hash bucket, and there is a free subslot in the secondary candidate bucket corresponding to the subslot of the secondary hash bucket, then the first signature value and the first address pointer are inserted into the free subslot of the secondary candidate bucket corresponding to the subslot of the secondary hash bucket.

[0047] In some embodiments, the method further comprises:

[0048] Get the second keyword;

[0049] Performing hash calculation on the second keyword to obtain an address of a hash bucket of the expanded first-level table and a second signature value;

[0050] Based on the address of the hash bucket, read the third signature value and the second address pointer in the hash bucket, wherein the second address pointer is used to store a subslot address of the second-level table;

[0051] If the second signature value is the same as the third signature value, the second data item of the subslot of the second-level table is obtained based on the second address pointer.

[0052] In some embodiments, the method further comprises:

[0053] Obtaining a second key-value pair, wherein the key-value pair includes a third keyword and a third data item;

[0054] Performing hash calculation on the third keyword to obtain an address of a hash bucket of the expanded first-level table and a fourth signature value;

[0055] Based on the address of the hash bucket, read the fifth signature value and the third address pointer in the hash bucket, wherein the third address pointer is used to store a subslot address of the second-level table;

[0056] If the fourth signature value is the same as the fifth signature value, obtaining the fourth keyword of the subslot of the second-level table based on the third address pointer;

[0057] If the third keyword is the same as the fourth keyword, the third data item is updated to the subslot of the second-level table pointed to by the third address pointer.

[0058] In some embodiments, the method further comprises:

[0059] If the second signature value is different from the third signature value, the current operation of obtaining the second data item is terminated;

[0060] If the fourth signature value is the same as the fifth signature value, and the third keyword is different from the fourth keyword, then the current operation of updating the third data item to the subslot of the second-level table is terminated;

[0061] or,

[0062] If the fourth signature value is different from the fifth signature value, the current operation of updating the third data item to the subslot of the second-level table is terminated.

[0063] In a second aspect, an embodiment of the present application provides a data insertion system for a hash table, wherein the hash table includes a first-level table and a second-level table, the first-level table includes a plurality of hash buckets, the hash bucket includes a plurality of sub-slots, the second-level table includes a plurality of sub-slots, and the system includes:

[0064] A hash table control unit, used for performing depth expansion on the first-level table to obtain an expanded first-level table;

[0065] A hash calculation unit, configured to obtain a first key-value pair, wherein the key-value pair includes a first keyword and a first data item, and obtain, according to the first keyword, a free subslot of a hash bucket in the extended first-level table and a first signature value corresponding to the first keyword;

[0066] A pointer control unit, used for acquiring a first address pointer, wherein the first address pointer is used for storing an address of a subslot of the second-level table;

[0067] A first-level table storage unit, used to store a first signature value and a first address pointer corresponding to a first keyword;

[0068] The second-level table storage unit is used to store the first keyword and the first data item.

[0069] In a third aspect, an embodiment of the present application provides an electronic device, including:

[0070] at least one processor, and

[0071] a memory communicatively coupled to at least one processor, wherein:

[0072] The memory stores instructions that can be executed by at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method for inserting data into a hash table according to the first aspect.

[0073] The beneficial effect of the embodiment of the present application is: different from the prior art, the embodiment of the present application provides a method for inserting data into a hash table, the hash table includes a first-level table and a second-level table, the first-level table includes multiple hash buckets, the hash bucket includes multiple subslots, and the second-level table includes multiple subslots. The method includes: deeply expanding the first-level table to obtain an expanded first-level table, obtaining a first key-value pair, wherein the key-value pair includes a first keyword and a first data item, according to the first keyword, obtaining a free subslot of a hash bucket in the expanded first-level table, a first signature value corresponding to the first keyword, obtaining a first address pointer, wherein the first address pointer is used to store the address of a subslot of the second-level table, storing the first signature value corresponding to the first keyword and the first address pointer in the free subslot of the hash bucket of the first-level table, and storing the first keyword and the first data item in the second-level table according to the first address pointer.

[0074] By deeply expanding the first-level table to obtain the expanded first-level table, obtaining the keyword and the data item, obtaining the address of the hash bucket based on the keyword, obtaining the free subslot of a hash bucket of the expanded first-level table and the signature value of the keyword, obtaining the address pointer, storing the address pointer and the signature value in the free subslot of the hash bucket, and storing the keyword and the data item in the subslot of the second-level table pointed to by the address pointer, the utilization rate of the hash table can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0075] One or more embodiments are exemplarily described by pictures in the corresponding drawings, and these exemplified descriptions do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings represent similar elements, and unless otherwise stated, the figures in the drawings do not constitute proportional limitations.

[0076] Figure 1 It is a flowchart of a method for inserting data into a hash table provided in an embodiment of the present application;

[0077] Figure 2 yes Figure 1 A detailed flow chart of step S101 in FIG.

[0078] Figure 3 yes Figure 1 A detailed flow chart of step S103 in FIG.

[0079] Figure 4 This is an example schematic diagram of inserting a first signature value and a first address pointer into a hash table provided by an embodiment of the present application;

[0080] Figure 5 yes Figure 3 A detailed flowchart of step S1309 in FIG.

[0081] Figure 6 This is another example schematic diagram of inserting a first signature value and a first address pointer into a hash table provided by an embodiment of the present application;

[0082] Figure 7 yes Figure 3 A detailed flowchart of step S1310 in FIG.

[0083] Figure 8 yes Figure 3 A detailed flowchart of step S1311 in FIG.

[0084] Fig. 9 This is another example schematic diagram of inserting a first signature value and a first address pointer into a hash table provided by an embodiment of the present application;

[0085] Fig.10 This is a schematic diagram of an example structure of a hash table provided in an embodiment of the present application;

[0086] Fig.11 is a schematic diagram of an example structure of a sub-slot in a first-level table provided in an embodiment of the present application;

[0087] Fig.12 This is an example schematic diagram of inserting an element into a hash table provided by an embodiment of the present application;

[0088] Fig.13 This is a schematic diagram of a data search process provided by an embodiment of the present application;

[0089] Fig.14 This is a schematic diagram of a data update process provided by an embodiment of the present application;

[0090] Fig.15 It is a structural diagram of a hash table data insertion system provided in an embodiment of the present application;

[0091] Fig.16 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application.

[0092] Description of Figure Numbers:

[0093] DETAILED DESCRIPTION

[0094] In order to make the purpose, technical solutions and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application.

[0095] It should be noted that, if there is no conflict, the various features in the embodiments of the present application can be combined with each other, all within the scope of protection of the present application. In addition, although the functional module division is performed in the device schematic diagram and the logical order is shown in the flow chart, in some cases, the steps shown or described can be performed in a sequence different from the module division in the device or the flow chart. Furthermore, the words "first", "second", "third", etc. used in this application do not limit the data and execution order, but only distinguish the same items or similar items with basically the same functions and effects.

[0096] In addition, the technical features involved in each embodiment of the present application described below can be combined with each other as long as they do not conflict with each other.

[0097] Before introducing the embodiments of the present application, a brief introduction to the prior art known to the inventors of the present application is first given to facilitate subsequent understanding of the embodiments of the present application.

[0098] Cuckoo hash is a commonly used data structure for fast table lookup and insertion. It stores key-value pairs by selecting hash buckets and subslots through hash values.

[0099] Since the input performance of the hash table is limited by the load factor, the load factor will gradually increase with the increase in the number of replacements of the hash table's subslots. When the load factor exceeds a certain value, the insertion performance of the hash table will decrease exponentially, resulting in reduced utilization of the hash table.

[0100] In response to the above problems, the present application provides a method for inserting data into a hash table, which deeply expands the first-level table to obtain an expanded first-level table, obtains keywords and data items, obtains the address of a hash bucket based on the keyword, obtains a free subslot of a hash bucket of the expanded first-level table and a signature value of the keyword, obtains an address pointer, stores the address pointer and the signature value in the free subslot of the hash bucket, and stores the keywords and data items in the subslot of the second-level table pointed to by the address pointer, thereby improving the utilization rate of the hash table.

[0101] The technical solution of this application is described in detail below with reference to the accompanying drawings:

[0102] See also Figure 1 , Figure 1 It is a flowchart of a method for inserting data into a hash table provided in an embodiment of the present application.

[0103] The method for inserting data into a hash table is applied to an electronic device. Specifically, the execution subject of the method for inserting data into a hash table is one or at least two processors of the electronic device.

[0104] like Figure 1 As shown, the data insertion method of the hash table includes:

[0105] Step S101: perform depth expansion on the first-level table to obtain an expanded first-level table.

[0106] In an embodiment of the present application, the hash table includes a first-level table, the first-level table includes multiple hash buckets, and each hash bucket includes multiple sub-slots.

[0107] In an embodiment of the present application, when data is inserted into a hash table, there will be a hash conflict, resulting in low utilization of the hash table. Before obtaining the key-value pairs, the first-level table needs to be expanded to expand the depth of the first-level table so that when the hash table performs an insertion operation, the search path of the free subslot is stabilized at a certain value.

[0108] Specifically, the first-level table is deeply expanded to obtain an expanded first-level table. Figure 2 .

[0109] See also Figure 2 , Figure 2 yes Figure 1 Detailed flowchart of step S101 in FIG.

[0110] like Figure 2 As shown, step S101 includes:

[0111] Step S111: Obtain the load factor.

[0112] In the embodiment of the present application, the load factor is a parameter used to measure the degree of filling of the hash table, that is, the ratio between the number of elements stored in the hash table and the total capacity of the hash table, load factor = number of stored elements / total capacity of the hash table.

[0113] Specifically, the load factor is obtained through simulation testing.

[0114] For example, a hash table can hold 1,000 entries, but when actually adding elements, only 800 entries can be added, so the load factor = 800 / 1,000% = 80%.

[0115] Step S112: Obtain the maximum depth of the hash table.

[0116] The maximum depth of the hash table is the entry size of the hash table.

[0117] Step S113: Based on the load factor and the maximum depth of the hash table, the depth of the first-level table is expanded.

[0118] Specifically, based on the load factor and the maximum depth of the hash table, the depth of the first-level table is extended, and the depth of the first-level table = (1 / σ)*D, σ is the load factor, and D is the maximum depth of the hash table.

[0119] Step S102: Obtain a first key-value pair, wherein the key-value pair includes a first keyword and a first data item.

[0120] Specifically, a first key-value pair is obtained, where the key-value pair includes a first keyword and a first data item.

[0121] Step S103: According to the first keyword, obtain a free subslot of a hash bucket of the expanded first-level table.

[0122] Specifically, the first keyword is hashed to obtain a free subslot of a hash bucket of the expanded first-level table. For specific steps, please refer to Figure 3 .

[0123] See also Figure 3 , Figure 3 yes Figure 1 Detailed flowchart of step S103 in FIG.

[0124] like Figure 3 As shown, step S103 includes:

[0125] Step S1301: According to the first keyword, obtain the address of the primary hash bucket, the address of the secondary hash bucket, and the first signature value corresponding to the first keyword, obtain the primary hash bucket according to the primary hash bucket address, and obtain the secondary hash bucket according to the secondary hash bucket address.

[0126] In the embodiment of the present application, the hash bucket includes a primary hash bucket and a secondary hash bucket, the primary hash bucket includes a plurality of sub-slots, and the secondary hash bucket includes a plurality of sub-slots.

[0127] Specifically, the hash function hash_function() is used to perform hash calculation on the first keyword to obtain the address of the primary hash bucket and the address of the secondary hash bucket, for example, (pri_bucket_addr, sec_bucket_addr) = hash_function(key), where key is the keyword, pri_bucket_addr is the primary hash bucket address, and sec_bucket_addr is the secondary hash bucket address.

[0128] Specifically, the address of the primary hash bucket and the secondary hash bucket address are input into the acquisition function to obtain the primary hash bucket and the secondary hash bucket. For example, the acquisition function is get_bucket(), (pri_bucket, sec_bucket) = get_bucket(pri_bucket_addr, sec_bucket_addr), pri_bucket_addr is the primary hash bucket address, sec_bucket_addr is the secondary hash bucket address, pri_bucket is the primary hash bucket, and sec_bucket is the secondary hash bucket.

[0129] Step S1302: Determine whether there is an idle subslot in the main hash bucket.

[0130] Specifically, it is determined by a detection function whether there is an idle subslot in the main hash bucket. If there is an idle subslot in the main hash bucket, the process jumps to step S1303; if there is no idle subslot in the main hash bucket, the process jumps to step S1304.

[0131] The detection function may be IsEmptySlot(), for example, IsEmptySlot(pri_bucket), where pri_bucket is the primary hash bucket.

[0132] Step S1303: insert the first signature value and the first address pointer into the free subslot of the main hash bucket.

[0133] The first address pointer stores the subslot address of the second-level table, and the subslot of the second-level table is used to store key-value pairs.

[0134] Specifically, when there is an idle subslot in the main hash bucket, the first signature value and the first address pointer are inserted into the idle subslot of the main hash bucket.

[0135] In an embodiment of the present application, the main hash bucket includes multiple sub-slots. When there are multiple free sub-slots in the main hash bucket, the first signature value and the first address pointer are inserted starting from the tail of the main hash bucket.

[0136] Step S1304: Determine whether there is an idle subslot in the secondary hash bucket.

[0137] Specifically, it is determined whether there is an idle subslot in the secondary hash bucket through a detection function. If there is an idle subslot in the secondary hash bucket, the process jumps to step S1305 . If there is no idle subslot in the secondary hash bucket, the process jumps to step S1306 .

[0138] The detection function may be IsEmptySlot(), for example, IsEmptySlot(sec_bucket), where sec_bucket is a secondary hash bucket.

[0139] Step S1305: insert the first signature value and the first address pointer into the free subslot of the secondary hash bucket.

[0140] Specifically, when there is an idle subslot in the secondary hash bucket, the first signature value and the first address pointer are inserted into the idle subslot of the secondary hash bucket.

[0141] In the embodiment of the present application, the secondary hash bucket includes multiple sub-slots. When the secondary hash bucket has multiple free sub-slots, the first signature value and the first address pointer are inserted starting from the tail of the secondary hash bucket.

[0142] In the embodiment of the present application, when detecting whether there are free subslots in the primary hash bucket and the secondary hash bucket, the detection can be performed simultaneously without limiting the order. When both the primary hash bucket and the secondary hash bucket have free subslots, one of the primary hash bucket and the secondary hash bucket is determined as the bucket corresponding to the first signature value and the first address pointer storage according to the new element priority principle. The specific content of the new element priority principle is described in Figure 5 A detailed introduction is given in.

[0143] See also Figure 4 , Figure 4 This is an example schematic diagram of inserting a first signature value and a first address pointer into a hash table provided by an embodiment of the present application.

[0144] like Figure 4 As shown, the Figure 4In the case where both the primary hash bucket and the secondary hash bucket have free subslots, a first signature value and a first address pointer are inserted into an example diagram of a hash table. Assume that the hash table has 5 domain segments, domain segment 0 to domain segment 4, each domain segment includes four hash buckets, and a hash bucket (bucket) has 4 subslots, a total of 80 slots, 20 buckets, bucket0 to bucket19, the maximum load of each domain segment (i.e., the number of subslots) is 16 slots, H (head) represents the head of the hash bucket, and T (tail) represents the tail of the hash bucket. After obtaining the first signature value and the first address pointer, after hash calculation, the primary hash bucket H1 is obtained as bucket2, and the secondary hash bucket H2 is obtained as bucket12. There are free subslots in both bucket2 and bucket12, the number of free subslots of bucket2 is 2, and the number of free subslots of bucket12 is 1, then the first signature value and the first address pointer are inserted into the free subslot of bucket2.

[0145] Step S1306: Determine whether there is an idle subslot in the candidate bucket corresponding to the main hash bucket.

[0146] In the embodiment of the present application, the primary hash bucket and the secondary hash bucket each include a plurality of sub-slots, and each sub-slot corresponds one-to-one to a candidate bucket.

[0147] Specifically, when there are no free subslots in either the primary hash bucket or the secondary hash bucket, obtain the alternative buckets corresponding to each subslot of the primary hash bucket, and determine whether there are free subslots in the alternative buckets corresponding to the primary hash bucket. If there are free subslots in the alternative buckets corresponding to the subslots of the primary hash bucket, jump to step S1307; if there are no free subslots in the alternative buckets corresponding to the subslots of the primary hash bucket, jump to step S1308.

[0148] In an embodiment of the present application, if the address of the main alternative bucket is known, then the address of each subslot of the main alternative bucket is known. Then, according to the address of the subslot in the main alternative bucket, the alternative bucket corresponding to the subslot of the main alternative bucket can be obtained, and each subslot corresponds to a one-to-one alternative bucket. For example, when there are four subslots, {pri_slot3, pri_slot2, pri_slot1, pri_slot0} = pri_bucket, pri_bucket is the main alternative bucket, pri_slot3, pri_slot2, pri_slot1, pri_slot0 are the four subslots of the main alternative bucket, then the main alternative buckets corresponding to pri_slot3, pri_slot2, pri_slot1, pri_slot0 are {sec_bucket3, sec_bucket2, sec_bucket1, sec_bucket0} respectively.

[0149] Step S1307: Determine the hash bucket according to the priority principle of the newly added element.

[0150] Specifically, when there is an idle subslot in the candidate bucket corresponding to the subslot of the main hash bucket, the hash bucket is determined according to the priority principle of the newly added element, and the first signature value and the first address pointer are inserted into the hash bucket. The specific content of the priority principle of the newly added element is in Figure 5 A detailed introduction is given in.

[0151] In an embodiment of the present application, since there are multiple subslots in the main hash bucket, each subslot corresponds to an alternative bucket. When there are free subslots in the alternative buckets corresponding to each subslot in the main hash bucket, it is necessary to determine an alternative bucket from all the alternative buckets corresponding to the main hash bucket to further insert the first signature value and the first address pointer into the hash bucket.

[0152] Step S1308: Determine whether there is an idle subslot in the candidate bucket corresponding to the secondary hash bucket.

[0153] Specifically, when there is no free subslot in the alternative bucket corresponding to the subslot of the main hash bucket, the alternative bucket corresponding to each subslot of the secondary hash bucket is obtained to determine whether there is a free subslot in the alternative bucket corresponding to the secondary hash bucket. If there is a free subslot in the alternative bucket corresponding to the subslot of the secondary hash bucket, jump to step S1309; if there is no free subslot in the alternative bucket corresponding to the subslot of the secondary hash bucket, jump to step S1310.

[0154] In an embodiment of the present application, if the address of the secondary alternative bucket is known, then the address of each subslot of the secondary alternative bucket is known. Then, according to the address of the subslot in the secondary alternative bucket, the alternative bucket corresponding to the subslot of the secondary alternative bucket can be obtained, and each subslot corresponds to a one-to-one alternative bucket. For example, when there are four subslots in the secondary alternative bucket, {sec_slot3, sec_slot2, sec_slot1, sec_slot0} = sec_bucket, sec_bucket is the secondary alternative bucket, sec_slot3, sec_slot2, sec_slot1, sec_slot0 are the four subslots of the secondary alternative bucket, then the secondary alternative buckets corresponding to {sec_slot3, sec_slot2, sec_slot1, sec_slot0} are {sec_bucket3, sec_bucket2, sec_bucket1, sec_bucket0} respectively.

[0155] Step S1309: Determine the first elimination sub-slot and the first candidate bucket according to the priority principle of the newly added element.

[0156] Among them, the newly added element priority principles include the principle of the largest number of free subslots in the hash bucket, the principle of the largest number of free subslots in the domain segment, and the tail priority principle.

[0157] Specifically, when there is an idle subslot in the candidate bucket corresponding to the subslot of the secondary hash bucket, a new element priority principle is added to determine the first elimination subslot and the first candidate bucket, wherein the first elimination subslot is a subslot in the secondary hash bucket, and the first candidate bucket is the candidate bucket corresponding to the first elimination subslot.

[0158] See also Figure 5 , Figure 5 yes Figure 3 A detailed flowchart of step S1309 in FIG.

[0159] like Figure 5 As shown, step S1309 includes:

[0160] Step S1391: Determine the third candidate bucket according to the principle that the number of free sub-slots in the hash bucket is the largest.

[0161] Specifically, the number of free subslots in the candidate bucket corresponding to each subslot of the secondary hash bucket is obtained, and the candidate bucket with the largest number of free subslots is used as the third candidate bucket.

[0162] For example, the secondary hash bucket has 4 subslots, and the numbers of free subslots in the candidate buckets corresponding to the 4 subslots are (1, 2, 3, 4) respectively, then the candidate bucket with 4 free subslots is used as the third candidate bucket.

[0163] In an embodiment of the present application, there may be a situation where the number of free subslots in the alternative bucket corresponding to each subslot of the secondary hash bucket is equal. For example, there are 4 subslots in the secondary hash bucket, and the numbers of free subslots in the alternative buckets corresponding to the 4 subslots are (4, 4, 3, 4) respectively. Then the alternative bucket with 4 free subslots will be used as the third alternative bucket. There are three third alternative buckets, and it is necessary to further select a third alternative bucket from them.

[0164] In the embodiment of the present application, the more free sub-slots a hash bucket has, the lighter the bucket load is.

[0165] Step S1392: Determine whether the number of the third candidate buckets is greater than the preset number.

[0166] Among them, the preset number is 1.

[0167] Specifically, determine whether the number of the third candidate buckets is greater than the preset number. If the number of the third candidate buckets is equal to the preset number, jump to step S1393; if the number of the third candidate buckets is greater than the preset number, jump to step S1394.

[0168] Step S1393: Determine the third secondary candidate bucket as the first secondary candidate bucket.

[0169] Specifically, when the number of third candidate buckets is equal to the preset number, the third secondary candidate bucket is used as the first secondary candidate bucket, that is, the candidate bucket with the largest number of free subslots among the multiple candidate buckets corresponding to the secondary hash bucket is used as the first secondary candidate bucket.

[0170] Step S1394: Determine the fourth candidate bucket based on the principle of the largest number of free subslots in the domain segment.

[0171] In an embodiment of the present application, the hash table includes multiple domain segments, and each domain segment includes multiple hash buckets.

[0172] Specifically, when the number of the third alternative buckets is greater than the preset number, the fourth alternative bucket is determined based on the principle of the largest number of free subslots in the domain segment. The fourth alternative bucket is an alternative bucket in the third alternative bucket, and the number of free subslots in the domain segment where the fourth alternative bucket is located is greater than the number of free subslots in the domain segment where the alternative bucket in the third alternative bucket is located.

[0173] For example, there are three third candidate buckets, and the domain segments where the three third candidate buckets are located are domain segment 1, domain segment 2, and domain segment 3 respectively. The number of free subslots in domain segment 1 is 7, the number of free subslots in domain segment 2 is 8, and the number of free subslots in domain segment 3 is 8. Then the third candidate buckets in domain segments 2 and 3 are used as the fourth candidate buckets. At this time, there is more than one candidate bucket, and it is necessary to further select an alternative bucket from multiple fourth candidate buckets.

[0174] In the embodiment of the present application, the more free sub-slots a domain segment has, the lighter the domain segment load is.

[0175] Step S1395: Determine whether the number of the fourth candidate barrels is greater than a preset number.

[0176] Specifically, determine whether the number of the fourth candidate buckets is greater than the preset number. If the number of the fourth candidate buckets is equal to the preset number, jump to step 1396. If the number of the fourth candidate buckets is greater than the preset number, jump to step 1397.

[0177] Step S1396: Determine the fourth secondary candidate bucket as the first secondary candidate bucket.

[0178] Specifically, when the number of the fourth secondary candidate buckets is equal to the preset number, the fourth secondary candidate bucket will be used as the first secondary candidate bucket, that is, the fourth secondary candidate bucket will be used as the first secondary candidate bucket, that is, the candidate bucket with the largest number of free subslots in the domain segment where the multiple candidate buckets corresponding to the secondary hash bucket are located will be used as the first secondary candidate bucket.

[0179] Step S1397: According to the tail priority principle, determine the fifth candidate bucket, and use the fifth candidate bucket as the first candidate bucket.

[0180] In the embodiment of the present application, each hash bucket in the hash table corresponds to a serial number one by one.

[0181] Among them, the tail priority principle means that the candidate bucket with the largest sequence number is selected first.

[0182] Specifically, according to the tail priority principle, the candidate bucket with the largest sequence number in the fourth candidate bucket is selected as the fifth candidate bucket, and the fifth candidate bucket is used as the first candidate bucket.

[0183] For example, there are two fourth candidate buckets with serial numbers 12 and 18 respectively, then the fourth candidate bucket with serial number 18 is selected as the fifth candidate bucket.

[0184] Step S1398: Use the subslot in the secondary hash bucket corresponding to the first candidate bucket as the first elimination subslot.

[0185] Specifically, the subslot in the secondary hash bucket corresponding to the first candidate bucket is used as the first elimination subslot.

[0186] For example, if the first candidate bucket is the candidate bucket corresponding to the first sub-slot in the secondary hash bucket, the first sub-slot in the secondary hash bucket is used as the first eliminated sub-slot.

[0187] In the embodiment of the present application, the first elimination subslot and the first alternative bucket are determined, and the elements in the first elimination subslot are moved to the free subslot of the first alternative bucket, and the first signature value and the first address pointer are added to the first elimination subslot.

[0188] See also Figure 6 , Figure 6 This is another example schematic diagram of inserting a first signature value and a first address pointer into a hash table provided in an embodiment of the present application.

[0189] like Figure 6 As shown, the Figure 6In the case where there are no free subslots in the primary hash bucket and the secondary hash bucket, and there are free subslots in the candidate bucket corresponding to the subslot of the primary hash bucket, the first signature value and the first address pointer are inserted into an example diagram of the hash table. Assume that there are 5 domains in the hash table, domain 0 to domain 4, each domain includes four hash buckets, and a hash bucket (bucket) has 4 subslots, a total of 80 slots, 20 buckets, bucket0 to bucket19, the maximum load of each domain (i.e., the number of subslots) is 16 slots, H (head) represents the head of the hash bucket, T (tail) represents the tail of the hash bucket, after obtaining the first signature value and the first address pointer, after hash calculation, the primary hash bucket H1 is obtained as bucket2, the secondary hash bucket H2 is obtained as bucket9, and there are no free subslots in bucket2 and bucket9, then the primary candidate bucket corresponding to the subslot of the primary hash bucket H1 (i.e., bucket2) is obtained, and the candidate bucket corresponding to bucket2 has free subslots, and bucket2 The candidate buckets corresponding to the subslots {slot3, slot2, slot1, slot0} are {bucket10, bucket8, bucket5, bucket4} respectively, among which the bucket loads of {bucket10, bucket8, bucket5, bucket4} are {1, 1, 1, 0} respectively, and the domain loads of {bucket10, bucket8, bucket5, bucket4} are {7, 7, 1, 1} respectively. The smaller the bucket load, the more free subslots there are in the hash bucket, and the smaller the domain load, the more free subslots there are in the domain. According to the priority principle of new elements, the hash bucket with the smallest bucket load is selected first. Bucket4 has the smallest bucket load. The subslot of bucket2 corresponding to bucket4 is slot0. Then slot0 is used as the removed subslot, and the elements of slot0 are moved to the free subslot of bucket4. Then the first signature value and the first address pointer are inserted into slot0 of bucket2.

[0190] Step S1310: According to the priority principle of culling elements, determine the corresponding second culling sub-slot and the second candidate bucket in the secondary candidate bucket.

[0191] Specifically, when there are no free subslots in the main hash bucket, the secondary hash bucket, the candidate bucket corresponding to the main hash bucket, and the candidate bucket corresponding to the secondary hash bucket, the corresponding second elimination subslot and the second candidate bucket in the secondary candidate bucket are determined according to the elimination element priority principle.

[0192] Among them, the element removal priority principle includes the tail-first removal principle and the head-first removal principle. The tail-first removal principle is to give priority to the subslot at the tail of the hash bucket to remove the subslot at the tail of the hash bucket. The head-first removal principle is to give priority to the subslot at the head of the hash bucket to remove the subslot at the head of the hash bucket.

[0193] See also Figure 7 , Figure 7 yes Figure 3 Detailed flowchart of step S1310 in .

[0194] like Figure 7 As shown, step S1310 includes:

[0195] Step S13111: Determine whether the element culling priority principle is the tail-first culling principle.

[0196] Specifically, when the priority principle for removing elements is not the tail-first removal principle, jump to step S13112; when the priority principle for removing elements is the tail-first removal principle, jump to step S13113.

[0197] Step S13112: Use the first sub-slot at the head of the secondary hash bucket as the second exclusion sub-slot.

[0198] The second elimination sub-slot is a sub-slot in the secondary hash bucket.

[0199] Specifically, when the element removal priority principle is the head first removal principle, the first subslot at the head of the secondary hash bucket is used as the second removal subslot, for example, the second removal subslot kickout_slot=TailSlot(sec_bucket)=sec_slot3, sec_bucket is the secondary hash bucket.

[0200] Step S13113: Use the first sub-slot at the tail of the secondary hash bucket as the second exclusion sub-slot.

[0201] Specifically, when the element removal priority principle is the tail first removal principle, the first subslot at the tail of the secondary hash bucket is used as the first removal subslot, for example, the first removal subslot kickout_slot=HeadSlot(sec_bucket)=sec_slot0, sec_bucket is the secondary hash bucket.

[0202] Step S13114: Use the candidate bucket corresponding to the second rejection sub-slot as the second candidate bucket.

[0203] Specifically, the candidate bucket corresponding to the second rejection sub-slot is used as the second candidate bucket.

[0204] Step S1311: All subslots in the secondary candidate hash bucket are cyclically shifted in sequence, covering the candidate hash bucket, and the number of eliminations is increased by one.

[0205] Specifically, all subslots in the secondary candidate hash bucket are shifted in sequence, covering the candidate hash bucket, and the number of eliminations is increased by one, that is, all subslots in the secondary hash bucket are moved by one position based on the second elimination subslot to obtain a new secondary hash bucket, and the new secondary hash bucket covers the original secondary hash bucket, and the elimination of the current secondary hash bucket is increased by 1. The secondary candidate hash bucket is the secondary hash bucket.

[0206] See also Figure 8 , Figure 8 yes Figure 3 A detailed flowchart of step S1311 in .

[0207] like Figure 8 As shown, step S1311 includes:

[0208] Step S13111: all subslots in the secondary candidate hash bucket are sequentially moved one position toward the tail of the secondary candidate hash bucket until all subslots in the secondary candidate hash bucket are shifted.

[0209] Specifically, all subslots in the secondary candidate hash bucket are moved one position to the tail of the secondary candidate hash bucket in sequence until all subslots in the secondary candidate hash bucket are shifted, that is, all subslots in the secondary hash bucket are moved one position to the tail of the secondary hash bucket in sequence until all subslots in the secondary hash bucket are shifted.

[0210] The subslot shift is achieved through CyclicShift().

[0211] For example, the secondary hash bucket sec_bucket has four sub-slots {slot3, slot2, slot1, slot0}, and the tail-first culling principle is selected, then slot3 is the culled sub-slot, and the secondary hash bucket after shifting is CyclicShift(sec_bucket)={slot2, sec_slot1, slot0, slot3}.

[0212] In the embodiment of the present application, when all the sub-slots in the secondary hash bucket are sequentially cyclically shifted, the second elimination sub-slot may also be moved toward the head of the secondary hash bucket.

[0213] Step S1312: Replace the second candidate bucket with the secondary candidate bucket.

[0214] Among them, the second candidate bucket is the candidate bucket corresponding to the second rejection sub-slot

[0215] Specifically, if all candidate buckets corresponding to the subslots in the secondary hash bucket have no free subslots, the second candidate bucket is used as the secondary candidate bucket, that is, the second candidate bucket is determined not to be a candidate bucket that currently needs to determine to remove the subslots.

[0216] Specifically, obtain the alternative buckets corresponding to all subslots of the second alternative bucket, determine whether there are free subslots in the alternative buckets corresponding to all subslots of the second alternative bucket, if there are free subslots, select a subslot from the second alternative bucket as a rejection subslot according to the rejection element priority principle, move the elements in the rejection subslot in the second alternative bucket to the alternative bucket corresponding to the rejection subslot in the second alternative bucket, move the elements in the second rejection subslot to the rejection subslot in the second alternative bucket, move the first rejection subslot to the second rejection subslot, and add the first signature value and the first address pointer to a rejection subslot.

[0217] For example, see Fig. 9 , Fig. 9 is another example schematic diagram of inserting the first signature value and the first address pointer into the hash table provided by an embodiment of the present application, such as Fig. 9 As shown, the Fig. 9In the case that there are no free subslots in the main hash bucket, the secondary hash bucket, the candidate bucket corresponding to the subslot of the main hash bucket, and the candidate bucket corresponding to the subslot of the secondary hash bucket, the first signature value and the first address pointer are inserted into an example diagram of the hash table. Assume that the hash table has 5 domains, domains 0 to 4, each domain includes four hash buckets, and a hash bucket (bucket) has 4 subslots (slots), a total of 80 slots, 20 buckets, bucket0 to bucket19, the maximum load of each domain (i.e., the number of subslots) is 16 slots, H (head) represents the head of the hash bucket, T (tail) represents the tail of the hash bucket, after obtaining the first signature value and the first address pointer, Through hash calculation, we get that the main hash bucket H1 is bucket2, and the secondary hash bucket H2 is bucket9. There are no free subslots in bucket2 and bucket9. Then we get the candidate buckets corresponding to the subslots of the main hash bucket H1 (i.e. bucket2). The candidate buckets corresponding to all the subslots of bucket2 {slot3, slot2, slot1, slot0} are {bucket10, bucket8, bucket5, bucket4} respectively. There are no free subslots in bucket10, bucket8, bucket5, and bucket4. We further get the candidate buckets corresponding to the subslots of the secondary hash bucket H1 (i.e. bucket9). The candidate buckets corresponding to all subslots {slot3, slot2, slot1, slot0} of t9 are {bucket2, bucket12, bucket15, bucket19} respectively. There are no free subslots in bucket2, bucket12, bucket15, and bucket19. Then, according to the priority principle of element removal, the subslot corresponding to the secondary hash bucket is determined. For example, if the priority principle of element removal is the tail-first removal principle, the subslot of the secondary hash bucket H1 (i.e., bucket9) is determined to be slot3. All subslots in bucket9 are cyclically shifted in sequence. After the shift, the order of all subslots in bucket9 is {slot2, s lot1, slot0, slot3}, re-obtain the candidate bucket corresponding to slot3, the candidate bucket corresponding to slot3 in bucket9 after the shift is bucket19, bucket19 has no free subslots, obtain the candidate buckets corresponding to each subslot of bucket19 after the shift, and find that the candidate bucket corresponding to the subslot of bucket19 has free subslots. The candidate buckets corresponding to each subslot of bucket19 after the shift are {bucket13, bucket11, bucket18, bucket7}, bucket13, bucket11, bucket18, bucket7 all have free subslots, then according to the priority principle of the newly added elements,Determine that the bucket loads of bucket13, bucket11, bucket18, and bucket7 are {1, 1, 0, 0}, and the domain loads are {9, 13, 7, 9}. Select bucket18 and bucket7 based on the principle of the largest number of free subslots (i.e., the smaller the bucket load), and select bucket18 based on the second priority principle (i.e., the smaller the domain load). Use the subslot slot1 of bucket19 corresponding to bucket18 as the removed subslot, move the elements of slot1 of bucket19 to the free subslot of bucket18, move the elements of slot3 in the shifted bucket9 to slot1 of bucket19, and move the first signature value and the first address pointer to slot3 in the shifted bucket9.

[0218] In an embodiment of the present application, free subslots are found by using the priority principles of removed elements and new elements to avoid loops in the removal path, thereby greatly improving the efficiency of each removal path and reducing the delay in adding elements.

[0219] In an embodiment of the present application, when a conflict occurs when inserting an element, it is necessary to find the next available subslot in the hash table to store the data. If the number of times a subslot is eliminated is not limited, the subslot will be repeatedly searched and eliminated in the hash table, resulting in an unavailable subslot to store the data, thus forming an infinite loop. Therefore, when determining to eliminate a subslot, it is also necessary to determine whether the number of times the currently selected subslot exceeds the maximum number of eliminations to avoid an infinite loop when inserting elements.

[0220] Specifically, before shifting the subslot in the hash bucket, it is necessary to determine whether the number of times the subslot is eliminated is greater than or equal to the elimination number threshold. If the number of times the subslot is eliminated is less than the elimination number threshold, and there is no free subslot in the currently selected alternative bucket, the current alternative is used as the secondary hash bucket, and jump to step S1308 to continue looking for free subslots.

[0221] For example, the second candidate bucket has no free subslots. In order to find an empty slot, the second candidate bucket is used as the candidate bucket that currently needs to determine the subslot to be removed, that is, the second candidate bucket is used as the secondary hash bucket in step S1308 to cyclically search for free subslots until a free subslot is found.

[0222] It should be noted that the secondary hash bucket refers to the currently selected hash bucket. For example, if there is no subslot in the second candidate bucket corresponding to the second subslot, it is necessary to find out whether there are any free subslots in the candidate buckets corresponding to all subslots in the second candidate bucket. The second candidate bucket is then used as the currently selected candidate bucket, that is, the second candidate bucket is used as the secondary hash bucket.

[0223] In an embodiment of the present application, a subslot is selected from the secondary hash bucket as the first elimination subslot according to the element elimination priority principle, the subslots in the secondary hash bucket are cyclically shifted, and a subslot is determined in the alternative bucket corresponding to the subslot of the shifted secondary hash bucket as the second elimination subslot, and the elements of the second elimination subslot are moved to an idle position, and the elements of the first elimination subslot are overwritten to the second elimination subslot, thereby making room for the elements and inserting them into the first elimination subslot. By cyclically shifting and reallocating elements, the space utilization of the hash table can be optimized to avoid the entire hash table from being unable to continue working due to the inability to insert elements into a certain subslot.

[0224] In addition, according to the priority principles of new elements and removed elements, selecting alternative buckets and removing subslots can reduce the impact on frequently accessed elements and avoid the problem of dead loops caused by frequent access to a subslot, so as to reduce the problem of reduced efficiency of hash table data insertion due to conflict processing.

[0225] Step S104: Obtain a first address pointer, where the first address pointer is used to store the address of a subslot of the second-level table.

[0226] In an embodiment of the present application, the hash bucket includes a second-level table, the second-level table includes multiple sub-slots, and the second-level table is used to store key-value pairs.

[0227] In an embodiment of the present application, each subslot of the first-level table is used to store a status bit St, a signature value Sig, and a first address pointer Ptr. The status bit St is used to indicate the status of the current subslot. When there is data, the subslot is in a valid state. When there is no data, the subslot is in an idle state.

[0228] Specifically, after obtaining the free sub-slot in the first-level table, a first address pointer is obtained, wherein the first address pointer is used to store a sub-slot address of the second-level table.

[0229] In the embodiment of the present application, each subslot in the second-level table corresponds to an address pointer.

[0230] Step S105: Store the first address pointer in an idle subslot of the hash bucket of the expanded first-level table.

[0231] Specifically, the first address pointer is stored in an idle subslot of the hash bucket of the expanded first-level table.

[0232] Step S106: According to the first address pointer, the first keyword and the first data item are stored in the second-level table.

[0233] Specifically, according to the subslot address of the second-level table pointed to by the first address pointer, the first keyword and the first data item are stored in the subslot of the second-level table.

[0234] See 10, Fig.10 This is a schematic diagram of an example structure of a hash table provided in an embodiment of the present application.

[0235] like Fig.10 As shown, the hash table includes a first-level table and a second-level table. The first-level table includes multiple hash buckets, each hash bucket includes multiple subslots. Assuming that the maximum depth D of the hash table is 16K and the load factor is 80%, the depth of the hash subslot in the first-level table is 20K. When the ratio of hash buckets to hash subslots is selected as 1:5, the depth of the hash bucket is 4K.

[0236] The second-level table includes multiple hash buckets, each hash bucket includes multiple subslots, and each subslot in the second-level table corresponds to a pointer. When no data is inserted into the second-level table, the pointer corresponding to the subslot is a null pointer, such as null pointers 0, 1, 2...D.

[0237] See 11, Fig.11 This is a schematic diagram of an example structure of a sub-slot in a first-level table provided in an embodiment of the present application.

[0238] like Fig.11 As shown, each subslot in the first-level table includes a state bit st, a signature value sig, and an address pointer ptr. For example, the state bit in subslot 0 is st0, the signature value is sig0, and the address pointer is ptr0.

[0239] In an embodiment of the present application, when inserting a key-value pair, it is necessary to modify the state of the status bit in the first-level table, including the valid state and the idle state, and add the calculated signature value to the subslot, and obtain an address pointer pointing to the subslot of the second-level table and store it in the subslot of the first-level table.

[0240] Please refer to 12, Fig.12 This is an example schematic diagram of inserting an element into a hash table provided by an embodiment of the present application.

[0241] like Fig.10 As shown, an element K needs to be inserted at present. According to the element K, the address and signature value of the main hash bucket H1 and the secondary hash bucket H2 are obtained, and a free subslot is found. Assuming that there is a free subslot in the main hash bucket H1, an address pointer is obtained, and the signature value and address pointer corresponding to the element K are inserted into the free subslot of the main hash bucket H1, and according to the address pointer, the element K is inserted into the subslot of the second-level table pointed to by the address pointer.

[0242] See also Fig.13 , Fig.13 It is a flowchart of a data search provided in an embodiment of the present application.

[0243] like Fig.13As shown in FIG. 1 , the data search process includes:

[0244] Step S1301: Obtain the second keyword.

[0245] The second keyword is used to obtain the address of the hash bucket in the first-level table.

[0246] Step S1302: perform hash calculation on the second keyword to obtain the address of a hash bucket of the expanded first-level table and the second signature value.

[0247] Specifically, a hash calculation is performed on the second keyword to obtain an address of a hash bucket of the first-level table and a second signature value.

[0248] Step S1303: Based on the address of the hash bucket, read the third signature value and the second address pointer of the hash bucket, wherein the second address pointer is used to store a subslot address of the second-level table.

[0249] Specifically, based on the address of the hash bucket of the first-level table, the second signature value and the second address pointer are obtained from the subslot of the hash bucket of the first-level table.

[0250] The second address pointer is used to store the subslot address of the second-level table.

[0251] Step S1304: Determine whether the second signature value is the same as the third signature value.

[0252] Specifically, determine whether the second signature value is the same as the third signature value. If the second signature value is the same as the third signature value, jump to step S1305. If the second signature value is not the same as the third signature value, end the current data acquisition operation.

[0253] Step S1305: Based on the second address pointer, obtain the second data item of the subslot of the second-level table.

[0254] Specifically, when the second signature value is the same as the third signature value, the second data item of the subslot of the second-level table is obtained based on the second address pointer. The second data item is the data item stored in the subslot of the second-level table pointed to by the second address pointer.

[0255] See also Fig.14 , Fig.14 It is a flowchart of a data update provided in an embodiment of the present application.

[0256] like Fig.14 As shown, the data update process includes:

[0257] Step S1401: Acquire a second key-value pair, wherein the key-value pair includes a third keyword and a third data item.

[0258] The second key-value pair is a key-value pair to be updated in the hash table.

[0259] Step S1402: perform hash calculation on the third keyword to obtain an address of a hash bucket of the expanded first-level table and a fourth signature value.

[0260] Specifically, a hash calculation is performed on the third keyword to obtain an address of a hash bucket of the expanded first-level table and a fourth signature value.

[0261] Step S1403: Based on the address of the hash bucket, read the fifth signature value and the third address pointer of the hash bucket, wherein the third address pointer is used to store a subslot address of the second-level table.

[0262] Specifically, based on the address of the hash bucket of the first-level table, the fifth signature value and the third address pointer in the subslot of the hash bucket of the first-level table are read, wherein the third address pointer is used to store a subslot address of the second-level table.

[0263] Step S1404: Determine whether the fourth signature value is the same as the fifth signature value.

[0264] Specifically, determine whether the fourth signature value is the same as the fifth signature value. If the fourth signature value is the same as the fifth signature value, jump to step S1405. If the fourth signature value is not the same as the fifth signature value, end the current data update operation.

[0265] Step S1405: Determine whether the third keyword is the same as the fourth keyword.

[0266] Specifically, when the third signature value is the same as the fourth signature value, based on the third address pointer, the fourth keyword of the subslot of the second-level table is obtained to determine whether the third keyword is the same as the fourth keyword. If the third keyword is the same as the fourth keyword, jump to step S1206; if the third keyword is not the same as the fourth keyword, end the current data update operation.

[0267] Step S1406: Update the third data item to the subslot of the second-level table pointed to by the third address pointer.

[0268] Specifically, when the third keyword is the same as the fourth keyword, the third data item is updated to the subslot of the second-level table pointed to by the third address pointer.

[0269] See also Fig.15 , Fig.15 It is a structural diagram of a hash table data insertion system provided in an embodiment of the present application.

[0270] like Fig.15As shown, the data insertion system 1500 of the hash table includes: a hash table control unit 1501, a hash calculation unit 1502, a pointer control unit 1503, a first-level table storage unit 1504, and a second-level table storage unit 1505.

[0271] The hash table control unit 1501 is used to perform depth expansion on the first-level table to obtain an expanded first-level table.

[0272] The hash calculation unit 1502 is used to obtain a key-value pair, where the key-value pair includes a keyword and a data item, and obtain a free subslot of a hash bucket of the expanded first-level table according to the keyword.

[0273] The pointer control unit 1503 is used to obtain the address of a subslot of the second-level table.

[0274] The first-level table storage unit 1504 is used to store the address pointer of the second-level table, wherein the address pointer is used to store the address of a subslot of the second-level table.

[0275] The second-level table storage unit 1505 is used to store key-value pairs.

[0276] In the embodiment of the present application, the units of the data insertion system of the hash table cooperate with each other to improve the utilization rate of the hash table.

[0277] See also Fig.16 , Fig.16 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application.

[0278] like Fig.16 As shown, the electronic device 1600 includes one or more processors 1601 and a memory 1602. Fig.16 A processor 1601 is taken as an example.

[0279] The processor 1601 and the memory 1602 may be connected via a bus or other means. Fig.16 The example of connecting through bus is taken in the following.

[0280] Processor 1601 is used to provide computing and control capabilities to control electronic device 1600 to perform corresponding tasks, for example, to control electronic device 1600 to perform a data insertion method for a hash table in any one of the above-mentioned method embodiments, the method comprising: deeply expanding a first-level table to obtain an expanded first-level table, obtaining a first key-value pair, wherein the key-value pair comprises a first keyword and a first data item, obtaining a free subslot of a hash bucket of the expanded first-level table according to the first keyword, obtaining a first address pointer, wherein the first address pointer is used to store a subslot address of a second-level table, storing the first address pointer in a free subslot of the hash bucket of the expanded first-level table, and storing the first keyword and the first data item in the second-level table according to the first address pointer.

[0281] The processor 1601 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), a hardware chip or any combination thereof; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a programmable logic device (PLD) or a combination thereof. The above-mentioned PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL) or any combination thereof.

[0282] The memory 1602, as a non-transient computer-readable storage medium, can be used to store non-transient software programs, non-transient computer executable programs and modules, such as program instructions / modules corresponding to the data insertion method of the hash table in the embodiment of the present application. The processor 1601 can implement the data insertion method of the hash table in any of the above method embodiments by running the non-transient software programs, instructions and modules stored in the memory 1602. Specifically, the memory 1602 may include a volatile memory (volatile memory, VM), such as a random access memory (random access memory, RAM); the memory 1602 may also include a non-volatile memory (non-volatile memory, NVM), such as a read-only memory (read-only memory, ROM), a flash memory (flash memory), a hard disk (hard diskdrive, HDD) or a solid-state drive (solid-state drive, SSD) or other non-transient solid-state storage devices; the memory 1602 may also include a combination of the above types of memories.

[0283] The memory 1602 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 1602 may optionally include a memory remotely arranged relative to the processor 1601, and these remote memories may be connected to the processor 1601 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0284] One or more modules are stored in the memory 1602, and when executed by one or more processors 1601, the data insertion method of the hash table in any of the above method embodiments is executed, for example, the above described method is executed. Figure 1 The steps shown.

[0285] In the embodiment of the present application, the electronic device 1600 may also have components such as a wired or wireless network interface, a keyboard, and an input / output interface for input and output. The electronic device 1600 may also include other components for realizing device functions, which will not be elaborated here.

[0286] The embodiment of the present application also provides a non-volatile computer-readable storage medium, such as a memory including a program code, and the program code can be executed by a processor to complete the data insertion method of the hash table in the above embodiment. For example, the non-volatile computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CDROM), a magnetic tape, a floppy disk, and an optical data storage device.

[0287] The embodiment of the present application also provides a computer program product, which includes one or more program codes, and the program codes are stored in a non-volatile computer-readable storage medium. The processor of the flash memory device reads the program code from the non-volatile computer-readable storage medium, and the processor executes the program code to complete the method steps of the data insertion method of the hash table provided in the above embodiment.

[0288] A person skilled in the art will appreciate that all or part of the steps for implementing the above embodiments may be accomplished by hardware or by hardware associated with a program code, and the program may be stored in a non-volatile computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.

[0289] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a general hardware platform, and of course, by hardware. Based on this understanding, the above technical solution can be essentially or in other words, the part that contributes to the relevant technology can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute various embodiments or certain parts of the embodiments.

[0290] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Under the concept of the present application, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other changes in different aspects of the present application as above, which are not provided in detail for the sake of simplicity. Although the present application has been described in detail with reference to the aforementioned embodiments, a person of ordinary skill in the art should understand that the technical solutions described in the aforementioned embodiments can still be modified, or some of the technical features can be replaced by equivalents. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for inserting data into a hash table, characterized in that: The hash table includes a first-level table and a second-level table, the first-level table includes a plurality of hash buckets, the hash buckets include a plurality of sub-slots, the second-level table includes a plurality of sub-slots, and the method includes: Performing depth expansion on the first-level table to obtain an expanded first-level table; Obtaining a first key-value pair, wherein the key-value pair includes a first keyword and a first data item; According to the first keyword, obtaining a free subslot of a hash bucket in the extended first-level table and a first signature value corresponding to the first keyword; Obtaining a first address pointer, wherein the first address pointer is used to store an address of a subslot of the second-level table; The first signature value and the first address pointer corresponding to the first keyword are stored in the free subslot of the hash bucket of the expanded first-level table; According to the first address pointer, the first keyword and the first data item are stored in the second level table.

2. The method according to claim 1, characterized in that The deep expansion of the first-level table includes: Get the load factor; Get the maximum depth of the hash table; Based on the load factor and the maximum depth of the hash table, the depth of the first-level table is extended, wherein the depth of the first-level table=(1 / σ)*D, σ is the load factor, and D is the maximum depth of the hash table.

3. The method according to claim 1, characterized in that The hash bucket includes a primary hash bucket and a secondary hash bucket, and obtaining, according to the first keyword, a free subslot of a hash bucket of the extended first-level table includes: According to the first keyword, an address of a primary hash bucket and an address of a secondary hash bucket are obtained, a primary hash bucket is obtained according to the primary hash bucket address, and a secondary hash bucket is obtained according to the secondary hash bucket address, wherein the primary hash bucket and the secondary hash bucket each include a plurality of subslots, each subslot corresponds to an alternative bucket, and the subslot is used to store the first signature value; When none of the primary hash bucket, the secondary hash bucket, and the candidate bucket corresponding to the primary hash bucket has an empty subslot, determining whether there is an empty slot in the candidate bucket corresponding to the secondary hash bucket; If there is an idle subslot in the candidate bucket corresponding to the secondary candidate bucket, determine the first rejection subslot and the first candidate bucket according to the priority principle of the newly added elements, wherein the first rejection subslot is one of the multiple subslots in the secondary candidate bucket, and the first candidate bucket is the candidate bucket corresponding to the first rejection subslot; Adding the elements of the first eliminated sub-slot into the first candidate bucket, and using the first eliminated sub-slot as the sub-slot for storing the first signature value; If there is no free subslot in the candidate bucket corresponding to the secondary candidate bucket, determine the second elimination subslot and the second candidate bucket corresponding to the secondary candidate bucket according to the elimination element priority principle, wherein the second elimination subslot is a subslot in the secondary candidate hash bucket, and the second candidate bucket is the candidate bucket corresponding to the second elimination subslot; After all subslots in the secondary candidate hash bucket are cyclically shifted in sequence, the candidate hash bucket is covered, and the number of eliminations is increased by one; The second candidate bucket is replaced with the secondary candidate bucket, and it is determined whether there is an empty slot in the secondary candidate bucket until there is an empty sub-slot in the candidate bucket corresponding to the secondary candidate bucket.

4. The method according to claim 3, characterized in that The hash table includes a plurality of domain segments, each of which includes a plurality of hash buckets, and the priority principle of the newly added elements includes a principle of the largest number of free subslots of the hash bucket, a principle of the largest number of free subslots of the domain segment, and a tail priority principle; The step of determining the first elimination subslot and the first candidate bucket according to the priority principle of the newly added element includes: According to the principle that the number of free sub-slots of the hash bucket is the largest, a third candidate bucket is determined, wherein the third candidate bucket is the candidate bucket with the largest number of free sub-slots among all the candidate buckets corresponding to the secondary hash bucket; If the number of the third candidate buckets is greater than the preset number, a fourth candidate bucket is determined according to the principle that the number of free subslots in the domain segment is the largest, wherein the fourth candidate bucket is a candidate bucket in the third candidate bucket; If the number of the fourth candidate buckets is greater than the preset number, a fifth candidate bucket is determined according to the tail priority principle, and the fifth candidate bucket is used as the first candidate bucket, wherein the fifth candidate bucket is one of the plurality of fourth candidate buckets; The subslot in the secondary hash bucket corresponding to the first candidate bucket is used as the first eliminated subslot.

5. The method according to claim 4, characterized in that The element removal priority principle includes a tail-first removal principle or a head-first removal principle; According to the priority principle of the removed elements, determining the corresponding second removal sub-slot and the second candidate bucket in the secondary candidate bucket includes: If the element removal priority principle is the tail-first removal principle, the first subslot at the tail of the secondary hash bucket is used as the second removal subslot; If the element removal priority principle is the header first removal principle, the first subslot of the header of the secondary hash bucket is used as the second removal subslot; The candidate bucket corresponding to the second rejection sub-slot is used as the second candidate bucket.

6. The method according to claim 4, characterized in that The sequentially cyclically shifting all subslots in the secondary candidate hash bucket includes: All subslots in the secondary candidate hash bucket are sequentially moved one position toward the tail of the secondary candidate hash bucket until all subslots in the secondary candidate hash bucket are completely shifted.

7. The method according to claim 4, characterized in that The method further comprises: If the number of the third candidate buckets is equal to the preset number, the third candidate bucket is used as the first candidate bucket; If the number of the fourth candidate buckets is equal to the preset number, the fourth candidate buckets are used as the first candidate buckets.

8. The method according to claim 3, characterized in that The method further comprises: If the primary hash bucket has an idle subslot and the secondary hash bucket has no idle subslot, the first signature value and the first address pointer are stored in the idle subslot of the primary hash bucket; If the primary hash bucket does not have an idle subslot, and the secondary hash bucket has an idle subslot, storing the first signature value and the first address pointer in the idle subslot of the secondary hash bucket; If both the primary hash bucket and the secondary hash bucket have free subslots, determine a hash bucket according to a priority principle of newly added elements, and insert the first signature value and the first address pointer into the hash bucket, wherein the hash bucket is one of the primary hash bucket and the secondary hash bucket; If there is no free subslot in the primary hash bucket and no free subslot in the secondary hash bucket, and there is a free subslot in the primary candidate bucket corresponding to the subslot of the primary hash bucket, inserting the first signature value and the first address pointer into the free subslot of the primary candidate bucket corresponding to the subslot of the primary hash bucket; If there is no free subslot in the main hash bucket, and there is no free subslot in the secondary hash bucket, there is no free subslot in the main candidate bucket corresponding to the subslot of the main hash bucket, and there is a free subslot in the secondary candidate bucket corresponding to the subslot of the secondary hash bucket, then the first signature value and the first address pointer are inserted into the free subslot of the secondary candidate bucket corresponding to the subslot of the secondary hash bucket.

9. The method according to claim 1, characterized in that: The method further comprises: Get the second keyword; Performing hash calculation on the second keyword to obtain an address of a hash bucket of the expanded first-level table and a second signature value; Based on the address of the hash bucket, read a third signature value and a second address pointer in the hash bucket, wherein the second address pointer is used to store a subslot address of the second-level table; If the second signature value is the same as the third signature value, a second data item of the subslot of the second-level table is obtained based on the second address pointer.

10. The method according to claim 9, characterized in that The method further comprises: Obtaining a second key-value pair, wherein the key-value pair includes a third keyword and a third data item; Performing hash calculation on the third keyword to obtain an address of a hash bucket of the expanded first-level table and a fourth signature value; Based on the address of the hash bucket, read a fifth signature value and a third address pointer in the hash bucket, wherein the third address pointer is used to store a subslot address of the second-level table; If the fourth signature value is the same as the fifth signature value, obtaining the fourth keyword of the subslot of the second-level table based on the third address pointer; If the third keyword is the same as the fourth keyword, the third data item is updated to the subslot of the second-level table pointed to by the third address pointer.

11. The method according to claim 10, characterized in that The method further comprises: If the second signature value is different from the third signature value, then the current operation of obtaining the second data item is terminated; If the fourth signature value is the same as the fifth signature value, and the third keyword is different from the fourth keyword, then the current operation of updating the third data item to the subslot of the second-level table is terminated; or, If the fourth signature value is different from the fifth signature value, the current operation of updating the third data item to the subslot of the second-level table is terminated.

12. A data insertion system for a hash table, characterized in that: The hash table includes a first-level table and a second-level table, the first-level table includes a plurality of hash buckets, the hash buckets include a plurality of sub-slots, the second-level table includes a plurality of sub-slots, and the system includes: A hash table control unit, used for performing depth expansion on the first-level table to obtain an expanded first-level table; A hash calculation unit, configured to obtain a first key-value pair, wherein the key-value pair includes a first keyword and a first data item, and obtain, according to the first keyword, a free subslot of a hash bucket in the extended first-level table and a first signature value corresponding to the first keyword; A pointer control unit, used for acquiring a first address pointer, wherein the first address pointer is used for storing an address of a subslot of the second-level table; A first-level table storage unit, used to store a first signature value and a first address pointer corresponding to the first keyword; The second-level table storage unit is used to store the first keyword and the first data item.

13. An electronic device, characterized in that: include: at least one processor, and a memory communicatively coupled to the at least one processor, wherein: The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the data insertion method for the hash table as described in any one of claims 1 to 11 above.