Hash table data insertion method and system and electronic equipment

By introducing a combination scheme of the main hash bucket and the secondary hash bucket into the hash table, combining the culling threshold and priority principles, the loopback kick phenomenon of the cuckoo hash table during data insertion is solved, and data addition efficiency and data table utilization are improved.

CN119988374AActive Publication Date: 2025-05-13DAPUSTOR CORP
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
CN202411994793.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-13
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

Existing cuckoo hash tables are prone to loopback kicking when data is inserted, resulting in low data addition efficiency and low data table utilization.

Method used

Using the combination scheme of the main hash bucket and the secondary hash bucket, the removal sub-slot and alternative bucket are determined by setting the culling times threshold and priority principles, and the sub-slot is shifted cyclically to find the free position to avoid loop kicking.

Benefits of technology

It improves the efficiency of data insertion, improves the utilization rate of data tables, and avoids the vicious cycle caused by loopback kicking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988374A_ABST
    Figure CN119988374A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to the technical field of data processing, and discloses a data insertion method and system of a hash table and electronic equipment, and the method comprises the steps: obtaining a new element, calculating the address of a main hash bucket and the address of an auxiliary hash bucket according to the new element, determining that no empty slot exists in the main hash bucket, the auxiliary hash bucket and a main alternative bucket corresponding to a sub-slot of the main hash bucket, and inserting the data into the main hash bucket according to the empty slot. If yes, whether an empty slot exists in the alternative bucket corresponding to the sub-slot of the auxiliary Hash bucket or not is judged, if yes, according to a newly added element priority principle, a rejected sub-slot in the auxiliary Hash bucket is determined, elements in the rejected sub-slot are added to the alternative bucket corresponding to the rejected sub-slot, new elements are inserted into the rejected sub-slot of the auxiliary Hash bucket, and if not, the rejected sub-slot of the auxiliary Hash bucket is removed. And according to a rejection element priority principle, determining a rejection sub-slot in the auxiliary Hash bucket, performing cyclic shift on the sub-slot in the auxiliary Hash bucket, and then executing a judgment step until an idle sub-slot exists in an alternative bucket corresponding to the auxiliary alternative bucket, so that the data adding efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a method, system and electronic device for inserting data into a hash table. Background Art

[0002] In the fields of computer algorithms, database management and information processing technology, data table utilization is a key indicator for measuring system performance and resource utilization efficiency. Especially in big data processing scenarios, data table utilization is directly related to the scale and speed of data that the system can process.

[0003] At present, cuckoo hashing is usually used to manage data tables. Cuckoo hashing is a data structure that stores key-value pairs by selecting hash buckets and subslots through hash values. However, when a new key-value pair needs to be added, if a full hash bucket is encountered, other hash buckets and subslots need to be randomly replaced. During the replacement process, occupied or previously visited hash buckets and subslots may be encountered, resulting in a loop kick phenomenon, which reduces the efficiency of data addition and the utilization of data tables. Summary of the invention

[0004] The embodiments of the present application provide a method, system and electronic device for inserting data into a hash table, which can improve the efficiency of adding data.

[0005] To solve the above technical problems, the present application provides the following technical solutions:

[0006] In a first aspect, an embodiment of the present application provides a method for inserting data into a hash table, wherein the hash table includes a primary hash bucket and a secondary hash bucket, and the method includes:

[0007] Set the removal count threshold and obtain the new elements to be inserted;

[0008] According to the new element, the address of the primary hash bucket and the address of the secondary hash bucket are obtained, the primary hash bucket is obtained according to the address of the primary hash bucket, and the secondary hash bucket is obtained according to the address of the secondary hash bucket, wherein the primary hash bucket and the secondary hash bucket each include multiple subslots, each subslot corresponds to an alternative bucket, and the subslot is used to store the new element;

[0009] When there are no free subslots in the primary hash bucket, the secondary hash bucket, and the candidate bucket corresponding to the primary hash bucket, it is determined whether there is an empty slot in the candidate bucket corresponding to the secondary hash bucket;

[0010] If there is an idle subslot in the candidate bucket corresponding to the secondary candidate bucket, the first elimination subslot and the first candidate bucket are determined according to the priority principle of the newly added elements, wherein the first elimination subslot is one of the multiple subslots in the secondary candidate bucket, and the first candidate bucket is the candidate bucket corresponding to the first elimination subslot;

[0011] Add the element of the first elimination sub-slot to the first candidate bucket, insert the new element into the first elimination sub-slot, and the addition is successful;

[0012] If there is no free subslot in the candidate bucket corresponding to the secondary candidate bucket, the second elimination subslot and the second candidate bucket corresponding to the secondary candidate bucket are determined according to the elimination element priority principle, wherein the second elimination subslot is a subslot in the secondary candidate hash bucket, and the second candidate bucket is the candidate bucket corresponding to the second elimination subslot;

[0013] After all subslots in the secondary candidate hash bucket are shifted in sequence, they are overwritten with the candidate hash bucket, and the number of removals is increased by one;

[0014] The second candidate bucket is replaced with the secondary candidate bucket, and it is determined whether there is an empty slot in the secondary candidate bucket, until there is an idle sub-slot in the candidate bucket corresponding to the secondary candidate bucket or the number of eliminations is equal to the elimination number threshold.

[0015] In some embodiments, the hash table includes a plurality of domain segments, each domain segment includes a plurality of hash buckets, and the priority principle of the newly added element includes the principle of the largest number of free subslots of the hash bucket, the principle of the largest number of free subslots of the domain segment, and the tail priority principle;

[0016] According to the priority principle of the newly added elements, the first elimination sub-slot and the first candidate bucket are determined, including:

[0017] According to the principle that the number of free sub-slots of the hash bucket is the largest, a third candidate bucket is determined, wherein the third candidate bucket is the candidate bucket with the largest number of free sub-slots among all candidate buckets corresponding to the secondary hash bucket;

[0018] If the number of the third candidate buckets is greater than the preset number, a fourth candidate bucket is determined according to the principle that the number of free subslots in the domain segment is the largest, wherein the fourth candidate bucket is a candidate bucket in the third candidate bucket;

[0019] If the number of the fourth candidate buckets is greater than the preset number, the fifth candidate bucket is determined according to the tail priority principle, and the fifth candidate bucket is used as the first candidate bucket, wherein the fifth candidate bucket is one of the multiple fourth candidate buckets;

[0020] The subslot in the secondary hash bucket corresponding to the first candidate bucket is used as the first elimination subslot.

[0021] In some embodiments, the element culling priority principle includes a tail-first culling principle or a head-first culling principle;

[0022] According to the priority principle of culling elements, the corresponding second culling sub-slot and the second candidate bucket in the secondary candidate bucket are determined, including:

[0023] If the element removal priority principle is the tail first removal principle, the first subslot at the tail of the secondary hash bucket is used as the second removal subslot;

[0024] If the element removal priority principle is the head first removal principle, the first subslot at the head of the secondary hash bucket is used as the second removal subslot;

[0025] The candidate bucket corresponding to the second rejection sub-slot is used as the second candidate bucket.

[0026] In some embodiments, all subslots in the secondary candidate hash bucket are sequentially shifted cyclically, including:

[0027] All subslots in the secondary candidate hash bucket are moved one position toward the tail of the secondary candidate hash bucket in sequence until all subslots in the secondary candidate hash bucket are shifted.

[0028] In some embodiments, the method further comprises:

[0029] If the number of the third candidate buckets is equal to the preset number, the third candidate bucket is used as the first candidate bucket;

[0030] If the number of the fourth candidate buckets is equal to the preset number, the fourth candidate bucket is used as the first candidate bucket.

[0031] In some embodiments, the method further comprises:

[0032] If there is an idle subslot in the primary hash bucket and there is no idle subslot in the secondary hash bucket, insert the new element into the idle subslot in the primary hash bucket;

[0033] If there is no free subslot in the primary hash bucket, but there is a free subslot in the secondary hash bucket, insert the new element into the free subslot of the secondary hash bucket;

[0034] If both the primary hash bucket and the secondary hash bucket have free subslots, the hash bucket is determined according to the priority principle of the newly added element, and the new element is inserted into the hash bucket, where the hash bucket is one of the primary hash bucket and the secondary hash bucket;

[0035] If there is no free subslot in the primary hash bucket and no free subslot in the secondary hash bucket, and there is a free subslot in the primary candidate bucket corresponding to the subslot of the primary hash bucket, then the new element is inserted into the free subslot of the primary candidate bucket corresponding to the subslot of the primary hash bucket;

[0036] If there is no free subslot in the primary hash bucket, and there is no free subslot in the secondary hash bucket, there is no free subslot in the primary candidate bucket corresponding to the subslot of the primary hash bucket, and there is a free subslot in the secondary candidate bucket corresponding to the subslot of the secondary hash bucket, then the new element is inserted into the free subslot of the secondary candidate bucket corresponding to the subslot of the secondary hash bucket.

[0037] In a second aspect, an embodiment of the present application provides a data insertion system for a hash table, wherein the hash table includes a primary hash bucket and a secondary hash bucket, and the system includes:

[0038] A hash bucket determination unit is used to obtain a new element to be inserted, and obtain an address of a primary hash bucket and an address of a secondary hash bucket according to the new element, wherein the primary hash bucket and the secondary hash bucket are used to store data, the primary hash bucket includes a plurality of sub-slots, the secondary hash bucket includes a plurality of sub-slots, each sub-slot in the primary hash bucket corresponds to a primary candidate bucket, and each sub-slot in the secondary hash bucket corresponds to a secondary candidate bucket;

[0039] The culling element priority judgment unit is used to determine, based on the address of the primary hash bucket and the address of the secondary hash bucket, that when there are no idle subslots in the primary hash bucket, the secondary hash bucket, all primary candidate buckets corresponding to the primary hash bucket, and all secondary candidate buckets corresponding to the secondary hash bucket, then determine, according to the culling element priority principle, a first culling subslot corresponding to the secondary hash bucket, wherein the first culling subslot is a subslot in the secondary hash bucket;

[0040] A new element priority judgment unit is added, which is used to circularly shift all subslots in the secondary hash bucket in order based on the first eliminated subslot to obtain the shifted secondary hash bucket, obtain the first candidate bucket corresponding to each subslot in the shifted secondary hash bucket, and if there is an idle subslot in the first candidate bucket, determine the second candidate bucket according to the new element priority principle, wherein the second candidate bucket is one of the first candidate buckets;

[0041] The culling and replacement unit is used to add the elements in the first culling subslot to the free subslots in the second candidate bucket, and overwrite the new elements to the first culling subslot.

[0042] In some embodiments, the system further comprises:

[0043] Hash table, used to store new elements;

[0044] The culled element random access memory is used to record the number of culled times for each subslot in the hash table.

[0045] In a third aspect, an embodiment of the present application provides an electronic device, including:

[0046] at least one processor, and

[0047] a memory communicatively coupled to at least one processor, wherein:

[0048] The memory stores instructions that can be executed by at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method for inserting data into a hash table according to the first aspect.

[0049] The beneficial effects of the embodiment of the present application are as follows: Different from the prior art, the embodiment of the present application provides a method for inserting data into a hash table, the hash table includes a primary hash bucket and a secondary hash bucket, and the method includes: setting a threshold of the number of eliminations, and obtaining a new element to be inserted;

[0050] According to the new element, the address of the main hash bucket and the address of the secondary hash bucket are obtained, the main hash bucket is obtained according to the address of the main hash bucket, and the secondary hash bucket is obtained according to the address of the secondary hash bucket, wherein the main hash bucket and the secondary hash bucket each include multiple subslots, each subslot corresponds to an alternative bucket, and the subslot is used to store the new element. When there are no free subslots in the main hash bucket, the secondary hash bucket, and the alternative bucket corresponding to the main hash bucket, it is determined whether there is an empty slot in the alternative bucket corresponding to the secondary hash bucket. If there is a free subslot in the alternative bucket corresponding to the secondary alternative bucket, the first elimination subslot and the first alternative bucket are determined according to the priority principle of the newly added element, wherein the first elimination subslot is one of the multiple subslots in the secondary alternative bucket, and the first alternative bucket is the alternative bucket corresponding to the first elimination subslot. Bucket, add the elements of the first eliminated subslot to the first alternative bucket, insert the new element into the first eliminated subslot, and add successfully. If there is no free subslot in the alternative bucket corresponding to the secondary alternative bucket, then determine the second eliminated subslot and the second alternative bucket corresponding to the secondary alternative bucket according to the elimination element priority principle, wherein the second eliminated subslot is a subslot in the secondary alternative hash bucket, and the second alternative bucket is the alternative bucket corresponding to the second eliminated subslot. After all the subslots in the secondary alternative hash bucket are cyclically shifted in sequence, the alternative hash bucket is overwritten, the elimination times are increased by one, and the second alternative bucket is replaced with the secondary alternative bucket. It is determined whether there is an empty slot in the secondary alternative bucket until there is a free subslot in the alternative bucket corresponding to the secondary alternative bucket or the elimination times are equal to the elimination times threshold.

[0051] By obtaining the new element, the address of the main hash bucket and the address of the secondary hash bucket are calculated according to the new element, and it is determined that there are no empty slots in each of the main candidate buckets corresponding to the main hash bucket, the secondary hash bucket, and the subslots of the main hash bucket. Then, it is determined whether there are empty slots in each candidate bucket corresponding to the subslots of the secondary hash bucket. If there are empty slots, the elimination subslot in the secondary hash bucket is determined according to the priority principle of the newly added elements, and the elements in the elimination subslot are added to the candidate bucket corresponding to the elimination subslot, and the new element is inserted into the elimination subslot of the secondary hash bucket. If there are no empty slots, the elimination subslot in the secondary hash bucket is determined according to the priority principle of the eliminated elements, and the subslots in the secondary hash bucket are cyclically shifted, and then the judgment step is executed until there are free subslots in the candidate bucket corresponding to the secondary candidate bucket, which can avoid duplication of elimination paths and improve the efficiency of data addition. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] One or more embodiments are exemplarily described by pictures in the corresponding drawings, and these exemplified descriptions do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings represent similar elements, and unless otherwise stated, the figures in the drawings do not constitute proportional limitations.

[0053] Figure 1 It is a flowchart of a method for inserting data into a hash table provided in an embodiment of the present application;

[0054] Figure 2 This is an example schematic diagram of inserting a new element into a hash table provided by an embodiment of the present application;

[0055] Figure 3 yes Figure 1 A detailed flowchart of step S1010 in FIG.

[0056] Figure 4 is another example schematic diagram of inserting a new element into a hash table provided by an embodiment of the present application;

[0057] Figure 5 yes Figure 1 A detailed flowchart of step S1011 in FIG.

[0058] Figure 6 yes Figure 1 A detailed flowchart of step S1012 in FIG.

[0059] Figure 7 is another example schematic diagram of inserting a new element into a hash table provided by an embodiment of the present application;

[0060] Figure 8 It is a structural diagram of a hash table data insertion system provided in an embodiment of the present application;

[0061] Fig. 9 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application.

[0062] Description of Figure Numbers:

[0063] DETAILED DESCRIPTION

[0064] In order to make the purpose, technical solutions and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application.

[0065] It should be noted that, if there is no conflict, the various features in the embodiments of the present application can be combined with each other, all within the scope of protection of the present application. In addition, although the functional module division is performed in the device schematic diagram and the logical order is shown in the flow chart, in some cases, the steps shown or described can be performed in a sequence different from the module division in the device or the flow chart. Furthermore, the words "first", "second", "third", etc. used in this application do not limit the data and execution order, but only distinguish the same items or similar items with basically the same functions and effects.

[0066] In addition, the technical features involved in each embodiment of the present application described below can be combined with each other as long as they do not conflict with each other.

[0067] Before introducing the embodiments of the present application, a brief introduction to the prior art known to the inventors of the present application is first given to facilitate subsequent understanding of the embodiments of the present application.

[0068] Cuckoo hash is a data structure that stores key-value pairs in the form of hash buckets and subslots. When a new key-value pair needs to be added, if a full hash bucket is encountered, a random principle is adopted to randomly select a subslot from the hash bucket for replacement. During the replacement process, occupied or previously visited hash buckets and subslots may be encountered, resulting in a loop kick phenomenon, which reduces the efficiency of data addition and the utilization of the data table.

[0069] In response to the above problems, the present application provides a method for inserting data into a hash table, which obtains a new element and calculates the address of the main hash bucket and the address of the secondary hash bucket according to the new element, determines whether there are empty slots in the main candidate bucket corresponding to the main hash bucket, the secondary hash bucket, and the subslot of the main hash bucket, then determines whether there are empty slots in the candidate bucket corresponding to the subslot of the secondary hash bucket, if there are empty slots, determines the elimination subslot in the secondary hash bucket according to the priority principle of the newly added elements, adds the elements in the elimination subslot to the candidate bucket corresponding to the elimination subslot, inserts the new element into the elimination subslot of the secondary hash bucket, if there are no empty slots, determines the elimination subslot in the secondary hash bucket according to the priority principle of the eliminated elements, performs a circular shift on the subslots in the secondary hash bucket, and then executes the judgment step until there are free subslots in the candidate bucket corresponding to the secondary candidate bucket, thereby improving the efficiency of data addition.

[0070] The technical solution of this application is described in detail below with reference to the accompanying drawings:

[0071] See also Figure 1 , Figure 1 It is a flowchart of a method for inserting data into a hash table provided in an embodiment of the present application.

[0072] The method for inserting data into a hash table is applied to an electronic device. Specifically, the execution subject of the method for inserting data into a hash table is one or at least two processors of the electronic device.

[0073] like Figure 1 As shown, the data insertion method of the hash table includes:

[0074] Step S1001: Set a threshold for the number of eliminations and obtain a new element to be inserted.

[0075] In the embodiment of the present application, a hash table is a data structure used to implement a mapping relationship between key-value pairs. The hash table maps a key to a position in an array through a hash function, thereby realizing fast search, insertion and deletion operations on records.

[0076] The hash table includes multiple hash buckets, and a hash bucket is a container or slot for storing key-value pairs.

[0077] Specifically, a new element to be inserted into the hash table is obtained, where the new element is an element inserted into the hash table for the first time.

[0078] The culling times threshold is the maximum culling times for each sub-slot.

[0079] In the embodiment of the present application, when inserting a new element, it is necessary to find an idle subslot from the hash bucket to store the new element. When the hash bucket has no idle subslot, it is necessary to find an idle subslot from the candidate bucket corresponding to the hash bucket, and select a subslot from the hash bucket for replacement. The more times the subslot is replaced, the lower the data insertion rate of the hash table will be. When no idle subslot is found, it will fall into a loop to find an idle subslot. In order to avoid the loop kick phenomenon, it is necessary to set a culling threshold to limit the number of cullings for each subslot.

[0080] Step S1002: According to the new element, the address of the primary hash bucket and the address of the secondary hash bucket are obtained. The primary hash bucket is obtained according to the primary hash bucket address, and the secondary hash bucket is obtained according to the secondary hash bucket address.

[0081] In the embodiment of the present application, the hash bucket includes a primary hash bucket and a secondary hash bucket, the primary hash bucket includes a plurality of sub-slots, and the secondary hash bucket includes a plurality of sub-slots.

[0082] Specifically, the new element includes a keyword and a data item. The hash function hash_function() is used to perform hash calculation on the keyword of the new element to obtain the address of the primary hash bucket and the address of the secondary hash bucket. For example, (pri_bucket_addr, sec_bucket_addr) = hash_function(key), where key is the keyword, pri_bucket_addr is the primary hash bucket address, and sec_bucket_addr is the secondary hash bucket address.

[0083] Specifically, the address of the primary hash bucket and the secondary hash bucket address are input into the acquisition function to obtain the primary hash bucket and the secondary hash bucket. For example, the acquisition function is get_bucket(), (pri_bucket, sec_bucket) = get_bucket(pri_bucket_addr, sec_bucket_addr), pri_bucket_addr is the primary hash bucket address, sec_bucket_addr is the secondary hash bucket address, pri_bucket is the primary hash bucket, and sec_bucket is the secondary hash bucket.

[0084] Step S1003: Determine whether there is an idle subslot in the main hash bucket.

[0085] Specifically, it is determined by a detection function whether there is an idle subslot in the main hash bucket. If there is an idle subslot in the main hash bucket, the process jumps to step S1004; if there is no idle subslot in the main hash bucket, the process jumps to step S1005.

[0086] The detection function may be IsEmptySlot(), for example, IsEmptySlot(pri_bucket), where pri_bucket is the primary hash bucket.

[0087] Step S1004: Insert the new element into the free subslot of the main hash bucket.

[0088] Specifically, when there is an idle subslot in the main hash bucket, the new element is inserted into the idle subslot of the main hash bucket.

[0089] In the embodiment of the present application, the main hash bucket includes multiple sub-slots. When there are multiple free sub-slots in the main hash bucket, new elements are inserted starting from the tail of the main hash bucket.

[0090] Step S1005: Determine whether there is an idle subslot in the secondary hash bucket.

[0091] Specifically, it is determined whether there is an idle subslot in the secondary hash bucket through a detection function. If there is an idle subslot in the secondary hash bucket, the process jumps to step S1006 . If there is no idle subslot in the secondary hash bucket, the process jumps to step S1007 .

[0092] The detection function may be IsEmptySlot(), for example, IsEmptySlot(sec_bucket), where sec_bucket is a secondary hash bucket.

[0093] Step S1006: Insert the new element into the free subslot of the secondary hash bucket.

[0094] Specifically, when there is an idle subslot in the secondary hash bucket, the new element is inserted into the idle subslot of the secondary hash bucket.

[0095] In the embodiment of the present application, the secondary hash bucket includes multiple sub-slots. When the secondary hash bucket has multiple free sub-slots, new elements are inserted starting from the tail of the secondary hash bucket.

[0096] In the embodiment of the present application, when detecting whether there are free subslots in the primary hash bucket and the secondary hash bucket, the detection can be performed simultaneously without limiting the order. When both the primary hash bucket and the secondary hash bucket have free subslots, one of the primary hash bucket and the secondary hash bucket is determined as the bucket corresponding to the new element storage according to the priority principle of the new element. The specific content of the priority principle of the new element is described in Figure 3 A detailed introduction is given in.

[0097] See also Figure 2 , Figure 2 This is an example schematic diagram of inserting a new element into a hash table provided in an embodiment of the present application.

[0098] like Figure 2 As shown, the Figure 2 An example diagram of inserting a new element into a hash table when there are free subslots in both the primary hash bucket and the secondary hash bucket. Assume that the hash table has 5 domains, domain 0 to domain 4, each domain includes four hash buckets, and a hash bucket (bucket) has 4 subslots, a total of 80 slots, 20 buckets, bucket0 to bucket19, the maximum load of each domain (i.e., the number of subslots) is 16 slots, H (head) represents the head of the hash bucket, T (tail) represents the tail of the hash bucket, after obtaining the new element key, after hash calculation, the primary hash bucket H1 is bucket2, the secondary hash bucket H2 is bucket12, there are free subslots in bucket2 and bucket12, the number of free subslots of bucket2 is 2, and the number of free subslots of bucket12 is 1, then the new element key is inserted into the free subslot of bucket2.

[0099] Step S1007: Determine whether there is an idle subslot in the candidate bucket corresponding to the main hash bucket.

[0100] In the embodiment of the present application, the primary hash bucket and the secondary hash bucket each include a plurality of sub-slots, and each sub-slot corresponds one-to-one to a candidate bucket.

[0101] Specifically, when there are no free subslots in either the primary hash bucket or the secondary hash bucket, obtain the alternative buckets corresponding to each subslot of the primary hash bucket, and determine whether there are free subslots in the alternative buckets corresponding to the primary hash bucket. If there are free subslots in the alternative buckets corresponding to the subslots of the primary hash bucket, jump to step S1008; if there are no free subslots in the alternative buckets corresponding to the subslots of the primary hash bucket, jump to step S1009.

[0102] In an embodiment of the present application, if the address of the main alternative bucket is known, then the address of each subslot of the main alternative bucket is known. Then, according to the address of the subslot in the main alternative bucket, the alternative bucket corresponding to the subslot of the main alternative bucket can be obtained, and each subslot corresponds to a one-to-one alternative bucket. For example, when there are four subslots, {pri_slot3, pri_slot2, pri_slot1, pri_slot0} = pri_bucket, pri_bucket is the main alternative bucket, pri_slot3, pri_slot2, pri_slot1, pri_slot0 are the four subslots of the main alternative bucket, then the main alternative buckets corresponding to pri_slot3, pri_slot2, pri_slot1, pri_slot0 are {sec_bucket3, sec_bucket2, sec_bucket1, sec_bucket0} respectively.

[0103] Step S1008: Determine a hash bucket according to the priority principle of the newly added element, and insert the new element into the hash bucket.

[0104] Specifically, when there is an idle subslot in the candidate bucket corresponding to the subslot of the main hash bucket, the hash bucket is determined according to the priority principle of the new element, and the new element is inserted into the hash bucket. The specific content of the priority principle of the new element is in Figure 3 A detailed introduction is given in.

[0105] In an embodiment of the present application, since there are multiple subslots in the main hash bucket, each subslot corresponds to an alternative bucket. When there are free subslots in the alternative buckets corresponding to each subslot in the main hash bucket, it is necessary to determine an alternative bucket from all the alternative buckets corresponding to the main hash bucket to further insert new elements into the hash bucket.

[0106] Step S1009: Determine whether there is an idle subslot in the candidate bucket corresponding to the secondary hash bucket.

[0107] Specifically, when there is no free subslot in the alternative bucket corresponding to the subslot of the main hash bucket, the alternative bucket corresponding to each subslot of the secondary hash bucket is obtained, and it is determined whether there is a free subslot in the alternative bucket corresponding to the secondary hash bucket. If there is a free subslot in the alternative bucket corresponding to the subslot of the secondary hash bucket, jump to step S1010; if there is no free subslot in the alternative bucket corresponding to the subslot of the secondary hash bucket, jump to step S1011.

[0108] In an embodiment of the present application, if the address of the secondary alternative bucket is known, then the address of each subslot of the secondary alternative bucket is known. Then, according to the address of the subslot in the secondary alternative bucket, the alternative bucket corresponding to the subslot of the secondary alternative bucket can be obtained, and each subslot corresponds to a one-to-one alternative bucket. For example, when there are four subslots in the secondary alternative bucket, {sec_slot3, sec_slot2, sec_slot1, sec_slot0} = sec_bucket, sec_bucket is the secondary alternative bucket, sec_slot3, sec_slot2, sec_slot1, sec_slot0 are the four subslots of the secondary alternative bucket, then the secondary alternative buckets corresponding to {sec_slot3, sec_slot2, sec_slot1, sec_slot0} are {sec_bucket3, sec_bucket2, sec_bucket1, sec_bucket0} respectively.

[0109] Step S1010: Determine the first elimination sub-slot and the first candidate bucket according to the priority principle of the newly added element.

[0110] Among them, the newly added element priority principles include the principle of the largest number of free subslots in the hash bucket, the principle of the largest number of free subslots in the domain segment, and the tail priority principle.

[0111] Specifically, when there is an idle subslot in the candidate bucket corresponding to the subslot of the secondary hash bucket, a new element priority principle is added to determine the first elimination subslot and the first candidate bucket, wherein the first elimination subslot is a subslot in the secondary hash bucket, and the first candidate bucket is the candidate bucket corresponding to the first elimination subslot.

[0112] See also Figure 3 , Figure 3 yes Figure 1 A detailed flowchart of step S1010 in FIG.

[0113] like Figure 3 As shown, step S1010 includes:

[0114] Step S11101: Determine the third candidate bucket according to the principle that the number of free sub-slots in the hash bucket is the largest.

[0115] Specifically, the number of free subslots in the candidate bucket corresponding to each subslot of the secondary hash bucket is obtained, and the candidate bucket with the largest number of free subslots is used as the third candidate bucket.

[0116] For example, the secondary hash bucket has 4 subslots, and the numbers of free subslots in the candidate buckets corresponding to the 4 subslots are (1, 2, 3, 4) respectively, then the candidate bucket with 4 free subslots is used as the third candidate bucket.

[0117] In an embodiment of the present application, there may be a situation where the number of free subslots in the alternative bucket corresponding to each subslot of the secondary hash bucket is equal. For example, there are 4 subslots in the secondary hash bucket, and the numbers of free subslots in the alternative buckets corresponding to the 4 subslots are (4, 4, 3, 4) respectively. Then the alternative bucket with 4 free subslots will be used as the third alternative bucket. There are three third alternative buckets, and it is necessary to further select a third alternative bucket from them.

[0118] In the embodiment of the present application, the more free sub-slots a hash bucket has, the lighter the bucket load is.

[0119] Step S11102: Determine whether the number of the third candidate buckets is greater than a preset number.

[0120] Among them, the preset number is 1.

[0121] Specifically, determine whether the number of the third candidate buckets is greater than the preset number. If the number of the third candidate buckets is equal to the preset number, jump to step S11103; if the number of the third candidate buckets is greater than the preset number, jump to step S11104.

[0122] Step S11103: Determine the third secondary candidate bucket as the first secondary candidate bucket.

[0123] Specifically, when the number of third candidate buckets is equal to the preset number, the third secondary candidate bucket is used as the first secondary candidate bucket, that is, the candidate bucket with the largest number of free subslots among the multiple candidate buckets corresponding to the secondary hash bucket is used as the first secondary candidate bucket.

[0124] Step S11104: Determine the fourth candidate bucket according to the principle of the largest number of free subslots in the domain segment.

[0125] In an embodiment of the present application, the hash table includes multiple domain segments, and each domain segment includes multiple hash buckets.

[0126] Specifically, when the number of the third alternative buckets is greater than the preset number, the fourth alternative bucket is determined based on the principle of the largest number of free subslots in the domain segment. The fourth alternative bucket is an alternative bucket in the third alternative bucket, and the number of free subslots in the domain segment where the fourth alternative bucket is located is greater than the number of free subslots in the domain segment where the alternative bucket in the third alternative bucket is located.

[0127] For example, there are three third candidate buckets, and the domain segments where the three third candidate buckets are located are domain segment 1, domain segment 2, and domain segment 3 respectively. The number of free subslots in domain segment 1 is 7, the number of free subslots in domain segment 2 is 8, and the number of free subslots in domain segment 3 is 8. Then the third candidate buckets in domain segments 2 and 3 are used as the fourth candidate buckets. At this time, there is more than one candidate bucket, and it is necessary to further select an alternative bucket from multiple fourth candidate buckets.

[0128] In the embodiment of the present application, the more free sub-slots a domain segment has, the lighter the domain segment load is.

[0129] Step S11105: Determine whether the number of the fourth candidate barrels is greater than a preset number.

[0130] Specifically, determine whether the number of the fourth secondary candidate buckets is greater than the preset number. If the number of the fourth secondary candidate buckets is equal to the preset number, jump to step 11106. If the number of the fourth secondary candidate buckets is greater than the preset number, jump to step 11107.

[0131] Step S11106: Determine the fourth secondary candidate bucket as the first secondary candidate bucket.

[0132] Specifically, when the number of the fourth secondary candidate buckets is equal to the preset number, the fourth secondary candidate bucket will be used as the first secondary candidate bucket, that is, the fourth secondary candidate bucket will be used as the first secondary candidate bucket, that is, the candidate bucket with the largest number of free subslots in the domain segment where the multiple candidate buckets corresponding to the secondary hash bucket are located will be used as the first secondary candidate bucket.

[0133] Step S11107: According to the tail priority principle, determine the fifth candidate bucket, and use the fifth candidate bucket as the first candidate bucket.

[0134] In the embodiment of the present application, each hash bucket in the hash table corresponds to a serial number one by one.

[0135] Among them, the tail priority principle means that the candidate bucket with the largest sequence number is selected first.

[0136] Specifically, according to the tail priority principle, the candidate bucket with the largest sequence number in the fourth candidate bucket is selected as the fifth candidate bucket, and the fifth candidate bucket is used as the first candidate bucket.

[0137] For example, there are two fourth candidate buckets with serial numbers 12 and 18 respectively, then the fourth candidate bucket with serial number 18 is selected as the fifth candidate bucket.

[0138] Step S11108: Use the subslot in the secondary hash bucket corresponding to the first candidate bucket as the first elimination subslot.

[0139] Specifically, the subslot in the secondary hash bucket corresponding to the first candidate bucket is used as the first elimination subslot.

[0140] For example, if the first candidate bucket is the candidate bucket corresponding to the first sub-slot in the secondary hash bucket, the first sub-slot in the secondary hash bucket is used as the first eliminated sub-slot.

[0141] In the embodiment of the present application, the first culling subslot and the first candidate bucket are determined, and then the elements in the first culling subslot are moved to the free subslot of the first candidate bucket, and the new elements are added to the first culling subslot.

[0142] See also Figure 4 , Figure 4 This is another example schematic diagram of inserting a new element into a hash table provided in an embodiment of the present application.

[0143] like Figure 4 As shown, the Figure 4 An example diagram of inserting a new element into a hash table when there are no free subslots in either the primary hash bucket or the secondary hash bucket, and there are free subslots in the alternative bucket corresponding to the subslot of the primary hash bucket. Assume that the hash table has 5 domains, domains 0 to 4, each domain includes four hash buckets, and a hash bucket (bucket) has 4 subslots, a total of 80 slots, 20 buckets, bucket0 to bucket19, and the maximum load of each domain (i.e., the number of subslots) is 16 slots. H (head) represents the head of the hash bucket, and T (tail) represents the tail of the hash bucket. After obtaining the key of the new element, after hash calculation, the primary hash bucket H1 is obtained as bucket2, and the secondary hash bucket H2 is obtained as bucket9. There are no free subslots in bucket2 and bucket9, so the primary alternative bucket corresponding to the subslot of the primary hash bucket H1 (i.e., bucket2) is obtained. There are free subslots in the alternative bucket corresponding to bucket2. All subslots of bucket2 { The alternative buckets corresponding to slot3, slot2, slot1, slot0} are {bucket10, bucket8, bucket5, bucket4} respectively, among which the bucket loads of {bucket10, bucket8, bucket5, bucket4} are {1, 1, 1, 0} respectively, and the domain loads of {bucket10, bucket8, bucket5, bucket4} are {7, 7, 1, 1} respectively. The smaller the bucket load, the more free subslots there are in the hash bucket, and the smaller the domain load, the more free subslots there are in the domain. According to the priority principle of new elements, the hash bucket with the smallest bucket load is selected first. Bucket4 has the smallest bucket load. The subslot of bucket2 corresponding to bucket4 is slot0. Then slot0 is used as the removed subslot, and the elements of slot0 are moved to the free subslot of bucket4. Then the new element key is inserted into slot0 of bucket2.

[0144] Step S1011: According to the priority principle of culling elements, determine the corresponding second culling sub-slot and the second candidate bucket in the secondary candidate bucket.

[0145] Specifically, when there are no free subslots in the main hash bucket, the secondary hash bucket, the candidate bucket corresponding to the main hash bucket, and the candidate bucket corresponding to the secondary hash bucket, the corresponding second elimination subslot and the second candidate bucket in the secondary candidate bucket are determined according to the elimination element priority principle.

[0146] Among them, the element removal priority principle includes the tail-first removal principle and the head-first removal principle. The tail-first removal principle is to give priority to the subslot at the tail of the hash bucket to remove the subslot at the tail of the hash bucket. The head-first removal principle is to give priority to the subslot at the head of the hash bucket to remove the subslot at the head of the hash bucket.

[0147] See also Figure 5 , Figure 5 yes Figure 1 A detailed flowchart of step S1011 in FIG.

[0148] like Figure 5 As shown, step S1011 includes:

[0149] Step S1111: Determine whether the element culling priority principle is the tail-first culling principle.

[0150] Specifically, when the priority principle for removing elements is not the tail-first removal principle, jump to step S1112; when the priority principle for removing elements is the tail-first removal principle, jump to step S1113.

[0151] Step S1112: Use the first sub-slot at the head of the secondary hash bucket as the second exclusion sub-slot.

[0152] The second elimination sub-slot is a sub-slot in the secondary hash bucket.

[0153] Specifically, when the element removal priority principle is the head first removal principle, the first subslot at the head of the secondary hash bucket is used as the second removal subslot, for example, the second removal subslot kickout_slot=TailSlot(sec_bucket)=sec_slot3, sec_bucket is the secondary hash bucket.

[0154] Step S1113: Use the first sub-slot at the tail of the secondary hash bucket as the second elimination sub-slot.

[0155] Specifically, when the element removal priority principle is the tail first removal principle, the first subslot at the tail of the secondary hash bucket is used as the first removal subslot, for example, the first removal subslot kickout_slot=HeadSlot(sec_bucket)=sec_slot0, sec_bucket is the secondary hash bucket.

[0156] Step S1114: taking the candidate bucket corresponding to the second rejection sub-slot as the second candidate bucket.

[0157] Specifically, the candidate bucket corresponding to the second rejection sub-slot is used as the second candidate bucket.

[0158] Step S1012: All subslots in the secondary candidate hash bucket are cyclically shifted in sequence, covering the candidate hash bucket, and the number of eliminations is increased by one.

[0159] Specifically, all subslots in the secondary candidate hash bucket are shifted in sequence, covering the candidate hash bucket, and the number of eliminations is increased by one, that is, all subslots in the secondary hash bucket are moved by one position based on the second elimination subslot to obtain a new secondary hash bucket, and the new secondary hash bucket covers the original secondary hash bucket, and the elimination of the current secondary hash bucket is increased by 1. The secondary candidate hash bucket is the secondary hash bucket.

[0160] See also Figure 6 , Figure 6 yes Figure 1 A detailed flowchart of step S1012 in FIG.

[0161] like Figure 6 As shown, step S1012 includes:

[0162] Step S1121: all subslots in the secondary candidate hash bucket are sequentially moved one position toward the tail of the secondary candidate hash bucket until all subslots in the secondary candidate hash bucket are completely shifted.

[0163] Specifically, all subslots in the secondary candidate hash bucket are moved one position to the tail of the secondary candidate hash bucket in sequence until all subslots in the secondary candidate hash bucket are shifted, that is, all subslots in the secondary hash bucket are moved one position to the tail of the secondary hash bucket in sequence until all subslots in the secondary hash bucket are shifted.

[0164] The subslot shift is achieved through CyclicShift().

[0165] For example, the secondary hash bucket sec_bucket has four sub-slots {slot3, slot2, slot1, slot0}, and the tail-first culling principle is selected, then slot3 is the culled sub-slot, and the secondary hash bucket after shifting is CyclicShift(sec_bucket)={slot2, sec_slot1, slot0, slot3}.

[0166] In the embodiment of the present application, when all the sub-slots in the secondary hash bucket are cyclically shifted in sequence, the second elimination sub-slot may also be moved toward the head of the secondary hash bucket.

[0167] Step S1013: Replace the second candidate bucket with the secondary candidate bucket.

[0168] Among them, the second candidate bucket is the candidate bucket corresponding to the second rejection sub-slot

[0169] Specifically, if all candidate buckets corresponding to the subslots in the secondary hash bucket have no free subslots, the second candidate bucket is used as the secondary candidate bucket, that is, the second candidate bucket is determined not to be a candidate bucket that currently needs to determine to remove the subslots.

[0170] Specifically, obtain the alternative buckets corresponding to all subslots of the second alternative bucket, determine whether there are free subslots in the alternative buckets corresponding to all subslots of the second alternative bucket, if there are free subslots, select a subslot from the second alternative bucket as a rejection subslot according to the rejection element priority principle, move the elements in the rejection subslot in the second alternative bucket to the alternative bucket corresponding to the rejection subslot in the second alternative bucket, move the elements in the second rejection subslot to the rejection subslot in the second alternative bucket, move the first rejection subslot to the second rejection subslot, and add the new element to a rejection subslot.

[0171] For example, see Figure 7 , Figure 7 is another example schematic diagram of inserting a new element into a hash table provided by an embodiment of the present application, such as Figure 7 As shown, the Figure 7An example diagram of inserting a new element into a hash table when there are no free subslots in the primary hash bucket, the secondary hash bucket, the candidate bucket corresponding to the subslot of the primary hash bucket, and the candidate bucket corresponding to the subslot of the secondary hash bucket. Assume that the hash table has 5 domains, domain 0 to domain 4, each domain includes four hash buckets, and a hash bucket (bucket) has 4 subslots (slots), a total of 80 slots, 20 buckets, bucket0 to bucket19, the maximum load of each domain (i.e., the number of subslots) is 16 slots, H (head) represents the head of the hash bucket, T (tail) represents the tail of the hash bucket, after obtaining the new element key, after hash calculation, the primary hash is obtained. Bucket H1 is bucket2, and the secondary hash bucket H2 is bucket9. There are no free subslots in bucket2 and bucket9. Then, the candidate buckets corresponding to the subslots of the primary hash bucket H1 (i.e., bucket2) are obtained. The candidate buckets corresponding to all the subslots of bucket2 {slot3, slot2, slot1, slot0} are {bucket10, bucket8, bucket5, bucket4} respectively. There are no free subslots in bucket10, bucket8, bucket5, and bucket4. Further, the candidate buckets corresponding to the subslots of the secondary hash bucket H1 (i.e., bucket9) are obtained. All the subslots of bucket9 {s lot3, slot2, slot1, slot0} corresponding to the candidate buckets are {bucket2, bucket12, bucket15, bucket19}, bucket2, bucket12, bucket15, bucket19 do not have free sub-slots, then according to the elimination element priority principle, determine the elimination sub-slot corresponding to the secondary hash bucket, for example, the elimination element priority principle is the tail first elimination principle, then determine the elimination sub-slot of the secondary hash bucket H1 (that is, bucket9) is slot3, all sub-slots in bucket9 are shifted in sequence, and the order of all sub-slots in bucket9 after the shift is {slot2, slot1 , slot0, slot3}, re-obtain the candidate bucket corresponding to slot3, the candidate bucket corresponding to slot3 in bucket9 after the shift is bucket19, bucket19 has no free subslots, obtain the candidate buckets corresponding to each subslot of bucket19 after the shift, and find that the candidate bucket corresponding to the subslot of bucket19 has free subslots. The candidate buckets corresponding to each subslot of bucket19 after the shift are {bucket13, bucket11, bucket18, bucket7}, bucket13, bucket11, bucket18, bucket7 all have free subslots, then according to the priority principle of the newly added elements,Determine that the bucket loads of bucket13, bucket11, bucket18, and bucket7 are {1, 1, 0, 0}, and the domain loads are {9, 13, 7, 9}. Select bucket18 and bucket7 based on the principle of the largest number of free subslots (i.e., the smaller the bucket load), and select bucket18 based on the second priority principle (i.e., the smaller the domain load). Use the subslot slot1 of bucket19 corresponding to bucket18 as the removed subslot, move the element of slot1 of bucket19 to the free subslot of bucket18, move the element of slot3 in bucket9 after the shift to slot1 of bucket19, and move the new element key to slot3 in bucket9 after the shift.

[0172] In an embodiment of the present application, free subslots are found by following the priority principles of removed elements and new elements to avoid loops in the removal path, thereby greatly improving the efficiency of each removal path and reducing the delay in adding new elements.

[0173] Step S1014: Determine whether the number of elimination times is greater than or equal to the elimination times threshold.

[0174] In an embodiment of the present application, when a conflict occurs when inserting a new element, it is necessary to find the next available subslot in the hash table to store the data. If the number of times the subslot is eliminated is not limited, the subslot will be repeatedly searched and eliminated in the hash table, resulting in an unavailable subslot to store the data, thus forming an infinite loop. Therefore, when determining to eliminate the subslot, it is also necessary to determine whether the number of times the currently selected subslot exceeds the maximum number of eliminations to avoid the problem of an infinite loop when inserting new elements.

[0175] Specifically, before shifting the subslot in the hash bucket, it is necessary to determine whether the number of times the subslot is eliminated is greater than or equal to the elimination number threshold. If the number of times the subslot is eliminated is less than the elimination number threshold, and there is no free subslot in the currently selected alternative bucket, the current alternative is used as the secondary hash bucket, and jump to step S1009 to continue looking for free subslots.

[0176] For example, the second candidate bucket has no free subslots. In order to find an empty slot, the second candidate bucket is used as the candidate bucket that currently needs to determine the subslot to be removed, that is, the second candidate bucket is used as the secondary hash bucket in step S1009 to cyclically search for free subslots until a free subslot is found.

[0177] It should be noted that the secondary hash bucket refers to the currently selected hash bucket. For example, if there is no subslot in the second candidate bucket corresponding to the second subslot, it is necessary to find out whether there are any free subslots in the candidate buckets corresponding to all subslots in the second candidate bucket. The second candidate bucket is then used as the currently selected candidate bucket, that is, the second candidate bucket is used as the secondary hash bucket.

[0178] In an embodiment of the present application, a subslot is selected from a secondary hash bucket as the first elimination subslot according to the principle of element elimination priority, the subslots in the secondary hash bucket are cyclically shifted, a subslot is determined in the alternative bucket corresponding to the subslot of the shifted secondary hash bucket as the second elimination subslot, the elements of the second elimination subslot are moved to an idle position, and the elements of the first elimination subslot are overwritten to the second elimination subslot, thereby making room for the new element and inserting it into the first elimination subslot. By cyclically shifting and reallocating elements, the space utilization of the hash table can be optimized, and the entire hash table can be prevented from being unable to continue working due to the inability to insert a new element into a subslot.

[0179] In addition, according to the priority principles of new elements and removed elements, selecting alternative buckets and removing subslots can reduce the impact on frequently accessed elements and avoid the problem of dead loops caused by frequent access to a subslot, so as to reduce the problem of reduced efficiency of hash table data insertion due to conflict processing.

[0180] See also Figure 8 , Figure 8 It is a structural diagram of a hash table data insertion system provided in an embodiment of the present application.

[0181] like Figure 8 As shown, the data insertion method system 800 of the hash table includes a hash bucket determination unit 801, a removal element priority judgment unit 802, a new element priority judgment unit 803, a removal replacement unit 804, a hash table 805, and a removal element random access memory 806.

[0182] The hash table includes a primary hash bucket and a secondary hash bucket.

[0183] The hash bucket determination unit 801 is used to obtain a new element to be inserted, and obtain an address of a primary hash bucket and an address of a secondary hash bucket according to the new element, wherein the primary hash bucket and the secondary hash bucket are used to store data, the primary hash bucket includes a plurality of sub-slots, the secondary hash bucket includes a plurality of sub-slots, each sub-slot in the primary hash bucket corresponds to a primary candidate bucket, and each sub-slot in the secondary hash bucket corresponds to a secondary candidate bucket;

[0184] The culling element priority judgment unit 802 is used to determine, based on the address of the primary hash bucket and the address of the secondary hash bucket, that there are no idle subslots in the primary hash bucket, the secondary hash bucket, all primary candidate buckets corresponding to the primary hash bucket, and all secondary candidate buckets corresponding to the secondary hash bucket, then determine the first culling subslot corresponding to the secondary hash bucket according to the culling element priority principle, wherein the first culling subslot is a subslot in the secondary hash bucket;

[0185] The newly added element priority judgment unit 803 is used to sequentially shift all subslots in the secondary hash bucket based on the first eliminated subslot to obtain the shifted secondary hash bucket, obtain the first candidate bucket corresponding to each subslot in the shifted secondary hash bucket, and if there is an idle subslot in the first candidate bucket, determine the second candidate bucket according to the newly added element priority principle, wherein the second candidate bucket is one of the first candidate buckets;

[0186] The culling and replacement unit 804 is used to add the element in the first culling sub-slot to the free sub-slot in the second candidate bucket, and overwrite the new element to the first culling sub-slot.

[0187] Hash table 805, used to store new elements;

[0188] The culled element random access memory 806 is used to record the number of culled times for each subslot in the hash table.

[0189] In the embodiment of the present application, the units of the data insertion system through the hash table cooperate with each other to improve the efficiency of data addition.

[0190] See also Fig. 9 , Fig. 9 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application.

[0191] like Fig. 9 As shown, the electronic device 900 includes one or more processors 901 and a memory 902. Fig. 9 A processor 901 is taken as an example.

[0192] The processor 901 and the memory 902 may be connected via a bus or other means. Fig. 9 The example of connecting through bus is taken in the following.

[0193] The processor 901 is used to provide computing and control capabilities to control the electronic device 900 to perform corresponding tasks, for example, to control the electronic device 900 to perform the data insertion method of the hash table in any one of the above method embodiments, the method comprising: obtaining a new element to be inserted, and obtaining an address of a primary hash bucket and an address of a secondary hash bucket according to the new element, wherein the primary hash bucket and the secondary hash bucket are used to store data, the primary hash bucket includes a plurality of subslots, and the secondary hash bucket includes a plurality of subslots. When it is determined based on the address of the primary hash bucket and the address of the secondary hash bucket that there are no free subslots in the primary hash bucket and the secondary hash bucket, a primary candidate bucket corresponding to each subslot of the primary hash bucket is obtained; if the primary candidate bucket corresponding to each subslot in the primary hash bucket does not have a free subslot, a secondary candidate bucket corresponding to each subslot of the secondary hash bucket is obtained, and according to the priority principle of the newly added element, the new element is inserted into the secondary candidate bucket corresponding to one of the subslots of the secondary hash bucket.

[0194] The processor 901 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), a hardware chip or any combination thereof; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a programmable logic device (PLD) or a combination thereof. The above-mentioned PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL) or any combination thereof.

[0195] The memory 902, as a non-transient computer-readable storage medium, can be used to store non-transient software programs, non-transient computer executable programs and modules, such as program instructions / modules corresponding to the data insertion method of the hash table in the embodiment of the present application. The processor 901 can implement the data insertion method of the hash table in any of the above method embodiments by running the non-transient software programs, instructions and modules stored in the memory 902. Specifically, the memory 902 may include a volatile memory (volatile memory, VM), such as a random access memory (random access memory, RAM); the memory 902 may also include a non-volatile memory (non-volatile memory, NVM), such as a read-only memory (read-only memory, ROM), a flash memory (flash memory), a hard disk drive (hard disk drive, HDD) or a solid-state drive (solid-state drive, SSD) or other non-transient solid-state storage devices; the memory 902 may also include a combination of the above types of memories.

[0196] The memory 902 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 902 may optionally include a memory remotely arranged relative to the processor 901, and these remote memories may be connected to the processor 901 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0197] One or more modules are stored in the memory 902, and when executed by one or more processors 901, the data insertion method of the hash table in any of the above method embodiments is executed, for example, the above described method is executed. Figure 1 The steps shown.

[0198] In the embodiment of the present application, the electronic device 900 may also have components such as a wired or wireless network interface, a keyboard, and an input / output interface for input and output. The electronic device 900 may also include other components for realizing device functions, which will not be described in detail here.

[0199] The embodiment of the present application also provides a non-volatile computer-readable storage medium, such as a memory including a program code, and the program code can be executed by a processor to complete the data insertion method of the hash table in the above embodiment. For example, the non-volatile computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CDROM), a magnetic tape, a floppy disk, and an optical data storage device.

[0200] The embodiment of the present application also provides a computer program product, which includes one or more program codes, and the program codes are stored in a non-volatile computer-readable storage medium. The processor of the flash memory device reads the program code from the non-volatile computer-readable storage medium, and the processor executes the program code to complete the method steps of the data insertion method of the hash table provided in the above embodiment.

[0201] A person skilled in the art will appreciate that all or part of the steps for implementing the above embodiments may be accomplished by hardware or by hardware associated with a program code, and the program may be stored in a non-volatile computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.

[0202] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a general hardware platform, and of course, by hardware. Based on this understanding, the above technical solution can be essentially or in other words, the part that contributes to the relevant technology can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute various embodiments or certain parts of the embodiments.

[0203] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Under the concept of the present application, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other changes in different aspects of the present application as above, which are not provided in detail for the sake of simplicity. Although the present application has been described in detail with reference to the aforementioned embodiments, a person of ordinary skill in the art should understand that the technical solutions described in the aforementioned embodiments can still be modified, or some of the technical features can be replaced by equivalents. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for inserting data into a hash table, characterized in that: The hash table includes a primary hash bucket and a secondary hash bucket, and the method includes: Set the removal count threshold and obtain the new elements to be inserted; According to the new element, an address of a primary hash bucket and an address of a secondary hash bucket are obtained, a primary hash bucket is obtained according to the primary hash bucket address, and a secondary hash bucket is obtained according to the secondary hash bucket address, wherein the primary hash bucket and the secondary hash bucket each include a plurality of subslots, each subslot corresponds to an alternative bucket, and the subslot is used to store the new element; When none of the primary hash bucket, the secondary hash bucket, and the candidate bucket corresponding to the primary hash bucket has an empty subslot, determining whether there is an empty slot in the candidate bucket corresponding to the secondary hash bucket; If there is an idle subslot in the candidate bucket corresponding to the secondary candidate bucket, determine the first rejection subslot and the first candidate bucket according to the priority principle of the newly added elements, wherein the first rejection subslot is one of the multiple subslots in the secondary candidate bucket, and the first candidate bucket is the candidate bucket corresponding to the first rejection subslot; Add the element of the first rejection sub-slot to the first candidate bucket, insert the new element into the first rejection sub-slot, and add successfully; If there is no free subslot in the candidate bucket corresponding to the secondary candidate bucket, determine the second elimination subslot and the second candidate bucket corresponding to the secondary candidate bucket according to the elimination element priority principle, wherein the second elimination subslot is a subslot in the secondary candidate hash bucket, and the second candidate bucket is the candidate bucket corresponding to the second elimination subslot; After all subslots in the secondary candidate hash bucket are cyclically shifted in sequence, the candidate hash bucket is covered, and the number of eliminations is increased by one; The second candidate bucket is replaced with the secondary candidate bucket, and it is determined whether there is an empty slot in the secondary candidate bucket, until there is an idle sub-slot in the candidate bucket corresponding to the secondary candidate bucket or the number of eliminations is equal to the threshold number of eliminations.

2. The method according to claim 1, characterized in that The hash table includes a plurality of domain segments, each of which includes a plurality of hash buckets, and the priority principle of the newly added elements includes a principle of the largest number of free subslots of the hash bucket, a principle of the largest number of free subslots of the domain segment, and a tail priority principle; The step of determining the first elimination subslot and the first candidate bucket according to the priority principle of the newly added element includes: According to the principle that the number of free sub-slots of the hash bucket is the largest, a third candidate bucket is determined, wherein the third candidate bucket is the candidate bucket with the largest number of free sub-slots among all the candidate buckets corresponding to the secondary hash bucket; If the number of the third candidate buckets is greater than the preset number, a fourth candidate bucket is determined according to the principle that the number of free subslots in the domain segment is the largest, wherein the fourth candidate bucket is a candidate bucket in the third candidate bucket; If the number of the fourth candidate buckets is greater than the preset number, a fifth candidate bucket is determined according to the tail priority principle, and the fifth candidate bucket is used as the first candidate bucket, wherein the fifth candidate bucket is one of the plurality of fourth candidate buckets; The subslot in the secondary hash bucket corresponding to the first candidate bucket is used as the first eliminated subslot.

3. The method according to claim 2, characterized in that The element removal priority principle includes a tail-first removal principle or a head-first removal principle; According to the priority principle of the removed elements, determining the corresponding second removal sub-slot and the second candidate bucket in the secondary candidate bucket includes: If the element removal priority principle is the tail-first removal principle, the first subslot at the tail of the secondary hash bucket is used as the second removal subslot; If the element removal priority principle is the header first removal principle, the first subslot of the header of the secondary hash bucket is used as the second removal subslot; The candidate bucket corresponding to the second rejection sub-slot is used as the second candidate bucket.

4. The method according to claim 2, characterized in that The sequentially cyclically shifting all subslots in the secondary candidate hash bucket includes: All subslots in the secondary candidate hash bucket are sequentially moved one position toward the tail of the secondary candidate hash bucket until all subslots in the secondary candidate hash bucket are completely shifted.

5. The method according to claim 2, characterized in that: The method further comprises: If the number of the third candidate buckets is equal to the preset number, the third candidate bucket is used as the first candidate bucket; If the number of the fourth candidate buckets is equal to the preset number, the fourth candidate buckets are used as the first candidate buckets.

6. The method according to claim 1, characterized in that The method further comprises: If there is an idle subslot in the primary hash bucket and there is no idle subslot in the secondary hash bucket, insert the new element into the idle subslot in the primary hash bucket; If there is no free subslot in the primary hash bucket and there is a free subslot in the secondary hash bucket, insert the new element into the free subslot in the secondary hash bucket; If both the primary hash bucket and the secondary hash bucket have free subslots, determine a hash bucket according to the priority principle of newly added elements, and insert the new element into the hash bucket, wherein the hash bucket is one of the primary hash bucket and the secondary hash bucket; If there is no free subslot in the primary hash bucket and no free subslot in the secondary hash bucket, and there is a free subslot in the primary candidate bucket corresponding to the subslot of the primary hash bucket, insert the new element into the free subslot of the primary candidate bucket corresponding to the subslot of the primary hash bucket; If there is no free subslot in the primary hash bucket, and there is no free subslot in the secondary hash bucket, there is no free subslot in the primary candidate bucket corresponding to the subslot of the primary hash bucket, and there is a free subslot in the secondary candidate bucket corresponding to the subslot of the secondary hash bucket, then the new element is inserted into the free subslot of the secondary candidate bucket corresponding to the subslot of the secondary hash bucket.

7. A data insertion system for a hash table, characterized in that: The hash table includes a primary hash bucket and a secondary hash bucket, and the system includes: A hash bucket determination unit, used to obtain a new element to be inserted, and according to the new element, obtain an address of a primary hash bucket and an address of a secondary hash bucket, wherein the primary hash bucket and the secondary hash bucket are used to store data, the primary hash bucket includes a plurality of subslots, the secondary hash bucket includes a plurality of subslots, each subslot in the primary hash bucket corresponds to a primary candidate bucket, and each subslot in the secondary hash bucket corresponds to a secondary candidate bucket; a culling element priority judgment unit, configured to determine, based on the address of the primary hash bucket and the address of the secondary hash bucket, when there are no idle subslots in the primary hash bucket, the secondary hash bucket, all primary candidate buckets corresponding to the primary hash bucket, and all secondary candidate buckets corresponding to the secondary hash bucket, then determine, according to the culling element priority principle, a first culling subslot corresponding to the secondary hash bucket, wherein the first culling subslot is a subslot in the secondary hash bucket; A newly added element priority judgment unit is used to sequentially shift all subslots in the secondary hash bucket based on the first eliminated subslot to obtain a shifted secondary hash bucket, obtain a first candidate bucket corresponding to each subslot in the shifted secondary hash bucket, and if there is an idle subslot in the first candidate bucket, determine a second candidate bucket according to a newly added element priority principle, wherein the second candidate bucket is one of the first candidate buckets of the first plurality; The culling and replacement unit is used to add the element in the first culling sub-slot to the free sub-slot in the second candidate bucket, and overwrite the new element to the first culling sub-slot.

8. The system according to claim 7, characterized in that The system further comprises: A hash table, used to store the new elements; The culled element random access memory is used to record the number of culled times for each subslot in the hash table.

9. An electronic device, characterized in that: include: at least one processor, and a memory communicatively coupled to the at least one processor, wherein: The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the data insertion method for the hash table as described in any one of claims 1 to 6 above.

Citation Information

Patent Citations

  • Hash table construction method and system for nonvolatile memory

    CN107153707A

  • Hash bucket searching method and device, Hash table storage method and device, and Hash table searching method and device

    CN110457535A

  • High-speed network elephant flow accurate measurement method and architecture

    CN111262756A

  • In-memory hash sorting construction method based on novel memory

    CN113254720A

  • Write optimization extensible Hash index structure based on nonvolatile memory and insertion, refreshing and deletion methods

    CN113342706A