A key-value pair storage method, device, equipment and medium
By combining extendable hash and perfect hash, the performance problem of hash index in read-intensive and read-skewed scenarios is solved, and more efficient key-value pair storage and read performance improvements are achieved.
Patent Information
- Application Number
- CN202210976474.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-08-15
AI Technical Summary
Existing hash indexes cannot fully utilize the performance advantages of persistent memory (PMEM) in read-intensive and read-skewed scenarios, and the additional overhead caused by hash collisions limits the index's read throughput.
By combining extendable hash and perfect hash, we judge the size relationship between the number of rehashings of the virtual bucket group of key-value pairs and the number of extensions of the local hash table, and perform corresponding operations, including the transfer of historical key-value pairs and the expansion of buckets, to reduce the additional overhead caused by hash collisions.
Improves the efficiency of key-value pair storage, improves the read performance of indexes in read-intensive and read-skewed scenarios, and reduces the overhead of maintaining index perfection.
Smart Images

Figure CN115309745B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a key-value pair storage method, device, equipment and medium. Background Art
[0002] At present, the emergence of persistent memory (PMEM) has significantly improved the performance of traditional hash indexes. A large number of existing hash indexes for PMEM mainly use hardware features such as NUMA (Non Uniform Memory Access), cache granularity alignment, and read-write asymmetry to optimize performance; a small number of existing studies focus on the algorithm and structure of the hash index itself, and commonly used optimization techniques include: fingerprint acceleration, aggressive concurrency control, load balancing, etc. In addition, existing studies usually ignore the optimization of read performance, and even sacrifice read performance in exchange for improved write performance, which leads to the inability of existing indexing schemes to fully utilize the performance advantages of PMEM in large-scale read-intensive and read-skewed scenarios. The reason is that the additional overhead caused by hash collisions severely limits the read throughput of the index. Hash collisions require multiple additional PMEM accesses for a query, resulting in a linear increase in the single query time with the number of probes, especially under negative query workloads (queried key-value pairs are not in the hash table).
[0003] As can be seen from the above, in the process of key-value pair storage, how to give full play to the hardware characteristics of PMEM and the inherent advantages of perfect hash index, so as to improve the efficiency of key-value pair storage, improve the reading performance of indexes in read-intensive and read-skewed scenarios, and reduce the overhead of maintaining index perfection are problems to be solved in this field. Summary of the invention
[0004] In view of this, the purpose of the present invention is to provide a key-value pair storage method, device, equipment and medium, which can give full play to the hardware characteristics of PMEM and the inherent advantages of perfect hash index, thereby improving the efficiency of key-value pair storage, improving the index reading performance in read-intensive and read-skewed scenarios, and reducing the overhead of maintaining index perfection. The specific scheme is as follows:
[0005] In a first aspect, the present application discloses a key-value pair storage method, comprising:
[0006] Determine a key-value pair storage bucket in the key-value pair storage bucket group, and determine whether the remaining capacity in the key-value pair storage bucket is less than the occupied capacity of the key-value pair to be stored;
[0007] If the remaining capacity in the key-value pair storage bucket is less than the occupied capacity of the key-value pair to be stored, determining the relationship between the number of rehashing of the pre-acquired key-value pair virtual bucket group and the number of extensions of the local hash table;
[0008] If the number of rehashing times is less than the number of extension times, a key-value pair transfer storage bucket group is determined, and the historical key-value pairs belonging to the key-value pair virtual bucket group are transferred and stored in the key-value pair transfer storage bucket group;
[0009] A target key-value pair storage bucket is determined, and the key-value pair to be stored is stored in the target key-value pair storage bucket.
[0010] Optionally, before determining the key-value pair storage bucket in the key-value pair storage bucket group, the method further includes:
[0011] Acquire key-value pair storage information, and calculate the key-value pair keywords in the key-value pair storage information to obtain a unit number and a guide index;
[0012] The location information of the key-value pair to be stored in the key-value pair virtual bucket group is determined based on the unit number and the guide index.
[0013] Optionally, determining the key-value pair storage bucket in the key-value pair storage bucket group includes:
[0014] Input the unit number and the guide index into a preset mapper to obtain a layer index and a bucket offset;
[0015] The number of layers of the key-value pair storage bucket group and the location information of the key-value pair storage bucket are determined based on the layer index and the bucket offset to obtain the key-value pair storage bucket.
[0016] Optionally, transferring and storing the historical key-value pairs in the key-value pair virtual bucket group to the key-value pair transfer storage bucket group includes:
[0017] Filter out the target historical key-value pairs to be transferred from all historical key-value pairs belonging to the key-value pair virtual bucket group;
[0018] The target historical key-value pair is transferred and stored in the key-value pair transfer storage bucket group.
[0019] Optionally, the step of selecting a target historical key-value pair to be transferred from all historical key-value pairs in the key-value pair virtual bucket group includes:
[0020] Determine the unit number of the key-value pair virtual bucket group, and obtain the unit numbers of all historical key-value pairs belonging to the key-value pair virtual bucket group;
[0021] Determine whether the unit number of the key-value pair virtual bucket group is consistent with the unit number of the historical key-value pair; if the unit number of the key-value pair virtual bucket group is inconsistent with the unit number of the historical key-value pair, use the historical key-value pair as the target historical key-value pair to be transferred.
[0022] Optionally, after determining the magnitude relationship between the number of rehashing times and the number of extension times, the method further includes:
[0023] If the number of rehashing operations is equal to the number of extension operations, obtaining the remaining capacity of all the key-value pair storage buckets in the key-value pair storage bucket group;
[0024] Based on the capacity of the key value to be stored and the remaining capacity of all the key-value pair storage buckets in the key-value pair storage bucket group, the key-value pair transfer storage bucket is screened out from the key-value pair storage bucket group so as to transfer and store the historical key-value pairs belonging to the key-value pair storage bucket to the key-value pair transfer storage bucket.
[0025] Optionally, the key-value pair storage method further includes:
[0026] If the occupied capacity of the key-value pairs to be stored is greater than the remaining capacity of all the key-value pair storage buckets in the key-value pair storage bucket group, the key-value pair storage bucket group is expanded according to a preset expansion method to obtain a new key-value pair storage bucket group, and the values of the rehashing times and the extension times are increased, and then the process jumps to the step of determining the key-value pair storage buckets in the key-value pair storage bucket group.
[0027] In a second aspect, the present application discloses a key-value pair storage device, comprising:
[0028] A key-value pair storage bucket determination module is used to determine a key-value pair storage bucket in the key-value pair storage bucket group, and determine whether the remaining capacity in the key-value pair storage bucket is less than the occupied capacity of the key-value pair to be stored;
[0029] A judgment module, configured to judge the relationship between the number of rehashing of the pre-acquired key-value pair virtual bucket group and the number of extensions of the local hash table if the remaining capacity in the key-value pair storage bucket is less than the occupied capacity of the key-value pair to be stored;
[0030] A historical key-value pair transfer module, configured to determine a key-value pair transfer storage bucket group if the number of rehashing times is less than the number of extension times, and transfer and store the historical key-value pairs belonging to the key-value pair virtual bucket group to the key-value pair transfer storage bucket group;
[0031] The target key-value pair storage module is used to determine a target key-value pair storage bucket and store the key-value pair to be stored in the target key-value pair storage bucket.
[0032] In a third aspect, the present application discloses an electronic device, comprising:
[0033] Memory, used to store computer programs;
[0034] The processor is used to execute the computer program to implement the aforementioned key-value pair storage method.
[0035] In a fourth aspect, the present application discloses a computer storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the steps of the aforementioned disclosed key-value pair storage method are implemented.
[0036] It can be seen that the present application provides a key-value pair storage method, including determining a key-value pair storage bucket in a key-value pair storage bucket group, and judging whether the remaining capacity in the key-value pair storage bucket is less than the occupied capacity of the key-value pair to be stored; if the remaining capacity in the key-value pair storage bucket is less than the occupied capacity of the key-value pair to be stored, then judging the relationship between the number of rehashing times of a pre-acquired key-value pair virtual bucket group and the number of extension times of a local hash table; if the number of rehashing times is less than the number of extension times, then determining a key-value pair transfer storage bucket group, and transferring and storing the historical key-value pairs belonging to the key-value pair virtual bucket group to the key-value pair transfer storage bucket group; determining a target key-value pair storage bucket, and storing the key-value pairs to be stored in the target key-value pair storage bucket. This application combines scalable hashing with perfect hashing, and performs corresponding operations by judging the relationship between the number of rehashing times of a virtual bucket group of a key-value pair and the number of extension times of a local hash table, thereby eliminating the extra overhead caused by hash collisions during queries by introducing perfect hashing, thereby releasing the read performance of the index, and giving full play to the hardware characteristics of PMEM and the inherent advantages of the perfect hash index, thereby improving the efficiency of key-value pair storage, improving the read performance of the index in read-intensive and read-skewed scenarios, and reducing the overhead of maintaining index perfection. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0038] Figure 1 A flow chart of a key-value pair storage method disclosed in this application;
[0039] Figure 2 A structural diagram of a key-value pair storage system disclosed in this application;
[0040] Figure 3 A specific flow chart of a key-value pair storage method disclosed in this application;
[0041] Figure 4 A schematic diagram of a specific process of a key-value pair storage method disclosed in this application;
[0042] Figure 5 A schematic diagram of a specific process of a key-value pair storage method disclosed in this application;
[0043] Figure 6 A schematic diagram of a specific process of a key-value pair storage method disclosed in this application;
[0044] Figure 7 A schematic diagram of the structure of a key-value pair storage device disclosed in this application;
[0045] Figure 8 A structural diagram of an electronic device provided for this application. DETAILED DESCRIPTION
[0046] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0047] At present, the emergence of persistent memory (PMEM) has significantly improved the performance of traditional hash indexes. A large number of existing hash indexes for PMEM mainly use hardware features such as NUMA (Non Uniform Memory Access), cache granularity alignment, and read-write asymmetry to optimize performance; a small number of existing studies focus on the algorithm and structure of the hash index itself, and commonly used optimization techniques include: fingerprint acceleration, aggressive concurrency control, load balancing, etc. Existing studies usually ignore the optimization of read performance, and even sacrifice read performance in exchange for improved write performance, which leads to the inability of existing indexing schemes to fully utilize the performance advantages of PMEM in large-scale read-intensive and read-skewed scenarios. The reason is that the additional overhead caused by hash collisions severely limits the read throughput of the index. Hash collisions require multiple additional PMEM accesses for a query, resulting in a linear increase in the single query time with the number of probes, especially under negative query workloads (queried key-value pairs are not in the hash table). As can be seen from the above, in the process of key-value pair storage, how to give full play to the hardware characteristics of PMEM and the inherent advantages of perfect hash index, so as to improve the efficiency of key-value pair storage, improve the reading performance of indexes in read-intensive and read-skewed scenarios, and reduce the overhead of maintaining index perfection are problems to be solved in this field.
[0048] See also Figure 1 As shown, an embodiment of the present invention discloses a key-value pair storage method, which may specifically include:
[0049] Step S11: determine a key-value pair storage bucket in the key-value pair storage bucket group, and determine whether the remaining capacity in the key-value pair storage bucket is less than the occupied capacity of the key-value pair to be stored.
[0050] In this embodiment, before determining the key-value pair storage bucket in the key-value pair storage bucket group, it also includes: obtaining key-value pair storage information, and calculating the key-value pair keyword in the key-value pair storage information to obtain a unit number and a guide index; determining the location information of the key-value pair to be stored in the key-value pair virtual bucket group based on the unit number and the guide index, and then inputting the unit number and the guide index into a preset mapper to obtain a layer index and a bucket offset, and determining the number of layers of the key-value pair storage bucket group and the location information of the key-value pair storage bucket based on the layer index and the bucket offset to obtain the key-value pair storage bucket.
[0051] It can be understood that this application is a scalable perfect hash index for hybrid storage architecture PMEM-DRAM (dynamic random access memory). A perfect hash is usually composed of an index structure and a key-value table. The position of the key-value pair in the key-value table is determined by querying the index. The key-value pair is hashed, mapped, and persisted once to establish a one-to-one correspondence with a unique physical bucket (no hash collision). Therefore, only one PMEM access is required to complete the query operation, such as Figure 2 As shown, the specific system structure of the present application is divided into three parts: the index structure in DRAM, the hash table in PMEM, and the functional components. Specifically, in the first part, the present invention uses Figure 1 The data structure named unit in is used as a fast index structure. Figure 1As shown, in order to reduce expensive movement overhead, each unit maintains three types of metadata. Among them, the lock field is used for concurrency control; LD indicates the number of times the current unit is rehashed; the guide array (GA) is an array that records the movement displacement of key-value pairs; in the second part, the present invention organizes physical buckets in layers, and each layer contains a fixed number of physical buckets, which are indexed by layer pointers. Each pointer in the layer pointer array points to the pointer of the first physical bucket of each layer. In addition, the layer pointer array also maintains GD to indicate the number of extensions of the hash table, and cooperates with LD to control overflow operations; the bucket metadata occupies the first 48 bytes of each bucket, and contains some fields that assist in performing basic operations (insertion, query, and deletion). Among them, the 4-byte version lock is used to ensure the concurrency of key-value pair operations; the used slot field is used to accumulate the number of used slots. The present invention uses the least significant 8 bits of each keyword to generate a unique fingerprint to speed up the search process. The bitmap field is used to check the validity of the key-value pair in the slot; the metadata also saves the unit number and guide index of each key-value pair to facilitate locating the virtual bucket corresponding to the key-value pair; when performing fault recovery, the status field indicates whether the physical bucket is in a consistent state. In addition, padding bytes are added to the end of the metadata to maintain the same access granularity as PMEM; in the third part, the hash and modulus components respectively calculate the unit number and guide index for locating the unique virtual bucket, and the mapper component is used to build a mapping relationship between the virtual bucket and the physical bucket.
[0052] Specifically, the present application receives the unit number and the displacement-added guide index as input through a mapper, outputs the layer index and the bucket offset, and determines a physical bucket by these two values, so that the virtual bucket is mapped to a unique physical bucket. It is worth noting that a physical bucket can be mapped to multiple virtual buckets, and a virtual bucket will only be mapped to one physical bucket. That is, first obtain the key-value pair storage information, determine the key-value pair keyword k in the key-value pair storage information, and then use the following formula to specifically calculate the unit number c and guide index g used to index the virtual bucket.
[0053] ;
[0054] ;
[0055] Where N represents the number of units and G represents the length of the index array.
[0056] Then, a mapper is used to map the virtual bucket to a unique physical bucket. The mapper receives the unit number s and the guide index g as input, and outputs the layer index l and the bucket offset o through the following formula.
[0057] ;
[0058] ;
[0059] ;
[0060] Among them, GA (Guide Array) represents the guide array, B represents the number of physical buckets in each layer, GD (Global Depth) represents the extension times, and LD (Local Depth) represents the rehashing times.
[0061] Step S12: If the remaining capacity in the key-value pair storage bucket is less than the occupied capacity of the key-value pair to be stored, then determine the magnitude relationship between the rehashing times of the pre-obtained key-value pair virtual bucket group and the extension times of the local hash table.
[0062] Step S13: If the rehashing times are less than the extension times, then determine the key-value pair transfer storage bucket group, and transfer and store the historical key-value pairs belonging to the key-value pair virtual bucket group to the key-value pair transfer storage bucket group.
[0063] In this embodiment, if the rehashing times are less than the extension times, then determine the key-value pair transfer storage bucket group, then screen out the target historical key-value pairs to be transferred from all the historical key-value pairs belonging to the key-value pair virtual bucket group, and transfer and store the target historical key-value pairs to the key-value pair transfer storage bucket group.
[0064] Among them, after determining the key-value pair transfer storage bucket group, the specific process is as follows: determine the unit number of the key-value pair virtual bucket group, and obtain the unit numbers of all historical key-value pairs belonging to the key-value pair virtual bucket group, then determine whether the unit number of the key-value pair virtual bucket group is the same as the unit number of the historical key-value pair. If the unit number of the key-value pair virtual bucket group is not the same as the unit number of the historical key-value pair, then use the historical key-value pair as the target historical key-value pair to be transferred.
[0065] In this embodiment, if the remaining capacity in the key-value pair storage bucket is less than the occupied capacity of the key-value pair to be stored, and the rehashing times are less than the extension times, that is, when the physical bucket is full and LD < GD, perform a local rehashing operation. Different from the full-table rehashing in previous studies, the present invention performs rehashing in units of physical bucket groups. In other words, in the rehashing stage, only the physical bucket groups that cause insertion failures are operated on.
[0066] Step S14: Determine the target key-value pair storage bucket, and store the key-value pair to be stored in the target key-value pair storage bucket.
[0067] After the target historical key-value pair is transferred and stored in the key-value pair transfer storage bucket group, the above calculation steps are repeated to obtain a target key-value pair storage bucket, and then the key-value pair to be stored is stored in the target key-value pair storage bucket.
[0068] In this embodiment, a key-value pair storage bucket in a key-value pair storage bucket group is determined, and it is determined whether the remaining capacity in the key-value pair storage bucket is less than the occupied capacity of the key-value pairs to be stored; if the remaining capacity in the key-value pair storage bucket is less than the occupied capacity of the key-value pairs to be stored, the relationship between the number of rehashing times of the pre-acquired key-value pair virtual bucket group and the number of extension times of the local hash table is determined; if the number of rehashing times is less than the number of extension times, a key-value pair transfer storage bucket group is determined, and the historical key-value pairs belonging to the key-value pair virtual bucket group are transferred and stored in the key-value pair transfer storage bucket group; a target key-value pair storage bucket is determined, and the key-value pairs to be stored are stored in the target key-value pair storage bucket. This application combines scalable hashing with perfect hashing, and performs corresponding operations by judging the relationship between the number of rehashing times of a virtual bucket group of a key-value pair and the number of extension times of a local hash table, thereby eliminating the extra overhead caused by hash collisions during queries by introducing perfect hashing, thereby releasing the read performance of the index, and giving full play to the hardware characteristics of PMEM and the inherent advantages of the perfect hash index, thereby improving the efficiency of key-value pair storage, improving the read performance of the index in read-intensive and read-skewed scenarios, and reducing the overhead of maintaining index perfection.
[0069] See also Figure 3 As shown, an embodiment of the present invention discloses a key-value pair storage method, which may specifically include:
[0070] Step S21: determine a key-value pair storage bucket in the key-value pair storage bucket group, and determine whether the remaining capacity in the key-value pair storage bucket is less than the occupied capacity of the key-value pair to be stored.
[0071] Step S22: if the remaining capacity in the key-value pair storage bucket is less than the occupied capacity of the key-value pairs to be stored, determine the relationship between the number of rehashing times of the pre-acquired key-value pair virtual bucket group and the number of extension times of the local hash table.
[0072] Step S23: If the number of rehashing times is equal to the number of extension times, the remaining capacity of all the key-value pair storage buckets in the key-value pair storage bucket group is obtained, and then the key-value pair transfer storage bucket is filtered out from the key-value pair storage bucket group based on the capacity of the key value to be stored and the remaining capacity of all the key-value pair storage buckets in the key-value pair storage bucket group, so as to transfer and store the historical key-value pairs belonging to the key-value pair storage bucket to the key-value pair transfer storage bucket.
[0073] Step S24: determine a target key-value pair storage bucket, and store the key-value pair to be stored in the target key-value pair storage bucket.
[0074] In this embodiment, if the occupied capacity of the key-value pairs to be stored is greater than the remaining capacity of all the key-value pair storage buckets in the key-value pair storage bucket group, the key-value pair storage bucket group is expanded according to a preset expansion method to obtain a new key-value pair storage bucket group, and the values of the rehashing times and the extension times are increased, and then the process jumps to the step of determining the key-value pair storage buckets in the key-value pair storage bucket group.
[0075] Specifically, taking doubling as an example, the key-value pair storage bucket group is expanded by doubling, and the unit and layer pointers can also be expanded. After the expansion is successful, the new layer pointer will point to the new physical bucket layer in sequence, and increase the number of rehashing times of the key-value pair virtual bucket group and the number of extensions of the local hash table (i.e., the values of LD and GD). After changing LD and GD, the mapping relationship between the unit for storing key-value pairs and the physical bucket group will also change. The preset expansion method in this application includes but is not limited to doubling.
[0076] In this embodiment, if the occupied capacity of the key-value pair to be stored is greater than the remaining capacity of all the key-value pair storage buckets in the key-value pair storage bucket group, it means that the insertion of the key-value pair to be stored has failed. After expansion, since the target key-value pair storage bucket for the key-value pair to be stored has been determined before hashing (the same key-value pair storage buckets at different layers), there is no need to re-establish the mapping relationship between the key-value pair to be stored and the target key-value pair storage bucket, thus eliminating unnecessary calculations, and then reinsert the key-value pair that failed to be inserted before.
[0077] like Figure 4 As shown, the specific operation steps are: add 1 to the LD in all cells (cells 0 and 2) pointing to virtual bucket group 1. When the LD in a cell is added by 1, the cell will be mapped to a physical bucket group at a different level, so that the key-value pair hashed to cell 2 will be mapped to the physical bucket of the first level (i.e., the target key-value pair storage bucket). Traverse the historical key-value pairs belonging to virtual bucket group 0, obtain the unit number of the historical key-value pairs, and if the unit number of the key-value pair virtual bucket group is equal to the unit number of the historical key-value pair (cell 0 in the legend), skip it; otherwise, once the unit number of this key-value pair virtual bucket group is not equal to the unit number of the historical key-value pair, transfer this key-value pair to the physical bucket group corresponding to another cell (cell 2 in the legend).
[0078] Obtain the remaining capacity of all the key-value pair storage buckets, and then determine whether the occupied capacity of the key value to be stored is greater than the remaining capacity of all the key-value pair storage buckets in the key-value pair storage bucket group. If the occupied capacity of the key-value pair to be stored is greater than the remaining capacity of all the key-value pair storage buckets in the key-value pair storage bucket group, obtain the remaining capacity of all the key-value pair storage buckets; select the target key-value pair storage bucket from all the key-value pair storage buckets according to the remaining capacity of all the key-value pair storage buckets, so as to transfer and store the historical key-value pairs in the key-value pair storage bucket to the target key-value pair storage bucket. Figure 5 As shown in Figure 1, when inserting a key-value pair into physical bucket 2 fails, an idle bucket is searched in the physical bucket group corresponding to the unit ( Figure 5 Move the historical key-value pairs in virtual bucket 2 to physical bucket 0, and store the key-value pairs to be stored in the key-value pair storage bucket. Then, after the move, update the guide array so that the moved key-value pairs can be relocated next time. Atomically increase , the increment is .
[0079] The basic operations of the hash table in this application (such as insertion, query and deletion) are as follows Figure 5As shown, hashing and modulo are first performed, that is, the key-value pair storage information is obtained, and the key-value pair keyword in the key-value pair storage information is calculated to obtain a unit number and a guide index, and then the location information of the key-value pair to be stored is determined based on the unit number and the guide index, so as to locate a unique virtual bucket, and then mapping is performed, that is, mapping the virtual bucket to a unique physical bucket, that is, inputting the unit number and the guide index into a preset mapper to obtain a layer index and a bucket offset, and then determining the location information of the key-value pair storage bucket based on the layer index and the bucket offset to obtain the key-value pair storage bucket. Insert the key-value pair into the physical bucket and persist it. The specific query operation is to traverse the historical key-value pairs belonging to the virtual bucket group No. 0 as mentioned above, obtain the unit number of the historical key-value pair, and if the unit number of the key-value pair virtual bucket group is equal to the unit number of the historical key-value pair (unit No. 0 in the legend), skip it; otherwise, once the unit number of this key-value pair virtual bucket group is not equal to the unit number of the historical key-value pair, transfer this key-value pair to the physical bucket group corresponding to another unit (unit No. 2 in the legend). The present invention uses SIMD to speed up the search and comparison process, which allows the search of the entire bucket to be completed with only one comparison. If the key-value pair is not in this physical bucket, then this key-value pair must not be in the hash table. When deleting a key-value pair, the present invention first queries this key-value pair in the hash table. If the key-value pair is found, the corresponding position of the bitmap is set to 0, indicating that the key-value pair at this position has expired, and then the bucket metadata is persisted to complete the deletion operation. In addition, if Figure 6 As shown, the present application expands the cell array, layer pointer and physical bucket layer in a doubling manner. After the expansion is successful, the new layer pointer will point to the newly generated physical bucket layer in sequence. Similar to the local rehashing operation, the LD and GD values of the cell are atomically increased. After changing the LD and GD according to the above formula, the mapping relationship between the cell and the physical bucket group will also change. Rehashing causes the insertion of a physical bucket group that failed. This is a little different from local rehashing. Since the destination physical bucket of the key-value pair has been determined before hashing (the same physical bucket at different layers), there is no need to re-establish the mapping relationship between the key-value pair and the physical bucket, eliminating unnecessary calculations, and then reinsert the key-value pair that failed to be inserted before. The present invention achieves an increase in read throughput of approximately 2.21 times, and the 99th percentile tail latency is only 1 / 3 of these solutions.
[0080] In this embodiment, a key-value pair storage bucket in a key-value pair storage bucket group is determined, and it is determined whether the remaining capacity in the key-value pair storage bucket is less than the occupied capacity of the key-value pair to be stored; if the remaining capacity in the key-value pair storage bucket is less than the occupied capacity of the key-value pair to be stored, then the relationship between the number of rehashing times of the pre-acquired key-value pair virtual bucket group and the number of extension times of the local hash table is determined; if the number of rehashing times is equal to the number of extension times, then the remaining capacity of all the key-value pair storage buckets in the key-value pair storage bucket group is obtained, and then the key-value pair transfer storage bucket is screened out from the key-value pair storage bucket group based on the capacity of the key value to be stored and the remaining capacity of all the key-value pair storage buckets in the key-value pair storage bucket group, so as to transfer and store the historical key-value pairs belonging to the key-value pair storage bucket to the key-value pair transfer storage bucket; a target key-value pair storage bucket is determined, and the key-value pair to be stored is stored in the target key-value pair storage bucket. This application combines scalable hashing with perfect hashing, and performs corresponding operations by judging the relationship between the number of rehashing times of a virtual bucket group of a key-value pair and the number of extension times of a local hash table, thereby eliminating the extra overhead caused by hash collisions during queries by introducing perfect hashing, thereby releasing the read performance of the index, and giving full play to the hardware characteristics of PMEM and the inherent advantages of the perfect hash index, thereby improving the efficiency of key-value pair storage, improving the read performance of the index in read-intensive and read-skewed scenarios, and reducing the overhead of maintaining index perfection.
[0081] See also Figure 7 As shown, an embodiment of the present invention discloses a key-value pair storage device, which may specifically include:
[0082] The key-value pair storage bucket determination module 11 is used to determine a key-value pair storage bucket in the key-value pair storage bucket group, and determine whether the remaining capacity in the key-value pair storage bucket is less than the occupied capacity of the key-value pair to be stored;
[0083] The judging module 12 is used to judge the relationship between the number of rehashing of the pre-acquired key-value pair virtual bucket group and the number of extensions of the local hash table if the remaining capacity in the key-value pair storage bucket is less than the occupied capacity of the key-value pair to be stored;
[0084] A historical key-value pair transfer module 13 is used to determine a key-value pair transfer storage bucket group if the number of rehashing times is less than the number of extension times, and transfer and store the historical key-value pairs belonging to the key-value pair virtual bucket group to the key-value pair transfer storage bucket group;
[0085] The target key-value pair storage module 14 is used to determine a target key-value pair storage bucket and store the key-value pair to be stored in the target key-value pair storage bucket.
[0086] In this embodiment, a key-value pair storage bucket in a key-value pair storage bucket group is determined, and it is determined whether the remaining capacity in the key-value pair storage bucket is less than the occupied capacity of the key-value pairs to be stored; if the remaining capacity in the key-value pair storage bucket is less than the occupied capacity of the key-value pairs to be stored, the relationship between the number of rehashing times of the pre-acquired key-value pair virtual bucket group and the number of extension times of the local hash table is determined; if the number of rehashing times is less than the number of extension times, a key-value pair transfer storage bucket group is determined, and the historical key-value pairs belonging to the key-value pair virtual bucket group are transferred and stored in the key-value pair transfer storage bucket group; a target key-value pair storage bucket is determined, and the key-value pairs to be stored are stored in the target key-value pair storage bucket. This application combines scalable hashing with perfect hashing, and performs corresponding operations by judging the relationship between the number of rehashing times of a virtual bucket group of a key-value pair and the number of extension times of a local hash table, thereby eliminating the extra overhead caused by hash collisions during queries by introducing perfect hashing, thereby releasing the read performance of the index, and giving full play to the hardware characteristics of PMEM and the inherent advantages of the perfect hash index, thereby improving the efficiency of key-value pair storage, improving the read performance of the index in read-intensive and read-skewed scenarios, and reducing the overhead of maintaining index perfection.
[0087] In some specific embodiments, the key-value pair storage bucket determination module 11 may specifically include:
[0088] An information acquisition module, used to acquire key-value pair storage information and calculate the key-value pair keywords in the key-value pair storage information to obtain a unit number and a guide index;
[0089] The position information determining module is used to determine the position information of the key-value pair to be stored in the key-value pair virtual bucket group based on the unit number and the guide index.
[0090] In some specific embodiments, the key-value pair storage bucket determination module 11 may specifically include:
[0091] A mapper output module, used for inputting the unit number and the guide index into a preset mapper to obtain a layer index and a bucket offset;
[0092] The key-value pair storage bucket determination module is used to determine the number of layers of the key-value pair storage bucket group and the location information of the key-value pair storage bucket based on the layer index and the bucket offset to obtain the key-value pair storage bucket.
[0093] In some specific embodiments, the historical key-value pair transfer module 13 may specifically include:
[0094] A screening module, used for screening out a target historical key-value pair to be transferred from all historical key-value pairs belonging to the key-value pair virtual bucket group;
[0095] The transfer module is used to transfer and store the target historical key-value pair to the key-value pair transfer storage bucket group.
[0096] In some specific embodiments, the historical key-value pair transfer module 13 may specifically include:
[0097] A unit number determination module, used to determine the unit number of the key-value pair virtual bucket group, and obtain the unit numbers of all historical key-value pairs belonging to the key-value pair virtual bucket group;
[0098] The judging module is used to judge whether the unit number of the key-value pair virtual bucket group is consistent with the unit number of the historical key-value pair. If the unit number of the key-value pair virtual bucket group is inconsistent with the unit number of the historical key-value pair, the historical key-value pair is used as the target historical key-value pair to be transferred.
[0099] In some specific embodiments, the determination module 12 may specifically include:
[0100] A module for obtaining a remaining capacity, configured to obtain the remaining capacity of all the key-value pair storage buckets in the key-value pair storage bucket group if the number of rehashing times is equal to the number of extension times;
[0101] A transfer storage module is used to filter out the key-value pair transfer storage bucket from the key-value pair storage bucket group based on the capacity of the key value to be stored and the remaining capacity of all the key-value pair storage buckets in the key-value pair storage bucket group, so as to transfer and store the historical key-value pairs belonging to the key-value pair storage bucket to the key-value pair transfer storage bucket.
[0102] In some specific embodiments, the target key-value pair storage module 14 may specifically include:
[0103] An expansion module is used to expand the key-value pair storage bucket group according to a preset expansion method to obtain a new key-value pair storage bucket group if the occupied capacity of the key-value pair to be stored is greater than the remaining capacity of all the key-value pair storage buckets in the key-value pair storage bucket group, and increase the values of the rehashing times and the extension times, and then jump to the step of determining the key-value pair storage buckets in the key-value pair storage bucket group.
[0104] Figure 8A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the key-value pair storage method performed by the electronic device disclosed in any of the aforementioned embodiments.
[0105] In this embodiment, the power supply 23 is used to provide working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device, and the communication protocol it follows is any communication protocol that can be applied to the technical solution of the present application, and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs and is not specifically limited here.
[0106] In addition, the memory 22, as a carrier for storing resources, can be a read-only memory, a random access memory, a disk or an optical disk, etc. The resources stored thereon include an operating system 221, a computer program 222 and data 223, etc. The storage method can be temporary storage or permanent storage.
[0107] Among them, the operating system 221 is used to manage and control the hardware devices and computer programs 222 on the electronic device 20 to realize the operation and processing of the data 223 in the memory 22 by the processor 21, which can be Windows, Unix, Linux, etc. In addition to including a computer program that can be used to complete the key-value pair storage method performed by the electronic device 20 disclosed in any of the aforementioned embodiments, the computer program 222 can further include a computer program that can be used to complete other specific tasks. In addition to data transmitted from an external device received by the key-value pair storage device, the data 223 can also include data collected by its own input and output interface 25, etc.
[0108] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0109] Furthermore, an embodiment of the present application also discloses a computer-readable storage medium, in which a computer program is stored. When the computer program is loaded and executed by a processor, the key-value pair storage method steps disclosed in any of the aforementioned embodiments are implemented.
[0110] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0111] The key-value pair storage method, device, equipment and storage medium provided by the present invention are introduced in detail above. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for general technical personnel in this field, according to the idea of the present invention, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.
Claims
1. A key-value pair storage method, characterized in that: include: Determine a key-value pair storage bucket in the key-value pair storage bucket group, and determine whether the remaining capacity in the key-value pair storage bucket is less than the occupied capacity of the key-value pair to be stored; If the remaining capacity in the key-value pair storage bucket is less than the occupied capacity of the key-value pair to be stored, determining the relationship between the number of rehashing of the pre-acquired key-value pair virtual bucket group and the number of extensions of the local hash table; If the number of rehashing times is less than the number of extension times, a key-value pair transfer storage bucket group is determined, and the historical key-value pairs belonging to the key-value pair virtual bucket group are transferred and stored in the key-value pair transfer storage bucket group; Determine a target key-value pair storage bucket, and store the key-value pair to be stored in the target key-value pair storage bucket; The transferring and storing the historical key-value pairs in the key-value pair virtual bucket group to the key-value pair transfer storage bucket group includes: selecting a target historical key-value pair to be transferred from all historical key-value pairs in the key-value pair virtual bucket group; transferring and storing the target historical key-value pair to the key-value pair transfer storage bucket group; The method of selecting a target historical key-value pair to be transferred from all historical key-value pairs belonging to the key-value pair virtual bucket group includes: determining the unit number of the key-value pair virtual bucket group, and obtaining the unit numbers of all historical key-value pairs belonging to the key-value pair virtual bucket group; judging whether the unit number of the key-value pair virtual bucket group is consistent with the unit number of the historical key-value pair, and if the unit number of the key-value pair virtual bucket group is inconsistent with the unit number of the historical key-value pair, taking the historical key-value pair as the target historical key-value pair to be transferred.
2. The key-value pair storage method according to claim 1, characterized in that: Before determining the key-value pair storage bucket in the key-value pair storage bucket group, the method further includes: Acquire key-value pair storage information, and calculate the key-value pair keywords in the key-value pair storage information to obtain a unit number and a guide index; The location information of the key-value pair to be stored in the key-value pair virtual bucket group is determined based on the unit number and the guide index.
3. The key-value pair storage method according to claim 2, characterized in that: The determining of the key-value pair storage bucket in the key-value pair storage bucket group includes: Input the unit number and the guide index into a preset mapper to obtain a layer index and a bucket offset; The number of layers of the key-value pair storage bucket group and the location information of the key-value pair storage bucket are determined based on the layer index and the bucket offset to obtain the key-value pair storage bucket.
4. The key-value pair storage method according to any one of claims 1 to 3, characterized in that: After determining the magnitude relationship between the number of rehashing times and the number of extension times, the method further includes: If the number of rehashing operations is equal to the number of extension operations, obtaining the remaining capacity of all the key-value pair storage buckets in the key-value pair storage bucket group; Based on the capacity of the key value to be stored and the remaining capacity of all the key-value pair storage buckets in the key-value pair storage bucket group, the key-value pair transfer storage bucket is screened out from the key-value pair storage bucket group so as to transfer and store the historical key-value pairs belonging to the key-value pair storage bucket to the key-value pair transfer storage bucket.
5. The key-value pair storage method according to claim 4, characterized in that: Also includes: If the occupied capacity of the key-value pairs to be stored is greater than the remaining capacity of all the key-value pair storage buckets in the key-value pair storage bucket group, the key-value pair storage bucket group is expanded according to a preset expansion method to obtain a new key-value pair storage bucket group, and the values of the rehashing times and the extension times are increased, and then the process jumps to the step of determining the key-value pair storage buckets in the key-value pair storage bucket group.
6. A key-value pair storage device, characterized in that: include: A key-value pair storage bucket determination module is used to determine a key-value pair storage bucket in the key-value pair storage bucket group, and determine whether the remaining capacity in the key-value pair storage bucket is less than the occupied capacity of the key-value pair to be stored; A judgment module, configured to judge the relationship between the number of rehashing of the pre-acquired key-value pair virtual bucket group and the number of extensions of the local hash table if the remaining capacity in the key-value pair storage bucket is less than the occupied capacity of the key-value pair to be stored; A historical key-value pair transfer module, configured to determine a key-value pair transfer storage bucket group if the number of rehashing times is less than the number of extension times, and transfer and store the historical key-value pairs belonging to the key-value pair virtual bucket group to the key-value pair transfer storage bucket group; A target key-value pair storage module is used to determine a target key-value pair storage bucket and store the key-value pair to be stored in the target key-value pair storage bucket; The transferring and storing the historical key-value pairs in the key-value pair virtual bucket group to the key-value pair transfer storage bucket group includes: selecting a target historical key-value pair to be transferred from all historical key-value pairs in the key-value pair virtual bucket group; transferring and storing the target historical key-value pair to the key-value pair transfer storage bucket group; The method of selecting a target historical key-value pair to be transferred from all historical key-value pairs belonging to the key-value pair virtual bucket group includes: determining the unit number of the key-value pair virtual bucket group, and obtaining the unit numbers of all historical key-value pairs belonging to the key-value pair virtual bucket group; judging whether the unit number of the key-value pair virtual bucket group is consistent with the unit number of the historical key-value pair, and if the unit number of the key-value pair virtual bucket group is inconsistent with the unit number of the historical key-value pair, taking the historical key-value pair as the target historical key-value pair to be transferred.
7. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the key-value pair storage method according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that: Used to store computer programs; wherein, when the computer program is executed by a processor, the key-value pair storage method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Write optimization extensible Hash index structure based on nonvolatile memory and insertion, refreshing and deletion methods
CN113342706A
Persistent memory dynamic hash indexing method, system and equipment and storage medium
CN114385636A