Storage-computation integrated indexing system and key-value pair storage system
By using the collaborative design of cross-point arrays and hash tables in memory, efficient indexing of key-values on data is achieved, solving the problem of low hash index efficiency in the prior art, and improving the parallelism and performance of the system.
Patent Information
- Application Number
- PCT/CN2024/073739
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-09
- Filing Date
- 2024-01-24
- Publication Date
- 2025-07-17
AI Technical Summary
The index operation of existing key-value storage systems is inefficient, especially the performance degradation caused by frequent hash collisions and memory fetches in hash indexes.
The integrated storage and computing indexing system is adopted, and the parallel data comparison capability of the cross-point array is used to complete the indexing operation of key-value data in memory. Combined with the hash table design, the in-situ indexing operation is realized, reducing the hash indexing steps and improving parallelism.
It significantly improves the index operation performance, reduces the hash indexing steps, especially the processing efficiency in hash conflicts, and improves the overall performance of the system and CPU cache utilization.
Smart Images

Figure CN2024073739_17072025_PF_FP_ABST
Abstract
Description
A storage and computing integrated index system and key-value pair storage system
Technical field
[0001] The present invention belongs to the intersection of storage and computing, and more specifically, relates to a storage-computing integrated indexing system and a key-value pair storage system. [Background Technology]
[0002] Key-value stores organize data using simple key-value pairs. Each data item consists of a unique key and associated data. These systems are widely used in high-performance, low-latency, scalable, and highly concurrent applications. Data indexes are a crucial component of high-performance key-value stores, enabling efficient data operations and enabling fast insertion, querying, updating, and deletion of data corresponding to a given key. While different index structures vary in their design, they typically fall into two broad categories: 1) Tree index structures, where the index is organized as a multi-level tree structure and operations are performed hierarchically from the root to the leaf nodes. Examples include B+ trees, log-structured merge trees, skip lists, and radix trees. 2) Hash index structures, where operations are performed by applying a hash function to a given key to obtain an index value. The index then locates the corresponding data within a linear list. Tree indexes require traversing multiple layers of tree nodes, and their query time complexity is typically O(logN), where N is the amount of data stored. Hash indexes, due to their flat storage structure (linear table), can provide query operations within the theoretical time complexity of O(1). However, hash indexes are inevitably affected by hash conflicts, that is, after the hash function is calculated, multiple keys are mapped to the same position in the linear table. Regardless of the conflict handling method used, such as the open chain method or the linear probing method, more data access and comparison will be introduced to handle the conflict.
[0003] Both the node traversal of tree indexes and the conflict resolution of hash indexes require more frequent comparisons and memory accesses, resulting in reduced indexing efficiency and impacting the performance of key-value storage systems. While researchers have proposed novel indexing structures, such as learned indexes, these structures are typically organized as trees, still requiring multiple rounds of computation, comparison, and memory access, leading to reduced indexing efficiency. The storage performance of key-value systems still needs to be further improved.
[0004] [Summary of the invention]
[0005] In response to the defects of the existing technology and the need for improvement, the present invention provides a storage and computing integrated indexing system and a key-value pair storage system, the purpose of which is to utilize the parallel data comparison capability of the crosspoint array to complete the indexing operation of key-value pair data in the memory, so as to improve the parallelism of the indexing operation and reduce the delay of the indexing operation.
[0006] To achieve the above objectives, according to one aspect of the present invention, there is provided a storage-computing integrated indexing system, comprising: a controller and a plurality of array clusters;
[0007] In the array cluster, each array cluster includes multiple arrays; the array is a cross-point array composed of resistive memory cells, and the content is addressable; in the array, each row is used to store a key and its valid flag; the value of the valid flag is f v When , it means that the corresponding row stores a valid key, and the value of the valid flag is f i , it means that the corresponding row does not store a valid key; f v ≠f i ; Initially, the valid flag bits of each row are f i ;Multiple array clusters can operate in parallel;
[0008] The controller is used to perform index operations on key-value pair data stored in the array; the index operations include: insert operations; the insert operations include:
[0009] (I1) For the key-value pair data to be inserted [k i ,v i ], in hash bucket B i Search the mapped array cluster for a valid flag f i If the search is successful, go to step (I2); otherwise, return the operation failure;
[0010] Hash bucket B i For key k i The hash value of the hash bucket corresponding to the hash table; in the hash table, each hash bucket is mapped to an array cluster, which is used to store the address of the array in the array cluster;
[0011] (I2) key-value pair data [k i ,v i ] key k i After storing in the allocated row, set its valid flag to f v , and the value v i Write to an array row.
[0012] Furthermore, the index operation also includes: a query operation; the query operation includes:
[0013] (S1) For the key k to be queried s , read hash bucket B in sequence s The array address in , and search the corresponding array for the key k s And the valid flag is f v Line L sIf the search is successful, go to step (S2); otherwise, return that the data does not exist;
[0014] Hash bucket B s For key k s The hash bucket corresponding to the hash value in the hash table;
[0015] (S2) From row L s Read the value v from the corresponding row used to store the value s and returns.
[0016] Furthermore, the index operation also includes: an update operation; the update operation includes:
[0017] (U1) For the key-value pair data to be updated [k u ,v u ], read hash bucket B in sequence u The array address in , and search the corresponding array for the key k u And the valid flag is f v Line L u If the search is successful, go to step (U2); otherwise, return that the data does not exist;
[0018] Hash bucket B u For key k u The hash bucket corresponding to the hash value in the hash table;
[0019] (U2) will be L u Updates the value stored in the row that stores the value to v u .
[0020] Furthermore, the index operation also includes: a delete operation; the delete operation includes:
[0021] (D1) For the key k to be deleted d , read hash bucket B in sequence d The array address in , and search the corresponding array for the key k d And the valid flag is f v Line L d If the search is successful, go to step (D2); otherwise, return that the data does not exist;
[0022] Hash bucket B d For key k d The hash bucket corresponding to the hash value in the hash table;
[0023] (D2) Move row L d The valid flag position is f i .
[0024] Furthermore, the controller is further configured to perform a hash expansion operation; the hash expansion operation includes:
[0025] (E1) Expand the capacity of the hash table to 2N and determine the correspondence between each hash bucket and array cluster; N is the current capacity of the hash table, and N is a power of 2;
[0026] (E2) determining the hash bucket corresponding to each valid key stored in each array cluster in the new hash table, and moving the valid key to the array cluster mapped by the corresponding hash bucket;
[0027] For any valid key k, the hash bucket number corresponding to its hash value in the original hash table is recorded as i, and the log2Nth bit in the hash value is determined to be 0. If so, the hash bucket number corresponding to the valid key k in the new hash table is determined to be i. Otherwise, the hash bucket number corresponding to the valid key k in the new hash table is determined to be i+N.
[0028] Furthermore, the insert operation also includes:
[0029] In the key k i At the same time as storing it in the assigned row, the key k i The high (H-log2N) bits of the hash value are stored as hash valid bits in an array row; the hash valid bits of the keys stored in the same array are stored in the same array and aligned by column; H represents the length of the hash value;
[0030] Furthermore, in the hash expansion operation, the log2Nth bit of the hash value of the valid key is obtained in the following ways:
[0031] Read the column of the hash valid bit corresponding to the log2Nth bit in the hash value from the array.
[0032] Furthermore, in the hash table, the size of each hash bucket is the same as the size of a CPU cache line.
[0033] Furthermore, in the hash table, different hash buckets are mapped to different array clusters.
[0034] According to another aspect of the present invention, a key-value pair storage system including the above-mentioned storage-computation-integrated indexing system provided by the present invention is provided.
[0035] In general, the above technical solutions conceived by the present invention can achieve the following beneficial effects:
[0036] (1) The storage-computation integrated indexing system provided by the present invention is implemented collaboratively based on a crosspoint array composed of resistive memory cells and a hash table. The content addressable function (parallel data comparison function) of the array is used to implement in-situ indexing operations, so that the key of a key-value pair is mapped to a hash bucket after the hash function is passed through. The in-situ indexing operation of the array is then used to complete the specific operation on the data. Due to the parallelism of the underlying hardware, the number of hash indexing steps is drastically reduced. In particular, when a hash conflict occurs, the present invention can significantly reduce the number of hash indexing steps. In general, the present invention can effectively improve the performance of indexing operations.
[0037] (2) In the preferred technical solution of the present invention, when performing hash expansion, based on the capacity relationship of the hash table before and after the expansion, only the corresponding bits in the hash value are used to determine the movement position of the data, so that the destination of all row data in the array can be completed within the index system, and the data movement in the memory can be completed without the participation of the CPU, thereby reducing the data movement between the CPU and the memory, improving the efficiency of hash expansion, reducing the impact on the index function during hash expansion, and further improving the overall index performance of the index system.
[0038] (3) In a further preferred technical solution of the present invention, hash valid bits are determined in advance to indicate the destination of valid keys in the array each time the hash is expanded, and these hash valid bits are stored in the array in a column-aligned manner. When the hash is expanded, the columns where the corresponding hash valid bits are located are read from the array in parallel according to the capacity of the current hash table, so as to determine the destination of each valid key in the array. The entire process is completed in the index system without the need for the CPU to participate in data movement, thereby further improving the efficiency of the hash expansion.
[0039] (4) In a further preferred embodiment of the present invention, the capacity of each hash bucket in the hash table is the same as a CPU cache line, thereby improving the CPU cache utilization.
[0040] (5) In a further preferred embodiment of the present invention, different hash buckets in the hash table are mapped to different array clusters, thereby maximizing the underlying parallelism of the hardware and further improving the overall indexing operation performance.
Brief Description of the Drawings
[0041] FIG1 is a schematic diagram of the structure of an existing resistive random access memory.
[0042] FIG. 2 is a schematic diagram of a cross-point array structure formed by an existing resistive random access memory.
[0043] FIG3 is a schematic diagram of a content addressable process of an array according to an embodiment of the present invention.
[0044] FIG4 is a schematic diagram of the hardware architecture of the storage-computing integrated indexing system provided in an embodiment of the present invention.
[0045] FIG5 is a schematic diagram of a software-hardware collaborative hash index provided by an embodiment of the present invention.
[0046] FIG6 is a schematic diagram of an insertion operation in an embodiment of the present invention.
[0047] FIG7 is a schematic diagram of a query operation in an embodiment of the present invention.
[0048] FIG8 is a schematic diagram of an update operation in an embodiment of the present invention.
[0049] FIG9 is a schematic diagram of a deletion operation in an embodiment of the present invention. [Specific implementation method]
[0050] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below may be combined with each other as long as they do not conflict with each other.
[0051] In the present invention, the terms "first", "second", etc. (if any) in the present invention and the drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0052] In order to solve the technical problem that existing indexing methods for key-value pair data all lead to frequent comparisons and memory accesses, resulting in a decrease in index operation efficiency, the present invention provides a storage and computing integrated indexing system and a key-value pair storage system. The overall idea is to implement in-situ indexing operations based on a hash table design on the basis of a content-addressable cross-point array, including insertion, query, update, deletion, etc., and support data movement within the memory, thereby simplifying the hash indexing steps and effectively improving the index operation performance.
[0053] Before explaining the technical solution of the present invention in detail, a brief introduction to the resistive memory cell and the cross-point array structure composed thereof is given as follows.
[0054] Figure 1 shows the structure of a resistive memory cell. Its basic structure consists of a resistive material, a bottom electrode, and a top electrode. The resistive material can be a metal oxide, for example. Its resistance can be reversibly switched between a high-resistance state and a low-resistance state under an applied electric field, enabling data to be written. Data can be read by applying a small voltage and measuring the outgoing current or the converted voltage without changing the cell state.
[0055] The crosspoint array composed of resistive memory cells is shown in Figure 2. A calculation voltage can be applied to the array word line (row) and the matrix-vector multiplication operation between the input calculation voltage vector and the conductivity matrix stored in the array can be obtained by reading the current flowing out of the bit line (column). The matrix-vector multiplication operation complexity implemented is only O(1), so it is widely used to accelerate applications dominated by matrix-vector multiplication, such as neural networks and graph computing. The crosspoint array can also implement parallel data comparison operations. This parallel data comparison function is also called content addressable function. Through the content addressable function, it is possible to determine the row in which the data stored in the array is the same as the input data. Figure 3 shows a schematic diagram of a feasible content addressable. In this implementation, two resistive memory cells are used to form a content addressable unit, and the corresponding two columns form an input comparison unit. The content stored in the content addressable unit can be determined based on the mapping relationship at the bottom of Figure 3. For example, if the logical value to be stored is 1, the data actually stored in the content addressable unit is (1,0). With the help of the input register, each 1-bit input data is converted into an input for two bit lines. The specific conversion rules are determined by the mapping relationship at the bottom of Figure 3. For example, when the input logic value "0" is converted to (0, 1) after entering the input comparison unit. The input of more bit lines enables parallel comparison of multiple bits of input data with the data stored in each row. If the data stored in a row of the array matches (is identical to) the given input data, the sense amplifier (SA) inputs a "1" signal; otherwise, it outputs a "0".
[0056] It should be noted that FIG3 shows only one method of implementing content addressability for a resistive memory cross-point array. Other implementation methods or methods of implementing content addressability for other non-volatile memories or even traditional DRAM, SRAM and other memories are all feasible.
[0057] In practical applications, key-value pair data is usually expressed as [key, value], where key represents the key and value represents the value corresponding to the key. In the following embodiments, a similar expression method will be adopted. Specifically, [k, v] is used to represent key-value pair data, where k and v represent the key and value respectively.
[0058] The following are examples.
[0059] Example 1:
[0060] A storage and computing integrated indexing system, as shown in FIG4 , includes: a controller and multiple array clusters;
[0061] Multiple array clusters can operate in parallel; each array cluster includes multiple arrays; the array is a cross-point array composed of resistive memory cells, and the content is addressable; in the array, each row is used to store a key and its valid flag; the valid flag is used to indicate whether the corresponding row stores a valid key; in this embodiment, when the value of the valid flag is 1, it indicates that the corresponding row stores a valid key, and when the value of the valid flag is 0, it indicates that the corresponding row does not store a valid key; initially, the valid flag of each row is 0.
[0062] Optionally, as shown in FIG4 , in this embodiment, the array specifically adopts a storage architecture similar to DRAM memory to construct storage-computing integrated indexing hardware to avoid complex on-chip interconnection structures. The input registers and sense amplifiers required for the content addressable function can be shared with the input registers and sense amplifiers used for the storage function. It is only necessary to modify the sense amplifier to support two reference value comparisons, one for normal read operations and the other for content addressable operations. It should be noted that the hardware organization form shown in FIG4 is only an optional implementation method of the embodiment of the present invention. In some other embodiments of the present invention, other organizational forms such as H-trees can also be used to complete the organization of the array.
[0063] The controller is used to perform indexing operations on the key-value pairs stored in the array. In this embodiment, the indexing operations specifically include insert operations, query operations, update operations, and delete operations. These operations are all completed within the array.
[0064] This embodiment stores the key and its valid flag in a row of the array, so that the input key and the valid flag can be compared in parallel. It is easy to understand that in the initialization state, all rows of the array have not yet been inserted with data, and the valid flag of each row is 0; after a row of data is deleted, its valid flag is also 0.
[0065] As shown in Figure 5, this embodiment designs a hash table that collaborates with both software and hardware based on the storage and computing integrated index hardware. In the hash table, each hash bucket is mapped to an array cluster for storing the addresses of the arrays in the array cluster; optionally, in this embodiment, the serial number of the array in the array cluster is used as the address of the array cluster, and each array address occupies 8B. The capacity of each hash bucket is the same as that of a CPU cache line, so that a hash bucket can store multiple array addresses to fill a CPU cache line, thereby improving the cache utilization of the CPU. In order to maximize the underlying parallelism of the hardware, in this embodiment, different hash buckets are mapped to different array clusters. Since different array clusters can operate in parallel, the system performance can be further improved.
[0066] In this embodiment, when a key value is mapped to a hash table, the hash bucket number mapped in the hash table is hash(k)%N, where hash(k) represents the hash value of key k, N represents the current capacity of the hash table, and % represents a remainder operation.
[0067] Based on the designed hash table, this embodiment maps the key k of the key-value pair [k, v] to a hash bucket after a hash function. Operations can then be performed on the arrays in the array cluster mapped to the hash bucket. Similar to traditional hash table-based indexing systems, the CPU calculates the hash value hash(k) for key k and locates the corresponding hash bucket in the hash table.
[0068] In this embodiment, the insert operation includes:
[0069] (I1) For the key-value pair data to be inserted [k i ,v i ], in hash bucket B i Search the mapped array cluster for a valid flag f i If the search is successful, go to step (I2); otherwise, return the operation failure;
[0070] Hash bucket B i For key k i The hash value of the hash bucket corresponding to the hash table; in the hash table, each hash bucket is mapped to an array cluster, which is used to store the address of the array in the array cluster;
[0071] (I2) key-value pair data [k i ,v i ] key k i After storing in the allocated row, set its valid flag to f v , and the value v i Write to an array row.
[0072] Figure 6 shows an example of an insertion operation. In this example, the key k to be inserted i =0011; The left side of Figure 6 shows searching for an empty row in the array, that is, searching for a row with a valid flag bit of 0, and the right side of Figure 6 shows writing data into the empty row found. It should be noted that if multiple empty rows are found in the array, a random empty row can be selected. In this embodiment, the first empty row is selected; if there is no empty row, the operation fails. In addition, because the value v data does not need to participate in data comparison, it can be simply stored in a storage unit instead of a content addressable unit. At the same time, the value v i can be stored in and key k i , the valid flags can be stored in the same array or in different arrays. You only need to obtain the matching row number of the key and then operate on the corresponding row of the array where the value is located.
[0073] In this embodiment, the query operation includes:
[0074] (S1) For the key k to be queried s , read hash bucket B in sequence s The array address in , and search the corresponding array for the key k s And the valid flag is f v Line L s If the search is successful, go to step (S2); otherwise, return that the data does not exist;
[0075] Hash bucket B s For key k s The hash bucket corresponding to the hash value in the hash table;
[0076] (S2) From row L s Read the value v from the corresponding row used to store the value s and returns.
[0077] Figure 7 shows an example of a query operation. In this example, the key k to be queried s =1010; the left side of Figure 7 shows the search for stored k s The right side of Figure 7 shows the required value v being read from the array row corresponding to the storage value of the row. s .
[0078] In this embodiment, the update operation includes:
[0079] (U1) For the key-value pair data to be updated [k u ,v u ], read hash bucket B in sequence u The array address in , and search the corresponding array for the key k u And the valid flag is f v Line L u If the search is successful, go to step (U2); otherwise, return that the data does not exist;
[0080] Hash bucket B u For key k u The hash bucket corresponding to the hash value in the hash table;
[0081] (U2) will be L u Updates the value stored in the row that stores the value to v u .
[0082] Figure 8 shows an example of an update operation. In this example, the key k in the key-value pair data to be updated is u =0100; the left side of Figure 8 shows the search for stored k uThe right side of FIG8 shows the row for updating the storage value corresponding to the row.
[0083] For the delete operation, only the valid flag of the matching row is cleared. Accordingly, in this embodiment, the delete operation includes:
[0084] (D1) For the key k to be deleted d , read hash bucket B in sequence d The array address in , and search the corresponding array for the key k d And the valid flag is f v Line L d If the search is successful, go to step (D2); otherwise, return that the data does not exist;
[0085] Hash bucket B d For key k d The hash bucket corresponding to the hash value in the hash table;
[0086] (D2) Move row L d The valid flag position is f i .
[0087] Figure 9 shows an example of a delete operation. In this example, the key k of the key-value pair data to be deleted is d =1010; the left side of Figure 9 shows the search for stored k d The right side of Figure 9 shows that the valid flag of the row is cleared.
[0088] This embodiment, based on a content-addressable array, designs a corresponding hash table, enabling in-situ indexing within the array. This eliminates a large number of CPU memory access operations, effectively reduces the number of indexing steps, and improves indexing efficiency. Furthermore, it fully utilizes the parallelism of the underlying hardware, further improving overall system performance.
[0089] As the number of key-value pairs that need to be indexed increases, the hash table will gradually be filled up. When the remaining space in the hash table is small, hash expansion is required. In the process of hash expansion, data movement is involved. In order to further improve the performance of the indexing system, this embodiment adopts in-memory data movement. During hash expansion, the calculation of the index value is converted from the original hash(k)%N to hash(k)%(2N). In this embodiment, the capacity N of the hash table is a power of 2. Therefore, after the data originally stored in the i-th hash bucket is moved to the new hash bucket, it may be moved to the i-th hash bucket or the i+N-th hash bucket of the new hash table, depending on whether the log2(N)-th bit of the hash value hash(k) is 0 or 1. In this embodiment, this bit is used as a hash valid bit. If this bit is 0, then after calculating the index value according to hash(k)%(2N), the array address of key k is still located in the i-th hash bucket. Otherwise, it is located in the i+N-th hash bucket. Based on this, in this embodiment, the controller is also used to perform hash expansion operations; the hash expansion operation includes:
[0090] (E1) Expand the capacity of the hash table to 2N and determine the correspondence between each hash bucket and array cluster; N is the current capacity of the hash table, and N is a power of 2;
[0091] (E2) determining the hash bucket corresponding to each valid key stored in each array cluster in the new hash table, and moving the valid key to the array cluster mapped by the corresponding hash bucket;
[0092] For any valid key k, the hash bucket number corresponding to its hash value in the original hash table is recorded as i, and the log2Nth bit in the hash value is determined to be 0. If so, the hash bucket number corresponding to the valid key k in the new hash table is determined to be i. Otherwise, the hash bucket number corresponding to the valid key k in the new hash table is determined to be i+N.
[0093] In practice, the high (H-log2N) bits of the key's hash value are used as hash valid bits, indicating the destination of valid keys in the array each time the hash is expanded. H represents the length of the hash value. To facilitate the use of this information, in this embodiment, these hash valid bits are stored in the array during data insertion. By using the array's parallel read function, a hash valid bit of all rows in the array can be read. This allows the destination of all rows of data in the array to be known in the memory, completing the data movement within the memory. Accordingly, in this embodiment, the insertion operation also includes:
[0094] In the key k i At the same time as storing it in the assigned row, the key k i The high (H-log2N) bits of the hash value are stored as hash valid bits in an array row; the hash valid bits of the keys stored in the same array are stored in the same array and aligned in columns;
[0095] Furthermore, in the hash expansion operation, the log2Nth bit of the hash value of the valid key is obtained in the following ways:
[0096] Read the column of the hash valid bit corresponding to the log2Nth bit in the hash value from the array.
[0097] In this embodiment, the data in-memory movement in hash expansion can be implemented for all valid keys stored in an array by supporting instructions similar to mv(ori_addr, to_addr0, to_addr1, bit_pos), where ori_addr represents the original array address, bit_pos represents the bit_pos-th hash valid bit in the hash value corresponding to the key; to_addr0 and to_addr1 represent the array address to which the data in the original array needs to be moved when the bit_pos-th hash valid bit is 0 and 1, respectively, and are determined according to the array clusters mapped by the i-th hash bucket and the i+N-th hash bucket in the new hash table, respectively; when the hash valid bit is 0, the valid key is moved from the array corresponding to ori_addr to the array corresponding to to_addr0; when the hash valid bit is 1, the valid key is moved from the array corresponding to ori_addr to the array corresponding to to_addr1.
[0098] Through the above-mentioned hash expansion operation, this embodiment can complete the hash expansion only by moving the in-memory data without the need for CPU participation, thereby further improving the efficiency of the hash expansion.
[0099] Example 2:
[0100] A key-value pair storage system including the storage-computation-integrated indexing system provided in the above-mentioned embodiment 1.
[0101] In the key-value pair storage system, the index system defines the key-value pair data storage method and completes the storage of key-value pair data. Based on the storage and computing integrated index system provided in the above embodiment 1, the key-value pair storage system provided in this embodiment can provide efficient data operations and effectively improve performance.
[0102] It is easy to understand that this embodiment also includes other modules that cooperate with the indexing system, such as monitoring and management, security and permission control, load balancing, backup and other modules to improve availability and reliability. The specific implementation methods of these modules are the same as those of traditional key-value storage systems and will not be repeated here.
[0103] It will be easily understood by those skilled in the art that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A computing-in-memory indexing system, characterized in that Comprising: A controller and multiple array clusters; The multiple array clusters can operate in parallel; each array cluster includes a plurality of arrays; the arrays are cross-point arrays composed of resistive memory cells and are content-addressable; in each of the arrays, each row is used to store a key and its valid flag bit; the value of the valid flag bit is f v When it is, it means that the corresponding row stores a valid key, and the value of the valid flag bit is f i When it is, it means that the corresponding row does not store a valid key; f v ≠f i ; At the initial moment, the valid flag bits of each row are all f i ; The controller is configured to perform an indexing operation on the key-value pair data stored in the array; The indexing operation includes: an insertion operation; the insertion operation includes: (I1) For the key-value pair data [k i , v i to be inserted, search for a row with a valid flag of f i in the array cluster mapped by the hash bucket B i . If the search is successful, go to step (I2); otherwise, return operation failure; Hash bucket B i is the hash bucket corresponding to the hash value of key k i in the hash table; in the hash table, each hash bucket is mapped to an array cluster for storing the addresses of the arrays in the array cluster. (I2) the key-value pair data [k i ,v i ] key i After storing in the allocated row, set its valid flag to f v , and the value v i Write to an array row.
2. The in-memory computing index system according to claim 1, characterized in that The indexing operation further includes: a query operation; the query operation includes: (S1) For the key k to be queried s , sequentially read the array addresses in the hash bucket B s , and search in the corresponding array for the row L s that stores the key k v and whose valid flag is f s . If the search is successful, go to step (S2); otherwise, return that the data does not exist. Hash bucket B s is the hash bucket corresponding to the hash value of key k s in the hash table; (S2) Read the value v from the corresponding row L s for storing the value and return it. s 3. The in-memory computing index system according to claim 1, wherein The indexing operation further includes: an update operation; the update operation includes: (U1) For the key-value pair data [k u , v u to be updated, sequentially read the array addresses in the hash bucket B u and search in the corresponding array for the row L u that stores the key k v and has a valid flag of f u . If the search is successful, go to step (U2); otherwise, return that the data does not exist. Hash bucket B u is the hash bucket corresponding to the hash value of key k u in the hash table; (U2) Update the value stored in row L u to value v for the row used to store the value u .
4. The in-memory computing index system according to claim 1, wherein The indexing operation further includes: a deletion operation; the deletion operation includes: (D1) For the key k to be deleted d , sequentially read the array addresses in the hash bucket B d , and search for the row L that stores the key k d and the valid flag bit is f v in the corresponding array. d If the search is successful, go to step (D2); otherwise, return that the data does not exist. Hash bucket B d for key k d is the hash bucket corresponding to the hash value of key k in the hash table; (D2) Set the valid flag of row L d to f i .
5. The in-memory computing index system according to any one of claims 1 to 4, characterized in that The controller is further configured to perform a hash expansion operation; the hash expansion operation includes: (E1) Expand the capacity of the hash table to 2N, and determine the correspondence between each hash bucket and the array cluster; N is the capacity of the current hash table, and N is a power of 2; (E2) Determine the hash bucket corresponding to each valid key stored in each array cluster in the new hash table, and move the valid key to the array cluster mapped by the corresponding hash bucket; For any valid key k, record the hash bucket number corresponding to its hash value in the original hash table as i, and determine whether the log2N-th bit in the hash value is 0. If so, determine that the hash bucket number corresponding to the valid key k in the new hash table is i; otherwise, determine that the hash bucket number corresponding to the valid key k in the new hash table is i + N.
6. The in-memory computing index system according to claim 5, wherein The insertion operation further includes: While storing the key k i into the allocated row, store the high (H - log2N) bits of the hash value of the key k i as hash valid bits into an array row; for keys stored in the same array, their hash valid Bits are stored in the same array and aligned by column; H represents the length of the hash value; Moreover, in the hash expansion operation, the method for obtaining the log2N-th bit in the hash value of the valid key includes: Read the column where the hash valid bit corresponding to the log2N-th bit in the hash value is located from the array.
7. The in-memory computing indexing system according to any one of claims 1 to 4, characterized in that In the hash table, the size of each hash bucket is the same as the size of a CPU cache line.
8. The in-memory computing index system according to any one of claims 1 to 4, characterized in that In the hash table, different hash buckets are mapped to different array clusters.
9. A key-value pair storage system including the in-memory computing indexing system according to any one of claims 1 to 8.
Citation Information
Patent Citations
Writing method for resistive random access memory based on crosspoint array
CN108053852A
Data indexing method and device and electronic equipment
CN113157689A
System and method for using hash table having set of frequently accessed buckets and set of non-frequently accessed buckets
CN114830108A
Data storage method and related equipment
CN115729847A
Hash table processing method and device and electronic equipment
CN116719813A