A method, device and equipment for index query based on database

By designing a hash-based index query method in the database and using the page index structure for index query, the problem of inefficient index query in read-only scenarios is solved, and efficient index query and good space utilization are achieved.

CN119396897BActive Publication Date: 2025-05-16BEIJING VOLCANO ENGINE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411412975.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-10
Publication Date
2025-05-16
Estimated Expiration
2044-10-10

AI Technical Summary

Technical Problem

The prior art lacks an optimization solution for indexes in read-only scenarios, resulting in inefficient index query.

Method used

A database-based index query method is designed, using hash values ​​to determine page identity and obtain the corresponding page index file. The page index file has a page index structure, including an entry array, a block bitmap and a length array, and supports key-value pair searches with O(1) complexity.

Benefits of technology

It realizes efficient index query in read-only scenarios, improves data query efficiency, and since storage is continuous, it ensures high space utilization and avoids additional data parsing work.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119396897B_ABST
    Figure CN119396897B_ABST
Patent Text Reader

Abstract

The present application discloses a method, device and equipment for index query based on a database, which optimizes indexes for read-only scenarios. The method includes: determining the corresponding page identifier based on the hash value of the target key in the data query request; obtaining a page index file with a page index structure corresponding to the page identifier; the page index structure includes an entry array, a block bitmap and a length array, the entry array includes the entry data of the page index file, and is sorted in sequence according to the order of the buckets and the order in the buckets; the block bitmap includes a first field and a second field corresponding to each block, the first field uses a first preset number of bits to save the status value of whether the bucket constituting the block is an empty bucket, and the second field saves the first number value of the entry data included in all blocks before the current block; the length array includes a unary code of a non-empty bucket, each unary code corresponds to the number of entry data included in a non-empty bucket; in the page index file, query whether there is a value corresponding to the target key.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of database technology, and in particular to a database-based index query method, device and equipment. Background Art

[0002] In databases, especially distributed databases, the index structure is crucial to the storage engine. The index can quickly filter out unnecessary data, find the target data, and improve data query efficiency.

[0003] HashMap (Hash-based Map) or HashTable (Hash Table) is a very efficient query structure based on Hash, which can find the corresponding Value according to the Key, making it an essential data structure for almost every development language. There are many mature and efficient algorithm implementations for HashMap, which usually provide general capabilities such as add / delete / query.

[0004] In some scenarios, once a read-only file is built, it will not be changed. Therefore, during the query phase, the index structure to be loaded can be read-only and does not need to support update operations such as add / delete. However, there is currently no mature solution for optimizing indexes for read-only scenarios. Summary of the invention

[0005] In view of this, embodiments of the present application provide a database-based index query method, apparatus, and device to optimize indexes for read-only scenarios.

[0006] To solve the above problems, the technical solutions provided in the embodiments of the present application are as follows:

[0007] In a first aspect, an embodiment of the present application provides an index query method based on a database, the method comprising:

[0008] Obtaining a data query request for a database, wherein the data query request includes a target key;

[0009] Determine a corresponding page identifier based on the hash value of the target key;

[0010] Obtain a page index file corresponding to the page identifier, wherein the page index file has a page index structure; the page index structure includes an entry array, a block bitmap, and a length array, wherein the entry array includes entry data of the page index file, wherein the entry data is sequentially sorted according to the order of the buckets in which they are located and the order within the buckets, and each of the entry data includes a key-value pair of an index; the block bitmap includes a first field and a second field corresponding to each block, wherein the first field is used to save a status value of whether a bucket constituting the block is an empty bucket using a first preset number of bits, and the second field is used to save a first quantity value of the entry data included in all blocks before the current block; the length array includes unary codes of non-empty buckets, and each of the unary codes corresponds to the quantity of entry data included in a non-empty bucket;

[0011] In the page index file, it is queried whether there is a value corresponding to the target key.

[0012] In a second aspect, an embodiment of the present application provides an index query device based on a database, the device comprising:

[0013] A first acquisition unit, configured to acquire a data query request for a database, wherein the data query request includes a target key;

[0014] A first determining unit, configured to determine a corresponding page identifier based on a hash value of the target key;

[0015] a second acquisition unit, configured to acquire a page index file corresponding to the page identifier, wherein the page index file has a page index structure; the page index structure includes an entry array, a block bitmap, and a length array, the entry array includes entry data of the page index file, the entry data are sequentially sorted according to the order of the buckets in which they are located and the order within the buckets, and each of the entry data includes a key-value pair of an index; the block bitmap includes a first field and a second field corresponding to each block, the first field is used to save a status value of whether the bucket constituting the block is an empty bucket using a first preset number of bits, and the second field is used to save a first quantity value of the entry data included in all blocks before the current block; the length array includes a unary code of a non-empty bucket, and each of the unary codes corresponds to the quantity of entry data included in a non-empty bucket;

[0016] The query unit is used to query whether there is a value corresponding to the target key in the page index file.

[0017] In a third aspect, an embodiment of the present application provides a database-based index query device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the database-based index query method as described above is implemented.

[0018] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which instructions are stored. When the instructions are executed on a terminal device, the terminal device executes the database-based index query method as described above.

[0019] It can be seen that the embodiments of the present application have the following beneficial effects:

[0020] In an embodiment of the present application, the entire index file can be divided into multiple page index files. After obtaining the target key in the data query request for the database, the hash value of the target key is first determined to determine the corresponding page identifier, thereby determining the page index file corresponding to the page identifier. The page index file has a page index structure, which is designed for read-only scenarios. In the page index structure, it includes an entry array, a block bitmap, and a length array. The entry array continuously stores the entry data of the page index file, and the entry data is sorted in the order of the buckets and the order in the bucket. N buckets form a block, and the block bitmap includes a first field and a second field corresponding to each block. The first field includes N bits, and the N bits correspond to the status value of whether the N buckets constituting the block are empty buckets. The second field corresponds to the first number value of the entry data included in all blocks before the current block. The length array continuously stores the unary codes of non-empty buckets, and each unary code corresponds to the number of entry data included in a non-empty bucket. Through the page index file, it can be queried whether there is a value corresponding to the target key. Since the page index structure is designed for read-only scenarios, the storage is continuous, which ensures a high space utilization rate. And it can be directly loaded from other storage media into the memory for use, without the need for additional data parsing, thus achieving the optimization of indexes for read-only scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 A flowchart of a database-based index query method provided in an embodiment of the present application;

[0022] Figure 2 A schematic diagram of the overall index architecture in an embodiment of the present application;

[0023] Figure 3 A schematic diagram of the logical structure of a page index in an embodiment of the present application;

[0024] Figure 4 A schematic diagram of the physical structure of a page index in an embodiment of the present application;

[0025] Figure 5 This is a schematic diagram of the process of generating the initial Chunk Bitmap in an embodiment of the present application;

[0026] Figure 6 This is a schematic diagram of the process of generating Chunk Bitmap in an embodiment of the present application;

[0027] Figure 7 This is a schematic diagram of the process of associating Entry data with Chunk in an embodiment of the present application;

[0028] Figure 8 A schematic diagram of an index query device provided in an embodiment of the present application;

[0029] Fig. 9 A schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0030] In order to make the above-mentioned purposes, features and advantages of the embodiments of the present application more obvious and easy to understand, the embodiments of the present application are further described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0031] In order to facilitate understanding and explanation of the technical solution provided by the embodiments of the present application, the background technology of the embodiments of the present application will be described below.

[0032] In (distributed) databases, the index structure is crucial to the storage engine. Through the index, you can quickly filter out unnecessary data, find the target data, and improve query efficiency.

[0033] In some scenarios, read-only files will not be changed after they are built. For example, a database based on LSM-Tree (Log-Structured Merge Tree) always appends data, and usually stores the data in a local disk or remote distributed file storage in a manner similar to SSTable files. The generated SSTable files will not be modified again, but will only be merged to generate new SSTable files. This opens the door to convenient index design. The index will not be changed after it is built with the read-only file. Therefore, in the query phase, the index structure to be loaded is read-only and does not need to support update operations such as add / delete.

[0034] In actual applications, during the data query phase, data needs to be loaded into memory, and the computing engine is used to complete data filtering, aggregation, analysis, and other computing processes. Since the process of loading data from the original storage medium into the memory has a large delay and poor efficiency, a memory cache system is usually maintained inside the computing engine. If the data to be loaded already exists in the cache, the process of loading the data into the memory can be avoided, thereby improving computing efficiency. In the cache system, the cache object is usually designed as a block / page of the original data (usually at the KB or MB level). Correspondingly, data is loaded from the data file in units of Pages, that is, each time a Page is read from the file into the memory for use by the computing engine. If the Page exists in the cache, the Page data in the cache can be used directly, otherwise it is read from the file and placed in the cache system for the next query.

[0035] In order to adapt to this Page organization method, the index data in the file also needs to be organized in Page mode. When querying the index, only part of the Page index data needs to be loaded, and the complete index data does not need to be loaded. In addition, after obtaining the Page index data, it needs to be parsed into an index object and the query interface capability of the index is attached before the index query can be performed. In order to ensure query efficiency, the process of parsing the Page index data needs to be lightweight enough.

[0036] Based on this, the database-based index query method, device and equipment provided in the embodiments of the present application realize the optimization of the index for read-only scenarios, and the index structure can solve the following problems: provide overall Key / Value search capability to ensure O(1) computational complexity; be able to load indexes on demand at Page granularity and ensure the lightweight parsing process; design an efficient HashMap data structure to ensure O(1) search and lightweight parsing while also ensuring high space utilization and avoiding space waste; based on the structure, design a corresponding construction algorithm; based on the structure, design a corresponding search algorithm to ensure O(1) computational complexity.

[0037] To facilitate understanding of the embodiments of the present application, a database-based index query method provided in the embodiments of the present application is described below in conjunction with the accompanying drawings.

[0038] See also Figure 1 As shown in FIG. 1 , this figure is a flowchart of a database-based index query method provided in an embodiment of the present application, such as Figure 2 As shown, the method may include S101-S104:

[0039] S101: Obtain a data query request for a database, where the data query request includes a target key.

[0040] The embodiment of the present application can be applied to a storage engine, which includes a database. When data query is needed in the database, the storage engine can obtain a data query request for the database, which includes a target key (target Key) for querying the value (Value) corresponding to the target key in the database.

[0041] S102: Determine a corresponding page identifier based on the hash value of the target key.

[0042] When querying data, you can use the index to quickly filter out unnecessary data and find the target data. For example, a HashMap index is pre-established, and the Key value of the HashMap index is the data value of a column in the data table, and the Value value is the corresponding row number. When querying the index, enter the required data value as the target Key to check whether there is a corresponding row number. If it exists, you can quickly output the row number to get the value corresponding to the target Key.

[0043] Since the index can be divided into multiple page index files, the index can logically include two levels of indexes. Figure 2 As shown, a schematic diagram of the overall index architecture in an embodiment of the present application is shown. The first-level index uses the number of Pages in the metadata (Meta). For any Key, it is divided into the corresponding page index files of the lower layer according to its Hash value. Among them, the metadata is the data describing the index, which may include parameters such as the number of Pages. The second-level index consists of multiple physical page index files (Page), and each Page has an independent HashMap index.

[0044] The corresponding page ID is determined based on the hash value of the target key, so as to further search the index in the page index file corresponding to the page ID to narrow the search scope.

[0045] S103: Obtain a page index file corresponding to the page identifier, wherein the page index file has a page index structure. The page index structure includes an entry array, a block bitmap, and a length array. The entry array includes entry data of the page index file, and the entry data is sorted in sequence according to the order of the buckets in which they are located and the order in the buckets. Each entry data includes a key-value pair of an index. The block bitmap includes a first field and a second field corresponding to each block. The first field is used to save the state value of whether the bucket constituting the block is an empty bucket using a first preset number of bits, and the second field is used to save the first number value of the entry data included in all blocks before the current block. The length array includes a unary code of a non-empty bucket, and each unary code corresponds to the number of entry data included in a non-empty bucket.

[0046] After determining the page identifier of the target key, the page index file corresponding to the page identifier can be obtained. In actual applications, it can be confirmed whether the page index file is in the memory first. If not, the page index file is loaded into the memory. If it is, the page index file is obtained from the memory. The page index file has a page index structure, which is an index structure optimized according to the read-only scenario in the embodiment of the present application.

[0047] See also Figure 3 As shown, a schematic diagram of the logical structure of the page index in the embodiment of the present application is shown. The page index file uses a HashMap index, and the HashMap index logically adopts a chain (Separate Chaining) conflict resolution method, that is, the Entry data on the same bucket (which may be empty) is organized in a logical chain manner, and all Entry data in a bucket can be traversed sequentially. Each Entry data is a Key-Value pair of the index.

[0048] See also Figure 4 FIG. 1 is a schematic diagram showing the physical structure of a page index in an embodiment of the present application. In an embodiment of the present application, the page index is physically composed of three parts:

[0049] The Entry Array stores all Entry data, and each Entry data is a Key / Value pair. Logically, the Entry data on the same bucket (chain) is physically stored in the Entry Array in the order of the bucket. The order of Entry data on different buckets (chains) is consistent with the order of the buckets (chains). That is, the Entry data in the front bucket is also at the front in the Entry Array.

[0050] Chunk Bitmap, a first preset number of N buckets constitutes a chunk. For example, the first preset number is 32, that is, every 32 buckets constitute a chunk. In the Chunk Bitmap, it includes a first field and a second field corresponding to each chunk, and each field includes a first preset number of bits. That is, each chunk uses 2 times the first preset number of bits to describe the chunk. The first field uses N bits to save the status value of whether the bucket constituting the chunk is an empty bucket, and the second field saves the first quantity value of the entry data included in all chunks before the current chunk. The first field and the second field corresponding to each chunk are arranged consecutively, and each chunk is arranged in order according to the chunk identifier. For example, in the Chunk Bitmap, a 64-bits Word is used to describe the chunk, and the first 32 bits of the Word are the first field, corresponding to the status bitmap of the 32 buckets in the chunk. A bit value (status value) of 1 indicates that the bucket is not empty, and a bit value (status value) of 0 indicates that the bucket is empty. Considering the conflict, the number of Entry data in the bucket may be greater than 1. The last 32 bits of the word are the second field, which is the prefixPopCount (first quantity value) of the uint32 integer, indicating the total number of Entry data of all previous Chunks. For example, the first quantity value of the second Chunk is 50, which means that the total number of Entry data in the first Chunk is 50. In addition, the first preset number is 8 as an example in the figure, and the first preset number can be set according to the actual situation. A block bitmap is composed of multiple 64-bits Words.

[0051] The Length Array uses unary coding to describe the chain length of each non-empty bucket, that is, the number of Entry data included in the non-empty bucket. For example: 1 represents a length of 1, 01 represents a length of 2, 0001 represents a length of 4, and so on. The entire Length Array is a flat storage representation of the unary coding of all non-empty buckets, and each Entry data occupies 1 bit in the Length Array.

[0052] In terms of space occupied by the page index structure, assuming that the total number of Entry data is M, the total space occupied by all Key / Value is S, so the number of bytes occupied by the Entry Array is S. Assuming that the load factor is 1 / 4, that is, there are 4*M buckets, the number of Chunks is 4*M / 32=M / 8, and the number of bytes occupied by the entire Chunk Bitmap is M / 8*8=M, (one Chunk occupies 8 bytes). In the Length Array, each Entry data occupies 1 bit, so the number of bytes occupied by the entire LengthArray is M / 8. In summary, the overall space occupied by the page index structure is S+M+M / 8=S+1.125M. That is, in addition to the space S that must be occupied, the additional space required for the HashMap index is 1.125N bytes. In the case of a lower load factor, the additional space occupied is already very low data. Ensure high space utilization and avoid space waste.

[0053] In addition, after reading the page index file from other storage media into the memory, only the memory address of the data and the total number of Entry data in the metadata are needed to complete the initialization of the page index object. Therefore, the page index file can be loaded into the memory on demand with the page as the granularity, and no additional data parsing work is required.

[0054] S104: In the page index file, check whether there is a value corresponding to the target key.

[0055] In the page index file, check whether the value corresponding to the target key exists. If it exists, the value corresponding to the target key will also be returned.

[0056] In an embodiment of the present application, the entire index file can be divided into multiple page index files. After obtaining the target key in the data query request for the database, the hash value of the target key is first determined to determine the corresponding page identifier, thereby determining the page index file corresponding to the page identifier. The page index file has a page index structure, which is designed for read-only scenarios. In the page index structure, it includes an entry array, a block bitmap, and a length array. The entry array continuously stores the entry data of the page index file, and the entry data is sorted in the order of the buckets and the order in the bucket. N buckets form a block, and the block bitmap includes a first field and a second field corresponding to each block. The first field includes N bits, and the N bits correspond to the status value of whether the N buckets constituting the block are empty buckets. The second field corresponds to the first number value of the entry data included in all blocks before the current block. The length array continuously stores the unary codes of non-empty buckets, and each unary code corresponds to the number of entry data included in a non-empty bucket. Through the page index file, it can be queried whether there is a value corresponding to the target key. Since the page index structure is designed for read-only scenarios, the storage is continuous, which ensures a high space utilization rate. And it can be directly loaded from other storage media into the memory for use, without the need for additional data parsing, thus achieving the optimization of indexes for read-only scenarios.

[0057] Based on the page index structure in the embodiment of the present application, a corresponding search algorithm is also provided. In a possible implementation, the specific implementation of S104 querying whether there is a value corresponding to the target key in the page index file may include:

[0058] A1: Determine the corresponding bucket ID based on the hash value of the target key.

[0059] According to the hash value of the target key, the bucket identifier BID of the bucket where the target key is located can be calculated, and the bucket identifiers BID of each bucket are arranged in order. Specifically, assuming that the number of buckets is X, the bucket identifier BID can be obtained by modulo X with the hash value of the target key.

[0060] A2: Determine the corresponding block ID based on the bucket ID.

[0061] Since N buckets form a chunk, the chunk ID can be obtained from the bucket ID. Taking N as 32 as an example, BID / 32 is rounded to the integer to obtain the chunk ID.

[0062] A3: Query whether the bucket corresponding to the bucket identifier is empty in the first field corresponding to the block identifier in the block bitmap.

[0063] The area corresponding to the Chunk can be determined from the Chunk Bitmap by the Chunk ID, thereby obtaining the N bits of the first field corresponding to the Chunk and the first quantity value in the second field. In the N bits of the first field, the bits corresponding to the bucket identifier can be obtained to determine whether the bucket is empty. For example, if the Chunk ID is 0, the first field and the second field corresponding to the Chunk are obtained from the first 64 bits of the block bitmap. Since it is the first Chunk, the first quantity value in the second field is 0. Then the corresponding bit value (status value) is obtained by the bucket identifier.

[0064] A4: If the bucket corresponding to the bucket identifier is empty, the value corresponding to the target key that does not exist is returned.

[0065] If the bit value is 0, it means that the bucket corresponding to the bucket identifier is empty, indicating that the value corresponding to the target key does not exist and is returned directly.

[0066] A5: If the bucket corresponding to the bucket identifier is not empty, determine the starting position and number of entries of the bucket corresponding to the bucket identifier in the entry array according to the block bitmap and the length array.

[0067] If the bit value is 0, it means that the bucket corresponding to the bucket identifier is not empty, and the starting position of the bucket identifier in the Entry Array and the number of entries can be calculated.

[0068] In a possible implementation, if the bucket corresponding to the bucket identifier is not empty, A5 may include, according to the block bitmap and the length array, determining the starting position and the number of entries of the bucket corresponding to the bucket identifier in the entry array:

[0069] B1: If the bucket corresponding to the bucket identifier is not empty, obtain the number of non-empty buckets before the bucket corresponding to the bucket identifier in the first field corresponding to the block identifier in the block bitmap and record it as the second quantity value, and obtain the first quantity value in the second field corresponding to the block identifier in the block bitmap and record it as the third quantity value.

[0070] If the bit value is 0, it means that the bucket corresponding to the bucket identifier is not empty. In the current Chunk, the number of non-empty buckets before the bucket identifier BID is the second quantity value B, which is less than 32. The first quantity value prefixPopCount of the current Chunk is taken as the third quantity value P.

[0071] B2: In the length array, starting from the position where the third quantity value is added by one, read the second number of unary codes, and accumulate the number of entry data corresponding to the second number of unary codes to obtain a fourth quantity value.

[0072] In the Length Array, starting from the P+1th bit position, B unary codes are read, the number of Entry data corresponding to these unary codes is accumulated, and the sum is recorded as the fourth quantity value S.

[0073] B3: Add the third data value and the fourth quantity value to obtain the starting position of the bucket corresponding to the bucket identifier in the entry array, and read the quantity of entry data corresponding to the next unary code as the number of entries.

[0074] The starting subscript of the bucket (BID bucket chain) corresponding to the bucket identifier in the Entry Array is P+S. In the LengthArray, continue to read the next unary code. The number of Entry data corresponding to the unary code is the number of Entry data (number of entries) included in the bucket (BID bucket chain).

[0075] A6: Query from the starting position of the entry array whether the value corresponding to the target key exists in the number of entries.

[0076] According to the starting index and number of entries of the bucket (BID bucket chain) in the Entry Array, read the Key of each Entry data in the chain in the bucket from the EntryArray in turn. Make an equal value judgment with the target Key. If the Key is equal, return the corresponding Value value. If the Entry data with the same Key is not found, the search fails and the value corresponding to the target key does not exist in the index.

[0077] Calculate the complexity of the index search process: Based on the target Key and Hash value and other information, the bucket identifier can be directly determined, and the corresponding Chunk information can be directly read in the Chunk Bitmap. Based on the bucket identifier and Chunk information, at most B unary encoded values ​​(B<32) are read in the LengthArray to calculate the starting position and length information of the bucket chain. This process is at a constant level and has nothing to do with the number of entries. In the bucket chain, search for matching Key values. Under the premise of Hash function balance, the length of the chain (that is, the number of entries in a bucket) is limited. This is also the theoretical basis of all HashMap algorithms. The search complexity within the bucket chain is at a constant level, which guarantees the computational complexity of O(1).

[0078] Therefore, based on the index structure of the embodiment of the present application, a corresponding index search algorithm can be implemented, and the overall computational complexity of the index search algorithm is guaranteed to be O(1).

[0079] Based on the page index structure in the embodiment of the present application, a corresponding index structure construction algorithm is also provided. In a possible implementation, the database-based index query method provided in the embodiment of the present application further includes:

[0080] C1: Determine the total amount of data for all indexed key-value pairs, and generate the number of pages based on the total amount of data and the page size parameter.

[0081] In the embodiment of the present application, the overall second-layer index can be constructed first, and then the page index file in the second-layer index can be constructed. First, the Key-Value of all indexes is collected to obtain all Entry data. The total data volume TotalSize of all Entry data is calculated. According to the expected PageSize (which can be pre-determined as a parameter), a suitable number of Pages Y is generated using TotalSize / PageSize.

[0082] C2: Divide the entire index into multiple entry data sets based on the hash value of the key in the index's key-value pair.

[0083] Traverse all Entry data, and assign each Entry data to a specific Page according to the Hash value of the Key value of each Entry data and the number of Pages. Thus, the entire index is divided into multiple Entry data sets, and each Entry data set corresponds to a Page.

[0084] C3: For each entry data set, create a corresponding page index file.

[0085] For each Page's Entry data set, a page index file provided in the embodiment of the present application is constructed.

[0086] The following further describes the algorithm for constructing the page index file. In a possible implementation, C1 may establish a corresponding page index file for each entry data set by:

[0087] D1: for a target entry data set, obtain a fifth quantity value of the entry data in the entry data set, and calculate a sixth quantity value of the bucket in the page index structure according to the fifth quantity value; the target entry data set is any entry data set.

[0088] For a certain Entry data set, the total number of Entry data included in it is recorded as the fifth data value. According to the load factor, the number of buckets in the Page can be obtained, which is recorded as the sixth quantity.

[0089] D2: construct a block bitmap according to the hash value of the entry data key in the target entry data set, the sixth quantity value and the first preset quantity.

[0090] The Entry data set is traversed, and the hash value of the Key value of each Entry data is calculated. According to the hash value of the Key value and the sixth quantity, it can be determined to which bucket each Entry data is assigned, so as to determine whether each bucket is empty. Each first preset number of buckets is divided into a Chunk, and a Chunk Bitmap can be constructed based on the status value of whether each bucket is an empty bucket and the number of Entry data included in each Chunk.

[0091] In a possible implementation, D2 may construct a block bitmap according to the hash value of the entry data key in the target entry data set, the sixth quantity value, and the first preset quantity, and specifically implement the following:

[0092] E1: Determine the entry data corresponding to each bucket according to the hash value of the entry data key in the target entry data set, and determine whether the bucket is an empty bucket state value.

[0093] Specifically, in the target entry data set, the hash value of the Key value of each Entry data is calculated. The bucket to which each Entry data should belong can be determined by each hash value and the sixth quantity value, and then whether each bucket is empty can be determined, and the status value of whether each bucket is empty can be obtained.

[0094] E2: All buckets are divided into multiple blocks based on the sixth quantity value and the first preset quantity to generate an initial block bitmap; the initial block bitmap includes a third field and a fourth field, the third field is used to use the first preset number of bits to save the status value of whether the bucket constituting the block is an empty bucket, and the fourth field is used to save the seventh quantity value of the entry data included in the current block.

[0095] According to the sixth quantity value (the total number of buckets) and the first preset quantity (the number of buckets in each block), all buckets in the target entry data set (Entry Pool) can be divided into multiple blocks, and the initial Chunk Bitmap is updated. Figure 5As shown, a schematic diagram of the process of generating the initial Chunk Bitmap in an embodiment of the present application is shown. Similar to the aforementioned description of the Chunk Bitmap, a first preset number N buckets constitute a chunk (Chunk), for example, the first preset number is 32, that is, every 32 buckets constitute a Chunk. The initial Chunk Bitmap includes a third field and a fourth field corresponding to each chunk, and each Chunk uses 2 times the first preset number of bits to describe the Chunk. The third field uses N bits to save the status value of whether the bucket constituting the block is an empty bucket, and the fourth field saves the seventh quantity value of the entry data included in the current block. The third field and the fourth field corresponding to each block are arranged consecutively, and each block is arranged in order according to the block identifier. For example, a 64-bits Word is used to describe the Chunk, and the first 32 bits of the Word are the third field, corresponding to the status bitmap of the 32 buckets in the Chunk. A bit value (status value) of 1 indicates that the bucket is not empty, and a bit value (status value) of 0 indicates that the bucket is empty. The last 32 bits of the word are the fourth field, which is the PopCount (the seventh quantity value) of the uint32 integer, indicating the total number of Entry data of the current Chunk. In addition, the first preset quantity is 8 for example, and the first preset quantity can be set according to actual conditions.

[0096] E3: Generate a first quantity value of entry data included in all blocks before the current block from the fourth field corresponding to each block in the initial block bitmap to construct a block bitmap.

[0097] See also Figure 6 As shown, a schematic diagram of the process of generating Chunk Bitmap in the embodiment of the present application is shown. The PopCount (the seventh quantity value) of each block in the initial fast bitmap is accumulated to generate the prefixPopCount of the next block, and then the third field in the initial Chunk Bitmap is used as the first field of the Chunk Bitmap to complete the construction of the Chunk Bitmap.

[0098] D3: Associate the entry data in the target entry data set with the corresponding blocks to generate an entry array and an entry array.

[0099] Finally, the Entry data in the target entry data set can be associated with the corresponding blocks and buckets to generate the Entry Array and Length Array.

[0100] In a possible implementation, D3 associates the item data in the target item data set with the corresponding block, generates an item array, and the specific implementation of the item array may include:

[0101] F1: Associate the entry data to the corresponding block according to the hash value of the entry data key in the target entry data set.

[0102] The hash value of the Key in the Entry data in the target entry data set, which can be used to associate the Entry data with the corresponding Chunk. Figure 7 As shown, a schematic diagram of the process of associating Entry data with Chunks in an embodiment of the present application is shown. In actual applications, the number of Entry data included in each Chunk can be obtained from the Chunk Bitmap, and an EntryIndex Array (entry data sequence number array) is established to save the subscripts of the Entry data in each Chunk. For example, the first Chunk includes 3 Entry data, and the first 3 positions in the Entry Index Array correspond to the first Chunk, the second Chunk includes 3 Entry data, and the third position in the Entry Index Array corresponds to the second Chunk, the third Chunk includes 5 Entry data, and the third position in the Entry Index Array corresponds to the third Chunk, and so on. Traverse the Entry data, and associate the Entry data with the corresponding Chunk according to the hash value of the Entry data Key, that is, write the subscript of the Entry data into the corresponding area of ​​the Entry Index Array of the corresponding Chunk. For example, the second Entry data belongs to the first Chunk, and its subscript is written into the first 3 positions corresponding to the first Chunk of the Entry Index Array. In this way, the Entry data included in each Chunk can be obtained, but the bucket corresponding to the Entry data in the Chunk has not yet been determined.

[0103] F2: Traverse each block and assign corresponding entry data to each bucket according to the hash value of the entry data key in the block.

[0104] For each Chunk, the Entry data in the block is allocated to the corresponding bucket according to the hash value of its Key.

[0105] F3: Write the entry data in each bucket into the entry array in sequence, and write the number of entry data included in the non-empty bucket into the length array in a unary encoding manner.

[0106] Then the number of Entry data included in the non-empty bucket can be written into the Length Array in a unary encoding manner, and the Entry data in the bucket can be written into the Entry Array in sequence, thus completing the construction of the Length Array and the Entry Array.

[0107] Based on the index structure of the embodiment of the present application, a corresponding index construction algorithm can be implemented.

[0108] Based on the database-based index query method provided in the above method embodiment, the embodiment of the present application also provides a database-based index query device, which will be described below in conjunction with the accompanying drawings.

[0109] See also Figure 8 As shown in FIG. 1 , this figure is a schematic diagram of the structure of a database-based index query device provided in an embodiment of the present application. Figure 8 As shown, the database-based index query device includes:

[0110] A first acquisition unit 801 is used to acquire a data query request for a database, wherein the data query request includes a target key;

[0111] A first determining unit 802 is used to determine a corresponding page identifier based on the hash value of the target key;

[0112] The second acquisition unit 803 is used to acquire the page index file corresponding to the page identifier, wherein the page index file has a page index structure; the page index structure includes an entry array, a block bitmap, and a length array, wherein the entry array includes entry data of the page index file, wherein the entry data is sequentially sorted according to the order of the buckets in which they are located and the order within the buckets, and each of the entry data includes a key-value pair of an index; the block bitmap includes a first field and a second field corresponding to each block, wherein the first field is used to save a status value of whether the bucket constituting the block is an empty bucket using a first preset number of bits, and the second field is used to save a first quantity value of the entry data included in all blocks before the current block; the length array includes a unary code of a non-empty bucket, and each of the unary codes corresponds to the quantity of entry data included in a non-empty bucket;

[0113] The query unit 804 is used to query whether there is a value corresponding to the target key in the page index file.

[0114] In a possible implementation, the query unit includes:

[0115] A first determination subunit, configured to determine a corresponding bucket identifier based on a hash value of the target key;

[0116] A second determining subunit, configured to determine a corresponding block identifier according to the bucket identifier;

[0117] A first query subunit, configured to query whether the bucket corresponding to the bucket identifier is empty in the first field corresponding to the block identifier in the block bitmap;

[0118] A return subunit, used for returning a value corresponding to the target key that does not exist if the bucket corresponding to the bucket identifier is empty;

[0119] A third determining subunit is configured to determine, if the bucket corresponding to the bucket identifier is not empty, a starting position and a number of entries of the bucket corresponding to the bucket identifier in the entry array according to the block bitmap and the length array;

[0120] The second query subunit is used to query whether the value corresponding to the target key exists in the number of entry data starting from the starting position in the entry array.

[0121] In a possible implementation manner, the third determining subunit is specifically configured to:

[0122] If the bucket corresponding to the bucket identifier is not empty, obtaining the number of non-empty buckets before the bucket corresponding to the bucket identifier in the first field corresponding to the block identifier in the block bitmap and recording it as a second quantity value, obtaining the first quantity value in the second field corresponding to the block identifier in the block bitmap and recording it as a third quantity value;

[0123] In the length array, starting from the bit where the third quantity value is plus one, read the second number of unary codes, and accumulate the number of entry data corresponding to the second number of unary codes to obtain a fourth quantity value;

[0124] The third data value and the fourth quantity value are added to obtain the starting position of the bucket corresponding to the bucket identifier in the entry array, and the quantity of entry data corresponding to the next unary code is read as the entry quantity.

[0125] In a possible implementation manner, the device further includes:

[0126] A second determining unit is used to determine the total data volume of the key-value pairs of all indexes, and generate the number of pages according to the total data volume and a page size parameter;

[0127] A partitioning unit, configured to partition all indexes into a plurality of entry data sets according to hash values ​​of keys in key-value pairs of the indexes;

[0128] The establishing unit is used to establish a corresponding page index file for each entry data set.

[0129] In a possible implementation manner, the establishing unit includes:

[0130] an acquisition subunit, configured to acquire, for a target entry data set, a fifth quantity value of entry data in the entry data set, and calculate a sixth quantity value of buckets in a page index structure according to the fifth quantity value; the target entry data set is any of the entry data sets;

[0131] A construction subunit, configured to construct a block bitmap according to a hash value of an entry data key in the target entry data set, the sixth quantity value, and a first preset quantity;

[0132] The generating subunit is used to associate the entry data in the target entry data set with the corresponding blocks, and generate an entry array and an entry array.

[0133] In a possible implementation, the construction subunit is specifically used for:

[0134] Determine the entry data corresponding to each bucket according to the hash value of the entry data key in the target entry data set, and determine whether the bucket is an empty bucket;

[0135] According to the sixth quantity value and the first preset quantity, all the buckets are divided into a plurality of blocks to generate an initial block bitmap; the initial block bitmap includes a third field and a fourth field corresponding to each block, the third field is used to save a status value of whether the bucket constituting the block is an empty bucket using the first preset number of bits, and the fourth field is used to save the seventh quantity value of the entry data included in the current block;

[0136] A first quantity value of entry data included in all blocks before the current block is generated from the fourth field corresponding to each block in the initial block bitmap to construct a block bitmap.

[0137] In a possible implementation, the generating subunit is specifically used for:

[0138] Associating the entry data to the corresponding block according to the hash value of the entry data key in the target entry data set;

[0139] Traverse each block and assign corresponding entry data to each bucket according to the hash value of the entry data key in the block;

[0140] The entry data in each bucket is written into the entry array in sequence, and the number of entry data included in the non-empty bucket is written into the length array in a unary encoding manner.

[0141] In addition, an embodiment of the present application further provides a computer program product, including computer program instructions. When the computer program instructions are executed on a computer, the computer executes any of the database-based index query methods described above.

[0142] In an embodiment of the present application, the entire index file can be divided into multiple page index files. After obtaining the target key in the data query request for the database, the hash value of the target key is first determined to determine the corresponding page identifier, thereby determining the page index file corresponding to the page identifier. The page index file has a page index structure, which is designed for read-only scenarios. In the page index structure, it includes an entry array, a block bitmap, and a length array. The entry array continuously stores the entry data of the page index file, and the entry data is sorted in the order of the buckets and the order in the bucket. N buckets form a block, and the block bitmap includes a first field and a second field corresponding to each block. The first field includes N bits, and the N bits correspond to the status value of whether the N buckets constituting the block are empty buckets. The second field corresponds to the first number value of the entry data included in all blocks before the current block. The length array continuously stores the unary codes of non-empty buckets, and each unary code corresponds to the number of entry data included in a non-empty bucket. Through the page index file, it can be queried whether there is a value corresponding to the target key. Since the page index structure is designed for read-only scenarios, the storage is continuous, which ensures a high space utilization rate. And it can be directly loaded from other storage media into the memory for use, without the need for additional data parsing, thus achieving the optimization of indexes for read-only scenarios.

[0143] Based on the database-based index query method provided in the above method embodiment, the present application also provides an electronic device, including: one or more processors; a storage device, on which one or more programs are stored, when the one or more programs are executed by the one or more processors, the one or more processors implement the database-based index query method described in any of the above embodiments.

[0144] Reference below Fig. 9 , which shows a schematic diagram of the structure of an electronic device 1300 suitable for implementing the embodiment of the present application. The terminal device in the embodiment of the present application may include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (portable android devices), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs (televisions), desktop computers, etc. Fig. 9 The electronic device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0145] like Fig. 9As shown, the electronic device 1300 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 1301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1302 or a program loaded from a storage device 1306 into a random access memory (RAM) 1303. In the RAM 1303, various programs and data required for the operation of the electronic device 1300 are also stored. The processing device 1301, the ROM 1302, and the RAM 1303 are connected to each other via a bus 1304. An input / output (I / O) interface 1305 is also connected to the bus 1304.

[0146] Typically, the following devices may be connected to the I / O interface 1305: an input device 1306 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1306 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1309. The communication device 1309 may allow the electronic device 1300 to communicate with other devices wirelessly or by wire to exchange data. Although Fig. 9 The electronic device 1300 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.

[0147] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 1309, or installed from the storage device 1306, or installed from the ROM 1302. When the computer program is executed by the processing device 1301, the above-mentioned functions defined in the method of the embodiment of the present application are executed.

[0148] The electronic device provided in the embodiment of the present application and the database-based index query method provided in the above embodiment belong to the same inventive concept. The technical details not fully described in this embodiment can be referred to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.

[0149] Based on the database-based index query method provided in the above method embodiment, an embodiment of the present application provides a computer-readable medium on which a computer program is stored, wherein when the program is executed by a processor, the database-based index query method as described in any of the above embodiments is implemented.

[0150] It should be noted that the computer-readable medium in the embodiment of the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In an embodiment of the present application, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device. In an embodiment of the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0151] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.

[0152] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0153] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device executes the database-based index query method.

[0154] The computer program code for performing the operation of the embodiment of the present application can be written in one or more programming languages ​​or a combination thereof, and the above-mentioned programming languages ​​include but are not limited to object-oriented programming languages-such as Java, Smalltalk, C++, and also include conventional procedural programming languages-such as "C" language or similar programming languages. The program code can be executed completely on the user's computer, partially on the user's computer, as an independent software package, partially on the user's computer and partially on the remote computer, or completely on the remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network-including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, using an Internet service provider to connect through the Internet).

[0155] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0156] The units involved in the embodiments described in this application may be implemented by software or hardware, wherein the name of the unit / module does not, in some cases, constitute a limitation on the unit itself.

[0157] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

[0158] In the context of the present application embodiment, machine-readable medium can be a tangible medium that can contain or store a program for use by an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable medium can include but is not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the above. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above.

[0159] It should be noted that the various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same or similar parts between the various embodiments can be referred to each other. For the system or device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description.

[0160] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0161] It should also be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0162] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0163] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A database-based index query method, characterized in that: The method comprises: Obtaining a data query request for a database, wherein the data query request includes a target key; Determine a corresponding page identifier based on the hash value of the target key; Obtain a page index file corresponding to the page identifier, wherein the page index file has a page index structure; the page index structure includes an entry array, a block bitmap, and a length array, wherein the entry array includes entry data of the page index file, wherein the entry data is sequentially sorted according to the order of the buckets in which they are located and the order within the buckets, and each of the entry data includes a key-value pair of an index; the block bitmap includes a first field and a second field corresponding to each block, wherein the first field is used to save a status value of whether a bucket constituting the block is an empty bucket using a first preset number of bits, and the second field is used to save a first quantity value of the entry data included in all blocks before the current block; the length array includes unary codes of non-empty buckets, and each of the unary codes corresponds to the quantity of entry data included in a non-empty bucket; In the page index file, it is queried whether there is a value corresponding to the target key.

2. The method according to claim 1, characterized in that The querying whether there is a value corresponding to the target key in the page index file includes: Determine a corresponding bucket identifier based on the hash value of the target key; Determine a corresponding block identifier according to the bucket identifier; querying in the first field corresponding to the block identifier in the block bitmap whether the bucket corresponding to the bucket identifier is empty; If the bucket corresponding to the bucket identifier is empty, return the value corresponding to the target key that does not exist; If the bucket corresponding to the bucket identifier is not empty, determining the starting position and the number of entries of the bucket corresponding to the bucket identifier in the entry array according to the block bitmap and the length array; Starting from the starting position in the entry array, query whether there is a value corresponding to the target key in the entry number of entry data.

3. The method according to claim 2, characterized in that If the bucket corresponding to the bucket identifier is not empty, determining the starting position and the number of entries of the bucket corresponding to the bucket identifier in the entry array according to the block bitmap and the length array, comprises: If the bucket corresponding to the bucket identifier is not empty, obtaining the number of non-empty buckets before the bucket corresponding to the bucket identifier in the first field corresponding to the block identifier in the block bitmap and recording it as a second quantity value, obtaining the first quantity value in the second field corresponding to the block identifier in the block bitmap and recording it as a third quantity value; In the length array, starting from the bit of the third quantity value plus one, read the second quantity value number of unary codes, and accumulate the number of entry data corresponding to the second quantity value number of unary codes to obtain a fourth quantity value; The third quantity value is added to the fourth quantity value to obtain the starting position of the bucket corresponding to the bucket identifier in the entry array, and the quantity of entry data corresponding to the next unary code is read as the entry quantity.

4. The method according to claim 1, characterized in that: The method further comprises: Determine the total data volume of the key-value pairs of all indexes, and generate the number of pages based on the total data volume and a page size parameter; Dividing the entire index into a plurality of entry data sets according to the hash value of the key in the key-value pair of the index; For each entry data set, a corresponding page index file is created.

5. The method according to claim 4, characterized in that The step of establishing a corresponding page index file for each entry data set includes: For a target entry data set, obtaining a fifth quantity value of entry data in the entry data set, and calculating a sixth quantity value of buckets in a page index structure according to the fifth quantity value; the target entry data set is any of the entry data sets; constructing a block bitmap according to the hash value of the entry data key in the target entry data set, the sixth quantity value and the first preset quantity; The entry data in the target entry data set are associated with corresponding blocks to generate an entry array and an entry array.

6. The method according to claim 5, characterized in that The step of constructing a block bitmap according to the hash value of the entry data key in the target entry data set, the sixth quantity value, and the first preset quantity includes: Determine the entry data corresponding to each bucket according to the hash value of the entry data key in the target entry data set, and determine whether the bucket is an empty bucket; According to the sixth quantity value and the first preset quantity, all the buckets are divided into a plurality of blocks to generate an initial block bitmap; the initial block bitmap includes a third field and a fourth field corresponding to each block, the third field is used to save a status value of whether the bucket constituting the block is an empty bucket using the first preset number of bits, and the fourth field is used to save the seventh quantity value of the entry data included in the current block; A first quantity value of entry data included in all blocks before the current block is generated from the fourth field corresponding to each block in the initial block bitmap to construct a block bitmap.

7. The method according to claim 5, characterized in that The step of associating the entry data in the target entry data set with the corresponding blocks to generate an entry array and an entry array includes: Associating the entry data to the corresponding block according to the hash value of the entry data key in the target entry data set; Traverse each block and assign corresponding entry data to each bucket according to the hash value of the entry data key in the block; The entry data in each bucket is written into the entry array in sequence, and the number of entry data included in the non-empty bucket is written into the length array in a unary encoding manner.

8. A database-based index query device, characterized in that: The device comprises: A first acquisition unit, configured to acquire a data query request for a database, wherein the data query request includes a target key; A first determining unit, configured to determine a corresponding page identifier based on a hash value of the target key; a second acquisition unit, configured to acquire a page index file corresponding to the page identifier, wherein the page index file has a page index structure; the page index structure includes an entry array, a block bitmap, and a length array, the entry array includes entry data of the page index file, the entry data are sequentially sorted according to the order of the buckets in which they are located and the order within the buckets, and each of the entry data includes a key-value pair of an index; the block bitmap includes a first field and a second field corresponding to each block, the first field is used to save a status value of whether the bucket constituting the block is an empty bucket using a first preset number of bits, and the second field is used to save a first quantity value of the entry data included in all blocks before the current block; the length array includes a unary code of a non-empty bucket, and each of the unary codes corresponds to the quantity of entry data included in a non-empty bucket; The query unit is used to query whether there is a value corresponding to the target key in the page index file.

9. A database-based index query device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the database-based index query method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, and when the instructions are executed on a terminal device, the terminal device executes the database-based index query method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Multi-table association query method and device based on Bitmap

    CN117762920A

  • Data file and data retrieving method

    JP2001043237A