Data storage method and data storage device for multiple types of indexes
By constructing a multi-type indexing method of multi-tree and Huffman coding, the storage mismatch problem of SSTable file structure under different resources and KV data distribution is solved, and flexible and efficient data storage is achieved.
Patent Information
- Application Number
- CN202410398364.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-03
- Publication Date
- 2025-10-14
AI Technical Summary
The existing SSTable file structure cannot effectively match different resources and different KV data distributions, resulting in inflexible and inefficient storage.
A data storage method with multiple types of indexes is adopted, including header identification segment, key data segment, value data segment and value index segment. By constructing a multi-tree and Huffman coding, flexible compression and matching storage of keys and values is achieved.
It achieves efficient storage under different resource and KV distribution conditions, is suitable for key common prefix scenarios and large data volume scenarios, and improves storage flexibility and efficiency.
Smart Images

Figure CN120780668A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data storage technology, and in particular provides a data storage method and a data storage device for multiple types of indexes. Background Art
[0002] In a classic LSM tree, an SSTable (static sorted table) is a file structure that stores ordered KV data in blocks. An SSTable stores sorted data (actually, key-value pairs sorted by key). Each SSTable can contain multiple data-storing files, called segments, and the data within each segment is sorted by key.
[0003] However, the problem is that the existing SSTable file structure cannot provide better dynamic matching for KV data, different key distributions, different value distributions, and running applicable resources.
[0004] Therefore, how to use a more suitable storage solution for different resources and different KV data distributions is the technical problem to be solved by the patent of this invention. Summary of the Invention
[0005] In order to solve the above technical problems, the present invention proposes a data storage method and data storage device for multiple types of indexes, specifically, adopting the following technical solutions:
[0006] In a first aspect, the present invention provides a data storage method for multiple types of indexes, including a first data storage structure, wherein the first data storage structure includes:
[0007] The header identification section is used to write the type data representing the file encoding method;
[0008] The key data segment bitrie is used to write the key of the KV key-value pair in sequence, build the key into a multi-branch tree, number each node in sequence, and record the current key number on the corresponding leaf node when adding a key. Then, the multi-branch tree is traversed to obtain the complete bitmap, labels, and leaf-value. Among them, the bitmap records the parent-child node relationship of the entire multi-branch tree, the labels are used to map the nodes of the entire multi-branch tree to specific characters for matching, and the leaf-value is used to calculate the complete match from the root node to the leaf node and then query the value;
[0009] The value data segment value-block stores the corresponding values in the order of the keys;
[0010] The value index segment value-index stores the offset of each value in the file in sequence.
[0011] As an optional embodiment of the present invention, in a data storage method for multiple types of indexes of the present invention, the process of writing data into the key data segment bitrie includes:
[0012] The keys of multiple KV key-value pairs written in sequence are decomposed into single characters, and all the single characters of each key are added to the multi-tree in sequence. At the same time, when a single character is added to the multi-tree, the sequence number of the current key is recorded in the leaf node. The root node of the multi-tree does not store specific characters and is used to index each key;
[0013] Encode each node on the multi-branch tree. If there are N leaf nodes, the encoding is N 1s and 1 0. If there is no leaf node, the encoding is 0.
[0014] Perform hierarchical traversal on the encoded multi-branch tree, write the encoding of each node into the bitmap according to the level, and write each leaf node into the labels;
[0015] According to the bitmap and labels, the corresponding relationship between key-serial number, key and leaf-serial number is obtained, sorted according to leaf-serial number, and the leaf-serial number and key-serial number data pairs are stored in the leaf-value structure.
[0016] As an optional embodiment of the present invention, in a data storage method of a multi-category index of the present invention, the formula for calculating the parent and child nodes of the current node based on the bitmap is:
[0017] FirstChild(x)=rank1(select0(x)+1)
[0018] Parent(x)=rank0(select1(x))
[0019] select0(x) indicates the position of the xth 0 in the bitmap;
[0020] select1(x) indicates the position of the xth 1 in the bitmap;
[0021] Rank0(x) indicates how many zeros appear from bitmap 0 to bitmap x;
[0022] Rank1(x) indicates how many 1s appear from bitmap 0 to bitmap x.
[0023] As an optional embodiment of the present invention, a data storage method of a multi-category index of the present invention includes a data search process based on a first data storage structure:
[0024] The first child node of the root node of the multitree in the bitrie of the key data segment is searched for the query key and compared. If the comparison is successful, the child nodes of the first child node are further searched until the query key is found in the leaf node of the first child node or the query key is not found, and the process is exited. If the comparison fails, the sibling nodes of the first child node are searched until the query key is found in the leaf node of the sibling node or the query key is not found, and the process is exited.
[0025] Get the leaf-number based on the key found in the bitrie query of the key data segment, query the leaf-value structure based on the leaf-number, and locate the value-number through binary search;
[0026] Search in the value index section array-index through value-serial number. Use value-serial number to directly locate the offset of value in the file, and get the value from the value data section value-block.
[0027] As an optional embodiment of the present invention, in a data storage method for multiple indexes of the present invention, comparing the first child node of the multitree root node in the bitrie of the query key data segment according to the query key includes:
[0028] First, the number of the corresponding node is queried based on the bitmap, and then the position of the number is directly located based on the information saved in the labels. The value is compared with the first character of the query key to obtain the comparison result.
[0029] As an optional embodiment of the present invention, a data storage method for multiple indexes of the present invention includes replacing a second data storage structure formed by a value data Huffman segment value-block of a first data storage structure with a value data Huffman segment value-block (Huffman), wherein the value data Huffman segment value-block (Huffman) data storage structure includes:
[0030] The value storage area uses the Huffman algorithm for value compression storage;
[0031] The decoding index area decode-index is used to store the index relationship between the compressed storage char file of value and the codec table.
[0032] As an optional embodiment of the present invention, in a data storage method for multiple types of indexes of the present invention, the value storage area uses the Huffman algorithm to perform value compression storage, which includes:
[0033] Construct Huffman tree and Huffman encoding and decoding table for the values written in sequence;
[0034] The values written sequentially are stored in the value storage area according to the Huffman encoding and decoding table. The encoding structure of the value storage area includes: a value encoding storage area encode-value for storing the Huffman encoding of the value, an alignment area align for aligning the Huffman encoding of the value stored in the value encoding storage area encode-value to a fixed length, and an alignment bit counting area align-num for writing the number of alignment bits used when aligning.
[0035] As an optional embodiment of the present invention, in a data storage method for multiple types of indexes of the present invention, the encoding structure of the decoding index area decode-index is:
[0036] char-block1 char-block2 char-block3 char-block4
[0037] Among them, a block encoding structure is:
[0038] decode-value char1 encode-len
[0039] For each character, according to the frequency statistics table of the Huffman algorithm, the decoding value of the decoding table is written into decode-value, the corresponding character is written into char, and then the encoding bit length of the character is written into encode-len, and the encoding and decoding information of each character is written in turn.
[0040] As an optional embodiment of the present invention, a data storage method of a multi-category index of the present invention includes a data search process based on a first data storage structure:
[0041] When querying, the data information stored in the bitrie of the key data segment can be used to determine whether the query key exists. If it exists, its value-sequence number can be obtained. The value-sequence number is searched in the value index segment value-index to obtain the offset of its value in the file. The value size is obtained based on the offset difference between the previous and next values.
[0042] Read the encoded data of value in the file;
[0043] According to the align-num, the valid bit information is known, and the decoding ends when the valid bit is decoded, and the decoded value is returned. In a second aspect, the present invention provides a data storage device for multiple types of indexes, including a first data storage module, the first data storage module including:
[0044] The header identification section is used to write the type data representing the file encoding method;
[0045] The key data segment bitrie is used to write the key of the KV key-value pair in sequence, build the key into a multi-branch tree, number each node in sequence, and record the current key number on the corresponding leaf node when adding a key. Then, the multi-branch tree is traversed to obtain the complete bitmap, labels, and leaf-value. Among them, the bitmap records the parent-child node relationship of the entire multi-branch tree, the labels are used to map the nodes of the entire multi-branch tree to specific characters for matching, and the leaf-value is used to calculate the complete match from the root node to the leaf node and then query the value;
[0046] The value data segment value-block stores the corresponding values in the order of the keys;
[0047] The value index segment array-index stores the offset of each value in the file in sequence.
[0048] Compared with the prior art, the present invention has the following beneficial effects:
[0049] The first data storage structure of the present invention can perform extreme compression on keys when there are a large number of common prefixes for the keys. It is suitable for scenarios with common prefixes for the keys and is more memory / disk friendly.
[0050] The second data storage structure of the present invention can use different compression algorithms to compress keys and values, is suitable for scenarios with large data volumes and the need for extreme compression, and is disk-friendly.
[0051] This invention uses a more appropriate storage solution for different resources and KV distributions: if the CPU is idle and the overall file size is relatively large, the second data storage structure is used for extreme compression; if the key is larger than the value in the file and the CPU is relatively idle, the first data storage structure can be used. These solutions are combined, differentiated by version number, and can be parsed according to different parsing formats, improving scalability. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 An overall coding structure diagram of a first data storage structure according to a third embodiment of the present invention;
[0053] Figure 2 The coding structure diagram of the key index segment hash-index of the third embodiment of the present invention;
[0054] Figure 3 Overall coding structure diagram of the fourth data storage structure of the fourth embodiment of the present invention;
[0055] Figure 4 An overall coding structure diagram of a first data storage structure according to a first embodiment of the present invention;
[0056] Figure 5 An example of constructing a multi-branch tree with a key in the first embodiment of the present invention;
[0057] Figure 6 In the first embodiment of the present invention Figure 5 Bitmap and labels corresponding to the multi-tree example;
[0058] Figure 7 An overall coding structure diagram of a second data storage structure according to a second embodiment of the present invention;
[0059] Figure 8 An example of a Huffman tree constructed in the second embodiment of the present invention;
[0060] Figure 9 An example of a decoding table constructed in the second embodiment of the present invention;
[0061] Figure 10 The encoding structure diagram of the value in the value-block in the second embodiment of the present invention;
[0062] Figure 11 An example of a codec table for the decode-index area in the second embodiment of the present invention;
[0063] Figure 12 An example of the frequency statistics of the characters obtained by sequentially arranging v1-v5 using the Huffman algorithm in Example 6 of the present invention;
[0064] Figure 13 Example 1 of the Huffman tree constructed in Example 6 of the present invention;
[0065] Figure 14 Example 2 of the Huffman tree constructed in Example 6 of the present invention;
[0066] Figure 15 A schematic structural diagram of an electronic device according to a seventh embodiment of the present invention;
[0067] Figure 16A schematic diagram of the computer readable recording medium of the seventh embodiment of the present application. DETAILED DESCRIPTION
[0068] In order to make the objects, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments of the present application.
[0069] Therefore, the following detailed description of the embodiments of the present application is not intended to limit the scope of the claimed application, but merely represents some embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.
[0070] It should be noted that the embodiments in the present application and the features and technical solutions in the embodiments can be combined with each other without conflict.
[0071] It should be noted that: similar reference numbers and letters represent similar items in the following drawings, therefore, once an item is defined in one drawing, it does not need to be further defined and explained in the subsequent drawings.
[0072] In the description of the present application, it should be noted that the terms "upper", "lower", and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the product of the present application is usually placed, or the orientation or positional relationship commonly understood by those skilled in the art, such terms are only for the convenience of describing the present application and simplifying the description, and are not intended to indicate or imply that the indicated device or element must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as limiting the present application. In addition, the terms "first", "second", and the like are only used to distinguish the description, and cannot be understood as indicating or implying relative importance.
[0073] Example 1
[0074] The data storage method of the plurality of categories index of the embodiment comprises a first data storage structure, the first data storage structure comprises:
[0075] A header identification section header for writing type data representing a file encoding method;
[0076] The key data segment bitrie is used to write the key of the KV key-value pair in sequence, build the key into a multi-branch tree, number each node in sequence, and record the current key number on the corresponding leaf node when adding a key. Then, the multi-branch tree is traversed to obtain the complete bitmap, labels, and leaf-value. Among them, the bitmap records the parent-child node relationship of the entire multi-branch tree, the labels are used to map the nodes of the entire multi-branch tree to specific characters for matching, and the leaf-value is used to calculate the complete match from the root node to the leaf node and then query the value;
[0077] The value data segment value-block stores the corresponding values in the order of the keys;
[0078] The value index segment value-index stores the offset of each value in the file in sequence.
[0079] The first data storage structure of this embodiment can perform extreme compression on the key when there are a large number of common prefixes for the key. It is suitable for scenarios with common prefixes for the key and is more memory / disk friendly. The matching process will traverse multiple layers and consume some CPU resources.
[0080] The first data storage structure of this embodiment is encoded as follows: Figure 4 As shown, the header identification segment header is a uint8 enumeration value, which represents the file encoding method. It can support up to 256 encoding methods and can be customized according to the file encoding method.
[0081] The encoding structure of the key data segment bitrie in this embodiment is as follows:
[0082] bitmap labels leaf-value
[0083] The data encoding process of the bitrie of the key data segment in this embodiment includes:
[0084] The keys of multiple KV key-value pairs written in sequence are decomposed into single characters, and all the single characters of each key are added to the multi-tree in sequence. At the same time, when a single character is added to the multi-tree, the sequence number of the current key is recorded in the leaf node. The root node of the multi-tree does not store specific characters and is used to index each key;
[0085] Encode each node on the multi-branch tree. If there are N leaf nodes, the encoding is N 1s and 1 0. If there is no leaf node, the encoding is 0.
[0086] Perform hierarchical traversal on the encoded multi-branch tree, write the encoding of each node into the bitmap according to the level, and write each leaf node into the labels;
[0087] According to the bitmap and labels, the corresponding relationship between key-serial number, key and leaf-serial number is obtained, sorted according to leaf-serial number, and the leaf-serial number and key-serial number data pairs are stored in the leaf-value structure.
[0088] like Figure 5 As shown in the figure, an example of building a multi-branch tree with 8 keys and 8 values (not shown in the figure).
[0089] These keys are broken down into characters and added to the multitree in sequence (the root node does not store specific characters but is used to index each key). At the same time, the sequence number of the current key is recorded in the leaf node when adding.
[0090] Leaf nodes Corresponding key Current key number D AD 0 E AE 1 F BF 2 K BGK 3 L BGL 4 H BH 5 I CI 6 J CJ 7
[0091] Each node is numbered. The number of leaf nodes is encoded as N 1s and 1 0. If there is no leaf node, it is numbered 0.
[0092] For example, node A has two leaf nodes, numbered 110.
[0093] The bitmap stores the result of the hierarchical traversal. The encoding of each node is written into the bitmap according to the hierarchy, and each node is also written into the labels.
[0094] The result of the above-mentioned hierarchical traversal is stored in a bitmap, which can be compressed and stored; the characters obtained by the hierarchical traversal are stored in this field in order for character matching. Figure 6 for Figure 5 The bitmap and labels corresponding to the example.
[0095] In a data storage method for multiple indexes of this embodiment, the formula for calculating the parent and child nodes of the current node based on the bitmap is:
[0096] FirstChild(x)=rank1(select0(x)+1)
[0097] Parent(x)=rank0(select1(x))
[0098] select0(x) indicates the position of the xth 0 in the bitmap;
[0099] select1(x) represents the position of the xth 1 in the bitmap;
[0100] rank0(x) represents how many 0s appear from the 0th bit to the xth bit of the bitmap;
[0101] rank1(x) represents how many 1s appear from the 0th bit to the xth bit of the bitmap.
[0102] For example, in the example of Figure 5 and Figure 6 , the data of the leaf-value corresponding to the example is:
[0103] FirstChild(B) = rank1(select0(2) + 1) = rank1(6 + 1) = rank1(7) = 6 = F
[0104] Parent(B) = rank0(select1(2)) = rank0(1) = 0 = root.
[0105] and Figure 5 , Figure 6 The data of the leaf-value corresponding to the example is:
[0106] key serial number key leaf-number 0 AD 4 1 AE 5 2 BF 6 3 BGK 11 4 BGL 12 5 BH 8 6 CI 9 7 CJ 10
[0107] According to the sorting of the leaf-serial numbers, the following corresponding relationship is obtained
[0108] leaf number 4 5 6 8 9 10 11 12 key serial number 0 1 2 5 6 7 3 4
[0109] The data is stored in the leaf-value structure, and the stored values are (4, 0), (5, 1), (6, 2),..., (12, 4).
[0110] The encoding structure of the value index section value-index of the embodiment is as follows:
[0111] v1.offset v2.offset v3.offset ...
[0112] The offset of each value in the file is stored in turn, and it is fixed-length.
[0113] The data storage method of the multi-class index of the embodiment includes a data searching process based on the first data storage structure:
[0114] According to the query key, the first child node of the root node of the multi-way tree in the key data section bitrie is queried for comparison, if the comparison is successful, the child node of the first child node is further queried until the query key is found in the leaf node of the first child node or is not found and exits, if the comparison fails, the sibling node of the first child node is queried until the query key is found in the leaf node of the sibling node or is not found and exits;
[0115] According to the key obtained by querying the key data section bitrie, the leaf-value structure is queried according to the leaf-serial number, and the value-serial number is located through binary search, such as querying Figure 5 And Figure 6 In the example, the AE obtains the leaf serial number 5, and then queries the leaf-value structure, each leaf serial number and key serial number are fixed length, and the leaf-serial number is ordered, and the value can be located through binary search, (5, 1);
[0116] The value-serial number is used to search in the value index section array-index, the value-serial number can directly locate the offset of the value in the file, and the value is obtained from the value data section value-block and returned.
[0117] Further, the embodiment described according to the query key, the first child node of the root node of the multi-way tree in the key data section bitrie is queried for comparison, including:
[0118] First, the number corresponding to the node is queried according to the bitmap, and then the position of the number is directly located according to the information saved by the labels, and the value is compared with the first character of the query key, so that the comparison result is obtained.
[0119] The embodiment also provides a data storage device of multiple category indexes, including a first data storage module, the first data storage module includes:
[0120] The header identification section header is used for writing type data representing the encoding mode of the file;
[0121] key data section bitrie, for sequentially writing key of KV key-value pair, building key into multi-way tree, sequentially numbering each node, while adding a key, recording serial number of current key on corresponding leaf node, then performing hierarchical traversal of multi-way tree to obtain complete bitmap, labels and leaf-value, wherein bitmap records parent-child node relationship of entire multi-way tree, labels are used to correspond nodes of entire multi-way tree to specific characters for matching, and leaf-value is used to calculate value after complete matching from root node to leaf node;
[0122] value data section value-block, sequentially storing value corresponding to key;
[0123] value index section array-index, sequentially storing offset of each value in file.
[0124] Example 2
[0125] The data storage method of the multi-type index of the embodiment includes replacing the value data section value-block of the first data storage structure with a second data storage structure composed of a value data Huffman section value-block (Huffman), and the data storage structure of the value data Huffman section value-block (Huffman) includes:
[0126] value storage area, value compression storage is performed by using a Huffman algorithm;
[0127] decode index section decode-index, for storing index relationship between compressed storage char file of value and coding table.
[0128] Generally, key and value have different distributions, the second data storage structure can use different compression algorithms to compress key and value, is suitable for a scenario in which data volume is large and extreme compression is required, is friendly to disk, and multi-level key lookup and value decoding consume more CPU, and if decoded cache is added, more memory is consumed.
[0129] The overall encoding structure of the second data storage structure of the embodiment is as shown in Figure 7 The header section header is the same as that of the first embodiment, and an enumerated value represents that the file encoding is of the second data storage structure type.
[0130] The bitrie is consistent with the key data segment bitrie of the first embodiment, and is used to store the index relationship between the key and the value sequence number.
[0131] The value-index is consistent with the value index segment value-index in the first embodiment.
[0132] The encoding structure of the Huffman segment value-block (Huffman) of the value data in the second data storage structure of this embodiment is as follows:
[0133] value1 value2 value3 ... decode-index
[0134] In a data storage method for multiple indexes of this embodiment, the value storage area uses the Huffman algorithm to perform value compression storage, including:
[0135] Construct Huffman tree and Huffman encoding and decoding table for the values written in sequence;
[0136] The values written sequentially are stored in the value storage area according to the Huffman encoding and decoding table. The encoding structure of the value storage area includes: a value encoding storage area encode-value for storing the Huffman encoding of the value, an alignment area align for aligning the Huffman encoding of the value stored in the value encoding storage area encode-value to a fixed length, and an alignment bit counting area align-num for writing the number of alignment bits used when aligning.
[0137] Construct Huffman and encoding and decoding tables, value example:
[0138] value serial number value 0 acd 1 aaaabc 2 aaaab 3 abc 4 cc
[0139] Count by the number of times a character appears:
[0140] count(a)=10, count(b)=3, count(c)=5, count(d)=1.
[0141] Constructing a Huffman tree Figure 8 As shown, construct the decoding table as Figure 9 shown.
[0142] A coding table is constructed based on the Hafuman tree. Each character corresponds to a number of bits. The decoding table is padded with 0 according to the bit position. The prefix is the coding bit, and the back is padded with 0 to make it 16 bits. Then it is converted to a uint16 integer based on the 16 bits.
[0143] The encoding structure of the value in the value-block is as follows Figure 10 As shown, it is divided into three parts: the value encoding storage area encode-value, the alignment area align (8-byte alignment), and the end alignment bit count area align-num, which is the number of alignment bits used for 8-byte alignment.
[0144] The encoding structure of the decoding index area decode-index in this embodiment is:
[0145] char-block1 char-block2 char-block3 char-block4
[0146] Among them, a block encoding structure is:
[0147] decode-value char1 encode-len
[0148] For each character, according to the frequency statistics table of the Huffman algorithm, the decoding value of the decoding table is written into decode-value, the corresponding character is written into char, and then the encoding bit length of the character is written into encode-len, and the encoding and decoding information of each character is written in turn.
[0149] The codec table of the decode-index area of the example is as follows Figure 11 shown.
[0150] A data storage method for multiple types of indexes in this embodiment includes a data search process based on a second data storage structure:
[0151] When querying, the data information stored in the bitrie of the key data segment can be used to determine whether the query key exists. If it exists, its value-sequence number can be obtained. The value-sequence number is searched in the value index segment value-index to obtain the offset of its value in the file. The value size is obtained based on the offset difference between the previous and next values.
[0152] Read the encoded data of value in the file;
[0153] The valid bit information is known according to align-num. Decoding is completed when the valid bit information is reached, and the decoded value is returned.
[0154] Convert the current 2 bytes into integers, compare them with the decoding table, and perform a binary search to find the range where the character falls. This represents the character in the corresponding range.
[0155] For example, the encoding result for aaaabc is Figure 9The corresponding bit is: 00001101_00000000, the first 2B is obtained, which is converted into int as 3328, and the corresponding interval is [0, 32768), indicating that the first character after decoding is a, the encoding length of a is 1 (known from encode-len), and the bit is left shifted by 1 to obtain a new bit, which is converted into an integer to continue decoding. Through the above process, the final decoded value is aaaabc.
[0156] The embodiment also provides a data storage device for multiple types of indexes, including a second data storage module, wherein the second data storage module includes:
[0157] A header identification section header, used for writing type data representing a file encoding mode;
[0158] A key data section bitrie, used for sequentially writing a key of a KV key-value pair, constructing the key into a multi-way tree, sequentially numbering each node, recording a serial number of the current key on a corresponding leaf node when a certain key is added, and then performing hierarchical traversal of the multi-way tree to obtain complete bitmap, labels and leaf-value, wherein the bitmap records parent-child node relationships of the whole multi-way tree, the labels are used to correspond nodes of the whole multi-way tree to specific characters for matching, and the leaf-value is used to query a value after complete matching from a root node to a leaf node;
[0159] A value data Huffman section value-block (Huffman), and a data storage structure of the value data Huffman section value-block (Huffman) includes: a value storage area, which is stored by using a Huffman algorithm for value compression; and a decode index area decode-index, which is used for storing an index relationship between a compressed storage char file of the value and a coding and decoding table. Generally, keys and values have different distributions, and the fourth data storage structure can compress the keys and the values by using different compression algorithms, is suitable for a scenario in which data volume is large and extreme compression is required, is friendly to a disk, and consumes more CPUs when multi-level key searching and value decoding are performed, and more memory is consumed when a decoded cache is added;
[0160] A value index section array-index, used for sequentially storing an offset offset of each value in a file.
[0161] Example 3
[0162] The embodiment also provides a data storage method for multiple types of indexes, including a third data storage structure, wherein the third data storage structure includes:
[0163] header, used for writing type data representing the encoding mode of the file;
[0164] KVS, used for sequentially writing prefix-compressed KV key-value pair data;
[0165] array-index, recording the starting position offset of each KV key-value pair in the file, and the total length of a certain KV key-value pair being determined by the difference between the starting position offset of the KV key-value pair and the starting position offset of the next KV key-value pair;
[0166] hash-index, used for recording the corresponding index relationship between the key and the sequence number of the KV key-value pair in the overall file.
[0167] The third data storage structure of the embodiment can be used as a default writing format, and can be competent for most scenarios, especially for scenarios in which the key-value data is discretely distributed and the total number of files is within 100G, and the cpu / memory / disk consumption is relatively low and more balanced.
[0168] The overall encoding of the third data storage structure of the embodiment is as shown in Figure 1 The header identification section header is a uint8 enumeration value, representing the encoding mode of the file, and can support up to 256 encoding modes, which can be customized according to the encoding mode of the file.
[0169] The encoding structure of the KV data section KVS in the third data storage structure of the embodiment is as follows:
[0170] kv1 kv2 kv3 ...
[0171] The KV data section KVS of the embodiment sequentially stores the prefix-compressed KV key-value pair. Specifically, if it is checked that the two adjacent keys have a common prefix, then the next key only records the position and suffix information of the previous key, and the prefix information is obtained from the previous one or more adjacent keys.
[0172] The encoding structure of a single KV in the KV data section KVS of the embodiment for the KV key-value pair with common prefix data is as follows:
[0173] keySize(2B) sharedHeader(1B) sharedLen(2B) suffix value
[0174] The keySize section is used for storing the length of the suffix data suffix in the key;
[0175] The sharedHeader section is used for setting the condition that the current KV key-value pair does not contain common prefix data or contains common prefix data and is the first one.
[0176] The sharedLen area is used to store the length of the common prefix data;
[0177] The suffix area is used to store suffix data, and the suffix area stores the complete key;
[0178] The value area is used to store the value of the KV key-value pair.
[0179] In this embodiment, the encoding structure of a single KV in the KV data segment KVS for a KV key-value pair that only stores suffix data is as follows:
[0180] keySize(2B) sharedHeader(1B) sharedPos(1B) suffix value
[0181] The keySize area is used to store the length of the suffix data in the key;
[0182] The sharedHeader area is used to set the current KV key-value pair to contain a common prefix data and is not the first one;
[0183] The sharedPost area is used to store the distance between keys with common prefix data;
[0184] The suffix area is used to store the suffix data of the key excluding the common prefix data;
[0185] The value area is used to store the value of the KV key-value pair.
[0186] Specifically, the sharedHeader area of this embodiment is used to set whether the current kv contains a common prefix, 0 means no common prefix (no common prefix with the previous and next keys), 1 means there is a common prefix and it is the first one, and 2 means there is a common prefix but it is not the first one.
[0187] The sharedLen area is used to store the length of the common prefix, and the maximum common length is 65535.
[0188] The sharedPost area is used to store the distance between keys with common prefix data. For example, if it is adjacent to the common prefix data, it is recorded as 1; if it is separated from the common prefix data by 1, it is recorded as 2. The KV index segment array-index is used to accurately locate its absolute position in the file.
[0189] suffix stores suffix data. If the key has the first common prefix data, the value is the complete key; if the key previously carries common prefix data, the value is the suffix information after removing the common prefix data.
[0190] As an optional implementation of this embodiment, the encoding structure of the KV index segment array-index of this embodiment is as follows:
[0191] kv1.offset kv2.offset kv3.offset ...
[0192] The array-index section of the KV index records the starting position of the KV key-value pair in the file. The total length of the KV key-value pair can be determined by the difference between it and the next kv.offset. Each offset is fixed-length, and uint32 allows a maximum single file size of 4GB.
[0193] The array-index section of the KV index provides an index of all KV key-value pairs. It can provide binary search to locate a specific element, and can also provide traversal requests. After locating a certain element, it can traverse the subsequent / previous elements in sequence.
[0194] As an optional implementation of this embodiment, the key index segment hash-index described in this embodiment includes:
[0195] Hash the key to get an unsigned integer, record the high bit of the unsigned integer to get the key shard, and record the item at bit 0 in the corresponding key shard. If there are multiple elements in the same key shard, they can be recorded in subsequent positions in sequence. The content of the item is the serial number of the KV key-value pair in the entire file.
[0196] Therefore, the encoding structure of the key index segment hash-index in this embodiment is as follows Figure 2 As shown, this section is the index section, designed to speed up single key lookups. Specifically, the key is hashed to obtain a uint32 (4B). The high-order bits of the uint32 (uint16, 2B) are recorded, resulting in 65,536 shards. The item is recorded at position 0 in the corresponding shard. If the same shard has multiple elements, they are recorded in subsequent positions. The content of the item is the sequence number corresponding to the key-value pair (the sequence number of the key-value pair in the entire file).
[0197] A data storage method for multiple indexes in this embodiment includes a data search process based on a third data storage structure:
[0198] Search the key index segment hash-index based on the query key. If the key shard of the query key exists in the key index segment hash-index, obtain the item in the key shard.
[0199] According to the item, the KV index section array-index is found, the starting position offset of the KV key-value pair in the KV index section array-index is located, the difference between the current offset and the starting position offset stored in the next position is obtained, and the KV data section KVS is obtained;
[0200] The keySize area and the shareHeader area are obtained by decoding the KV data section KVS, and it is judged whether the current key is complete according to the value of the sharedHeader area;
[0201] If the judgment result is yes, the current key and the query key are compared, if they are consistent, the value is returned, and if they are inconsistent, the process is rolled back to find the key index section hash-index according to the query key and continues to find until the query key does not exist in the element of the current key shard, and then the process is returned.
[0202] If the judgment result is no, the common prefix data of the current key is found, the complete key is obtained by splicing the common prefix data of the current key and the suffix data of the current key, the current complete key and the query key are compared, if they are consistent, the value is returned, and if they are inconsistent, the process is rolled back to find the key index section hash-index according to the query key and continues to find until the query key does not exist in the element of the current key shard, and then the process is returned.
[0203] Further, the decoding of the KV data section KVS obtains the keySize area and the shareHeader area, and the judgment of whether the current key is complete according to the value of the sharedHeader area includes:
[0204] When the value of the sharedHeader area indicates that the current key does not contain common prefix data or contains common prefix data and is the first one, the current key is a complete key;
[0205] When the value of the sharedHeader area indicates that the current key contains common prefix data and is not the first one, the common prefix data is further found;
[0206] According to the value of the sharedPost area of the current key, the position of the common prefix data is obtained, then the KV key-value pair corresponding to the common prefix data is obtained, the common prefix data is obtained through the sharedLen area of the KV key-value pair with the common prefix data, and the complete key is obtained by splicing the common prefix data and the suffix data of the suffix area of the current key.
[0207] The data finding process based on the third data storage structure is specifically shown as follows:
[0208] When querying a single key, first search the key index segment hash-index, take the hash based on the query key to obtain a uint32, obtain the high bit, search the corresponding shard in memory, and then perform a binary search within the current shard (process-1-1).
[0209] The process of binary search is to obtain the serial number value of the KV key-value pair based on the binary index, and use the serial number value to search the KV index segment array-index. The serial number value is the nth element of the KV index segment array-index. Because each element of the KV index segment array-index is fixed-length, its starting address in the KV index segment array-index can be directly located, and the fixed-length bytes are taken. The byte corresponds to the offset of the KV in the entire file. The KV segment is obtained by the current offset and the offset stored in the next position. Decode the first three bytes to get keySize and shareHeader. If the sharedHeader value is 0 / 1, it means that the current key is complete. Get the current key and compare it with the searched key. If they are consistent, return value. If they are inconsistent, fall back to (process-1-1) and continue searching until the key does not exist in the element of the current shard and then return. If the sharedHeader value is 2, the common prefix data needs to be found. The search steps are as follows: according to the 4th byte value of the kv field, the location of the common prefix data is obtained (the current key is in the subscript of array-index, subtracting this value is the location of the common prefix data), and then the KV key-value pair corresponding to the common prefix data is obtained. The common prefix data is obtained through sharedLen, and then it is concatenated with the suffix data of the current KV to obtain the complete key. The comparison and fallback process is the same as when sharedHeader is 0 / 1.
[0210] This embodiment also provides a data storage device for multiple types of indexes, including a third data storage module, wherein the third data storage module includes:
[0211] The header identification section is used to write the type data representing the file encoding method;
[0212] The KV data segment KVS is used to write prefix-compressed KV key-value pairs in sequence;
[0213] The KV index segment array-index records the starting offset of each KV key-value pair in the file. The total length of a KV key-value pair is determined by the difference between the starting offset of the KV key-value pair and the starting offset of the next KV key-value pair.
[0214] key index section hash-index, used to record the corresponding index relationship between key and the serial number of KV key-value pair in the overall file.
[0215] Example 4
[0216] The data storage method of the embodiment of the present application is a multi-category index data storage method, comprising a fourth data storage structure, wherein the fourth data storage structure comprises:
[0217] header identification section header, used to write type data representing the encoding mode of the file;
[0218] KV data block section block, used to store a plurality of compressed KV key-value pair data;
[0219] KV data block index section array-index(block), used to record the starting position information of each KV data block section block.
[0220] The fourth data storage structure of the embodiment of the present application is more space-saving than the third data storage structure in the third embodiment, because the KV data block section block is used to store a plurality of compressed KV key-value pair data. However, decoding the entire KV data block section block will consume more CPU, or using cache to store the decompressed data will consume more memory. Therefore, the fourth data storage structure is suitable for scenarios where the data volume exceeds 100G, and is friendly to disk.
[0221] The overall encoding structure of the fourth data storage structure of the embodiment of the present application is shown in Figure 3 The header identification section header in the fourth data storage structure represents a file encoded by the fourth data storage structure, and the specific enumeration value is the same as that in the first, second and third embodiments.
[0222] The KV data block section block in the embodiment of the present application stores a segment of KV key-value pair. This part of data is stored in the file after compression, and the original content can be obtained by obtaining the data and performing inverse decoding when reading. The KV data block section block in the embodiment of the present application comprises:
[0223] max-key section, used to write the maximum key (max-key) in the KV data block section block;
[0224] header section, used to write the length of the maximum key (max-key);
[0225] KV data section KVS, used to sequentially write the prefix-compressed KV key-value pair data;
[0226] The KV index segment array-index records the starting offset of each KV key-value pair in the file. The total length of a KV key-value pair is determined by the difference between the starting offset of the KV key-value pair and the starting offset of the next KV key-value pair.
[0227] The key index segment hash-index is used to record the corresponding index relationship between the key and the KV key-value pair in the entire file;
[0228] The data in the KV data segment KVS, the KV index segment array-index and the key index segment hash-index are compressed as a whole and stored in a file.
[0229] Specifically, the encoding structure of the KV data block segment block in this embodiment is as follows:
[0230] header max-key kvs array-index hash-index
[0231] The header indicates the maximum key (max-key) length, and other information is of fixed length; it is not compressed.
[0232] max-key indicates the largest key in the KV key-value pairs stored in the KV data block segment block; it is not compressed.
[0233] KVS represents the KV key-value pair data storage segment, which has the same meaning as the KV data segment KVS in Example 3.
[0234] Array-index represents the file offset of each KV key-value pair, which has the same meaning as the array-index of the KV index segment in Example 3.
[0235] The hash-index represents the shard information of the KV key-value pair and can locate the key more quickly than a binary search of the entire file.
[0236] KVS+array-index+hash-index as a whole is compressed and stored in a file.
[0237] A data storage method for multiple indexes in this embodiment includes a data search process based on the fourth data structure:
[0238] Get the KV data block segment block information from the KV data block index segment array-index(block);
[0239] Get the maximum key (max-key) in the max-key area and the length of the maximum key (max-key) in the header area according to the KV data block segment block;
[0240] Determine whether the query key is greater than the maximum key (max-key);
[0241] If the query key is greater than the maximum key (max-key), then the KV data block segment block is searched backward in binary search; if the query key is less than the maximum key (max-key), then the KV data block segment block is searched forward in binary search;
[0242] Until the KV data block segment where the query key is located is greater than the maximum key (max-key) of the previous KV data block segment and smaller than the maximum key (max-key) of the current KV data block segment, continue searching within the current KV data block segment;
[0243] Decompress the current KV data block segment block to obtain the original data information of the KV data segment KVS, the KV index segment array-index and the key index segment hash-index. The query key is queried based on the original data information. If there is a query result, value is returned. If there is no query result, empty is returned.
[0244] This embodiment also provides a data storage device for multiple types of indexes, including a fourth data storage module, wherein the fourth data storage module includes:
[0245] The header identification section is used to write the type data representing the file encoding method;
[0246] The KV data block segment is used to store multiple compressed KV key-value pairs;
[0247] The KV data block index segment array-index (block) is used to record the starting position information of each KV data block segment block.
[0248] Example 5
[0249] The first data storage structure, second data storage structure, third data storage structure, and fourth data storage structure of Examples 1 to 4 of the present invention can be used individually or in combination. A variety of different data storage structure types can be used, which are distinguished by version numbers. The file can be parsed according to different parsing formats, which has better scalability.
[0250] The present invention uses a more suitable storage solution for different resources and different KV distributions. If the CPU is idle and the overall file is relatively large, the second data storage structure is used for extreme compression; if the key ratio in the file is larger than the value, and the CPU is relatively idle, the first data storage structure can be used; if it is relatively balanced, the third data storage structure is selected; on this basis, if the file exceeds 100G, the fourth data storage structure is considered for data compression to reduce the file's disk occupancy. The combined use of these solutions has stronger adaptability and higher innovation than a single storage solution.
[0251] The data storage method of a multi-type index in this embodiment can be expanded to more writing modes, and can be distinguished and expanded through the header identification segment header.
[0252] Example 6
[0253] In the second embodiment of the present invention, the Huffman segment value-block (Huffman) of the value data adopts the Huffman algorithm for value compression storage. The Huffman algorithm is a character encoding scheme that scans the frequency of each character in the entire sequence, and then encodes the high-frequency characters into fewer bits and the low-frequency characters into more bits, and then encodes each character in the entire sequence according to the encoding table, and then the encoding table is also written somewhere for decoding.
[0254] See also Figure 12 As shown, the frequency of the characters obtained by sequentially arranging v1-v5 is counted and sorted. The combination with the lowest frequency is formed into a binary tree, and then the frequency of the parent node of the binary tree is the sum of the frequencies of the leaf nodes. Then find the combination with the lowest frequency and finally get the Huffman tree as shown below. Figure 13 shown.
[0255] Then encode, where the encoding table is:
[0256] character coding Frequency a 0 10 c 10 5 b 110 3 d 111 1
[0257] According to this encoding table, the character sequence consisting of the five strings above is bit-encoded. Given that one character requires 8 bits to store, the bit length before encoding is 19 (characters) * 8 (8 bits per character) = 152. After encoding, the bit length is 1 * 10 + 2 * 5 + 3 * 3 + 3 * 1 = 32, demonstrating significant compression.
[0258] However, if the distribution of individual characters is quasi-exponential, the constructed Huffman tree may be relatively high (256 characters may have a maximum height of 255 layers). In this case, if a character is in the middle or lower part of the tree, its encoding length will be 32 bits, or even more, up to 255 bits. Compared with the original method of storing a single character with only 8 bits, the encoding length is significantly expanded. This embodiment addresses this problem by controlling the encoding length within a reasonable range when the Huffman tree height is too high, and the encoding and decoding efficiency is very high.
[0259] Assume that the result of character frequency statistics is:
[0260]
[0261]
[0262] Constructing a Huffman tree Figure 14 As shown, generate the coding table:
[0263]
[0264]
[0265] As can be seen from the figure, the length of many characters exceeds 8 bits. It is known that one character only requires 8 bits to store, but the height of the constructed Huffman matrix exceeds 8 layers, which makes the encoding length of some characters exceed 8 bits, and some even reach 17 bits, or more.
[0266] The improvement plan is that if the encoding length is 16 bits or less, it is considered a reasonable length, and the characters within this length are stored using the corresponding encoding. Above 16 bits, use the prefix + suffix encoding method. The prefix uses all 0{16 0s} or the prefix uses all 1{16 1s} (if the leaf node of the Huffman tree is on the right); the suffix is the original character itself. The prefix and suffix are concatenated to obtain the encoded bit information. For example, if the length of character a exceeds 16 bits, the encoding is: 0{16 0s}+a (itself occupies 8 bits), a total of 24 bits; the same is true for character b, the encoding is: 0{16 0s}+b (itself occupies 8 bits). In this way, when encoding characters with more than 16 bits, a maximum of 24 bits are used. The number of bits can be relaxed to 32 bits. If it exceeds 24 bits, the prefix uses 24 bits (fixed to all 0 or all 1), and the last 8 bits store the original character; the corresponding encoding is used for storage within 24 bits;
[0267] It can be seen from this that when a character string is compressed using a classic Huffman tree, the height of the constructed Huffman tree may be very high, and the encoding length may exceed the character itself by several times. After using the improved solution of this embodiment, the encoding length can be limited to 3 bytes or 4 bytes. In the worst case, the original Huffman tree may use 32 bytes (the highest level is 255, and the encoding length is 256 bits, that is, 32 bytes) to store the original character of 1 byte. In this scenario, the classic Huffman tree encoding takes up extra space, and the decoding speed increases with the increase of the encoding length, so the decoding speed is also lower. The technical solution of this embodiment reduces the character encoding length, thereby taking up less storage space, and the decoding table is smaller than before (because those characters exceeding the limited height do not need to be stored in the decoding table). The decoding speed corresponding to this encoding is also very fast, and the original character can be directly obtained when encountering a common prefix.
[0268] The key point of this embodiment is to stop using the classic Huffman tree encoding for characters with excessive height (for example, >16) and longer encoding lengths. The improved encoding method is: prefix + character itself, so that the encoding length can be controlled within 3 bytes (24 bits) or 4 bytes (32 bits). The improved scheme limits the encoding length to a reasonable range. This optimization is the key point. At the same time, it is easy to know that there will be no additional consumption for current encoding and decoding.
[0269] Example 7
[0270] The following describes an electronic device embodiment of the present invention, which can be considered a specific physical implementation of the method and apparatus embodiments of the present invention described above. Details described in the electronic device embodiment of the present invention should be considered supplementary to the above-mentioned method or apparatus embodiments; details not disclosed in the electronic device embodiment of the present invention can be implemented with reference to the above-mentioned method or apparatus embodiments.
[0271] Figure 15 This is a structural diagram of an electronic device according to an embodiment of the present invention. The electronic device includes a processor and a memory, wherein the memory is used to store a computer executable program. When the computer program is executed by the processor, the processor executes a multi-category index data storage method according to the above embodiment.
[0272] like Figure 15 As shown, the electronic device is implemented as a general-purpose computing device. The processor may be one or multiple processors working in concert. The present invention also does not exclude distributed processing, meaning that the processors may be dispersed across different physical devices. The electronic device of the present invention is not limited to a single entity but may also be the sum of multiple physical devices.
[0273] The memory stores a computer executable program, usually machine readable code. The computer readable program can be executed by the processor to enable the electronic device to perform the method of the present application, or at least some steps of the method.
[0274] The memory includes volatile memory, such as random access memory (RAM) and / or cache memory, and can also include non-volatile memory, such as read only memory (ROM).
[0275] Optionally, the electronic device further comprises an I / O interface for data exchange between the electronic device and external devices. The I / O interface can be one or more of several types of bus structures, including memory bus or memory controller, peripheral bus, graphics acceleration port, processing unit, or local bus using any of the bus structures.
[0276] It should be understood that Figure 15 The electronic device shown is only an example of the present application, and the electronic device of the present application can also include elements or components not shown in the above examples. For example, some electronic devices also include a display unit such as a display screen, and some electronic devices also include human-computer interaction elements such as buttons and keyboards. As long as the electronic device can execute the computer readable program in the memory to realize the method of the present application or at least some steps of the method, it can be considered as an electronic device covered by the present application.
[0277] Figure 16 is a schematic diagram of a computer readable recording medium according to an embodiment of the present application. As shown in Figure 16 The computer readable recording medium stores a computer executable program, which, when executed, realizes one or more of the above-mentioned data storage methods of the present application. The computer readable recording medium can include a data signal propagating in a baseband or as a carrier wave part of a carrier wave, which carries the readable program code. Such a propagating data signal can take many forms, including but not limited to electromagnetic signals, optical signals or any suitable combination of the above. The readable recording medium can also be any readable medium other than the readable recording medium, which can send, propagate or transmit programs for use by or in conjunction with an instruction execution system, device or apparatus. The program code contained on the readable recording medium can be transmitted by any suitable medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above.
[0278] The program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, and the like, as well as conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0279] Through the above description of the implementation mode, it is easy for those skilled in the art to understand that the present invention can be implemented by hardware capable of executing a specific computer program, such as the system of the present invention, and the electronic processing unit, server, client, mobile phone, control unit, processor, etc. contained in the system. The present invention can also be implemented by computer software that executes the method of the present invention, such as control software executed by a microprocessor, an electronic control unit, a client, a server, etc. However, it should be noted that the computer software that executes the method of the present invention is not limited to being executed by one or a specific hardware entity, and it can also be implemented in a distributed manner by unspecified specific hardware. For computer software, the software product can be stored in a computer-readable recording medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), or it can be distributed and stored on a network, as long as it enables an electronic device to execute the method according to the present invention.
[0280] The above embodiments are only used to illustrate the present invention and are not intended to limit the technical solutions described in the present invention. Although this specification has described the present invention in detail with reference to the above embodiments, the present invention is not limited to the above specific implementation methods. Therefore, any modification or equivalent replacement of the present invention; and all technical solutions and improvements thereof that do not depart from the spirit and scope of the invention are included in the scope of the claims of the present invention.
Claims
1. A data storage method for multiple types of indexes, characterized in that: The first data storage structure includes: The header identification section is used to write the type data representing the file encoding method; The key data segment bitrie is used to write the key of the KV key-value pair in sequence, build the key into a multi-branch tree, number each node in sequence, and record the current key number on the corresponding leaf node when adding a key. Then, the multi-branch tree is traversed to obtain the complete bitmap, labels, and leaf-value. Among them, the bitmap records the parent-child node relationship of the entire multi-branch tree, the labels are used to map the nodes of the entire multi-branch tree to specific characters for matching, and the leaf-value is used to calculate the complete match from the root node to the leaf node and then query the value; The value data segment value-block stores the corresponding values in the order of the keys; The value index segment value-index stores the offset of each value in the file in sequence.
2. The data storage method of a multi-category index according to claim 1, characterized in that: The process of writing data into the key data segment bitrie includes: The keys of multiple KV key-value pairs written in sequence are decomposed into single characters, and all the single characters of each key are added to the multi-tree in sequence. At the same time, when a single character is added to the multi-tree, the sequence number of the current key is recorded in the leaf node. The root node of the multi-tree does not store specific characters and is used to index each key; Encode each node on the multi-branch tree. If there are N leaf nodes, the encoding is N 1s and 1 0. If there is no leaf node, the encoding is 0. Perform hierarchical traversal on the encoded multi-branch tree, write the encoding of each node into the bitmap according to the level, and write each leaf node into the labels; According to the bitmap and labels, the corresponding relationship between key-serial number, key and leaf-serial number is obtained, sorted according to leaf-serial number, and the leaf-serial number and key-serial number data pairs are stored in the leaf-value structure.
3. The data storage method of a multi-category index according to claim 2, characterized in that: The formula for calculating the parent and child nodes of the current node based on the bitmap is: FirstChild(x)=rank1(select0(x)+1) Parent(x)=rank0(select1(x)) select0(x) indicates the position of the xth 0 in the bitmap; select1(x) indicates the position of the xth 1 in the bitmap; Rank0(x) indicates how many zeros appear from bitmap 0 to bitmap x; Rank1(x) indicates how many 1s appear from bitmap 0 to bitmap x.
4. A method for storing data of multiple indexes according to any one of claims 1 to 3, characterized in that: The data search process includes: The first child node of the root node of the multitree in the bitrie of the key data segment is searched for the query key and compared. If the comparison is successful, the child nodes of the first child node are further searched until the query key is found in the leaf node of the first child node or the query key is not found, and the process is exited. If the comparison fails, the sibling nodes of the first child node are searched until the query key is found in the leaf node of the sibling node or the query key is not found, and the process is exited. Get the leaf-number based on the key found in the bitrie query of the key data segment, query the leaf-value structure based on the leaf-number, and locate the value-number through binary search; Search in the value index section array-index through value-serial number. Use value-serial number to directly locate the offset of value in the file, and get the value from the value data section value-block.
5. The method for storing data of multiple indexes according to claim 4, characterized in that: The comparison of the first child node of the multitree root node in the query key data segment bitri e according to the query key includes: First, the number of the corresponding node is queried based on the bitmap, and then the position of the number is directly located based on the information saved in the labels. The value is compared with the first character of the query key to obtain the comparison result.
6. A method for storing data of multiple types of indexes according to any one of claims 1 to 3, characterized in that: A second data storage structure is formed by replacing the value data segment value-block of the first data storage structure with the value data Huffman segment value-block (Huffman), wherein the data storage structure of the value data Huffman segment value-block (Huffman) includes: The value storage area uses the Huffman algorithm for value compression storage; The decoding index area decode-index is used to store the index relationship between the compressed storage char file of value and the codec table.
7. The method for storing data of multiple indexes according to claim 6, characterized in that: The value storage area uses the Huffman algorithm to perform value compression storage, including: Construct Huffman tree and Huffman encoding and decoding table for the values written in sequence; The values written sequentially are stored in the value storage area according to the Huffman encoding and decoding table. The encoding structure of the value storage area includes: a value encoding storage area encode-value for storing the Huffman encoding of the value, an alignment area align for aligning the Huffman encoding of the value stored in the value encoding storage area encode-value to a fixed length, and an alignment bit counting area align-num for writing the number of alignment bits used during alignment in the alignment area align.
8. The method for storing data of multiple indexes according to claim 7, characterized in that: The encoding structure of the decoding index area decode-index is: Among them, a block encoding structure is: For each character, according to the frequency statistics table of the Huffman algorithm, the decoding value of the decoding table is written into decode-value, the corresponding character is written into char, and then the encoding bit length of the character is written into encode-len, and the encoding and decoding information of each character is written in turn.
9. The method for storing data of multiple indexes according to claim 8, characterized in that: The data search process includes: When querying, the data information stored in the key data segment bitri e can be used to determine whether the query key exists. If it exists, its value-sequence number can be obtained. The value-sequence number is searched in the value index segment value-index to obtain the offset of its value in the file. The value size is obtained based on the offset difference between the previous and next values. Read the encoded data of value in the file; The valid bit information is known according to align-num. Decoding is completed when the valid bit information is reached, and the decoded value is returned.
10. A data storage device with multiple indexes, characterized in that: The first data storage module includes: The header identification section is used to write the type data representing the file encoding method; The key data segment bitrie is used to write the key of the KV key-value pair in sequence, build the key into a multi-branch tree, number each node in sequence, and record the current key number on the corresponding leaf node when adding a key. Then, the multi-branch tree is traversed to obtain the complete bitmap, labels, and leaf-value. Among them, the bitmap records the parent-child node relationship of the entire multi-branch tree, the labels are used to map the nodes of the entire multi-branch tree to specific characters for matching, and the leaf-value is used to calculate the complete match from the root node to the leaf node and then query the value; The value data segment value-block stores the corresponding values in the order of the keys; The value index segment array-index stores the offset of each value in the file in sequence.