Efficient Use of TRIE Data Structure in a Database
By adopting bitmap and hierarchical structure in the trie data structure, the problem of inefficient storage space when storing database indexes is solved, efficient key operation and query performance is achieved, and it is suitable for indexing requirements of various cardinalities.
Patent Information
- Application Number
- CN201880093497.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-03-15
- Filing Date
- 2018-09-19
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2038-09-19
AI Technical Summary
In the prior art, trie data structures have problems with low storage space efficiency when storing database indexes, especially when the alphabet cardinality is large or trie is sparsely filled, and multi-dimensional query and low-cardinality standard query cannot be effectively processed.
Using a bitmap-based trie data structure, and allowing lazy calculations through a hierarchical trie structure, the first-level trie is used as an associative array to realize reverse indexing, and the second-level trie is used as a collection to store document lists, realizing a general indexing solution called "converged indexing".
This solution can frequently update indexes, is suitable for low cardinality and high cardinality keys, and saves space without compression and decompression, provides a constant time complexity of O(|M|) for key operations, which is better than tree-based methods, and query time is independent of database filling.
Smart Images

Figure CN112219199B_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to the efficient use of a trie data structure in databases and information retrieval systems, and to querying such systems with high performance. Background Art
[0002] Databases and information retrieval systems are used to process structured and unstructured information. Generally, structured data is the domain of databases (such as relational databases), while unstructured information is the domain of information retrieval systems (such as full-text search). A database engine is part of a database management system (or other application) that stores and retrieves data. For information retrieval systems, this function is performed by a search engine.
[0003] Indexing is used to improve the performance of a database or an information retrieval system. In the absence of an index for a query (also known as looking up or accessing by "key"), it would be necessary to scan the entire database or repository to deliver a result, which would be too slow to be useful.
[0004] Database indexing is comparable to the index provided by a book: To find a specific keyword, a user does not have to read the entire book, but can look up a certain keyword in the index, which contains references to the pages related to that keyword. This is also the basic principle behind search engines: For any given search term, they quickly find the documents containing that search term by consulting an appropriate index. An example query for an information retrieval system is full-text search, for which the terms (words) are stored as keys, and the document IDs are also stored as keys (see the description of FIG. 70 below). For example, if the term "apple" can be found in documents with IDs 10 and 33, then the two composite keys ("apple", 10) and ("apple", 33) are stored in the index. These terms are stored as keys so that searches such as "apple" and "pear" can be performed.
[0005] The classical method of indexing is to use a so-called B-tree index. The B-tree was invented by R. Bayer and E. McCreight in 1972 and remains the main data structure for this purpose. However, B-tree indexes have major drawbacks. For example, the time to access data increases logarithmically with the amount of data. An increase in the data size by one order of magnitude approximately doubles the access time, and an increase in the data size by two orders of magnitude quadruples the access time, and so on. In addition, B-tree indexes do not help to improve the performance of so-called low-cardinality criteria queries. For example, creating an index on the attribute "gender" with values "male", "female", and "unknown" does not improve performance compared to scanning and filtering all records. Finally, since B-tree indexes cannot be effectively joined (combined), multi-dimensional queries, i.e., queries involving multiple conditions / attributes, are difficult to process effectively.
[0006] M. Boehm et al., "Efficient In-Memory Indexing with Generalized Prefix-Trees", in: T. et al. (eds.), BTW. LNI, Vol. 180, pp. 227-246, Kaiserslautern, Germany (2011), proposed using a "trie" or trie data structure to store and process database indexes. A trie is a tree data structure and is sometimes referred to as a "radix tree" or "prefix tree". The word trie originated from the word "reTRIEval". Instead of storing keys inside nodes, the path to a node in a trie defines the key associated with it, where the root node represents the empty key. More specifically, each node is associated with a key part, and the value of this key part (sometimes referred to as the "value of the node" in this article) can be indicated by a pointer from the parent node. This value is selected from a predefined alphabet of possible values. The path from the root node in a trie to another node (e.g., to a leaf node) defines the key (or "key prefix" in the case of an internal node, i.e., a node that is not a leaf node) associated with the node, where the key is the concatenation of the key parts associated with the nodes on the path. The time complexity of a trie does not depend on the number of keys present in the trie, but rather on the key length.
[0007] Space-saving trie data structure
[0008] One way to implement a trie data structure is to store nodes separately in memory, where each node includes: node type information indicating whether the node is a leaf node; an array of pointers to its child nodes (at least in the case where the node is not a leaf node); and in the case of a leaf node, a possible payload, such as a value associated with the key. Using this method, traversing from node to node is a constant-time operation because the corresponding child pointers correspond to array entries representing the corresponding key parts. However, since the corresponding arrays include and store many empty entries, memory usage can be very inefficient for nodes with only a few child nodes. This is especially true when using a larger alphabet because they require storing a large number of child pointers.
[0009] To use memory space more effectively, implementations of trie data structures that avoid storing null pointers have been developed. Such solutions can include storing only a list of non-null pointers and their corresponding key-part values, rather than an array containing all possible pointers. A disadvantage of this approach is that for traversing from node to node, the list in the corresponding parent node must be scanned, or a binary search must be performed if the list is sorted by the value of the key part. Additionally, since specific pointers for specific values of the key part need to be identified, the associated values of the key part must also be stored, which reduces memory space efficiency.
[0010] Other ways of storing trie data structures in a compact format have been developed. For example, Dr. Bagwell, “Fast And Space Efficient Trie Searches”, Technical Report, EPFL, Switzerland (2000), discloses a bitmap-based trie data structure. This trie data structure uses a bitmap to mark all non-null pointers of a parent node. In particular, the set bits in the bitmap mark the valid (non-null) branches. Each parent node also includes one or more pointers, where each pointer is associated with a set bit in the bitmap and points to a child node of the parent node. The value of the key part of the child node is determined by the value of the bit (set) in the bitmap included in the parent node, where the bit is associated with the pointer pointing to that child node.
[0011] The pointers have a predetermined length or size and can be stored in the same or reverse order as the bits set in the bitmap. Based on the number of least significant bits set in the bitmap, the memory address of the pointer associated with the bits set in the bitmap can be easily calculated. Since the number of least significant bits set can be efficiently calculated using simple bit operations and the CTPOP (count trailing ones) operation for determining the number of set bits, the determination of the corresponding pointer and thus the next child node is fast. For example, such a count trailing ones method is available in the Java programming language and is called “Long.bitCount()”. The CTPOP itself can be implemented very efficiently using “bit-hacks”, and many modern CPUs even provide CTPOP as an internal instruction.
[0012] Since the bitmap indicates the alphabet values associated with the valid branches, only the existing (non-null) pointers need to be stored. Thus, memory usage can be reduced. On the other hand, based on the rank of its associated bit among the set bits in the bitmap, the address of a specific pointer can be easily determined in constant time. Finally, the value of the bit associated with the pointer pointing to a child node represents the value of the key part of the child node in an efficient manner. Thus, compared with based on Figure 4Compared with the trie data structure of the list shown in [reference], the bitmap-based trie data structure provides a more efficient way to handle memory allocation and process the trie.
[0013] However, the inventors have found that in many application scenarios, the memory space is used inefficiently. This is especially true when the trie is sparsely populated and / or degenerates into a chain of nodes (each node having a single child pointer and wasting space). Therefore, it is desirable to further reduce the storage space required to store the trie used by a database application or an information retrieval system, and in particular to reduce the amount of memory required to store the pointers and / or bitmaps used to implement the trie, without significantly increasing the speed required to traverse the trie.
[0014] Key Encoding
[0015] Typically, the radix of the alphabet of the trie used in the prior art is relatively large, such as 256, in order to be able to accommodate characters from a larger alphabet (such as Unicode). Using 256 different values, 8 bits (2 8 = 256) or 1 byte can be encoded. However, in the case where the trie uses a bitmap as described above, the width of the bitmap increases with the radix of the alphabet. For example, an alphabet with a radix of 256 requires a bitmap size of 256 bits. Such a large bitmap may be space-inefficient, especially when the trie is sparsely populated, because 256 bits need to be allocated for each node, although only a small portion of them may be used. In addition, since the bit width of the registers of modern computers is typically only 64 bits, a large bitmap with, for example, 256 bits cannot be processed in a time-saving manner. If the trie is to accommodate characters from an even larger alphabet, the efficiency problem will increase further. Therefore, the radix of the alphabet that can store its characters in the prior art trie in a space-saving and time-saving manner is limited.
[0016] In addition, the trie is only used to store predefined data types in the prior art, where the data types are typically limited to primitive data types such as numbers or characters. Such a use imposes constraints on the keys that can be stored in the trie and makes the trie data structure less suitable for storing database or information retrieval system indexes.
[0017] Accordingly, another object of the present invention is to provide a trie that can be used in a flexible manner, and / or a method of using a trie in a flexible manner. In fact, the object of the present invention is to provide a trie data structure that can be used as an index for a general database or an information retrieval system, and a method of using the trie data structure as an index for a general database or an information retrieval system. In addition, the object of the present invention is to provide a trie that stores keys of a database index in such a way that database queries involving more than one data item can be processed in a time-saving manner.
[0018] Query execution
[0019] Queries executed in a database are typically expressed in a logical algebra (e.g., SQL). For query execution, the query must be translated into a physical algebra: a physical query execution plan (QEP). The query is rewritten, optimized, and a QEP is prepared so that a query execution engine (QEE) executes the QEP generated by the previous steps on the database.
[0020] For processing, a QEP typically includes a set of operators aimed at producing a query result. Most databases represent a QEP as a tree, where nodes are operators, leaf nodes are data sources, and edges are relationships between operators in a producer-consumer form.
[0021] Many database engines follow an iterator-based execution model in the QEE, where an operator implements the following methods: Open (to make the operator ready to produce data), Next (to produce a new data unit on demand of an operator consumer), and Close (to complete execution and release resources). Invoking one of these operations starting from the root operator propagates it to its operator children, and so on, until the data source (and leaf nodes) is reached. In this way, in the query execution plan operator tree, control flows from the consumer down to the producer, while data flows from the producer up to the consumer. Since each operator does not require global knowledge, this approach provides a concise design and encapsulation.
[0022] However, such an execution model has serious drawbacks. First, since each operator does not have global knowledge, optimizations that would be beneficial from a global perspective cannot be applied. Additionally, since the query plan is static, it may be difficult to apply adaptive optimizations during query execution. As a result, there is a strong reliance on the query optimizer to create a good QEP involving complex algorithms. And finally, the iterator method delivers only one data unit, e.g., recording each operator call. This method is inefficient for operators that combine large sub-result sets that themselves return small result sets.
[0023] As an example of one of these drawbacks, if the query to be executed on a database is an expression such as "A intersect (B union C)", where A returns a short list of record IDs, e.g., (1, 2, ... 10), but B and C return long lists, e.g., (1, 2, ... 10,000) and (20,000, ... 50,000). If the query optimizer has no predictive ability and does not rewrite the query as (A intersect B) union (A intersect C) before its execution, the query is very inefficient.
[0024] Accordingly, there is a desire for a more time-efficient method of performing database or information retrieval system queries, which includes performing operations such as intersection (AND), union (OR), and difference (AND NOT) on two or more sets of keys stored in the database or information retrieval system, or on the input set or result key set of the database or information retrieval system.
[0025] In addition, the range query performance of existing technology databases degrades with the size of the index (the number of records included in the database and indexed by the index). Accordingly, there is a desire for a more time-efficient method of performing range queries, which can scale well as the index size increases. SUMMARY OF THE INVENTION
[0026] The present invention provides an indexing solution for database applications, where B-trees and derivatives remain the main strategy, and an indexing solution for information retrieval applications, where inverted indexes are typically used. An inverted index is an index where the terms list the documents containing the index. The present invention can utilize a hierarchical trie structure to allow lazy evaluation when the trie is processed level by level. Thus, the present invention can implement an inverted index using the trie on the first level as an associative array, where the terms are the keys stored in the trie. On the second level (i.e., the leaf nodes of the associative array), the present invention can use the trie as a set to implement the document list (i.e., the set of IDs). Thus, the present invention allows the use of a general solution instead of B-trees and inverted indexes, and can thus be referred to as a "confluence index".
[0027] The index data structure and query processing model of the present invention are based on a bitwise trie and have the potential to replace existing technology index structures: the data structure can be updated frequently, is suitable for keys with low and high cardinality, and saves space without compression / decompression (the data structure is compact or even succinct in some cases). Additionally, since it is based on a trie, it inherits an O(|M|) constant time complexity for insertion, update, deletion, and query operations by key (where |M| is the key length). This is better than the typical O(log n) complexity of tree-based methods (where n is the number of keys in the index). Then, it can provide query times independent of database population. In a preferred embodiment, the result of set operators in query processing is "displayed" as a trie. This allows functional composition and lazy evaluation. As a side effect, the physical algebra directly corresponds to the logical algebra, thus simplifying the task of creating a suitable physical execution plan for a given query, as typically done in a query optimizer.
[0028] A first embodiment of the present invention is a trie for an electronic database application or information retrieval system, the trie including one or more nodes, wherein each parent node included in the trie, preferably each parent node having more than one child node, includes a bitmap and one or more pointers, wherein each pointer is associated with a bit set in the bitmap and points to a child node of the parent node. The trie is characterized in that each parent node included in the trie, preferably each parent node having only one child node, does not include a pointer to the child node, and / or the child node is stored in a predefined position in the memory relative to the parent node.
[0029] According to a second embodiment, in the first embodiment, the child node of a parent node having only one child node is stored in a position in the memory directly behind the parent node.
[0030] According to a third embodiment, in the first or second embodiment, a node, preferably each child node, is associated with a key part, and a path from a root node in the trie to another node, particularly to a leaf node, defines a key, which is a concatenation of the key parts associated with the nodes in the path.
[0031] Terminal optimization
[0032] A fourth embodiment of the present invention is a trie for an electronic database application or an information retrieval system, the trie including one or more nodes, wherein each parent node included by the trie, preferably each parent node having more than one child node, includes a bitmap; a node, preferably each child node, is associated with a key portion; and the value of the key portion of the child node, preferably the value of the key portion of at least each child node whose parent node has more than one child node, is determined by the value of the bit (set) in the bitmap included by the parent node, depending on which bit the child node is associated with. The trie is characterized in that a node, preferably each node having only one child node and all its descendant nodes having at most one child node, is marked as a terminal branch node, and the value of the key portion associated with the terminal branch node, preferably each descendant node of the terminal branch node, preferably each descendant node, is not determined by the value of the bit (set) in the bitmap included by the parent node of the descendant node.
[0033] According to a fifth embodiment, in the fourth embodiment, the terminal branch node has more than one descendant node.
[0034] According to a sixth embodiment, in the fourth or fifth embodiment, the parent node of the terminal branch node has more than one child node.
[0035] According to a seventh embodiment, in the fourth to sixth embodiments, the mark of the terminal branch node is a bitmap with no set bits.
[0036] According to an eighth embodiment, in the seventh embodiment, the bitmap of the terminal branch node has the same length or format as the bitmap included by the parent node having more than one child node.
[0037] According to a ninth embodiment, in any one of the fourth to eighth embodiments, the terminal branch node, preferably each terminal branch node included by the trie, and / or the descendant nodes of the terminal branch node, preferably each descendant node, do not include a pointer to its child node, and / or the child node is stored in a predefined position in the memory relative to the parent node, preferably in a position directly behind the parent node in the memory.
[0038] According to a tenth embodiment, in any one of the fourth to ninth embodiments, the value of the key portion associated with the terminal branch node, preferably each descendant node of the terminal branch node, preferably each descendant node, is included by the parent node of the descendant node.
[0039] According to the 11th embodiment, in any one of the 4th to 10th embodiments, the values of the key portions associated with the terminal branch nodes, preferably the descendant nodes of each terminal branch node, preferably all descendant nodes, are stored continuously after the terminal branch nodes.
[0040] According to the 12th embodiment, in any one of the 4th to 11th embodiments, compared with the bitmaps included by the parent nodes having more than one child node, the encoding of the values of the key portions associated with the terminal branch nodes, preferably the descendant nodes of each terminal branch node, preferably each descendant node, requires less memory space.
[0041] According to the 13th embodiment, in any one of the 4th to 12th embodiments, the values of the key portions associated with the terminal branch nodes, preferably the descendant nodes of each terminal branch node, preferably each of the descendant nodes, are encoded as binary numbers.
[0042] According to the 14th embodiment, in any one of the 4th to 13th embodiments, the bitmaps included by the parent nodes, preferably each parent node having more than one child node, have 32, 64, 128, or 256 bits, and the key portions associated with the terminal branch nodes, preferably the descendant nodes of each terminal branch node, preferably each of the descendant nodes, are encoded by 5, 6, 7, or 8 bits respectively.
[0043] According to the 15th embodiment, in any one of the 4th to 14th embodiments, the values of the key portions associated with the terminal branch nodes, preferably the descendant nodes of each of the terminal branch nodes, preferably each of the descendant nodes, are encoded as integer values.
[0044] According to the 16th embodiment, in any one of the 4th to 14th embodiments, the descendant nodes of the terminal branch nodes as parent nodes, preferably each descendant node as a parent node, do not include a bitmap in which the set bits determine the values of the key portions associated with their child nodes.
[0045] According to the 17th embodiment, in any one of the 4th to 16th embodiments, the parent nodes included by the trie, preferably at least each parent node having more than one child node, include one or more pointers, wherein each pointer is associated with a bit set in the bitmap included by the parent node and points to a child node of the parent node.
[0046] According to the 18th embodiment, in any one of the 4th to 17th embodiments, the trie is a trie according to any one of the 1st or 2nd embodiments.
[0047] Bitmap Compression
[0048] The 19th embodiment of the present invention is a trie for an electronic database application or an information retrieval system. The trie includes one or more nodes, where each node, preferably each parent node having more than one child node, includes a bitmap in the form of a logical bitmap and a plurality of pointers. Each pointer is associated with a bit set in the logical bitmap and points to a child node of the node. The trie is characterized in that the logical bitmap is divided into a plurality of sections and encoded by a header bitmap and a plurality of content bitmaps; each section is associated with a bit in the header bitmap; and for each section of the logical bitmap in which one or more bits are set, the bit associated with the section in the header bitmap is set and the section is stored as a content bitmap.
[0049] According to the 20th embodiment, in the 19th embodiment, for each section of the logical bitmap in which no bit is set, the bit associated with the section in the header bitmap is not set and the section is not stored as a content bitmap.
[0050] According to the 21st embodiment, in any one of the 19th or 20th embodiments, each of the sections is contiguous.
[0051] According to the 22nd embodiment, in any one of the 19th to 21st embodiments, all sections have the same size.
[0052] According to the 23rd embodiment, in the 22nd embodiment, the size of the section is one byte.
[0053] According to the 24th embodiment, in any one of the 19th to 23rd embodiments, the number of sections stored as content bitmaps is equal to the number of bits set in the header bitmap.
[0054] According to the 25th embodiment, in any one of the 19th to 24th embodiments, the size of the header bitmap is one byte.
[0055] According to the 26th embodiment, in any one of the 19th to 25th embodiments, the content bitmaps are stored in predefined positions in the memory relative to the header bitmap.
[0056] According to the 27th embodiment, in any one of the 19th to 26th embodiments, the content bitmaps of the logical bitmap are stored in an array, a list, or in contiguous physical or virtual memory locations.
[0057] According to the 28th embodiment, in any one of the 19th to 27th embodiments, the content bitmaps are stored in the same or reverse order, where the set bits associated with their sections are arranged in the header bitmap.
[0058] According to the 29th embodiment, in any one of the 19th to 28th embodiments, among all the set bits in the head bit map, the rank of the content bit map in all the content bit maps of the logic bit map corresponds to the rank of the set bit associated with the section of the content bit map.
[0059] According to the 30th embodiment, in any one of the 19th to 29th embodiments, the pointers included in the node, preferably each pointer of the node and / or each pointer of each node that is not a leaf node, are encoded in the encoding manner defined for the logic bit map in any one of the 19th to 29th embodiments.
[0060] According to the 31st embodiment, in any one of the 19th to 30th embodiments, the trie is the trie in any one of the 1st to 18th embodiments.
[0061] According to the 32nd embodiment, in any one of the 19th to 31st embodiments, each child node of the node is preferably associated with a key part, and the path from the root node to another node, especially to a leaf node, in the trie defines a key, which is a concatenation of the key parts associated with the nodes in the path.
[0062] A key including control information
[0063] The 33rd embodiment of the present invention is a trie for an electronic database application or an information retrieval system, the trie including one or more nodes, wherein each child node of the node is preferably associated with a key part; the path from the root node to another node, especially to a leaf node, in the trie defines a key associated with the node, and the key is a concatenation of the key parts associated with the nodes on the path. The trie is characterized in that the key includes control information and content information.
[0064] According to the 34th embodiment, in the 33rd embodiment, the key includes one or more key parts containing content information, and for each of the key parts, the control information includes a data type information element that specifies the data type of the content information included in the key part.
[0065] According to the 35th embodiment, in the 34th embodiment, each key part preferably includes a data type information element that specifies the data type of the content information included in the key part.
[0066] According to the 36th embodiment, in the 35th embodiment, the data type information element is located by the content information element, preferably before the content information element.
[0067] According to the 37th embodiment, in the 34th embodiment, the data type information elements are located together and are preferably arranged in the same or reverse order as the content information elements for which they specify the data type.
[0068] According to the 38th embodiment, in the 37th embodiment, the control information is located before the content information in the key.
[0069] According to the 39th embodiment, in any one of the 34th to 38th embodiments, the key includes two or more key parts, which include content information of different data types.
[0070] According to the 40th embodiment, in any one of the 34th to 39th embodiments, at least one of the data types is a fixed-size data type.
[0071] According to the 41st embodiment, in the 40th embodiment, the fixed-size data type is an integer, a long integer, a double-precision floating point, or a time / date primitive.
[0072] According to the 42nd embodiment, in any one of the 34th to 41st embodiments, at least one of the data types is a variable-size data type.
[0073] According to the 43rd embodiment, in the 42nd embodiment, the variable-size data type is a string, preferably a Unicode string, or a variable-precision integer.
[0074] According to the 44th embodiment, in any one of the 34th to 43rd embodiments, the information of the key part is included by two or more key parts.
[0075] According to the 45th embodiment, in any one of the 34th to 44th embodiments, the data type of the content information included by the key part is a variable-size data type, and the end of the content information element is marked by a specific bit or a specific symbol in a specific one of the key parts that include the key part.
[0076] According to the 46th embodiment, in any one of the 34th to 45th embodiments, the control information includes information identifying the last key part.
[0077] According to the 47th embodiment, in any one of the 33rd to 46th embodiments, the control information includes information about whether the trie is for storing a dynamic set or an associative array.
[0078] According to the 48th embodiment, in any one of the 33rd to 47th embodiments, the trie is a trie according to any one of the 1st to 32nd embodiments.
[0079] According to the 49th embodiment, in any one of the 33rd to 48th embodiments, nodes, preferably at least each parent node having more than one child node, include a bitmap and a plurality of pointers, wherein each pointer is associated with a bit set in the bitmap and points to a child node of the node.
[0080] Interleaved multi-key
[0081] The 50th embodiment of the present invention is a trie for an electronic database application or an information retrieval system, the trie including one or more nodes, wherein nodes, preferably each child node, is associated with a key part; a path from a root node in the trie to another node, particularly to a leaf node, defines a key associated with the node, and the key is a concatenation of key parts associated with the nodes on the path. The trie is characterized in that two or more data items are encoded in the key, at least one or two of the data items, preferably each, consisting of two or more components; and the key contains two or more consecutive sections, at least one or two of the sections, preferably each, including two or more components of the data items encoded in the key.
[0082] According to the 51st embodiment, in the 50th embodiment, each section of the key, preferably each section, contains at least one component and / or at most one component from each data item encoded in the key.
[0083] According to the 52nd embodiment, in any one of the 50th or 51st embodiments, for two or more sections of the key, preferably all sections, components belonging to different data items are sorted in the same sequence within the section.
[0084] According to the 53rd embodiment, in any one of the 50th to 52nd embodiments, the order of the sections including components of the data item corresponds to the order of the components within the data item.
[0085] According to the 54th embodiment, in any one of the 50th to 53rd embodiments, the key part associated with a child node, preferably each child node, corresponds to a part of a component of the data item.
[0086] According to the 55th embodiment, in any one of the 50th to 53rd embodiments, the key part associated with a child node, preferably each child node, corresponds to a component of the data item, and / or a component of the data item, preferably each component of each data item, preferably corresponds to the key part associated with a child node of the trie.
[0087] According to the 56th embodiment, in any one of the 50th to 53rd embodiments, the key portion associated with the child node, preferably with each of the child nodes, corresponds to more than one component of the data item.
[0088] According to the 57th embodiment, in any one of the 50th to 56th embodiments, two or more, preferably all, of the data items of the key have the same number of components.
[0089] The type of the data item and its components
[0090] According to the 58th embodiment, in any one of the 50th to 57th embodiments, two or more data items represent geographical location data.
[0091] According to the 59th embodiment, in any one of the 50th to 58th embodiments, the data item represents longitude, latitude, or an index, or a string, or a combination of two or more of these.
[0092] According to the 60th embodiment, in any one of the 50th to 59th embodiments, the component of the data item is a group of bits of the binary encoding of the data item.
[0093] According to the 61st embodiment, in the 60th embodiment, the group of bits includes 6 bits.
[0094] According to the 62nd embodiment, in any one of the 50th to 61st embodiments, the data item is a number.
[0095] According to the 63rd embodiment, in the 62nd embodiment, the data item is an integer, a long integer, or a double long integer.
[0096] According to the 64th embodiment, in any one of the 62nd or 63rd embodiments, the data item is a 64-bit integer.
[0097] According to the 65th embodiment, in any one of the 62nd to 64th embodiments, the component of the data item is a digit.
[0098] According to the 66th embodiment, in the 65th embodiment, the digit has a predefined radix, preferably 64.
[0099] According to the 67th embodiment, in any one of the 50th to 66th embodiments, the data item is a string.
[0100] According to the 68th embodiment, in the 67th embodiment, the component of the data item is a single character.
[0101] According to the 69th embodiment, in any one of the 50th to 68th embodiments, the data item is a byte array.
[0102] According to the 70th embodiment, in any one of the 50th to 69th embodiments, the trie is a trie according to any one of the 1st to 49th embodiments.
[0103] According to the 71st embodiment, in the 70th embodiment, when subordinate to the 34th embodiment, the data item corresponds to the key part or the content information included by the key part.
[0104] According to the 72nd embodiment, in any one of the 33rd to 71st embodiments, the node, preferably at least each parent node having more than one child node, includes a bitmap and a plurality of pointers, wherein each pointer is associated with a bit set in the bitmap and points to a child node of the node.
[0105] Overall trie features
[0106] Bitmap and memory details
[0107] According to the 73rd embodiment, in any one of the 1st to 32nd, or 49th or 72nd embodiments, the bitmap is stored in the memory as an integer of a predefined size.
[0108] According to the 74th embodiment, in any one of the 1st to 32nd, or 49th or 72nd or 73rd embodiments, the size of the bitmap is 32, 64, 128 or 256 bits.
[0109] According to the 75th embodiment, in any one of the 1st to 32nd, or 49th or 72nd to 74th embodiments, the trie is suitable for storage and processing on a target computer system, and the size of the bitmap is equal to the bit width of the CPU register, system bus, data bus and / or address bus of the target computer system.
[0110] According to the 76th embodiment, in any one of the 1st to 32nd, or 49th or 72nd to 75th embodiments, the bitmap and / or pointer and / or node of the trie are stored in an array, preferably in a long integer array or byte array, a list or in consecutive physical or virtual memory locations.
[0111] According to the 77th embodiment, in any one of the 1st to 16th, or 26th to 27th or 76th embodiments, the memory is or includes physical memory or virtual memory, preferably consecutive memory.
[0112] Pointer
[0113] According to the 78th embodiment, in any one of the 1st to 32nd, or 49th, or 72nd to 77th embodiments, the number of pointers included by each parent node, preferably each parent node having more than one child node, is equal to the number of bits set in the bitmap included by the parent node.
[0114] According to the 79th embodiment, in any one of the 1st to 32nd, or 49th, or 72nd to 78th embodiments, the rank of a pointer among all the pointers of a parent node corresponds to the rank of the associated set bit of the pointer among all the set bits in the bitmap of the parent node.
[0115] According to the 80th embodiment, in any one of the 1st to 32nd, or 49th, or 72nd to 79th embodiments, the pointers are stored in the same or reverse order as the bits set in the bitmap.
[0116] According to the 81st embodiment, in any one of the 1st to 33rd, or 49th, or 72nd to 80th embodiments, the pointers included by a parent node point to the bitmap included by the child node.
[0117] According to the 82nd embodiment, in any one of the 1st to 33rd, or 49th, or 72nd to 81st embodiments, the number of pointers included by the leaf node, preferably each leaf node of the trie, is zero.
[0118] Key part
[0119] According to the 83rd embodiment, in any one of the 3rd or 33rd to 82nd embodiments, the value of the key part of a child node, preferably the value of the key part of at least each child node whose parent node has more than one child node, is determined by the value of the bit(s) set in the bitmap included by the parent node, depending on which bit the child node is associated with.
[0120] According to the 84th embodiment, in any one of the 4th to 18th or the 83rd embodiment, the maximum amount of different values available for the key part is defined by the size of the bitmap.
[0121] According to the 85th embodiment, in any one of the 4th to 18th, or the 83rd or 84th embodiments, the size of the bitmap defines the possible alphabet of the key part.
[0122] According to the 86th embodiment, in any one of the 3rd or 32nd to 85th embodiments, each key part in the trie can store a value of the same predefined size.
[0123] According to the 87th embodiment, in the 86th embodiment, the predefined size corresponds to a 5-bit, 6-bit, 7-bit, or 8-bit value.
[0124] Key encoding for range queries
[0125] According to the 88th embodiment, in any of the foregoing embodiments, the encoding of the value of a data item, preferably the encoding of the values of all data items, is obtained by converting the data type of the data item into an offset binary representation containing an unsigned integer.
[0126] According to the 89th embodiment, in the 88th embodiment, the integer is a long integer.
[0127] According to the 90th embodiment, in any of the 88th or 89th embodiments, if the data type of the data item is a floating-point number, the encoding is obtained by converting the data type of the data item into an offset binary representation.
[0128] According to the 91st embodiment, in any of the 88th to 90th embodiments, if the data type of the data item is a two's complement signed integer, the encoding is obtained by converting the data type of the data item into an offset binary representation.
[0129] Others
[0130] According to the 92nd embodiment, in any of the foregoing embodiments, the trie stores a dynamic set or an associative array.
[0131] Boolean operations on the trie
[0132] The 93rd embodiment of the present invention is a method for retrieving data from an electronic database or an information retrieval system, including the following steps: obtaining representations of two or more input tries; evaluating at least part of the resulting trie, where the resulting trie is a combination of the input tries using logical operations; and providing as output: a representation of the resulting trie, a set or subset of keys and / or other data items associated with the nodes of the resulting trie, or a set of keys or values derived from the keys associated with the nodes of the resulting trie.
[0133] According to the 94th embodiment, in the 93rd embodiment, the keys and / or other data items associated with the leaf nodes of the resulting trie are provided as output.
[0134] According to the 95th embodiment, in the 93rd or 94th embodiment, the set of keys provided as output is provided in the trie.
[0135] According to the 96th embodiment, in any of the 93rd or 94th embodiments, the set of keys provided as output is provided by a cursor or an iterator.
[0136] According to the 97th embodiment, in any one of the 93rd to 96th embodiments, where a part of the trie is a node and / or a branch of the trie.
[0137] According to the 98th embodiment, in any one of the 93rd to 97th embodiments, where the representation of the trie is a physical trie data structure or a handle to a trie operator capable of evaluating the trie.
[0138] According to the 99th embodiment, in any one of the 93rd to 98th embodiments, where the representation of the trie represents the root node of the trie and implements the trie node interface, where the trie node interface represents a trie node and provides methods or functions for querying and / or traversing the child nodes of the trie node it represents.
[0139] According to the 100th embodiment, in any one of the 93rd to 99th embodiments, where evaluating the trie or a part of the trie includes traversing the trie or a part of the trie.
[0140] According to the 101st embodiment, in the 100th embodiment, where evaluating the trie or a part of the trie includes providing the following information: the traversed node exists in the trie.
[0141] According to the 102nd embodiment, in any one of the 100th or 101st embodiments, where at any given moment during the evaluation of the trie or a part of the trie, at least those parts of the physical data structure implementing the trie or the part of the trie, which need to be implemented at that moment in order to traverse the trie or the part of the trie.
[0142] According to the 103rd embodiment, in any one of the 100th to 102nd embodiments, where at any given moment, less than the entire physical data structure of the trie or the part of the trie is implemented, and preferably only those parts of the physical data structure of the trie or the part of the trie, which need to be implemented at that moment in order to traverse the trie or the part of the trie, are implemented.
[0143] According to the 104th embodiment, in any one of the 93rd to 102nd embodiments, where evaluating the trie or a part of the trie includes implementing the physical data structure of the trie or the part of the trie as a whole.
[0144] According to the 105th embodiment, in any one of the 93rd to 104th embodiments, where at least one input trie is a virtual trie, which is at least partially evaluated during the step of evaluating at least a part of the result trie.
[0145] According to the 106th embodiment, in the 105th embodiment, at least at the moment of obtaining the representation of the virtual trie, there is no physical data structure of the virtual trie as a whole.
[0146] According to the 107th embodiment, in any one of the 105th or 106th embodiments, the physical data structure of the virtual trie or a part of the virtual trie as a whole is not implemented.
[0147] According to the 108th embodiment, in any one of the 105th to 107th embodiments, at least those parts of the virtual trie are evaluated, which are required for at least a part of the evaluation result trie.
[0148] According to the 109th embodiment, in any one of the 105th to 108th embodiments, at least some parts of the virtual trie that are not required for at least a part of the evaluation result trie are not evaluated, and preferably, only those parts of the virtual trie that are required for at least a part of the evaluation result trie are evaluated.
[0149] According to the 110th embodiment, in any one of the 105th to 109th embodiments, at least a part of the virtual trie part is evaluated after at least a part of the result trie has been evaluated.
[0150] According to the 111th embodiment, in any one of the 105th to 110th embodiments, the parts of the result trie and the virtual trie are evaluated in an interleaved manner.
[0151] According to the 112th embodiment, in any one of the 105th to 111th embodiments, two or more input tries are high-order input tries, and the result trie is a high-order result trie; and a representation of at least one virtual trie is provided as the output of a low-order combinatorial operator that performs the following steps: obtaining the representations of two or more low-order input tries; at least partially evaluating a low-order result trie that is a combination of the low-order input tries using low-order logical operations; and providing a representation of the low-order result trie as the output.
[0152] According to the 113th embodiment, in the 112th embodiment, the step of providing a representation of the low-order result trie as the output is performed before the step of at least partially evaluating the low-order result trie is completed.
[0153] According to the 114th embodiment, in any one of the 112th or 113th embodiments, the low-order logical operation is the same logical operation as the high-order logical operation.
[0154] According to the 115th embodiment, in any one of Embodiments 112 or 113, where the low-order logical operation is a logical operation different from the high-order logical operation.
[0155] According to the 116th embodiment, in any one of Embodiments 112 to 115, where at least one low-order input trie is a virtual special case as defined in any one of Embodiments 105 to 111, with necessary adaptations as necessary, which is represented as being provided as the output of an even lower-order combinatorial operator.
[0156] According to the 117th embodiment, in any one of Embodiments 93 to 116, where the input trie stores a set of keys stored in an electronic database or an information retrieval system or a resulting set of keys derived from one or more sets of keys stored in an electronic database or an information retrieval system.
[0157] According to the 118th embodiment, in any one of Embodiments 93 to 117, where the trie contains one or more nodes, each child node is associated with a key part, and the path from the root node to another node in the trie defines the key associated with that node, which is the concatenation of the key parts associated with the nodes on the path.
[0158] According to the 119th embodiment, in any one of Embodiments 93 to 118, where the combination of two or more input tries using a logical operation is the following trie, which stores the set of keys obtained when combining the corresponding sets of keys stored in the corresponding input tries using the logical operation.
[0159] According to the 120th embodiment, in any one of Embodiments 93 to 119, if the logical operations are different, the parent node of the resulting trie is the parent node of the first input trie, and the leaf nodes of the parent node of the resulting trie are the combination of the set of child nodes of the corresponding parent node in the first input trie and the set of child nodes of any corresponding parent node in the other input tries using the logical operation, and if the logical operations are not different, the set of child nodes of each node in the resulting trie is the combination of the set of child nodes of the corresponding node in the input tries using the logical operation; and if the keys associated with the nodes of different tries are the same, two or more nodes of different tries correspond to each other.
[0160] According to the 121st embodiment, in any one of Embodiments 93 to 120, where the set of child nodes of the nodes of the resulting trie is determined by combining the sets of child nodes of the nodes of the input tries corresponding to the nodes of the resulting trie using a logical operation.
[0161] Merge step
[0162] According to the 122nd embodiment, in any one of the 93rd to 122nd embodiments, wherein the step of evaluating a result trie or a part of the result trie includes: performing a combination function on the root node of the result trie; wherein performing the combination function on the input node of the result trie includes: determining a subset of child nodes of the input node of the result trie by using the logical operation to combine subsets of child nodes of nodes of an input trie corresponding to the input node of the result trie; and performing the combination function on each child node determined for the input node of the result trie.
[0163] According to the 123rd embodiment, in any one of the 93rd to 122nd embodiments, the step of evaluating a result trie or a part of the result trie is performed using depth - first traversal, breadth - first traversal, or a combination thereof.
[0164] According to the 124th embodiment, in the 123rd embodiment, the step of evaluating a result trie or a part of the result trie by depth - first traversal includes: performing the combination function on one of the child nodes of the input node and traversing a sub - trie formed by the child node before performing the combination function on the next sibling node of the child node.
[0165] According to the 125th embodiment, in any one of the 123rd or 124th embodiments, the step of evaluating a result trie or a part of the result trie by breadth - first traversal includes: performing the combination function on each child node determined for the input node of the result trie, and determining a subset of child nodes of each child node determined for the input node of the result trie before performing the combination function on any grand - child nodes of the input node of the result trie.
[0166] bitmap
[0167] According to the 126th embodiment, in any one of the 93rd to 125th embodiments, nodes in the input trie, preferably at least all parent nodes in the input trie, include a bitmap.
[0168] According to the 127th embodiment, in the 126th embodiment, the value of the key part of a child node in the trie is determined by the value of a bit set in the bitmap included by the parent node of the child node, depending on which bit the child node is associated with.
[0169] According to the 128th embodiment, in any one of the 126th or 127th embodiments, wherein the subsets of the child nodes of the nodes of the input trie corresponding to the nodes of the result trie are combined using a logical operation, preferably including: combining the bitmaps of each node of the input trie corresponding to the nodes of the result trie bitwise using a logical operation, thereby obtaining a combined bitmap.
[0170] According to the 129th embodiment, in the 128th embodiment, wherein the child nodes of the nodes of the result trie are determined and / or evaluated based on the combined bitmap, preferably only based on the combined bitmap.
[0171] According to the 130th embodiment, in any one of the 128th or 129th embodiments, wherein the step of combining the bitmaps is performed before determining and / or evaluating the child nodes of the nodes of the result trie and / or before traversing the child nodes of the nodes of the result trie for the evaluation of the result trie.
[0172] Logical operation
[0173] In the 131st embodiment, in any one of the 93rd to 130th embodiments, the logical operation is an intersection, union, difference, or exclusive OR.
[0174] According to the 132nd embodiment, in the 131st embodiment, using the logical operation includes combining using an AND Boolean operation, an OR Boolean operation, an AND NOT Boolean operation, or an XOR Boolean operation.
[0175] According to the 133rd embodiment, in any one of the 93rd to 132nd embodiments and the 105th embodiment, using the logical operation includes combining the bitmaps of the nodes using a bitwise AND Boolean operation, a bitwise OR Boolean operation, a bitwise AND NOT Boolean operation, or a bitwise XOR Boolean operation.
[0176] Range query
[0177] The 134th embodiment is a method for retrieving data from an electronic database or an information retrieval system by performing a range query on a set of keys stored in the electronic database or the information retrieval system or a result key set derived from one or more sets of keys stored in the electronic database or the information system, the method including the steps of: obtaining a definition of one or more ranges; and performing the method for retrieving data from the electronic database or the information retrieval system according to any one of the 93rd to 133rd embodiments, wherein one input trie is an input set trie that stores the set of keys or the result key set to search the one or more ranges; another input trie is a range trie that stores all values included in the one or more ranges for which the definition has been obtained; and the logical operation is an intersection.
[0178] According to the 135th embodiment, in the 134th embodiment, a range is a set of discrete ordered values, including all values between a first value and a second value of a certain data type.
[0179] According to the 136th embodiment, in the 135th embodiment, the range includes the first value and / or the second value.
[0180] According to the 137th embodiment, in any one of the 134th to 136th embodiments, in the steps of performing the method of any one of Embodiments 93 to 133, the method of any one of Embodiments 93 to 111 is performed, wherein at least one of the input tries of the virtual trie is a range trie.
[0181] Single-item input set trie
[0182] According to the 138th embodiment, in the 137th embodiment, the definition of the one or more ranges includes the definition of one or more ranges for the one data item.
[0183] Multi-item input set trie
[0184] According to the 139th embodiment, in any one of the 134th to 137th embodiments, the key associated with the leaf node of the input set trie encodes two or more data items of a specific data type.
[0185] According to the 140th embodiment, in the 139th embodiment, the definition of the one or more ranges includes the definition of one or more ranges for one or more of the data items.
[0186] Obtain a multi-item range trie from a single-item range trie
[0187] According to the 141st embodiment, in any one of the 139th or 140th embodiments, the range trie is a multi-item range trie, which is obtained by combining single-item range tries for each data item encoded by the key associated with the leaf node of the input set trie, wherein the single-item range trie for a data item stores all values included in one or more ranges of the data item.
[0188] According to the 142nd embodiment, in the 141st embodiment, the combination of the single-item range tries is performed within the function that implements the combination of the input set trie and the multi-item range trie.
[0189] According to the 143rd embodiment, in the 141st embodiment, the combination of the single-range tries is performed by the following function, which provides the multi-range trie as input to a function that implements the combination of the input set trie and the multi-range trie.
[0190] According to the 144th embodiment, in any one of the 141st to 143rd embodiments, the single-range trie is the following virtual trie, which is at least partially evaluated during the operation of combining the single-range tries, preferably a virtual trie as defined in any one of Embodiments 105 to 111 (with necessary modifications).
[0191] According to the 145th embodiment, in any one of the 141st to 144th embodiments, for each data item for which no range definition is obtained, the single-range trie stores the entire range of possible values of the data item.
[0192] According to the 146th embodiment, in any one of the 141st to 144th embodiments, the multi-range trie stores all combinations of the values of the data items stored in the single-range tries.
[0193] Structure of the range trie
[0194] According to the 147th embodiment, in any one of the 141st to 146th embodiments, the range trie has the same structure or format as the input set trie.
[0195] According to the 148th embodiment, in the 147th embodiment, the key pairs associated with the leaf nodes of the range trie encode data items of the same data type as the keys associated with the leaf nodes of the input set trie.
[0196] According to the 149th embodiment, in any one of the 147th or 148th embodiments, in the range trie, data items of a specific data type or components of such data items are encoded in nodes of the same level as the corresponding data items or components of the data items in the input set trie.
[0197] Output of the range query
[0198] According to the 150th embodiment, in any one of the 134th to 149th embodiments, the method provides the set of keys associated with the leaf nodes of the input set trie as output.
[0199] According to the 151st embodiment, in any one of the 139th to 149th embodiments, the method provides a simplified set of item keys as output, which encodes a subset of the data items encoded by the keys associated with the leaf nodes of the input set trie.
[0200] According to the 152nd embodiment, in the 151st embodiment, before providing the output, a set of simplified item keys obtained from different branches of the input set trie related to data items not encoded in the simplified item key and being the result of the combination of the input set trie and the range trie is merged.
[0201] According to the 153rd embodiment, in either the 151st or the 152nd embodiment, before providing the output, a set of simplified item keys obtained as a result of the operation of combining the input set trie and the range trie is written into a newly created trie, thereby eliminating duplicate keys.
[0202] Fuzzy search
[0203] The 154th embodiment is a method for retrieving data from an electronic database or an information retrieval system by performing approximate string matching, the method comprising the steps of: obtaining a search string; constructing a matching trie that stores a set of approximate strings including the search string and / or variants of the search string; evaluating at least a portion of a result trie that is a combination of the matching trie and a storage trie using an intersection operation, the storage trie storing a set of strings stored in the electronic database or the information retrieval system; providing strings and / or other data items associated with a result set of nodes of the result trie as output; wherein a trie includes one or more nodes, each child node is associated with a key portion, and a path from a root node in the trie to another node defines a key associated with the node, the key being a concatenation of key portions associated with nodes on the path.
[0204] According to the 155th embodiment, in the 154th embodiment, strings included in the set of strings stored in the matching trie are permanently stored in the electronic database or the information retrieval system.
[0205] According to the 156th embodiment, in any one of the 154th or 155th embodiments, strings included in the set of strings stored in the matching trie represent entries in the electronic database or the information retrieval system.
[0206] According to the 157th embodiment, in any one of the 154th to 156th embodiments, one or more child nodes in the matching trie have more than one parent node.
[0207] According to the 158th embodiment, in any one of the 154th or 157th embodiments, one or more child nodes in the matching trie have only one parent node.
[0208] According to the 159th embodiment, in any one of the 154th to 158th embodiments, the result trie is the following trie that stores the key set obtained when performing an intersection operation on the key set stored in the storage trie and the key set stored in the matching trie.
[0209] According to the 160th embodiment, in any one of the 154th to 158th embodiments, the set of child nodes of each node in the result trie is the intersection of the sets of child nodes of the corresponding nodes in the matching trie and the storage trie, where if the same key is associated with nodes in different tries, the nodes in the different tries correspond to each other.
[0210] According to the 160th embodiment, in any one of the 154th to 160th embodiments, the matching trie is a virtual trie that is at least partially evaluated or dynamically generated during the step of evaluating at least part of the result trie.
[0211] According to the 161st embodiment, in the 160th embodiment, a physical data structure of the virtual trie or a part of the virtual trie as a whole is not implemented.
[0212] According to the 162nd embodiment, in any one of the 160th or 161st embodiments, at least those parts of the virtual trie that are required for evaluating at least part of the result trie are evaluated.
[0213] According to the 163rd embodiment, in any one of the 160th to 162nd embodiments, at least some parts of the virtual trie that are not required for evaluating at least part of the result trie are not evaluated, and preferably only those parts of the virtual trie that are required for evaluating at least part of the result trie are evaluated.
[0214] According to the 164th embodiment, in any one of the 160th to 163rd embodiments, at least part of the virtual trie is evaluated after at least part of the result trie has been evaluated.
[0215] According to the 165th embodiment, in any one of the 160th to 164th embodiments, the parts of the result trie and the parts of the virtual trie are evaluated in an interleaved manner.
[0216] According to the 166th embodiment, in any one of the 154th to 165th embodiments, evaluating a trie or a part of a trie includes traversing the trie or the part of the trie.
[0217] According to the 167th embodiment, in any one of the 154th to 166th embodiments, wherein evaluating the trie or a portion of the trie includes providing the following information: the traversed node exists in the trie.
[0218] According to the 168th embodiment, in any one of the 154th to 167th embodiments, wherein at any given moment during the evaluation of the trie or a portion of the trie, at least those portions of the physical data structure implementing the trie or the portion of the trie that are required to be implemented at that moment for traversing the trie or the portion of the trie.
[0219] According to the 169th embodiment, in any one of the 154th to 168th embodiments, wherein at any given moment, less than the entire physical data structure of the trie or the portion of the trie is implemented, and preferably, only those portions of the physical data structure of the trie or the portion of the trie that are required to be implemented at that moment for traversing the trie or the portion of the trie are implemented.
[0220] According to the 170th embodiment, in any one of the 154th to 169th embodiments, wherein evaluating the trie or a portion of the trie includes: implementing the physical data structure of the trie or the portion of the trie as a whole.
[0221] According to the 171st embodiment, in any one of the 154th to 170th embodiments, the data item provided in the output represents a data unit that contains a string associated with a node of the result set of the nodes of the result trie, preferably a document identifier.
[0222] According to the 172nd embodiment, in any one of the 154th to 171st embodiments, the storage trie is an index trie or a physical index trie, and preferably, the strings included by the document and the corresponding document identifier are stored as, for example, two key parts (string, long).
[0223] According to the 173rd embodiment, in any one of the 154th to 172nd embodiments, the matching trie includes a set of matching nodes, each matching node is associated with one or more keys corresponding to one of the strings from the approximate string set, and the result set of the nodes is a set of nodes of the result trie corresponding to the set of matching nodes in the matching trie, wherein if the key associated with the node of the result trie is the same as the key associated with the node of the matching trie, the node of the result trie corresponds to the node of the matching trie.
[0224] According to the 174th embodiment, in any one of the 154th to 173rd embodiments, it further includes the step of obtaining a number N, wherein the variants of the search string include a set of strings that can be obtained by at most N single-character insertions, deletions, and / or replacements of the search string.
[0225] According to the 175th embodiment, in any one of the 154th to 174th embodiments, the step of building the matching trie includes: building a finite automaton representing the set of approximate strings; and deriving the matching trie from the finite automaton.
[0226] According to the 176th embodiment, in the foregoing embodiments, the transitions between two states of the finite automaton, preferably each transition is associated with a specific character, preferably a character included in the search string, or a wildcard or an empty string.
[0227] According to the 177th embodiment, in any one of the 175th or 176th embodiments, the step of building the finite automaton includes: building a non-deterministic finite automaton representing the set of approximate strings; and deriving a deterministic finite automaton from the non-deterministic finite automaton; and wherein the matching trie is derived from the deterministic finite automaton.
[0228] According to the 178th embodiment, in the foregoing embodiments, the transitions between two states of the deterministic finite automaton, preferably each transition is associated with a specific character, preferably a character included in the search string or a wildcard.
[0229] According to the 179th embodiment, in any one of the 154th to 178th embodiments, the nodes in the matching trie and the storage trie, preferably at least all the parent nodes, include bitmaps, and the value of the key part of the child nodes in the trie is determined by the value of the bits (set) in the bitmap included in the parent node of the child node, depending on which bit the child node is associated with.
[0230] According to the 180th embodiment, in the foregoing embodiments, wherein determining the child nodes of the nodes of the result trie includes: intersecting the set of child nodes of the nodes of the matching trie and the set of child nodes of the nodes of the storage trie corresponding to the nodes of the result trie.
[0231] According to the 181st embodiment, in the foregoing embodiments, wherein intersecting the set of child nodes of the nodes of the matching trie with the set of child nodes of the nodes of the storage trie includes: combining the bitmaps of the nodes of the matching trie and the nodes of the storage trie using a bitwise AND operation to obtain a combined bitmap.
[0232] According to the 182nd embodiment, in the foregoing embodiments, the children of the result trie node are determined and / or evaluated based on the combined bitmap, preferably based only on the combined bitmap.
[0233] According to the 183rd embodiment, in any one of the 181st or 182nd embodiments, the step of combining the bitmaps is performed before determining and / or evaluating the children of the nodes of the result trie and / or before the evaluation of the result trie traverses to the children of the nodes of the result trie.
[0234] According to the 184th embodiment, in any one of the 171st and 179th to 183rd embodiments, the step of deriving the matching trie from the finite automaton includes: obtaining an enhanced finite automaton by associating the transitions between two states of the finite automaton, preferably each transition via the encoding of a specific character or wildcard associated with the transition, where the encoding consists of or is represented by one or more bitmaps, the length and / or format of which is equal to that of the bitmap included in the parent node of the matching trie, and where the matching trie is derived from the enhanced finite automaton.
[0235] According to the 185th embodiment, in the foregoing embodiments, for the encoding of a specific character, exactly one bit is set in each of the bitmaps included or represented by the encoding.
[0236] According to the 186th embodiment, in any one of the 184th or 185th embodiments, for the encoding of a wildcard, the bits of all valid character encodings are set in the bitmap included or represented by the encoding, or all the bits of the valid character encodings except for the encoding of the specific character associated with the state from which the transition departs.
[0237] According to the 187th embodiment, in any one of the 179th to 186th embodiments, the characters stored in the matching trie, the storage trie, or the result trie are encoded by the key part of the corresponding trie with a quantity of M>1, preferably 5>M.
[0238] According to the 188th embodiment, in the foregoing embodiment, the step of deriving the matching trie from the finite automaton includes: replacing a transition between two states of the finite automaton by using one or more sequences of M - 1 levels of intermediate states and M transitions linking the two states via the M - 1 intermediate states, preferably each transition, or associating a transition between two states of the finite automaton, preferably each transition, with one or more sequences of M - 1 levels of intermediate states and M transitions linking the two states via the M - 1 intermediate states to obtain a complete finite automaton representing the approximate string set, wherein each of the M transitions in the sequence is associated with an intermediate encoding, the intermediate encoding being composed of or represented by a bitmap, the length and / or format of which is equal to the bitmap included in the parent node of the matching trie, and wherein the matching trie is derived from the complete finite automaton.
[0239] According to the 189th embodiment, in the 188th embodiment, if a transition between two states of the finite automaton is associated with a specific character, the concatenation of the bitmaps included in or represented by the intermediate encodings associated with the M transitions of the sequence is the encoding of the specific character, and exactly one bit is set in each of the bitmaps.
[0240] According to the 190th embodiment, in any one of the 188th or 189th embodiments, if a transition between two states of the finite automaton is associated with a wildcard, the concatenation of the bitmaps included in or represented by the intermediate encodings associated with the M transitions of the sequence includes: the following encoding, wherein the bits of all valid character encodings are set in the bitmap included in or represented by the encoding, or all the bits of the valid character encodings except for the bit of the encoding of the specific character associated with the state from which the transition departs, and / or one or more encodings including one or more parts of the encoding of the specific character and one or more parts of the following encoding, wherein the bits of all valid character encodings are set in the bitmap included in or represented by the encoding, or all the bits of the valid character encodings except for the bit of the encoding of the specific character associated with the state from which the transition departs.
[0241] According to the 191st embodiment, in any one of the 184th to 186th, or in any one of the 188th to 190th embodiments, the enhanced finite automaton or the complete finite automaton is respectively represented by a data structure including a plurality of rows or stored therein, each row representing a state of the enhanced finite automaton or the complete finite automaton and including a tuple for each transition departing from the state, each tuple including the encoding associated with the transition and a reference to the state at the end of the transition.
[0242] According to the 192nd embodiment, in the foregoing embodiments, for each state at the end of the conversion, the data structure includes information as to whether the state is a matching state, preferably encoded as a bit in each reference to the state.
[0243] According to the 193rd embodiment, in any one of the 191st or 192nd embodiments, the data structure includes a row for each of the states of the augmented finite automaton or the complete finite automaton from which the transition departs.
[0244] Data structure
[0245] According to the 194th embodiment, in any one of the 93rd to 193rd embodiments, the trie is a trie according to any one of the 1st to 92nd embodiments.
[0246] Inventions of different categories
[0247] The 195th embodiment is a computer-implemented method of using a trie according to any one of the 1st to 92nd embodiments in an electronic database application or an information retrieval system, in particular for storing keys or keys and values, for storing result keys or query keys and values, or for storing input keys or query keys and values.
[0248] The 196th embodiment is a computer-implemented method of generating a trie according to any one of the 1st to 92nd embodiments.
[0249] The 197th embodiment is a non-transitory computer-readable medium having stored thereon a trie according to any one of the 1st to 92nd embodiments.
[0250] The 198th embodiment is an electronic data stream representing a trie according to any one of the 1st to 92nd embodiments.
[0251] The 199th embodiment is an electronic database or an information retrieval system that stores keys or keys and values, result keys or query keys and values, or input keys or query keys and values by means of a trie according to any one of the 1st to 92nd embodiments.
[0252] The 200th embodiment of the present invention is a computer program, in particular a database application program or an information retrieval system program, including instructions for performing the method according to any one of the 93rd to 196th embodiments.
[0253] The 201st embodiment of the present invention is a data processing device or system including one or more processors and a memory, the data processing device or system being configured to perform the method according to any one of the 93rd to 196th embodiments.
[0254] The 202nd embodiment of the present invention is preferably a non - transitory computer - readable medium having stored thereon the computer program of the 201st embodiment. BRIEF DESCRIPTION OF THE DRAWINGS
[0255] Hereinafter, the present invention will be described in more detail by combining preferred embodiments and referring to the accompanying drawings, where
[0256] Figure 1 an example of a trie data structure used in the prior - art database is shown;
[0257] Figure 2 another example of a trie data structure used in the prior art is shown;
[0258] Figure 3 shows Figure 2 a prior - art implementation of a trie data structure of
[0259] Figure 4 shows Figure 2 another prior - art implementation of a trie data structure of
[0260] Figure 5 shows Figure 2 another prior - art implementation of a trie data structure of
[0261] Figure 6 shows how to determine the child pointers of a node in a prior - art trie;
[0262] FIG. 7A to FIG. 7C shows how leaf nodes according to the present invention are stored in a map and how a "last" bitmap is stored in a set in a trie;
[0263] Figure 8 provides an example of a trie where the first space optimization according to the present invention is effective;
[0264] Fig. 9 shows the first space optimization according to the present invention;
[0265] Fig.10 provides an example of a trie where the second space optimization according to the present invention is effective;
[0266] Fig.11 shows the general storage configuration of the second space optimization according to the present invention;
[0267] Fig.12 shows the second space optimization according to the present invention;
[0268] Fig.13Shows how to access the nodes of a spatially optimized trie according to the present invention in a unified manner;
[0269] FIG. 14A to FIG. 14C Shows inserting a key into a spatially optimized trie according to the present invention;
[0270] Fig.15 Shows the results of an experiment conducted to measure the efficiency of the first and second spatial optimizations according to the present invention;
[0271] Fig.16 Shows the third spatial optimization according to the present invention;
[0272] Fig.17 Shows the results of an experiment conducted to measure the efficiency of the third spatial optimization according to the present invention;
[0273] Fig.18 Shows the first way of arranging control and content information in a key to be stored in a trie according to the present invention;
[0274] Fig.19 Shows the second way of arranging control and content information in a key to be stored in a trie according to the present invention;
[0275] Fig. 20 Shows an example of key encoding according to the present invention;
[0276] Fig.21 Shows a trie storing keys encoded according to the present invention;
[0277] Fig. 22 Shows an example of a two-dimensional key with values (X = 12, Y = 45);
[0278] Fig.23 Gives an example of how to store the key shown in Fig.21 in a trie according to the present invention;
[0279] Fig.24 Shows the stages of database query processing in the prior art;
[0280] Fig.25 Shows an exemplary query execution plan used in a prior art database;
[0281] Fig.26 Shows operator control and data flow according to the query execution model in a prior art database;
[0282] Fig. 27 Shows tuple-at-a-time processing according to the query execution model of a prior art database;
[0283] Fig.28 Shows the operator - at - a - time processing of the query execution model according to the prior art database;
[0284] Fig.29 Shows the trie control and data flow of the query execution model according to the present invention;
[0285] Fig.29A Shows another trie control and data flow of the query execution model according to the present invention;
[0286] Fig.30 Shows an example of applying the intersection operator to two tries according to the present invention;
[0287] Fig.31 Shows two example tries on which the intersection operation according to the present invention will be performed;
[0288] Fig.32 Shows the bit - wise AND operation performed on the bitmaps of the two example tries at levels 1 and 2 Fig.31 ;
[0289] Fig.33 Shows the branch jumps during the intersection operation of the two tries Fig.31 ;
[0290] Fig.34 Shows the result trie of the intersection operation of the two tries Fig.31 ;
[0291] Fig.35 Shows an example of applying the union operator to two tries according to the present invention;
[0292] Fig.36 Shows the result trie of the union operation on the two tries Fig.35 ;
[0293] Fig.37 Shows an example of applying the difference operator to two tries according to the present invention;
[0294] Fig.37A Shows three tries combined according to the data flow shown Fig.29A ;
[0295] Fig.37B Shows the result of the combination of the tries shown Fig.37A ;
[0296] Fig.38Illustrates the interaction of the intersection operator, input set trie, and range trie in range queries according to the present invention;
[0297] Fig.39 Illustrates an example of a method for performing a range query according to the present invention;
[0298] Fig.40 Illustrates an example of a method for performing a two-dimensional range query according to the present invention;
[0299] Fig.41 Illustrates how to combine in an interleaved manner Fig.40 one-dimensional range tries;
[0300] Fig.42 Illustrates due to Fig.41 the obtained interleaved range trie;
[0301] Fig.43 Illustrates how to combine in a non - interleaved manner Fig.40 one-dimensional range tries;
[0302] Fig.44 Illustrates due to Fig.43 the obtained non - interleaved range trie;
[0303] Fig.45 Illustrates which nodes of the input set trie are accessed during a range query if the input set trie is stored in a non - interleaved manner Fig.40 ;
[0304] Fig.46 Illustrates which nodes of the input set trie are accessed during a range query if the input set trie is stored in an interleaved manner Fig.40 ;
[0305] Fig.47 Illustrates a two - dimensional range query with one - dimensional output according to the present invention;
[0306] Fig.48A Illustrates a non - deterministic finite automaton to match the search string "abc" with a maximum edit distance of 2;
[0307] Fig.48B Illustrates a deterministic finite automaton for matching "abc" when the edit distance is 1;
[0308] Fig.48C Illustrates Fig.48B the enhancement of the transitions of the automaton, where the encoding scheme of Unicode characters and Unicode strings described with reference to Fig.21 is used;
[0309] Fig.48D Shows the top of the results of a matching trie, where each 10-bit Unicode character is represented by two key parts;
[0310] Fig.48E Illustrates a data structure of an array of arrays, which can be used to represent the states of a complete finite automaton, from which a Fig.48D matching trie can be derived;
[0311] Fig.49 Illustrates the specification of an experimental geographical data search query for all locations within a small rectangle in the Munich area;
[0312] Fig.50 Illustrates the remaining records loaded into the database to execute the first series of experimental geographical data search queries;
[0313] Fig.51 Illustrates the remaining records loaded into the database to execute the second series of experimental geographical data search queries;
[0314] Fig.52 Illustrates the temporary result set and the final result set determined by the experimental geographical data search query;
[0315] Fig.53 Illustrates the performance measurement results of the prior art method;
[0316] Fig.54 Illustrates an abstract view of a non-interleaved two-dimensional index trie used in the standard indexing method according to the present invention;
[0317] Fig.55 Illustrates the use of a multi-OR operator in the standard indexing method according to the present invention;
[0318] Fig.56 Provides an overview of the components used in the standard indexing method according to the present invention;
[0319] Fig.57 Illustrates the performance measurement results of the standard indexing method according to the present invention;
[0320] Fig.58 Illustrates an example of a first-level index for a variable-precision indexing method according to the present invention;
[0321] Fig.59 Illustrates an example of a second-level index for a variable-precision indexing method according to the present invention;
[0322] Fig.60 Illustrates the performance measurement results of the variable-precision indexing according to the present invention;
[0323] Fig.61Shows how parts of binomial keys stored in an interleaved manner according to the present invention are combined with each other or with parts of a range trie in a two-dimensional indexing method;
[0324] Fig.62 Shows the performance measurement results of the two-dimensional indexing method according to the present invention;
[0325] Fig.63 Shows the structure of an index storing three item keys for an indexing method for single indexing according to the present invention;
[0326] Fig.64 Provides an overview of the components used in the indexing method for single indexing according to the present invention;
[0327] Fig.65 Shows the performance measurement results of the indexing method for single indexing according to the present invention;
[0328] Fig.66 Shows the experimental results where query performance was measured with an increasing result size;
[0329] Fig.67 The results are shown in a graph with a logarithmic x-axis Fig.66 of;
[0330] Fig.68 Shows the experimental results where index performance was measured on an increasing index size;
[0331] Fig.69 Shows the space requirements for different indexing methods;
[0332] Figures 70A-70F In an information retrieval application, the index performance, index size, and query performance of a database using a prior art index are compared with a database using an index according to the present invention. Detailed Description
[0333] Detailed description of the preferred embodiments of the present invention
[0334] Figure 1Shows an example of a trie data structure used in a database according to the prior art. It shows a trie data structure 101, where each child node (i.e., all nodes except the root node) is associated with a key part, and the value of the key part is indicated by a pointer from the parent node and selected from the alphabet {0…9}, i.e., the nodes at each level except the root node are associated with a decimal digit. The path from the root node in the trie to another node defines the key associated with that node, which is the concatenation of the key parts associated with the nodes on the path. The trie 101 “stores” keys with values “007” and “042” because the leaf node 107 of the trie 101 is associated with a key with the value “007” and the leaf node 108 of the trie 101 is associated with a key with the value “042”.
[0335] The root node 102 located on the first level 110 has a child node 104 associated with a key part with the value “0”. Thus, there is a pointer 103 from the root node 102 to the child node 104 located on the second level 111 of the trie 101, which indicates the value “0”. From the child node 104 on the second level, two different pointers point to the nodes 105, 106 located on the third level 112, and from each of these nodes 105, 106, an additional pointer points to the leaf nodes 107, 108 respectively. Thus, the concatenation of the key parts of the nodes on the path from the root node to the leaf nodes results in keys with values “007” and “042”.
[0336] Figure 2 Shows another example of a trie data structure used in the prior art. The trie data structure of this example is similar to the Figure 1 trie data structure, but has a larger alphabet of possible values for the key parts associated with the nodes. In particular, Figure 2 shows a 256 - ary trie data structure, where the representation of the key part value has a size of 8 bits (or 1 byte).
[0337] Figure 2The trie data structure stores two different keys ("0000FD" and "002A02" with two different keys associated with their leaf nodes) and has four levels 210 to 214. The root node 202 located on the first level 210 has a pointer to the child node 204 associated with the key partial value "00". The child node 204 includes two pointers to child nodes 205, 206, which are associated with the key partial values "00" and "2A" respectively. Each of the child nodes 205, 206 located in the third level 212 includes one pointer to a leaf node 207, 208 respectively. The leaf nodes 207, 208 are located in the fourth level 213. The leaf node 207 is associated with the key partial value "FD", and the leaf node 208 is associated with the key partial value "02".
[0338] Figure 3 illustrates Figure 2 A prior art implementation of the trie data structure. The nodes with associated pointers are fully allocated in memory. Thus, in this case, even a null pointer occupies memory space. For example, the root node 302 allocates 256 pointers in its array 306, and only one pointer 303 is non - null, which is associated with the key partial value "00" and points to the child node 304. The child nodes are implemented in the same way.
[0339] Figure 4 depicts Figure 2 An exemplary implementation of the trie data structure, which provides a known solution for more efficiently using memory space and avoiding storing null pointers. The solution includes storing only a list of non - null pointers and their corresponding key partial values instead of an array containing all possible pointers. For example, the root node 402 with one child node associated with the key partial value "00" includes a list with one entry. This list entry includes the key partial value 403 and the associated pointer 404 to the child node 405. The other child nodes are implemented accordingly.
[0340] Figure 5 illustrates Figure 2 Another exemplary implementation of the trie data structure, which provides a known solution to the allocation problem based on a bitmap for marking all non - null pointers of the parent nodes. In other words, the set bits in the bitmap mark the valid (non - null) branches. Each parent node also includes one or more pointers, where each pointer is associated with a set bit in the bitmap and points to a child node of the parent node.
[0341] To determine the pointer of a child node, the number of previous child pointers must be calculated. As Figure 6As shown in [the figure], the offset for finding the pointer is the number of least significant bits set in the bitmap before the target position. This compact trie node structure eliminates the need to store nil pointers and at the same time allows for very fast access.
[0342] Figure 5 For the example of the trie data structure, the alphabet base is 256, which results in a bitmap size of 256 bits. In other words, each bitmap can identify 256 pointers for 256 different key part values (or child nodes). With 256 different values, 8 bits (2 8 ) or 1 byte can be encoded.
[0343] The root node 502 has one bit set in its bitmap, representing the key part value "00". Thus, the root node 502 includes only one pointer 503 to the child node 504, which is associated with the key part having the value "00". The child node 504 has two bits 505, 507 set in its bitmap, i.e., the bits representing the key part values "2A" and "00". Thus, the child node 504 includes two pointers 507 and 508, which point to the corresponding child nodes 509, 510.
[0344] The pointer 506 associated with the bit having the value "2A" in the bitmap is addressed by calculating how many least significant bits are set starting from the bit 505 representing the key part value "2A". In this case, only one least significant bit is set, i.e., the bit 506, so it can be determined that there is an offset of one pointer, and the pointer we are looking for is the second pointer included by the child node 504.
[0345] General features of a preferred embodiment of the trie according to the present invention
[0346] Similar to all tries, the trie or trie data structure according to the present invention includes one or more nodes. As in the prior art trie described above with reference to Figures 1 to 6 each child node of the trie of the preferred embodiment is preferably associated with a key part, where the path from the root node in the trie to another node, particularly to a leaf node, defines a key, which is the concatenation of the key parts associated with the nodes in the path. The root node is not associated with a key part, and it will be understood that the trie may include additional nodes that are not associated with key parts, e.g., because they are used for other purposes. For example, the number of entries in a subtree can be stored in such a node to avoid or accelerate counting operations, which typically require traversing the entire tree.
[0347] In a preferred embodiment of a trie according to the present invention, each node, preferably at least each parent node having more than one child node, includes a bitmap and a plurality of pointers. Each pointer is associated with a bit set in the bitmap and points to a child node of the node. Typically, a bit is "set" in the bitmap if the value of the bit is "1". However, in a particular embodiment, a bit may be considered "set" if the value of the bit is "0". A bit in the bitmap is considered "set" here if the value of the bit in the bitmap corresponds to a value associated with the concept of a valid branch marked by the bit in the bitmap, as explained in the prior art trie shown above with reference to Figure 5 as shown.
[0348] Preferably, the bitmap is stored in memory as an integer of a predefined size. Additionally, the size of the bitmap is preferably 32, 64, 128, or 256 bits. The operating performance of storing and processing the trie in the target computer system can be improved by selecting the size of the bitmap such that it is equal to the bit width of the CPU's registers, system bus, data bus, and / or address bus of the target computer system.
[0349] For example, as mentioned above, the memory address of a pointer associated with a bit set in the bitmap can be calculated based on the number of least significant bits set in the bitmap. This determination can be made very efficiently using simple bit operations and the CTPOP (Count Trailing Ones Pop) operation for determining the number of set bits. Many modern CPUs even provide CTPOP as an internal instruction. However, since in modern CPUs, a long integer is 64 bits wide, CTPOP only works on 64 bits. This means that for a prior art trie using a 256-bit bitmap, this operation is performed at most four times (4 x 64 = 256). Alternatively, the prior art trie stores the total bit count of the previous bitmap along with the first three bits Figure 1 stored. Then the number of least significant bits can be calculated as the CTPOP of the last bit group + the bit count of the previous bit group.
[0350] Since the system bit width is 64 bits in most computer systems currently, a 64-bit bitmap size is currently the most preferred size and is used by the inventor in an exemplary embodiment of his invention. This results in a radix-64 trie, which means that each node can store symbols of an alphabet of 64 symbols, that is, it can encode 6 bits (2 6 = 64). As will be explained below, a trie according to an embodiment of the present invention can use several nodes and their associated key parts to store information included by the original data type. For example, to store a key represented by a 64-bit long integer, a radix-64 trie with 11 levels is required (11 * 6 bits >= 64).
[0351] The bitmap and / or pointer can be stored in, for example, an array, in a list, or in a continuous physical or virtual memory location. Note that whenever the term "memory" is used in this article, it can refer to physical or virtual memory, preferably continuous memory. In a preferred embodiment, a long integer (64 bits) is used to represent the bitmap, and is also used to represent each of the child pointers. Instead of allocating nodes separately in memory, the nodes are stored in an array of long integers, and instead of having a memory pointer for the node, the current node is assigned to the array by an index. The child pointer can be an index of a node position in an array. When traversing the trie, the offset for finding the index of the child node based on the current node index is the number of least significant bits set in the bitmap before the target position for the bitmap plus 1.
[0352] The preferred embodiment applies to several such arrays. One part of the pointer, e.g. the lower part, is an index into the array, while another part of the pointer, e.g. the upper part, is a reference to the array. This is done for memory management reasons, since it is not always possible to allocate an array of arbitrarily large size. For example, in Java, the size of an array is limited to a 32-bit integer, and this results in an array size of 2. 31 (positive values only) = 2,147,483,648. However, many real applications need to contain arrays of 16MB or more, which corresponds to 2 million entries for an array of 64-bit long integers.
[0353] Similar to Figure 5 and Figure 6 In a prior art trie of , a parent node, preferably at least each parent node having more than one child node, typically includes a number of pointers equal to the number of bits set in the bitmap included by the parent node. For example, Figure 5 The node 504 in has a bitmap with two bits 505, 506 set, and includes two pointers 507, 508. The rank of a pointer among all pointers of the parent node preferably corresponds to the rank of the associated set bit of the pointer among all set bits in the bitmap of the parent node. Figure 5 In the parent node 504 of , the first pointer 507 corresponds to the first set bit 505 of the bitmap, and the second pointer 508 corresponds to the second set bit 506. The pointers are typically stored in the same or opposite order as the bits set in the bitmap. Each of them preferably points to (the address of) the bitmap included by the child node, as shown in the following example: Fig. 9 and Fig.12 As shown in, for example, the starting address of this bitmap.
[0354] Similar to Figure 5 or Figure 6In the prior art trie, the value of the key part of a child node of the trie of the preferred embodiment, preferably the value of the key part of at least each child node whose parent node has more than one child node, is determined by the value of the bit (set) in the bitmap included by the parent node, depending on which bit the child node is associated with. Thus, typically, the maximum number of different values available for the key part is defined by the size of the bitmap, and / or the size of the bitmap defines the possible alphabet of the key part.
[0355] In a preferred embodiment of the present invention, each key part in the trie is capable of storing a value of the same predefined size, such as a 5-bit value (if the size of the bitmap is 32 bits), a 6-bit value (if the size of the bitmap is 64 bits), a 7-bit value (if the size of the bitmap is 128 bits), or an 8-bit value (if the size of the bitmap is 256 bits). The alphabet of the characters represented by the node or key part is the set of all possible bit groups of that size. For example, in the case where the key part is capable of storing a 6-bit value, the alphabet is the set of all bit groups including 6 bits.
[0356] The trie data structure according to the present invention can be used to implement: a key-value mapping (also known as an "associative array"), where the value is stored in the leaf node; and a key set (also known as a "dynamic set"), where no data is stored in the leaf node. In the case where each key has only one value, a mapping is used to look up the value of a given key, while a set is used to determine whether a given set contains a given key. For both cases, set operations on keys (such as union, intersection, or difference) are also frequently required operations.
[0357] FIG. 7A to FIG. 7C Shows how to store leaf nodes in a mapping in a compact manner and how to store the "last" bitmap in a set, and enables all values to be accessed in constant time. Fig. 7A Shows how to store leaf nodes for a key-value mapping for a value data type with a large and / or variable size (such as a long string or text). As Fig. 7A shown, the value is stored separately. If the size of the value is large compared to the size of the pointer, the space required for an additional pointer idx pointing to the value can be ignored.
[0358] In the case where the value data type is of a fixed size (such as a date or an integer), it is more efficient to store the value "inline", as Figure 7BAs shown, for example, directly behind the bitmap of its parent node. For example, according to a preferred embodiment of the present invention having inlining, a long-to-long mapping can be effectively implemented by a trie data structure, that is, a mapping where the key is of long integer type and the value is also of long integer type. Inlining only works for data types with fixed-size values because in this case, the position can be calculated, for example, as CTPOP(bitmap & (bitpos - 1)) * size. In contrast, for data types with variable-size values, the position must be determined through all value entries.
[0359] Figure 7C It shows how the inlining of storing the "last" bitmap can be used for a set. Since there is no value, the "last" bitmap itself is inlined without the need for a pointer. Note that in the terminology used herein, these "last" bitmaps at the physical level are part of the parent node of the leaf node, but the bits set in these bitmaps indicate the values of the key parts of the leaf nodes at the logical level.
[0360] The trie data structure stores keys in an ordered manner and thus allows traversing the keys in order. For example, a 64-bit long integer key starting from the most significant 6 bits (or 4 bits since 64 = 4 + 10 * 6) to the least significant 6 bits can be stored. In this way, the integer is treated as an unsigned long integer. For signed integers typically using two's complement encoding, to have the correct order, it must be converted to offset binary representation, for example, by adding 2 for a 64-bit long integer 64-1 .. Floating-point numbers are handled in a similar way. Therefore, encoding the value of a data item of a key (such as a floating-point number or a two's complement signed integer) can include converting the data type of the data item to an offset binary representation containing an unsigned integer (e.g., an unsigned long integer).
[0361] Space-saving trie data structure
[0362] In many application scenarios, the use efficiency of memory space is very low. This is especially true when the trie is sparsely populated and / or degenerates into a chain of nodes, where each node has only a single child pointer.
[0363] Linked node optimization
[0364] The inventors found in empirical studies that for any key, many nodes in a trie in typical application scenarios have only a single child node. This is because many keys share common prefixes, infixes, or suffixes. In this case, the prior art trie degenerates into a chain of nodes with a single child pointer, and the space efficiency of the prior art trie data structure is low.
[0365] When a node has only a single child node, i.e., when only a single bit is set in the bitmap included by the parent node, the first space optimization of the present invention eliminates the child pointer. This method is referred to herein as "chained node optimization". Figure 8 An example of a trie in which the chained node optimization is effective is shown in Figure 8 . The leaf nodes 841, 842, 843 at level 4 of the trie share a common prefix, which includes: the root node 800 at level 0; the node 810 at level 1, which is the only child node of the root node 800; the node 820 at level 2, which is the only child node of the node 810; and the node 830 at level 3, which is the only child node of the node 280.
[0366] The first space optimization of the present invention is applied to a trie including one or more nodes, wherein each parent node included by the trie, preferably each parent node having more than one child node, includes a bitmap and one or more pointers, wherein each pointer is associated with a bit set in the bitmap and points to a child node of the parent node. The optimization is achieved by the fact that each parent node included by the trie, preferably each parent node having only one child node, does not include a pointer to the child node, and / or the child node is stored in a predefined position in the memory relative to the parent node. Preferably, the child node of a parent node having only one child node is stored in a position in the memory directly behind the parent node.
[0367] In Fig. 9 An example of the first space optimization according to the present invention is shown, which shows a part of the trie 910 before the chained node optimization and a part of the trie 920 corresponding to the trie 910 after the application of the chained node optimization. As in the preferred embodiment described above, the nodes in the tries 910, 920 are stored in a long integer array, which is indicated by the expression "long[]" on the left side of the illustration of the parts of the tries 910, 920.
[0368] The trie 910 includes a first node 911, which has only a single child node, as indicated by the (64-bit wide) bitmap of the node 911, wherein only one bit is set (1) and all other bits are not set (0). Similar to in the prior art trie, the node 911 thus includes a single pointer 914, which is a long integer, pointing to the child node (node 912) of the node 911. The node 912 has two child nodes, in Fig. 9is not shown in the figure, as indicated by two bits set in the (64-bit wide) bitmap at node 912 and by two pointers 915, 916 included by node 912. As indicated by the three dots ("…") between node 911 and node 912, node 912 will typically not be stored in the memory location directly behind node 911, but can be stored anywhere in the long integer array.
[0369] Trie 920 also includes a first node 921 having only a single child node, as indicated by the (64-bit wide) bitmap of node 921, where only one bit is set (1) and all other bits are not set (0). However, compared to node 911 in trie 910, node 921 in trie 920 does not include a pointer to the child node (node 922) of node 921. Instead, node 922 is stored in the memory location directly behind node 921, and preferably, but alternatively, can be stored anywhere in the long integer array as long as the position relative to the parent node 921 in the memory is predefined. For example, the child node 922 can be stored directly in front of the parent node 921, or there can be another data object of a fixed length between the parent node 921 and the child node 922. Similar to node 912 of trie 910, node 912 of trie 920 has two child nodes, which are Fig. 9 not shown in the figure, and whose position in the memory (in the long integer array) is indicated by two pointers 925 and 926 included by node 922.
[0370] In Fig. 9 the example embodiment of, each of the bitmap and pointers included by the node is represented by a long integer, i.e., by the same data type or data types of the same length. According to the chained node optimization of the present invention, the memory space required to store a node with a single child node is reduced by 50%.
[0371] Terminal Optimization
[0372] The second space optimization of the present invention provides a more compact representation of the trie in memory, where the "end" of the trie includes a chain or string of single nodes, i.e., many keys do not have a common suffix. Fig.10 An example of a trie for which the second space optimization is effective is shown in Fig.10Each of the leaf nodes 1041, 1042, 1043, 1044, 1045 in the trie is part of an independent (non-common) suffix. Each suffix includes a node 1021, 1022, 1023, 1024, 1025 at level 3 of the trie, a node 1031, 1032, 1033, 1034, 1035 at level 4 of the trie, and a leaf node at level 5. Each of the nodes 1021, 1022, 1023, 1024, 1025 at level 2 and the nodes 1031, 1032, 1033, 1034, 1035 at level 4 has only a single child node.
[0373] According to the second space optimization, the node at the start of the string of a single node is marked as a "terminal branch node". In Fig.10 it, the terminal branch nodes are the nodes 1021, 1022, 1023, 1024, 1025 at level 3 of the trie. The values of the key parts of the remaining nodes in the string are stored continuously only in their "local" or literal encoding, rather than being determined by the values of the bits (sets) in the bitmap included by their parent nodes. This method is referred to herein as "terminal optimization".
[0374] Thus, the second space optimization of the present invention is applicable to a trie including one or more nodes, wherein each parent node included by the trie, preferably at least each parent node having more than one child node, includes a bitmap; the nodes, preferably each child node, is associated with a key part; and the value of the key part of the child node (preferably at least each child node whose parent node has more than one child node) is determined by the value of the bits (sets) in the bitmap included by the parent node, depending on which bit the child node is associated with.
[0375] In Dr. Bagwell's "Fast And Space Efficient Trie Searches", technical report, EPFL, Switzerland (2000), a method called "tree tail compression" that independently allocates nodes in memory uses pointers to string nodes for reference, or directly stores the numerical value of the terminal string in the terminal branch node. However, since the offset and node type (nodes with bitmaps or nodes with character / pointer lists) must be stored in the node, this method does not save space.
[0376] Terminal optimization according to the present invention overcomes this problem by marking nodes with a bitmap in which no bits are set, preferably each node in the trie that has only one child and all of whose descendant nodes have at most one child as a terminal branch node. The present invention makes use of the special quality of the bitmaps included in the standard nodes of the preferred embodiment, namely that they always have at least one bit set. This is because a node with an all-zero bitmap would be a node without children, but a node without children does not need to be represented in memory. Thus, a special meaning can be attributed to a bitmap with no bits set, and the length or format of the bitmap of a terminal branch node can be the same as that of the bitmap included in a parent node with more than one child.
[0377] The value of the key part associated with the descendant nodes (preferably each descendant node) of a terminal branch node (preferably each terminal branch node) is not determined by the value of the bit(s) set in the bitmap included in the parent node of the descendant node. Instead, the value of the key part is encoded such that it represents less memory space compared to the bitmap included in a parent node with more than one child. Typically, the value of the key part will be encoded as a binary number (digit), such as an integer value. For example, in cases where the bitmaps included in the standard nodes have 32 bits, 64 bits, 128 bits, or 256 bits respectively, the key part associated with the descendant nodes of a terminal branch node is encoded by 5 bits, 6 bits, 7 bits, or 8 bits respectively.
[0378] Fig.11 A general storage configuration of terminal optimization according to a preferred embodiment of the present invention is shown in. A 64-bit wide bitmap in which no bits are set is followed by a plurality of (usually at least two) 6-bit key parts encoded as binary numbers. However, in the preferred embodiment, each 6-bit key part is stored in an 8-bit block (one byte). Since only 6 out of the available 8 bits are used, this wastes some space, but converting from 6 bits out of 8 bits to 8 bits out of 8 bits encoding and back increases the complexity of the implementation and reduces performance. Measurements made by the inventors have shown that the wasted space is acceptable and the improvement in space by storing 8 bits out of 8 bits is very limited.
[0379] A terminal branch node and / or its descendant nodes do not need to include a pointer to its one child (if any), because the child node can be stored in memory at a predefined position relative to the parent node, preferably directly behind the parent node, as Fig.12 shown in. Additionally, the values of the key parts associated with the descendant nodes (preferably each terminal branch node) of a terminal branch node (preferably all descendant nodes) are stored consecutively after the terminal branch node. Finally, as can be seen in Fig.12As observed in, in the string of a single node, only for a terminal branch node must it have a bitmap to mark the node as a terminal branch node, while the descendant nodes without terminal branch nodes do not need to include a bitmap, especially a bitmap where the set bits determine the values of the key parts associated with its child nodes.
[0380] As can be seen from Fig. 14C and as will become apparent from the examples shown below and discussed below, only when a terminal branch node has more than one descendant node, compared to the chain node optimization, the terminal optimization according to the most preferred embodiment saves more space. Additionally, if the first node in a series of single nodes has been marked as a terminal branch node, the maximum space savings can be achieved, such that the parent node of the terminal branch node has more than one child node.
[0381] In Fig.12 a second space optimization according to the present invention is shown, which shows a part of trie 1210 before terminal optimization and a part of trie 1220 corresponding to trie 1210 after applying terminal optimization. The nodes of trie 1210 and trie 1220 are again stored in a long integer array.
[0382] Trie 1210 includes a first node 1211, which has only a single child node, as indicated by the bitmap of node 1111, where only one bit is set (1) and all other bits are not set (0). The value of the set bit is "60", as can be seen from the fact that it is the fourth bit from the left in a 64-bit wide bitmap, where the rightmost bit has a value of "0" and the leftmost bit has a value of "63". Node 1211 includes a single pointer 1213, which is a long integer, pointing to the child node (node 1212) of node 1211. As indicated by the three dots ("…") between node 1211 and node 1212, node 1112 will typically not be stored in the memory location directly behind node 1111, but can be stored anywhere in the long integer array. Node 1212 also has a child node, as indicated by the second bit from the right set in the 64-bit wide bitmap of node 1212, and the bit value is "01". However, since the child node of node 1212 is a leaf node, node 1212 does not include a pointer to its child node, but a pointer to leaf node part 1214, which can be a mapped pointer or value (see Fig. 7A and Figure 7B ), or an empty set (see Figure 7C ). The leaf nodes are not represented in the memory. Note that no specific "leaf node indicator" is required, because in the preferred embodiment, the depth of the trie or the length of the keys stored in the trie is known.
[0383] As can be observed, node 1211 is a terminal branch node because it has only one child node 1212 and all its descendant nodes (1212 and the child nodes of 1212) have at most one child node (node 1212 has one child node and the child node of 1212 has zero child nodes). As a result of the application of terminal optimization, trie 1220 is obtained from trie 1210. The node 1221 of trie 1220 corresponding to node 1211 of trie 1210 has been marked as a terminal branch node by providing it with a 64-bit wide bitmap with no set bits. The value of the key portion associated with its child node 1222 (corresponding to child node 1212 of trie 1211) is not determined by the value of the bits (set) in the bitmap of node 1221. The value of the key portion is encoded as a binary number included in node 1221, such as an integer, as Fig.12 indicated by the number "60" in 6 . And thus, it requires much less memory space compared to the 64-bit wide bitmap included in a parent node having more than one child node. For example, the value 60 can be encoded as the binary number "111100".
[0384] The terminal branch node 1221 does not include a pointer to its child node 1222. Instead, the child node 1222 is stored in a predefined position in memory relative to its parent node 1221, i.e., directly behind the parent node. The node 1222, which is a descendant node of the terminal branch node 1221, does not include a bitmap, nor does it include a pointer to its child node, but only includes a binary number encoding the value of the key portion associated with the child node of node 1222, as indicated by the number "01" in Fig.12 . For example, the value 01 can be encoded as the binary number "000001". The representation of node 1222 is followed by the leaf node portion 1224 in memory. Similarly, the leaf node portion 1224 can be a mapped pointer or value (see Fig. 7A and Figure 7B ), or an empty set (see Figure 7C ).
[0385] In Fig.12 's example embodiment, where each of the bitmap and pointer included in a node is represented by a long integer, the terminal optimized trie 1220 requires 64 + 6 + 6 = 76 bits to store nodes 1221 and 1222. In contrast, the unoptimized trie 1210 requires 64 + 64 + 64 = 192 bits to store the same information (nodes 1211 and 1212, not counting the leaf node portion 1214 that exists in both trie 1210 and trie 1220).
[0386] Similar to the preferred embodiment, a long integer array is used to store the trie, and the terminal optimization according to the present invention suffers from alignment loss. In the worst case, a 6-bit key part is stored in a 64-bit long integer. However, experiments show that, on average, 50% of the space used to store the descendant nodes of the terminal branch nodes is occupied. In addition, terminal optimization still requires less space compared to optimizing the storage of several individual child nodes using pointers or using chained nodes.
[0387] Now, reference will be made to Fig.13 an overview of a method for accessing standard nodes, nodes optimized by chained node optimization, and nodes optimized by terminal optimization in a unified manner. The main methods of node access used in the query execution model in the example embodiment are: getBitSet(), which returns a bitmap with bits set for all non-empty sub-pointers of the trie node; and getChildNode(bitNum), which returns the child node of the given node branch specified by the bit number. Both of these methods are provided by the interface CDBINode, where CDBI stands for "Confluent Database Index". Another interface, CDBINodeMem, provides an object-oriented access to the data model through the CDBINode interface.
[0388] The difficulty that must be overcome is how to handle these three cases in a unified central location without having to handle them separately in many places in the code. According to the solution discovered by the inventors, and as Fig.13 shown in, not only are nodes referenced via node pointers (indexes), but also via a base node index ("nodeRef") and an index within the node ("idxInNode"). A chain without a pointer and a terminal are regarded as a single node, whose base node index points to the starting point of the terminal node or the chain. In this way, the memory space optimization does not add significant complexity to the "fetch" operation and thus does not degrade performance.
[0389] Since nodeRef always points to the first bitmap, it is used to detect the three cases in the implementations of getBitSet() and getChildNode(bitNum):
[0390] - If the bitmap has more than one bit set, it belongs to a regular node. getBitSet() returns the bitmap; getChildNode(bitNum) determines the idx (pointer to the child node) and returns a new CDBINodeMem, setting nodeRef to it (and idxInNode to 0).
[0391] - If the bit map is 0 (bit not set), it belongs to a terminal branch node. getBitSet() converts the 6-bit value literally stored at the idxInNode position into a bit map and returns it; getChildNode(bitNum) returns a new CDBINodeMem with the same nodeRef and idxInNode + 1.
[0392] - If the bit map has one bit set, it belongs to a chained node. getBitSet() returns the bit map at the idxInNode position; getChildNode(bitNum) again returns a new CDBINodeMem with the same nodeRef and idxInNode + 1. The CDBINodeMem can also be used in lightweight mode by not creating a new time-consuming child node, but just updating the nodeRef and idxInNode (gotoChildNode method), i.e., by modifying the existing object acting as a proxy.
[0393] FIG. 14A to FIG. 14C Shows the trie growth when inserting two keys into an empty trie. The key part values of the first key are [00, 02, 02], and the key part values of the second key are [00, 03, 00, 01]. Fig.14A Shows an empty trie, which is represented by a root index pointer with a value of 0. Fig. 14B Shows the trie after adding the first key with key part values [00, 02, 02]. Since there is no previous entry, the terminal optimization is used to store the first key. Fig. 14C Shows the trie after adding the second key with key part values [00, 03, 00, 01]. The existing terminal node has been split. Since the matching prefix ("00") has only one child node, it is stored as a chained node, followed by a regular node with two child nodes ("02" and "03"). The remainder of the first key ("02") is stored as a node with a leaf node part (storing it using the terminal optimization would require more space). The remainder of the second key ("00" and "01") is again stored using the terminal optimization.
[0394] To measure the space requirements of the data structures according to various embodiments of the present invention, experiments were conducted in which a random set of long integers with the full range of long integer values was stored. Fig.15 Shows the measurement results, where the x-axis indicates the number of entries loaded into the trie index on a logarithmic scale, and the y-axis indicates the number of bytes required to store one entry on average.
[0395] It can be observed that the chained node optimization alone reduces the space requirement by approximately 40%, and the terminal optimization alone reduces it by approximately 60 - 75%. Compared to the terminal optimization alone, the combined chained node and terminal optimization does not provide a visible space improvement (the graph overlaps with the terminal optimization case). However, empirical measurements made by the inventors indicate that these two optimizations are still worth applying together. When the chained node optimization is applied in addition to the terminal optimization, the performance is improved because fewer pointers have to be followed, and the data locality is better, and the memory hierarchy (CPU cache) is respected.
[0396] Bitmap Compression
[0397] The third space optimization of the present invention provides a more compact representation of the trie in the memory of a sparsely filled trie. By grouping and efficiently storing sections of the same value (e.g., sections with a value of 0 in the case of sparsely filled nodes or sections with a value of 1 in the case of densely filled nodes), the memory space of a bitmap (e.g., a bitmap indicating the key part values of the child nodes) can be reduced. This third space optimization is referred to herein as "bitmap compression".
[0398] The third space optimization of the present invention is applied to a trie including one or more nodes, where each node, preferably at least each parent node having more than one child node, includes a bitmap in the form of a logical bitmap and a plurality of pointers, where each pointer is associated with a bit set in the logical bitmap and points to a child node of the node. As mentioned with respect to other aspects of the present invention, the logical bitmap can correspond to a bitmap including key part values. The optimization is achieved by the fact that the logical bitmap is divided into a plurality of sections and encoded by a header bitmap and a plurality of content bitmaps, where each section is associated with a bit in the header bitmap, and where for each section of the logical bitmap in which one or more bits are set, the bit associated with the section in the header bitmap is set and the section is stored as a content bitmap.
[0399] Storing the logical bitmap using the header bitmap and a plurality of content bitmaps can significantly reduce the storage space required by omitting the content bitmaps for sections of the logical bitmap in which no bits are set. In other words, only the content bitmaps (i.e., sections of the logical bitmap) having at least one set bit are stored in the memory. In the worst case, where each section of the logical bitmap has at least one set bit, the memory usage increases slightly due to the need to store an additional header bitmap. However, nodes are typically sparsely filled, and thus less memory is typically required when using bitmap compression.
[0400] In Fig.16An embodiment of bitmap compression according to the present invention is shown, which shows a section of trie 1601 at the upper part, and its parent node includes a logical bitmap 1602 without bitmap compression. At its lower part, a section of trie 1611 is shown, and its parent node is obtained after bitmap compression has been applied to the parent node of trie 1601. The parent node of trie 1611 includes a header bitmap 1612 and two content bitmaps 1613, 1614 generated by bitmap compression.
[0401] The parent nodes in trie 1601 and trie 1611 also include pointers 1603 to 1605. In trie 1601, each pointer is associated with bits 1606 to 1608 set in the logical bitmap 1602. In trie 1611, pointers 1603 to 1605 are associated with bits 1616 to 1618 set in the content bitmaps.
[0402] By dividing the logical bitmap 1602 into sections 1621 (e.g., 8 bits) and storing the sections 1622, 1623 in which at least one bit is set as content bitmaps 1613, 1614, the logical bitmap in trie 1601 is converted into the header bitmap 1612 and content bitmaps 1613, 1614 in the lower part 1611. The sections in which no bits are set are not stored as content bitmaps. Each bit in the header bitmap 1612 represents a different section of the logical bitmap. Content bitmaps 1613, 1614 are referenced by the corresponding bits 1619, 1620 set in the header bitmap 1612. Content bitmaps 1613, 1614 can be stored in the same order (not shown) or in the reverse order, where the set bits 1619, 1620 associated with their sections are arranged in the header bitmap 1612. In other words, the rank of the content bitmaps in all content bitmaps of the logical bitmap can correspond to the rank of the set bits in all set bits in the header bitmap that are associated with the sections of the content bitmaps. In this way, the content bitmaps can be easily addressed when processing the trie. Also, the sections of the logical bitmap are preferably all contiguous in the memory. Therefore, the entire logical bitmap can be represented contiguously in the memory by the header bitmap followed by multiple content bitmaps.
[0403] In a preferred embodiment, all sections have the same size. Sections of the same size allow for efficient processing of compressing and decompressing the logical bitmap, because no further information about the structure of the sections is required. Also, the size of the header bitmap can be the same as the size of the sections.
[0404] Different structures can be used for storing the header bitmap and the content bitmaps. The header bitmap and the content bitmaps of the logical bitmap can be stored in an array, a list, or in contiguous physical or virtual memory locations. When the size of the header bitmap and the content bitmaps is one byte, as Fig.16 As shown in, the bit map can be stored in a byte array instead of in a long integer array, as is done in trie 1601 or Fig. 9 and Figure 12 to Figure 1 in the trie of 4, which reduces alignment loss. The content bit map is preferably stored in a predefined position in the memory relative to the header bit map. Storing the content bit map and the header bit map close to each other can improve processing efficiency. In Fig.16 it, the content bit map is stored directly behind the header bit map in the memory.
[0405] The bitmap compression described above can also be applied to pointers, such as pointers for referencing child nodes, and can also be applied to inline leaf node bitmaps. This will typically further improve space efficiency but will degrade performance because variable-size encoding makes it necessary to traverse the pointer when calculating the offset of a particular pointer.
[0406] Bitmap compression can be combined with other aspects of the present invention. For example, in combination with pointer reduction, terminal branch nodes of the present invention that are not set according to (logical) bitmap markers can be encoded as only a header bitmap without any content bitmap.
[0407] Fig.17 illustrates the experimental results described above with reference to Fig.15 However, in addition to chained node and terminal optimizations, bitmap compression was also applied. As can be observed from the comparison with Fig.15 when chained node or terminal optimizations are not applied, the space savings achieved by bitmap compression are approximately 40%, and additionally when only chained node optimization is applied, the savings are approximately 50 - 70%, and additionally when only terminal optimization or both chained node and terminal optimizations are applied, the savings are approximately 30 - 60%.
[0408] Key Encoding
[0409] The present invention provides ways to store keys of different primitive data types having fixed or variable sizes (such as strings), as well as composite keys including two or more items of primitive data types in a trie.
[0410] Keys including control information
[0411] In a preferred embodiment of the present invention, keys can be encoded in a flexible manner such that iteration can be performed, for example, through a cursor even without prior knowledge of the number, data type, or length of the components stored in the key.
[0412] These embodiments are applied to tries used in database applications or information retrieval systems, e.g., a trie or trie data structure according to one of the embodiments of the trie and trie data structure described above. The trie includes one or more nodes, where a node, preferably each child node, is associated with a key part, and the path from the root node in the trie to another node, particularly to a leaf node, defines a key associated with that node, where the key is a concatenation of the key parts associated with the nodes on the path. The above-mentioned flexibility is achieved by the fact that, in addition to content information, the key also includes control information.
[0413] The key typically will include one or more key parts, where each key part includes content information that is part of the overall content information included by the key. For each of the key parts, the control information preferably includes a data type information element that specifies the data type of the content information included by the key part.
[0414] In principle, there are two ways to arrange the control information and content information associated with the key parts. The first way is shown in Fig.18 where the key part or preferably each key part 10, 20, 30 includes a data type information element ( Fig.18 the shaded element in ) and a content information element. The data type information element specifies the data type of the content information element. This means that the control information and content information are distributed across different key parts 10, 20, 30. In this case, the key's data type information element or each data type information element is typically located by the content information element associated with the data type information element in the key, preferably (directly) before the data type information element. Thus, the data type information element is similar to a prefix or header element of the key parts 10, 20, 30.
[0415] Fig.19 The second way of arranging the control information and content information associated with the key parts is shown in, where the data type information elements ( Fig.19 the shaded elements in ) are located together and preferably arranged in the same or opposite order as the content elements for which they specify their data types. The control information such as the data type information element is preferably located before the content information in the key, as a prefix or header element of the key. Research conducted by the inventors has shown that the second method of storing the data type information elements is preferable because keys with the same data type have the same prefix (starting node in the trie), and thus the number of nodes required in the trie is reduced, resulting in space savings.
[0416] The data type of the content information associated with a key part can have a fixed size, such as in the case of an integer, long integer, double-precision floating point, or time / date primitive, or it can have a variable size, such as in the case of a string, for example a Unicode string or a variable-precision integer. In some embodiments of the present invention, a key includes two or more key parts, which include content information of different (primitive) data types.
[0417] As will be explained below with reference to Fig. 20 and Fig.21 the control information can include information identifying the last key part (e.g., by the high-order status of the data type information element). Alternatively, the key part count can be stored separately. In addition, the control information can include information about whether the trie is for storing a dynamic set or an associative array.
[0418] The content information of a key part can be contained in a single key part, but typically it is contained in two or more key parts. For a key part of fixed size, the number of key parts required to contain the content information included by the key part is typically known. In the case where the data type of the content information included by the key part is a variable-size data type, the end of the content information element can be marked by a specific symbol, such as a null-terminated string, where the last character of the string is a null character ('\0', called NUL in ASCII). Alternatively and preferably, it can be marked by a specific bit in a specific one of the key parts containing the key part, as will be explained below with reference to Fig.21 the Unicode string.
[0419] Although, as mentioned above, the content information of a key part will typically be contained in two or more key parts, a key part preferably does not contain the content information of two or more key parts. In other words, the content information of a key part is aligned with the boundary of the key part. Similarly, a key part preferably does not contain the information of two or more control information elements, such as data type information elements, key part counts, or information about whether the trie is for storing a dynamic set or an associative array. This approach makes the implementation easier and more efficient, generally without significant alignment loss. In addition, as will be explained below, it allows the content information of different key parts to be stored in an interleaved manner.
[0420] In Fig. 20An example of key encoding according to the present invention is shown. The keys to be stored in the trie include control information 2010 and content information 2020. The keys include several key parts containing content information, and for each of the key parts, the control information 2010 includes data type information elements 2012, 2013 specifying the data type of the content information included by the key part. In addition, the control information 2010 includes information 2011 on whether the trie is for storing a dynamic set or an associative array. If the trie is used to store an associative array, the leaf nodes of the trie associated with the keys will typically include a leaf node value 2030 or a pointer to the leaf node value.
[0421] As mentioned above, in a preferred embodiment of the present invention, each parent node in the trie includes a 64-bit wide bitmap, and thus each key part in the trie is capable of storing a 6-bit value. The information 2011 on whether the trie is for storing a dynamic set or an associative array is stored by the first key part, and thus 6 bits are used for this information. In fact, for this yes / no information, 1 bit would be sufficient, but for the alignment reasons mentioned above, the entire key part capable of storing a 6-bit value is used. This information is encoded in node 2041, including a 64-bit wide bitmap (where the corresponding bit is set) and a pointer (idx) to the corresponding child node of node 2041.
[0422] Each of the data type information elements 2012, 2013 is also stored by a key part, and its value is encoded in the bitmaps of nodes 2042, 2043. 5 bits are used for the data type information, which allows 32 different type identifiers. The 6th bit (e.g., the high bit of the key part) that can be stored by the corresponding key part is used to indicate whether the key part associated with the data type information element is the last key part in the key. In Fig. 20 the example, the high bit of the data type information element 2013 is set to indicate that the key part associated with the data type information element 2013 is the last key part in the key.
[0423] The content information included by each of the key parts is also decomposed into generally 6-bit values 2021, 2022, and each of the values is stored by a key part. Due to space reasons, the nodes whose bitmaps are used to encode (6-bit) values 2021, 2022 are not shown in Fig. 20 For example, in the case where the key part includes a 32-bit integer value, this 32-bit value is stored by six key parts, the first key part stores a 2-bit value, and the last five key parts each store a 6-bit value (32 = 2 + 6 + 6 + 6 + 6 + 6).
[0424] Fig.21A trie storing keys encoded according to the present invention is shown. Each parent node of the trie includes a 64-bit wide bitmap. The key includes a first key part and a second key part. The first key part includes a 32-bit integer with a value of "100", and the second key part includes a string with a value of "ab". The control information of the key includes (1) a dynamic set identifier, (2) an integer identifier, and (3) a string type identifier with a tag of the last type. The content information of the key includes (1) a 32-bit integer value "100" and (2) a string value "ab" encoded in Unicode.
[0425] The dynamic set identifier is a 6-bit number with a value of 0 (0x00). Thus, Fig.21 the bitmap of the root node 2100 of the trie has the bit with the value "0" set, and the second-level node associated with this bit is associated with the key part with the value "0". The integer identifier is a 6-bit number with a value of 5 (0x05). Thus, the bitmap of the second-level node has the bit with the value "5" set, and the third-level node associated with this bit is associated with the key part with the value "5". The string type identifier with a tag of the last key part is a 6-bit number with a value of 39 (0x27). Thus, the bitmap of the third level has the bit with the value "39" set, and the fourth-level node associated with this bit is associated with the key part with the value "39".
[0426] The integer value "100" is encoded in 32-bit binary as "00 000000 000000 000000 000001100100". Thus, the key parts associated with the nodes on levels 5 to 10 for storing the integer value "100" are associated with the values 0 (0x00), 0 (0x00), 0 (0x00), 0 (0x00), 1 (0x01), and 36 (0x24) respectively.
[0427] The string value "ab" is encoded as the Unicode value of the character "a", followed by the Unicode value of the character "b". Each Unicode character is stored using 2 - 4 key parts, depending on the Unicode value, which may require 10, 15, or 21 bits. The encoding scheme for Unicode characters used in the preferred embodiment of the present invention is as follows:
[0428] 10-bit Unicode character: 00xxxx xxxxxx
[0429] 15-bit Unicode character: 010xxx xxxxxx xxxxxx
[0430] 21-bit Unicode character: 011xxx xxxxxx xxxxxx xxxxxx
[0431] Marking the last character in a string by setting a high bit results in the following encoding scheme for the last character:
[0432] 10-bit Unicode character: 10xxxx xxxxxx
[0433] 15-bit Unicode character: 110xxx xxxxxx xxxxxx
[0434] 21-bit Unicode character: 111xxx xxxxxx xxxxxx xxxxxx
[0435] The value of the Unicode character 'a' is 97 (0x61) and its 10-bit Unicode encoding is '0001100001'. According to the encoding scheme used in the preferred embodiment, the Unicode character 'a' is encoded as '000001100001'. The value of the Unicode character 'b' is 98 (0x62) and its 10-bit Unicode encoding is '0001100010'. According to the encoding scheme used in the preferred embodiment, the Unicode character 'b' is encoded as '100001100010' with the high bit set because 'b' is the last character in the string with the value 'ab'. Thus, the key parts associated with the nodes at levels 11 to 14 used to store the string value 'ab' are associated with the values 1 (0x01), 33 (0x21), 33 (0x21), and 34 (0x22) respectively.
[0436] Interleaved multi-item keys
[0437] Embodiments of the present invention provide a way to store data in a trie such that queries involving more than one data item can be performed in a more efficient manner. The method of the present invention is particularly useful for: storing keys or keys and values in a database or information retrieval system such that they can be queried more efficiently, storing the result keys or keys and values of database or information retrieval system queries, or storing the input keys or keys and values for database queries such that the queries can be performed more efficiently.
[0438] The method of storing data in the present invention uses a trie, such as a trie having the data structure described above, which trie includes nodes, where nodes, preferably each child node, is associated with a key part, and where the path from the root node in the trie to another node defines the key associated with that node, and where the key is the concatenation of the key parts associated with the nodes on the path. To achieve performance improvement in queries involving multiple data items, two or more data items are encoded in the key, and at least one or two, preferably each of the data items, consists of two or more components. The key includes two or more consecutive sections, and at least one or two, preferably each of the sections, includes two or more components of the data items encoded in the key. An "item" is sometimes referred to herein as a "dimension", and it may correspond to the so-called "key part" or "content information of the key part" above.
[0439] Fig. 22 An example of a two-dimensional key with values (X = 12, Y = 45) is shown, that is, a key encoding two data items X and Y. Both of these data items consist of two components, namely "1" and "2" in the case of item X, and "4" and "5" in the case of item Y. The key includes two consecutive sections S1 and S2. Both sections include exactly one component of each of the data items x and y encoded in the key: section S1 includes the first component "1" of data item X and the first component "4" of data item y; section S2 includes the second component "2" of data item X and the second component "5" of data item Y. It can be said that in the preferred embodiment, the key encodes multiple data items in an interleaved manner, such as X1Y1X2Y2.
[0440] According to a preferred embodiment, the encoding of the key is such that each section of the key, preferably each of the sections, contains at least and / or at most one component from each of the data items encoded in the key. For example, Fig. 22 both sections S1 and S2 of the key shown in contain exactly one component of each of the data items X = 12 and Y = 45 encoded in the key.
[0441] Furthermore, for two or more sections of the key, preferably all sections, the components belonging to different data items are sorted in the same order within that section. For example, in Fig. 22 the two sections of the key shown, the components are arranged in such a sequence: the components of data item X are ranked first, and the components of data item Y are ranked second.
[0442] Furthermore, the order of the sections including the components of the data items preferably corresponds to the order of the components within the data items. For example, in Fig. 22Among the keys shown, section S1 is located before section S2, which corresponds to the order of the components they include, in their respective entries: S1 includes the component "1" of entry X, which is before the component "2" of entry X included by section S2. S1 also includes the component "4" of entry Y, which is before the component "5" of entry Y included by section S2.
[0443] Two or more data entries of the key, preferably all data entries, have the same number of components. For example, in Fig. 22 the entries X and Y encoded in the key shown both have two components. However, the data entries of the key can also have different numbers of components. For example, it may be the case where one entry is a 64-bit integer and the other entry is a 32-bit integer. In the case where the key encodes the data entries in a strictly regular interleaved manner, such as X1Y1X2Y2, the data entry with the fewer number of components can be padded, for example resulting in X1Y1X2Y2X3*X4*. Alternatively, the interleaving method may have to be modified, for example resulting in X1Y1X2Y2X3X4.
[0444] Fig.23 An example of how to store the key shown in Fig. 22 in a trie is given. In a preferred embodiment, the key portion associated with a child node, preferably each of the child nodes, corresponds to a component of the data entry. In other words, the components of the data entry, preferably each data entry, preferably each component, correspond to the key portion associated with a child node of the trie. In the example of Fig.23 , the second-level node 2302 is associated with the key portion 1 corresponding to the first component of entry X, the third-level node 2303 is associated with the key portion 4 corresponding to the first component of entry Y, the fourth-level node 2304 is associated with the key portion 2 corresponding to the second component of entry X, and the fifth-level node 2305 is associated with the key portion 5 corresponding to the second component of entry Y. Although this is not preferred, the key portion associated with a child node can also only correspond to a part of the component of the data entry, or correspond to more than one component of the data entry.
[0445] In Fig. 22 and 23 's example, the data entries X and Y are 2-digit decimal numbers, and the components are decimal digits. According to the present invention, other examples of data entries encoded in the key stored by the trie are geographical location data, such as longitude or latitude, indices, any kind of numbers or data types, such as integers, long integers or double-long integers, 32-bit integers or 64-bit integers, strings, byte arrays, or combinations of two or more of these.
[0446] In the case where the data entry is a number, the components of the data entry can be digits (similar to Fig. 22 and Fig.23 In the example). In the case where the data item is a string, the component of the data item can be a single character. In the case where the data item is a byte array, the component of the data item can be a single byte.
[0447] However, in the preferred embodiment, the component of the data item is a group of bits of the binary encoding of the data item, and this group of bits preferably includes 6 bits. This is because, as explained above, in the preferred embodiment, the value of the key part of the child node is determined by the value of the bit set in the bitmap included in the parent node, depending on which bit the child node is associated with. Therefore, the size of the bitmap defines the possible alphabet of the key part. For example, in the case where the size of each bitmap is 64 bits, the number of different values available for the key part of the node is 2 6 . This means that a group of bits of the binary encoding including 6-bit data items can be represented by the key part associated with the node. In the case of using a 32-bit bitmap, a group including 5 bits can be represented, and so on.
[0448] For example, in the case where the data item is a 64-bit long integer and each component is a 6-bit group of the binary encoding of the integer, the data item has 64 / 6 = 11 components. In the case where the data item is a character encoded in Unicode, it may have 2 to 4 6-bit components, as explained above. In the case where the data item is a string composed of several characters, the components in the preferred embodiment are still 6-bit groups, that is, similar to any other data item, the string has the same type of components (6-bit groups). The number of components of the string corresponds to the number of components of a single character multiplied by the number of characters in the string.
[0449] Instead of regarding the components of the preferred embodiment as groups of bits, such as 6-bit groups, they can also be regarded as digits with a predefined radix or base, such as 64.
[0450] As will become apparent from the description of the range queries with reference to Fig.45 and Fig.46 as well as Figure 61 to Figure 65 , the interleaved way of storing keys with multiple data items can greatly improve the performance of range queries involving multiple data items. In addition, those skilled in the art will understand that interleaved storage can improve the performance of bidirectional search and search involving the "NOT" operator, which involves multiple data items.
[0451] Set operations
[0452] Embodiments of the present invention provide a time-saving method for performing queries in a database or information retrieval system, including operations such as intersection (Boolean AND), union (Boolean OR), difference (Boolean AND NOT), and exclusive OR (Boolean XOR) on two or more sets of keys stored in the database or information retrieval system, or the result key set of a database or information retrieval system query. These operations are herein referred to as "set operations" or "logical operations".
[0453] Still, most databases use the Volcano processing model, which means "one tuple at a time". However, for modern CPU architectures that take into account multi-level caches and in-memory databases, this is not efficient. Since all the operators in the physical execution plan run tightly interleaved, the combined instructions of the operators may take up too much space to fit into the instruction cache, and the combined state of the operators may take up too much space to fit into the data cache. Therefore, some database applications use a one-operator-at-a-time model or a combination of both, a vectorized execution model. The index data structure and unified step-by-step processing model according to the present invention result in a very lean instruction footprint regarding accessing the index trie and operator implementation.
[0454] Fig.24 An illustrative flowchart showing the various stages of database query processing according to the prior art is shown. The query is parsed 2401, rewritten 2402, and optimized 2403, and a query execution plan QEP is prepared and refined 2404 so that a query execution engine (QEE 2405) can execute the QEP generated by the previous steps on a database 2406.
[0455] Fig.25 An exemplary QEP used in a prior art database is shown. The QEP is represented by a tree, where the parent nodes 2501, 2502 are operators and the leaf nodes 2503, 2504 are data sources. In a first operation, the result from a table scan 2504 is sorted by a sorting operator 2502. In a second operation, the sorted result is merged and joined with the result of an index scan 2503.
[0456] Fig.26 An operator control and data flow in an iterator-based execution model of a prior art database is shown, where the operator implements the following methods: Open (prepare the operator to produce data), Next (produce a new data unit according to the needs of the operator's consumers), and Close (complete the execution and release resources). Each call to the Next method produces a new tuple.
[0457] The iterator-based execution model provides a unified model for operators, which is independent of the data model of the data source (database table or database index) and unifies the intermediate results of operator nodes. However, each operator call delivers only one data unit, such as a record. This approach is not efficient for operators that combine large sub-result sets, where the sub-result sets themselves return smaller result sets.
[0458] Tuples can be passed to other operators, such as Fig. 27 shown in, which shows a single tuple processing according to the query execution model of the prior art database. Starting from the root operator 2701, calls to next() will propagate to its operator children 2702, 2703, and so on, until the data source (and the leaf node of the tree representation) is reached. In this way, in the query execution plan operator tree, control flows from the consumer down to the producer, and data flows from the producer up to the consumer.
[0459] The single tuple processing model has smaller intermediate results and thus lower memory requirements. The operators in the execution plan run closely interleaved and can be very complex. However, a large number of function calls and the combined state of all operators result in a large amount of function call overhead and instruction and data cache misses because their footprint is usually too large to fit into the CPU cache.
[0460] Both the iterator-based execution model and the single tuple processing model require a complex query optimizer to be efficient.
[0461] Fig.28 Shows a single operator processing according to the query execution model of the prior art database, which immediately returns all tuples as the result of the operator in a suitable data structure such as a list. The single operator approach is cache-efficient and has tight loops with low function call overhead, but each operator produces a large intermediate result, which may not fit into the data cache or even into the main memory, which may render this approach useless.
[0462] The present invention solves the problems of the prior art with a novel execution model in which all data sources and preferably all operators are tries. Two or more input tries are combined according to the corresponding logical operations (set operations) to obtain a set of keys associated with the nodes of the corresponding result trie.
[0463] Then, the database query provides as output a representation of the resulting trie, a set of keys associated with the nodes of the resulting trie, or a subset of the keys associated with the nodes of the resulting trie, in particular the keys associated with the leaf nodes of the resulting trie, or a set of keys or values derived from the keys associated with the nodes of the resulting trie. Alternatively, it can provide other data items associated with the nodes of the resulting trie, such as document identifiers. The set of keys provided as output can be provided in the trie, for example in a physical trie data structure. It should be noted that the concept of "resulting trie" is used herein to define the set of keys that need to be obtained when combining input tries using logical operations. However, during the combination of the input tries, it is not necessarily required to form the resulting trie in a physical trie data structure, and the output set of keys can also be provided, for example, by a cursor or an iterator.
[0464] In general, the combination of two or more input tries using logical operations results in the following (resulting) trie, which stores the set of keys obtained when combining the corresponding sets of keys stored in the respective input tries using logical operations. In particular, if the logical operation is the difference set (AND NOT), then the parent node in the resulting trie is the parent node in the first input trie, and the leaf nodes of the parent node in the resulting trie are the AND NOT combination of the set of children of the corresponding parent node in the first input trie and the set of children of the corresponding parent node in the other input tries, if any. If the logical operation is not the difference set, for example if the logical operation is the intersection (AND), union (OR), or exclusive or (XOR), then the set of children of each node in the resulting trie is the combination of the logical operations (such as AND, OR, or XOR) using the sets of children of the corresponding nodes in the input tries. In this case, since each child node is associated with a key part, and the path from the root node in the trie to another node defines the key associated with that node, which is the concatenation of the key parts associated with the nodes on the path, two or more nodes from different tries "correspond" to each other if the keys associated with them are the same.
[0465] In a preferred embodiment, for its consumers, each operator itself appears again as a trie (including the case where the consumer is another operator using logical operations) until the root operator node in the query execution plan tree, for example, by implementing the corresponding trie (node) interface. Such a trie node interface represents a trie node and provides methods or functions for querying or traversing the child nodes of the trie node it represents, and the representation of the trie represents the root node of the trie. Thus, instead of using an iterator interface, a node interface can be used. This allows for functional composition and a simpler and more concise software architecture, and it further improves the performance of the database engine because the lower implementation complexity directly results in fewer function call overheads and fewer data and instruction cache misses.
[0466] Preferably, the data structure for implementing the trie is as described above. In particular, it is advantageous if the nodes in the input trie, preferably at least all the parent nodes in the input trie, include bitmaps, and the value of the key part of the child nodes in the trie is determined by the value of the bits (set) in the bitmap included in the parent node (depending on which bit the child node is associated with). Generally, the set of child nodes of the nodes of the result trie is determined by combining the sets of child nodes of the nodes of the input trie using logical operations. In such an embodiment, the combination of the sets of child nodes of the corresponding nodes of the input trie can be easily performed by bitwise combining the bitmaps of each node of the corresponding nodes of the input trie using logical operations such as bitwise AND, bitwise OR, bitwise AND NOT, or bitwise XOR. A combined bitmap is obtained, and the result of the combination is performed based on the combined bitmap. In other words, the child nodes of the nodes of the result trie are determined (evaluated) preferably only based on the combined bitmap. As a result, the step of combining the bitmap is performed before determining the child nodes of the nodes of the result trie and before the evaluation of the result trie traverses to the child nodes of the nodes of the result trie.
[0467] Thus, the physical algebra in the implementation of the trie directly corresponds to the logical algebra of set operations. While in the prior art, bitmaps are only used in the trie to reduce the memory space required for pointers, the present invention utilizes bitmaps to perform set operations on the trie.
[0468] As mentioned above, an exemplary embodiment of the trie node interface named "CDBINode" has the following main methods: getBitSet() – returns a bitmap that has bits set for all non-empty child pointers of the trie node; and getChildNode(bitNum) – returns the child node of the given node branch specified by the bit number.
[0469] Fig.29 shows the trie control and data flow of the query execution model according to the present invention. Operator 2903 performs a set operation on two input basic data sets 2901, 2902, such as intersection, union, difference, or exclusive OR. Each of the two sets of basic data is provided in an input trie, which is stored, for example, in a physical trie data structure, and the two input tries are combined by operator 2901 according to the corresponding set operation. By implementing the same trie (node) interface as the input tries, for its users, operator 2903 appears as a trie itself. Executor 2904 calls the above-mentioned trie node methods getBitSet and getChildNode to traverse the result provided by operator 2903, as indicated by the dashed arrow between the executor and the operator. When performing the set operation, operator 2903 calls the same methods getBitSet and getChildNode to traverse input tries 2901, 2902, as indicated by the dashed arrow between the operator and the input tries. In the data flow direction indicated by the solid arrow, the bitmap and sub-trie nodes are passed from input tries 2901, 2902 to operator 2903 and from operator 2903 to executor 2904.
[0470] As will be understood, by using the same or different logical operators, one or more of the input tries for higher-order set operations can be the output of another (lower-order) set operation on a trie. As described above, such an output of a lower-order set operation on a trie can be provided as a physical trie data structure. However, in a preferred embodiment, the input trie (which is the output of a lower-order set operation) for a higher-order set operation is provided as a virtual trie. A virtual trie can be understood as being able to evaluate the lower-order operators of a trie to be used in a higher-order operation. In this case, at the moment when the higher-order operator receives a representation of its input trie and at least one input trie is provided as the output of a lower-order operator, the physical data structure of that input trie does not exist as a whole and may in fact not be implementable at all.
[0471] In Fig.29AThis is shown in which the high-order operator 2906 performs a high-order set operation on, for example, an input basic data set 2905 provided in a physical Trie data structure and the output of the low-order operator 2903. The high-order operator 2905 evaluates at least part of the high-order result trie, which is a combination of the input basic data 2905 and the output of the low-order operator 2903 using the high-order set operation. The low-order operator 2903 evaluates at least part of the low-order result trie, which is a combination of two input basic data sets 2901, 2902 provided in a physical trie data structure, for example, using a low-order set operation. The low-order set operation can be the same set operation as the high-order set operation or a different set operation. As used herein, the expression "part of a trie" may refer to the nodes and / or branches of a trie.
[0472] At least part of the low-order result trie is calculated only in the step of evaluating at least part of the high-order result trie. The part of the low-order result trie is evaluated according to the requirements of the high-order operator, so the part of the high-order result trie and the part of the low-order input result trie are evaluated in an interleaved manner. Therefore, at least part of the low-order result trie is calculated only after at least part of the high-order result trie has been calculated.
[0473] It can be understood that when the high-order operator 2905 evaluates the high-order result trie, it may not be necessary to access all parts of the low-order result trie. Therefore, at least some (preferably all) of the parts of the low-order result trie that are not required for evaluating the high-order result trie are not evaluated. However, clearly at least those parts of the low-order result trie that are required for evaluating the high-order result trie need to be evaluated.
[0474] Fig.30 An example of applying an intersection (Boolean AND) operator 3003 to two input tries 3001, 3002 is shown. For readability, the input tries are in octal, but in a preferred embodiment, the trie data structure described above is used. The input trie 3001 includes three leaf nodes associated with the keys "13", "14", and "55". The input trie 3002 also includes three leaf nodes associated with the keys "13", "15", and "64". The trie 3005 depicts the result trie obtained by the AND combination of the input tries 3001, 3002.
[0475] As can be observed, the set of child nodes of each node in the resulting trie is the AND combination of the sets of child nodes of the corresponding nodes in the input tries. For example, the root node of the resulting trie 3005 has one child node associated with the key "1". This one child node is obtained when taking the intersection of the set of child nodes of the root node of the input trie 3001 ("1", "5") and the set of child nodes of the root node of the input trie 3002 ("1", "6"). Additionally, the node associated with the key "1" also has a child node associated with the key "13". In fact, the node "13" is obtained when taking the intersection of the set of child nodes of the node "1" in the input trie 3001 ("13", "14") and the set of child nodes of the corresponding node "1" in the input trie 3002 ("13", "15").
[0476] Thus, the resulting trie of the intersection operation (here the resulting trie 3005) includes all nodes and only includes the nodes included by each of the input tries (here the input tries 3001, 3002). In particular, the set of leaf nodes of the resulting trie (here the node associated with the key "13") includes all leaf nodes and only includes the leaf nodes included by each input trie and all input tries.
[0477] The algorithm performed by the preferred embodiment of the intersection operator 3003 can be described in pseudocode as follows:
[0478] 1. nodeA = the root node of trie A
[0479] 2. nodeB = the root node of trie B
[0480] 3. getBitSet of nodeA -> 00100010
[0481] 4. getBitSet of nodeB -> 01000010
[0482] 5. Bitwise AND -> 00100010
[0483] 6. For all set bits
[0484] nodeA = getChildNode of nodeA
[0485] nodeB = getChildNode of nodeB
[0486] If leaf node
[0487] Perform bitwise AND
[0488] Otherwise
[0489] Recursion (Step 3)
[0490] In this preferred embodiment, all trie nodes include the bitmap as described above. Further, the trie is formed by nodes implementing an interface including the getBitSet and getChildNode methods as described above. A bitwise AND operation is performed between the bitmaps of corresponding nodes of two input tries to determine the set of child nodes common to the two corresponding nodes.
[0491] As will be referred to below Figure 31 to Figure 34 As will be shown, the solution of the present invention utilizes a hierarchical trie structure. Since the trie is processed level by level, it allows for lazy evaluation. Thus, the performance of set operations can be significantly improved.
[0492] Fig.31 Two example input tries 3110, 3120 are shown on which an intersection operation will be performed. Similar to all tries, the input tries 3110 and 3120 have one (root) node at level 1. The input trie 3110 has two nodes at level 2, as indicated by the fact that the fourth and seventh bits are set in the bitmap 3111 included in the root node of the input trie 3110. The input trie 3120 has one node at level 2, indicated by the seventh bit set in the bitmap 3121 included in the root node of the input trie 3120. The two input tries have further sub-tries at level 3 or deeper levels, as indicated by the triangles depending on the bits set in the corresponding bitmaps of the nodes at level 2.
[0493] Fig.32 A bitwise AND operation performed on the bitmaps of corresponding nodes of the input tries 3110 and 3120 at levels 1 and 2 is shown. The root nodes at level 1 of the input tries always correspond to each other, and thus their bitmaps are combined with the first bitwise AND operation. Since the fourth and seventh bits are set in the bitmap 3111 of the root node of the input node 3110, and (only) the seventh bit is set in the bitmap 3121 of the root node of the input node 3120, (only) the seventh bit is set in the combined bitmap 3201. At level 2, the node depending on the seventh bit of the bitmap 3111 of the root node of the input trie 3110 corresponds to the node depending on the seventh bit of the bitmap 3121 of the root node of the input trie 3120, while the node depending on the second bit of the bitmap 3111 of the root node of the input trie 3110 has no corresponding node in the input trie 3120. Thus, (only) the bitmaps 3113 and 3122 are combined using the bitwise AND operation at level 2, and the third and sixth bits are set in the combined bitmap 3202.
[0494] Fig.33 shows branch skipping during the intersection operation of input tries 3110 and 3120. Since performing a bitwise AND operation on the bitmaps 3111 and 3121 of the root nodes of the two input tries results in a bitmap 3201 in which only the seventh bit is set, depending on the fourth bit of the bitmap 3111 of the root node of input trie 3110, the intersection operation does not need to traverse the branches of input trie 3110. This is indicated by the "X" in the fourth position of the bitmap 3111 and the dashed line used to draw the skipped branches in Fig.33 which. Similarly, at level 2, the seventh bit is set only in the bitmap 3113 of the node in input trie 3110 and not in the corresponding bitmap 3122 of the node in input trie 3120. Therefore, depending on the seventh bit of the bitmap 3113, the intersection operation does not need to traverse this branch.
[0495] Finally, Fig.34 shows the resulting trie of the intersection operation on tries 3110 and 3120. The bitmaps 3201 and 3202 associated with the nodes of the resulting trie at levels 1 and 2 correspond respectively to the bitmaps of the combinations calculated by the Fig.32 bitwise AND operation shown in.
[0496] Figure 31 to Figure 34 The example of shows that, generally speaking, during the process of set operations on combined input tries, it is not necessary to traverse some branches of the input tries, which can significantly improve performance. The combined bitmaps can be used to determine which branches need to be further traversed and which branches can be skipped. This method can be called "result prediction" or "tree pruning".
[0497] Fig.35 shows an example of applying the union (Boolean OR) operator 3503 to two input tries 3001, 3002 of Fig.30 . The illustration of the resulting trie 3605 obtained by the OR combination of input tries 3001, 3002 is shown in Fig.36 which.
[0498] As can be observed, the set of child nodes of each node in the resulting trie is the OR combination of the sets of child nodes of the corresponding node in the input trie. For example, the root node of the resulting trie 3605 has three child nodes, which are associated with the keys "1", "5", and "6". These three child nodes are obtained when forming the union of the set of child nodes of the root node of the input trie 3001 ("1", "5") and the set of child nodes of the root node of the input trie 3002 ("1", "6"). The node associated with the key "1" also has three child nodes, which are associated with the keys "13", "14", and "15". In fact, these nodes are obtained when forming the union of the set of child nodes of node "1" in the input trie 3001 ("13", "14") and the set of child nodes of the corresponding node "1" in the input trie 3002 ("13", "15"). Finally, each of the nodes in the resulting trie associated with the keys "5" and "6" respectively has one child node associated with the keys "55" and "64" respectively, which are the child nodes of the corresponding nodes in the input tries 3001 and 3002.
[0499] Thus, the resulting trie of the union operation (here it is the resulting trie 3605) includes all the nodes included in any of the input tries, here the input tries 3001, 3002. In particular, the set of leaf nodes of the resulting trie (here the nodes associated with "13", "14", "15", "55", and "64") includes all the leaf nodes included by any of the input tries.
[0500] The algorithm performed by the preferred embodiment of the union operator 3503 can be described in pseudocode as follows:
[0501] 1. nodeA = the root node of trie A
[0502] 2. nodeB = the root node of trie B
[0503] 3. getBitSet of nodeA -> 00100010
[0504] 4. getBitSet of nodeB -> 01000010
[0505] 5. bitwise OR -> 01100010
[0506] 6. For all set bits
[0507] If the bit is set in both nodeA and nodeB
[0508] nodeA = getChildNode of nodeA
[0509] nodeB = getChildNode of nodeB
[0510] If leaf node
[0511] Perform bitwise OR
[0512] Otherwise
[0513] Recursion (step 3)
[0514] If the bit is set only in nodeA
[0515] nodeA = getChildNode of nodeA
[0516] Recursively call trieA only (skip bitwise OR)
[0517] If the bit is set only in nodeB
[0518] nodeB = getChildNode of nodeB
[0519] Recursively call trieB only (skip bitwise OR)
[0520] Similarly, all trie nodes include bitmaps as described above, and the trie is formed by nodes implementing an interface including the getBitSet and getChildNode methods as described above. A bitwise OR operation is performed between the bitmaps of the corresponding nodes of the two input tries to determine the set of child nodes included in any of the two corresponding nodes. If the bit is set only in the bitmap of one of the two corresponding nodes, the child trie depending on that one node is added to the resulting trie, which is indicated by "Recursively call trieA only" / "Recursively call trieB only" in the pseudocode above.
[0521] Fig.37 Shows an example of applying the difference set (Boolean AND NOT) operator 3703 to two input tries 3001, 3002. Trie 3705 is an illustration of the resulting trie obtained from the AND NOT combination of the input tries 3001, 3002. Fig.30 As can be observed, all the parent nodes of the resulting trie 3705 correspond to the parent nodes of the first input trie 3001. The leaf nodes depending on the parent nodes of the resulting trie 3705 are the AND NOT combination of the set of child nodes of the corresponding parent nodes in the first input trie 3001 and the set of child nodes of any corresponding parent nodes in the input trie 3002.
[0522]
[0523] For example, the root node of the resulting trie 3705 has two child nodes associated with the keys "1" and "5", which are themselves parent nodes. These two nodes correspond to the two child nodes of the root node of the input trie 3002, which are themselves parent nodes. The node associated with the key "1" has a child node, which is a leaf node and is associated with the key "14". This leaf node is obtained when forming the difference set of the set of child nodes ("13", "14") of the node "1" of the first input trie 3001 and the set of child nodes ("13", "15") of the corresponding node "1" of the input trie 3002. Finally, the node in the resulting trie associated with the key "5" has a child node, which is associated with the key "55" and corresponds to the child node of the node with the key "5" of the first input trie 3001. This node with the key "5" has no corresponding node in the input trie 3002.
[0524] Thus, the resulting trie of the difference set operation (here the resulting trie 3605) includes all the parent nodes included in the first input trie (here the input trie 3001). The set of leaf nodes of the resulting trie (here the nodes associated with the keys "14" and "55") includes all the leaf nodes of the first input trie (here the trie 3001) minus the leaf nodes of the second input trie (here the trie 3002).
[0525] The algorithm performed by the preferred embodiment of the difference set operator 3703 can be described in pseudocode as follows:
[0526] 1. nodeA = the root node of trie A
[0527] 2. nodeB = the root node of trie B
[0528] 3. getBitSet of nodeA -> 00100010
[0529] 4. getBitSet of nodeB -> 01000010
[0530] 5. bit set of nodeA -> 00100010
[0531] 6. For all set bits
[0532] If the bit is set in both nodeA and nodeB
[0533] nodeA = getChildNode of nodeA
[0534] nodeB = getChildNode of nodeB
[0535] If a leaf node
[0536] Perform bitwise NAND
[0537] Otherwise
[0538] Recursion (step 3)
[0539] If a bit is set only in nodeA
[0540] nodeA = getChildNode of nodeA
[0541] Recurse only trieA (skip bitwise NAND)
[0542] Similarly, all trie nodes form and implement the interfaces as described above. If a bit is set in the bitmaps of the corresponding nodes of input tries 3001 and 3002, recursion exists on both tries. If a bit is set only in the bitmap of the node of the first input trie 3001, the child trie of that node is added to the result trie, which is indicated by "Recurse only trieA" in the above pseudocode. Bits set only in the bitmap of the node of trie 3002 are ignored. If the child nodes of the corresponding nodes of the two input tries are leaf nodes, only a bitwise AND NOT operation is performed between the bitmaps of the corresponding nodes of the two input tries.
[0543] The execution of the operator includes a recursive descent at the trie level (in the preferred embodiment, each level is a digit of radix / base 64). At each level, the bitmap of each node is used as a result prediction and then iterated through the prediction bits. Thus, evaluating the result trie includes performing a combination function on the root node of the result trie. Performing the combination function for an input node of the result trie includes: determining the set of child nodes of the input node of the result trie (which may also be empty) by combining the sets of child nodes of the nodes of the input tries corresponding to the input node of the result trie; using logical operations; and performing the combination function for each of the determined child nodes of the input node of the result trie. As already mentioned above, the root node and / or input nodes of the result trie do not have to be physically generated.
[0544] The steps of evaluating a result trie can be performed using depth - first traversal, breadth - first traversal, or a combination thereof. Evaluating a result trie in depth - first traversal includes: performing a combination function on one of the child nodes of an input node and traversing the sub - trie formed by that child node before performing the combination function on the next sibling node of that child node. Evaluating a result trie in breadth - first traversal includes: performing a combination function on each of the child nodes determined for the input node of the result trie; and determining a set of child nodes for each of the child nodes determined for the input node of the result trie before performing the combination function on any grand - child nodes of the input node of the result trie.
[0545] One or more of the input tries for set operations can be virtual tries, i.e., tries that are dynamically generated as needed during the operation of combining input tries. Typically, only those parts of the virtual trie are dynamically generated that are required to combine the input tries using logical operations. There are several scenarios for applications of virtual input tries. One is the implementation of database - range queries, which will be described later. Another scenario is that one or more of the input tries for higher - order set operations are the output of lower - order set operations on tries, as described above. Fig.29A which was described above.
[0546] A virtual trie is a trie whose physical data structure does not exist as a whole when, for example, an input trie that is a set operator receives the virtual trie. A virtual trie can be understood as an operator capable of evaluating a trie. Preferably, the operator represents the root node of the trie and implements a trie - node interface, where the trie - node interface represents a trie node and provides methods or functions for querying or traversing the child nodes of the trie node it represents.
[0547] "Evaluating" a trie or a part of a trie can also be referred to as "rendering" or "composing" the trie or the part of the trie, and generally includes traversing the trie or the part of the trie and providing the following information: the nodes traversed exist in the trie. It can include implementing the physical data structure of the trie or the part of the trie as a whole. However, in a preferred embodiment, at any given moment, less than the entire physical data structure of the trie or the part of the trie is implemented, and preferably, only those parts of the physical data structure of the trie or the part of the trie that need to be implemented at that moment for traversing the trie or the part of the trie are implemented. Obviously, at any given moment during the evaluation of a trie or a part of a trie, at least those parts of the physical data structure of the trie or the part of the trie that need to be implemented at that moment for traversing the trie or the part of the trie are implemented.
[0548] In a preferred embodiment, in the case where the virtual trie implements the trie (node) interface, as described above, the methods getBitSet and getChildNode are used to evaluate the trie or a part thereof. For example, during a higher-order intersection operation where one of the input tries is a virtual trie, the higher-order intersection operator accesses and the virtual range trie delivers, "on the fly" via an application programming interface (API), the components of the virtual trie required to evaluate the higher-order result trie. For example, the bitmap of the current node of the virtual trie can be accessed via the getBitSet() method introduced above, while the child node can be accessed via the getChildNode() method described above. The API returns, on the one hand, the corresponding bitmap and, on the other hand, an object (instead of a child node of a real trie) that will provide the corresponding bitmap for the next recursion. For the higher-order operator, the virtual trie appears just like a trie implemented in a physical trie data structure.
[0549] Generally, a representation of the virtual trie is provided to an operator that uses the virtual trie as an input for some operation. During the step of performing the operation, at least a part of the virtual trie is evaluated (on demand). In cases where not all parts of the virtual trie need to be accessed to perform the operation, at least some of the parts that are not required to perform the operation are not evaluated, and preferably, only those parts of the virtual trie that are required to perform the operation are evaluated. This is for Fig.29A the case shown in Fig.37A and Fig.37B which is illustrated where the virtual trie is a lower-order set operator 2903 and a representation of this virtual trie is provided as an input trie to a higher-order set operator 2906 that uses the virtual trie to evaluate at least a part of the higher-order result trie. The higher-order result trie is a combination of the input trie (i.e., the virtual trie) of the set operator 2906 and the basic data 2905 using logical operations. In Fig.37A and Fig.37B the case, the higher-order set operation is an AND operation and the lower-order set operation is an OR operation.
[0550] Fig.37AThree tries A, B, and C are shown. In the preferred embodiment, as described above, the trie is base 32, base 64. Thus, the bitmap referring to the pointers to the child nodes can include 32 or 64 bits. For simplicity, this example shows a quaternary trie. The tries are represented by the trie root node 3701 for trie A, 3702 for trie B, and 3703 for trie C. Each trie node includes bitmaps 3704, 3705, 3706 as described above.
[0551] Fig.37B Shows how to perform the ((A or B) and C) evaluation without implementing or materializing (A or B) as an intermediate result, i.e., creating a physical trie data structure. The figure shows two levels of expressions and their evaluation. For the first level, the bitmaps of the root nodes of trie A (3704) and trie B (3705) are combined by a bitwise OR operation to obtain bitmap 3709. Bitmap 3709 is combined by a bitwise AND operation with the bitmap 3706 of the root node of trie 3706. Since there is no match with Trie C, the resulting bitmap 3710 prunes Trie B for further consideration. The resulting physical Trie data structure is represented by the root node 3712. For the second level, the bitmaps 3707 and 3708 of trie A and trie C are combined together – there is no match for trie B here – resulting in bitmap 3711.
[0552] From Fig.37A and 37 the example of B, it can be seen that the fact of providing the result of the lower-order OR operation as a virtual input trie to the higher-order AND operator instead of as a physical trie data structure for a comprehensive evaluation enables "lazy evaluation". During the evaluation of the higher-order result trie of the higher-order AND operation, the lower-order result trie of the lower-order OR operation is evaluated. Thus, at least part of the lower-order result trie is evaluated after at least part of the higher-order result trie has been evaluated. In fact, parts of the higher-order result trie and parts of the lower-order result trie are evaluated in an interleaved manner, where parts of the lower-order result trie are evaluated according to the requirements of the higher-order operator. This only allows the evaluation of those parts of the lower-order result trie that are necessary for the evaluation of the higher-order result trie.
[0553] Note that for some operations, it is not necessary to actually implement a physical result trie data structure, for example if only the count (cardinality) of the result is needed. As will also be understood by those skilled in the art, higher-order operators can evaluate only parts of a higher-order result trie, for example if the higher-order result trie is used as input to an even higher-order operator. Similarly, the input trie for a lower-order operator does not have to be a physical trie data structure, but can again be a virtual trie, for example the output of an even lower-order operator. Finally, note that an operator typically obtains a representation of the input trie, where the representation of the trie can be a handle to a physical trie data structure, like a pointer. In the case of a virtual trie, the representation can be a handle to a trie operator capable of evaluating the trie.
[0554] Range query
[0555] The combination of two or more input tries through an intersection operation (Boolean AND) can be advantageously used to perform a range query in an efficient manner. A "range" can be described as a set of discrete ordered values that includes all values between a first value and a second value of some data type, where the first value and / or the second value may or may not be included in the range. A range query returns the keys in a key set whose values correspond to (match) one or more specified value ranges.
[0556] The range query according to the present invention is performed through an intersection operation of tries, where the trie stores a key set to be searched for one or more ranges (hereinafter referred to as "input set trie"), and the trie stores all values included in one or more ranges (hereinafter referred to as "range trie"). The trie is preferably implemented as described above. The key set to be searched is typically associated with the nodes of the input set trie, and the values (keys) indicating the value ranges to be matched are typically associated with the nodes of the range trie, particularly with the leaf nodes of the range trie. The key set to be searched is typically a key set stored in a database or the result of a database query or an input key set, or a key set stored in an information retrieval system or the result of an information retrieval system query or a set of input keys.
[0557] Obtain definitions of one or more ranges for performing a query through user input or other means. This definition is used to generate a range trie, where the values associated with the nodes (typically leaf nodes) of the range trie correspond to the values included in the one or more ranges. In the next step, the input set trie is combined with the range trie through the intersection operation as described above to obtain a set of keys associated with the nodes of the result trie. Finally, the set of keys associated with the nodes of the result trie, or a subset of the keys associated with the nodes of the result trie, particularly the keys associated with the leaf nodes of the result trie, or a set of keys or values derived from the keys associated with the nodes of the result trie is obtained as output.
[0558] Fig.38 An input set trie 3801 storing a set of keys to be searched within one or more ranges and a range trie 3802 generated to store all values included within the one or more ranges are shown. Both the trie 3801 and 3802 are combined by an intersection operator 3803 to obtain the set of keys in the input set trie 3801 whose values lie within the one or more ranges stored by the range trie 3802. As described above, not only the input set trie and the range trie, but also the intersection operator itself can implement the corresponding trie (node) interface, as indicated by the three triangles in Fig.38 In the case where the trie node interface has the getBitSet() and getChildNode(bitNum) methods as presented above, the intersection operator 3803 can call these operations to traverse the input set trie 3801 and the range trie 3802, as indicated by the dashed arrows in Fig.38 In the data flow direction indicated by the solid arrows, the bitmaps and child trie nodes are passed from the set trie 3801 and the range trie 3802 to the intersection operator 3803.
[0559] Fig.39Shows an example of a method for performing a range query according to the present invention. The leaf nodes of the input set trie 3901 are associated with three keys with values of "13", "14", and "55". The range trie 3902 is generated to include all leaf nodes associated with keys whose values are within the range or [14..56]. At each level of the range trie, the nodes at the end of the range (in the example of nodes 1 and 5 of the range trie 3902) contribute partially to the result. The nodes between the ends of the range (in the example of nodes 2, 3, and 4 of the range trie 3902) and their dependent sub-tries contribute fully to the result. The AND combination of the input trie 3901 and the range trie 3902 by the intersection operator 3903 obtains the set of keys associated with the nodes of the result trie 3905. The executor 3904 can output, for example, the values of the keys associated with the leaf nodes of the result trie 3905, that is, "14" and "55". This output corresponds to the values of all keys associated with the leaf nodes of the input set trie 3901 that are within the range of [14..56].
[0560] The algorithm executed by the intersection operator 3903 can be described in pseudocode as follows:
[0561] 1. nodeA = the root node of trieA (input set trie)
[0562] 2. nodeB = the root node of trie B (range trie)
[0563] 3. getBitSet of nodeA -> 00100010
[0564] 4. getBitSet of nodeB -> 01111110
[0565] 5. Bitwise AND -> 00100010
[0566] 6. For all set bits
[0567] nodeA = getChildNode of nodeA
[0568] nodeB = getChildNode of nodeB
[0569] If leaf node
[0570] Perform bitwise AND
[0571] Otherwise
[0572] Recursion (step 3)
[0573] Range trie 3902 is an example showing that a range trie can include many nodes. Thus, materializing range trie 3902 and all its nodes would be expensive in terms of time and memory space. For this reason, a range trie can be implemented as a virtual trie, as described above, i.e., a trie that is dynamically generated as needed during an intersection operation. This is indicated by the content of the triangle of range trie 3902 drawn with a dashed line in Fig.39 During an intersection operation, the intersection operator accesses and the virtual range trie delivers "on the fly" via an application programming interface (API) the components of the virtual trie required to evaluate the resulting trie of the intersection operation. For example, the bitmap of the current node of the virtual trie can be accessed via the getBitSet() method introduced above, while the child nodes can be accessed via the getChildNode() method. The API returns on the one hand the corresponding bitmap and on the other hand an object (instead of the child nodes of a real trie) that will provide the corresponding bitmap for the next recursion. To the intersection operator, the virtual trie appears like a real trie implemented in a physical trie data structure.
[0574] Typically, at least some parts of the virtual trie that are not required to evaluate the resulting tire of the intersection operation are not evaluated, and preferably only those parts of the virtual range trie are dynamically generated that are required to combine the input set trie and the range trie via the intersection operation. In the
[0575] example of Fig.39 due to branch skipping as explained above, the virtual sub-tries associated with bits 2, 3, 4, and 5 set in the bitmap of the root node of the range trie are not accessed during the intersection operation, and for this reason, none of their nodes have ever been physically generated (rendered or integrated).
[0576] In some embodiments of the range query according to the present invention, similar to the Fig.39 embodiment shown in
[0577] a key associated with a leaf node of the input set trie encodes a data item of a specific data type. These embodiments can also be referred to as "one-dimensional" range queries. In a one-dimensional range query, the definition of one or more ranges includes the definition of one or more ranges of a data item.
[0578] In other embodiments of the range query according to the present invention, the key associated with the leaf node of the input set trie encodes two or more data items of a specific data type. In this case, the definition of one or more ranges includes the definition of one or more ranges of one or more of the data items. Such an embodiment may also be referred to as a "multidimensional" range query. Although in principle range queries can be performed on each dimension and the intersection of the results can be performed, it is not efficient because too many tries and too many operators would be involved.
[0579] Fig.40 An example of an efficient multi-item or multidimensional range query process according to the present invention is shown. The input set trie 4001 is a two-dimensional or two-item trie, where each of its leaf nodes is associated with a key that specifies a value pair (x, y). Such a value pair can be used, for example, to specify the longitude and latitude of geographical location data, or a composite key such as (longitude, ID) or (latitude, ID), as will be explained below with reference to Figure 54 to Figure 57 As explained. In this example, each of the dimensions x and y contains two digits in the octal system for better readability, but in the preferred embodiment, the trie data structure described above is used. As will become apparent from the following explanation, according to the interleaved encoding of multiple data items in one key, the x and y dimensions are stored in an interleaved manner in the input set trie 4001.
[0580] The input set trie 4001 has a root node at level 1. The bitmap associated with this root node indicates the value of the first digit of the x dimension. Bits "1" and "5" are set in the bitmap associated with the root node, which indicates that the first digit of the x dimension of the keys stored in the input set trie is "1" or "5". The bitmap associated with the nodes at level 2 indicates the value of the first digit of the y dimension. Bits "3" and "4" are set in the bitmap associated with the nodes at level 2, which depend on the bit "1" of the root node, and bit "6" is set in the bitmap associated with the nodes at level 2, which depends on the bit "5" of the root node. This indicates that the first digit of the y dimension of the keys whose x dimension starts with "1" is "3" or "4", and the first digit of the y dimension of the keys whose x dimension starts with "5" is "6". The bitmap associated with the nodes at level 3 indicates the value of the second digit of the x dimension. The bits set in these bitmaps indicate the existence of keys with x dimensions of "12" and "16" whose y dimension starts with "3", keys with x dimensions of "12" and "15" whose y dimension starts with "4", and keys with x dimensions of "56" whose y dimension starts with "6". The bitmaps of the nodes at level 4 and the nodes at level 5 are not shown in Fig.40 and are only indicated by small triangles that depend on the nodes at level 4.
[0581] A range trie for multi - item or multi - dimensional range queries can be a multi - item range trie obtained by means of a single - item or one - dimensional range trie for each combination of data items encoded by keys associated with the leaf nodes of the input set trie. The single - item range trie of the data item stores all values included in one or more ranges of the data item. As described above, the single - item range trie can be a virtual range trie. This means that only those parts of the virtual single - item trie are dynamically generated which are necessary for combining the single - item range tries to obtain the multi - item range trie or for combining the input set trie and the single - item trie by an intersection operation.
[0582] In Fig.40 the example of, two one - dimensional input ranges are obtained for a two - dimensional range query: the range [15..55] for the x - dimension and the range [30, 31] for the y - dimension. A (virtual) range trie 4002 is created which stores the range [15..15] for the x - dimension, and a (virtual) range trie 4003 is created which stores the range [30..31] for the y - dimension.
[0583] In some multi - dimensional range queries, for some data items (dimensions), a definition of the range cannot be obtained. For example, the user can only specify the range [15..55] for the x - dimension in order to perform range query processing on Fig.40 the two - dimensional input set trie 4001, but there is no definition of the range for the y - dimension. This means that the user is interested in all keys stored in the input set trie 4001 whose x - dimension values are between 15 and 55, regardless of their y - dimension values. Dimensions for which the range is not specified are skipped or ignored. This can be achieved by creating (virtual) single - item range tries even for data items for which no range definition is obtained, which store the entire range of possible values of the data item. Such a trie can also be referred to as a "wildcard trie". For example, in one embodiment, the MatchAll trie is a trie that implements the above - mentioned trie (node) interface ("CDBINode"). Invoking the getBitSet() method on any of its nodes will always return a bitmap with all bits set.
[0584] A multi - item or multi - dimensional range trie obtained from the combination of single - item or one - dimensional range tries typically stores all combinations of the values of the data items stored in the single - item (one - dimensional) range tries. For example, if the range for the x - dimension is [11..13] and the range for the y - dimension is [7..8], the combined two - dimensional range trie will store the keys for the (x, y) value pairs (11, 7), (11, 8), (12, 7), (12, 8), (13, 7), and (13, 8).
[0585] Fig.41 shows how to combine Fig.40 different parts of the one-dimensional range tries 4002 and 4003 of Fig.42 to obtain the interleaved two-dimensional range trie shown in Fig.42 As can be observed in Fig.40 , the x-dimension and y-dimension in the combined two-dimensional range trie are stored in the same interleaved manner as the input set trie of Fig.43 shows how to combine Fig.40 the one-dimensional range tries 4002 and 4003 of Fig.44 to obtain the non-interleaved two-dimensional range trie shown in
[0586] In a preferred embodiment of the present invention, and this is the case for both one-dimensional and multi-dimensional range queries, the range trie has the same structure or format as the input set trie. Thus, in the case where the multi-dimensional input set trie stores data items in an interleaved manner, the multi-dimensional range trie preferably uses interleaved storage, and in the case where the input set trie stores data items in a non-interleaved manner, the multi-dimensional range trie preferably does not use interleaved storage either. Furthermore, the keys associated with the leaf nodes of the range trie preferably encode data items of the same data type as the keys associated with the leaf nodes of the input set trie. Finally, in the range trie, data items of a certain data type or components of such data items are preferably encoded in nodes at the same level as the corresponding data items or components of the data items in the input set trie.
[0587] In some embodiments of multi-item or multi-dimensional range query processing, combining single-item or one-dimensional range tries to obtain a multi-item or multi-dimensional range is performed by the following function, which provides the multi-item range trie as input to a function (such as an intersection operator) that implements the combination of the input set trie and the multi-item range trie. This is shown in the example of Fig.64 where the combination is performed by the interleaving operator 6404.
[0588] In other embodiments, the combination of the single-item or one-dimensional range trie is performed in a function or operator that implements the intersection of the input set trie and the range trie to obtain a multi-item or multi-dimensional range trie. In this case, the multi-dimensional range trie will only exist conceptually. In practice, the function or operator that implements the combination of the input set trie and the range trie accesses the (virtual) one-dimensional range tries as if they together formed a (virtual) multi-dimensional range trie. If there are dimensions with unspecified ranges, the function or operator that implements the intersection of the input set trie and the range trie will skip or ignore these dimensions, for example by creating a wildcard trie as described above.
[0589] In Fig.40 the example of, the combination is performed by the two-dimensional intersection (AND) operator 4004. The algorithm performed by the intersection operator 4004 can be described in pseudocode as follows:
[0590] 1. nodeA = the root node of trieA (input set trie)
[0591] 2. nodeB = the root node of trie B (range of the x dimension)
[0592] 3. nodeC = the root node of trie C (range of the y dimension)
[0593] 4. getBitSet of nodeA -> 00100010
[0594] 5. getBitSet of nodeB -> 00111110
[0595] 6. Bitwise AND -> 00100010
[0596] 7. For all set bits
[0597] nodeA = the getChildNode of nodeA,
[0598] 8. getBitSet of nodeA -> 00011000
[0599] 9. getBitSet of nodeC -> 00001000
[0600] 10. Bitwise AND -> 00001000
[0601] 11. For all set bits
[0602] Get the child node of child node A,
[0603] nodeB = the getChildNode of nodeB
[0604] nodeC = getChildNode of nodeC
[0605] If leaf node
[0606] Perform bitwise AND
[0607] Otherwise
[0608] Recursion (Step 4)
[0609] Conceptually create a (virtual) multi - dimensional range trie, where, if one - dimensional range tries are actually combined to obtain a (virtual) multi - dimensional range trie, at least if the multi - dimensional range trie has the same structure or format as the input set trie, the set of child nodes of each node in the resulting trie will be the result of the AND combination of the set of child nodes of the corresponding nodes in the input set trie and the multi - dimensional range trie.
[0610] For example, Fig.40 The resulting trie 4006 is shown, which stores the x - dimension and y - dimension in an interleaved manner similar to the input set trie 4001. The resulting trie 4006 has a node 4007, which is associated with a key having a value of (x = 1, y = 3). This node represents the entire set of child nodes of node 4108, which is associated with a key having a value of x = 1. The node in the input set trie 4001 corresponding to node 4007 of the resulting trie 4006 is node 4009, which is the node associated with the key having a value of x = 1. Fig.41 Shown in Fig.40 If the one - dimensional range tries 4002 and 4003 are actually combined, the resulting multi - dimensional range trie 4100 that would be obtained is shown, and it has the same structure and format as the input set trie 4001. The node in the multi - dimensional range trie 4100 corresponding to node 4007 of the resulting trie 4006 is node 4101, which is the node associated with the key having a value of x = 1. The AND combination of the set of child nodes (nodes 4010 and 4011) of node 4009 in the input set trie 4001 and the set of child nodes (node 4102) of node 4101 in the multi - dimensional range trie 4100 results in node 4007 of the resulting trie 4006.
[0611] In many cases, storing the different dimensions or items of a multi - dimensional or multi - item input set trie in an interleaved manner will result in more efficient range queries, as will now be explained with reference to Fig.45 and Fig.46 explained.
[0612] Fig.45 Shown storing with Fig.40the input set trie 4501 has the same values as the input set trie 4001, but in a non-interleaved manner. Similar to the input set trie 4001, the input set trie 4501 has a root node at level 1, and the bitmap associated with this root node indicates the value of the first digit of the x dimension. However, the bitmap associated with the nodes at level 2 of the input set trie 4501 indicates the value of the second digit of the x dimension, rather than the value of the first digit of the y dimension in the interleaved input set trie 4001. The bitmap associated with the nodes at level 3 indicates the value of the first digit of the y dimension, and the bitmap of the nodes at level 4 ( Fig.45 not shown) indicates the value of the second digit of the y dimension.
[0613] When performing the AND combination of the input set trie 4501 with the corresponding (also non-interleaved) two-dimensional range trie according to the present invention, the nodes traversed for a two-dimensional range query with range X = [15..55] and Y = [30..31] are shaded in Fig.45 (nodes at level 5 are not shown). Although only one node (x = 16, y = 3) in level 4 is a shaded node, at the three higher levels, a total of six nodes need to be traversed (visited): the root node x = 1, x = 5, x = 15, x = 16, and x = 53.
[0614] In contrast, Fig.46 shows Fig.40 the interleaved input set trie 4001, where again the nodes traversed for a two-dimensional range query with range X = [15..55] and Y = [30..31] according to the present invention are shaded (nodes at level 5 are not shown). One node (x = 16, y = 3) in level 4, which is the same as Fig.45 before, is a shaded node, but at the three higher levels, only four nodes need to be traversed in total: the root node x = 1, x = 5, and (x = 1, y = 3).
[0615] The reason for this is that although in the non-interleaved input set trie 4501, all x values falling within the range [15..55] are determined down to the last (second) digit, in the interleaved input set trie 4001, by looking at the first digit of the y dimension, nodes that are not worth traversing can be eliminated more quickly. The chance of eliminating nodes by looking at the first digit of another dimension is higher than the chance of eliminating nodes by looking at another digit of the same dimension. As will be understood, the more digits there are in different dimensions, the higher the performance gain of interleaved storage will be.
[0616] As mentioned above, range query processing can provide, as output, the set of keys associated with the leaf nodes of the input set trie, e.g., in the case of a one-dimensional range query or if the user is interested in all dimensions of the multi-dimensional keys stored in the input set trie. Alternatively, range query processing can provide, as output, a reduced item key set that encodes a subset of the data items encoded by the keys associated with the leaf nodes of the input set trie. Fig.47 and Fig.48 An example thereof is shown in
[0617] In Fig.47 a two-dimensional or binomial trie input set trie 4701 is the same as the input set trie 4001 of Fig.40 where each of its leaf nodes is associated with a key of a specified value pair (x, y). The user in the example is interested in all x values stored in the input set trie. Thus, although not shown in detail, the one-dimensional range trie 4702 is a wildcard trie that stores the entire range of possible values for the x dimension, and the one-dimensional range trie 4703 is a wildcard trie that stores the entire range of possible values for the y dimension.
[0618] Similar to Fig.40 the one-dimensional range tries 4002 and 4003 of Fig.40 the one-dimensional range tries 4702 and 4703 are combined, at least conceptually, to obtain a binomial or two-dimensional range trie. Since both one-dimensional range tries are wildcard tries, the two-dimensional range trie contains all possible value pairs (x, y). The two-dimensional range trie is then combined with the two-dimensional input set trie 4701 by a two-dimensional intersection (AND) operator 4704. Similar to
[0619] the two-dimensional intersection (AND) operator 4004 of Fig.40 the intersection operator 4704 combines the two-dimensional input set trie with the two-dimensional range trie according to the intersection operation to obtain the set of keys associated with the nodes of the corresponding result trie. Further, as explained above, the set of child nodes of each node in the result trie is the AND combination of the set of child nodes of the corresponding node in the two-dimensional input set trie and the set of child nodes of the corresponding node in the two-dimensional range trie. Since the two-dimensional range trie stores all possible value pairs (x, y), the result trie is the same as the input set trie.
[0619] Compared with Fig.40 the intersection operator 4004 of Fig.47The intersection operator 4704 does not provide the key set (x, y) associated with the leaf nodes of the input set trie as output. Instead, the intersection operator 4704 only provides the set of values for the x dimension as output, in this example, the set of all x values stored in the input set trie.
[0620] In the case where the range query processing provides a simplified item key set as output, similar to in Fig.47 the simplified item key set obtained from different branches of the input set trie related to data items not encoded in the simplified item key may contain duplicates. For example, as can be seen in Fig.47 there are at least two leaf nodes in the input set trie 4701 with their x value being 12, namely at least one leaf node with the first digit of the y value being 3 (12, 3…) and at least one leaf node with the first digit of the y value being 4 (12, 4…). To eliminate duplicate keys before providing the output, especially if the output set is used in an upstream operator, the simplified item key sets obtained from different branches of the input set trie related to data items not encoded in the simplified item key are merged before providing the output. For two dimensions, the merge can be performed on each level of the input set trie in an acceptable and efficient manner. However, it is usually more efficient to write the simplified item key set into a newly created trie (for example, where the input set trie has more than two dimensions), thus automatically eliminating duplicates.
[0621] In Fig.47 the example, to eliminate duplicate x values in the output set, the key set of x values obtained by combining the input set trie 4701 with a two-dimensional range trie through an intersection operation is written into a newly created one-dimensional trie, whose nodes are associated with keys having only x values. This newly created one-dimensional trie is shown in Fig.48 The values of the keys associated with its leaf nodes correspond to the set of all x values stored in the input set trie 4701.
[0622] The algorithm performed by the two-dimensional intersection operator 4004 that only outputs x values can be described in pseudocode as follows:
[0623] 1. nodeA = the root node of trieA (input set trie)
[0624] 2. nodeB = the root node of trie B (wildcard trie for the x dimension)
[0625] 3. nodeC = the root node of trie C (wildcard trie for the y dimension)
[0626] 4. getBitSet of nodeA -> 00100010
[0627] 5. getBitSet of nodeB -> 11111111
[0628] 6. Bitwise AND -> 00100010
[0629] 7. For all set bits
[0630] nodeA = getChildNode of nodeA,
[0631] 8. getBitSet of nodeA -> 00011000
[0632] 9. getBitSet of nodeC -> 11111111
[0633] 10. Bitwise AND -> 00011000
[0634] 11. For all set bits
[0635] Get the child node of child nodeA,
[0636] nodeB = getChildNode of nodeB
[0637] nodeC = getChildNode of nodeC
[0638] If leaf node
[0639] Execute bitwise AND
[0640] Write the result key (only x) to the output trie
[0641] Otherwise
[0642] Recursion (step 4)
[0643] In Fig.47 's example, only the value of the x dimension is provided as the output. As will be understood, when only the value of the y dimension is provided as the output, the same principle applies. In addition, since the one-dimensional range tries 4702 and 4703 are both wildcard tries, all x values stored in the input set trie 4701 are provided as the output. As will be understood, in the case where the one-dimensional range tries 4702 and / or 4703 store only a subset of all possible x and / or y values, only a subset of all x values stored in the input set trie 4701 can be provided as the output. For example, if the one-dimensional range trie 4702 is the same as Fig.40 's range trie 4002 and only x values in the range [15..55] are stored, the output provided by the intersection operator 4004 will not include the value of x = 12. As another example, if both one-dimensional range tries 4702 and 4703 are the same as Fig.40 If the corresponding range tries 4002 and 4003 are the same, storing only x values in the range [15..55] and accordingly storing only y values in the range [30..31], then the output provided by the intersection operator 4004 will include only the value x = 16.
[0644] Fuzzy search
[0645] A common requirement of text retrieval applications is to provide approximate string matching - also known as a fuzzy search capability. That is, to find strings that approximately match a pattern rather than exactly match it.
[0646] A typical measure of this "fuzziness" (the difference between two character sequences) is the Levenshtein distance. The Levenshtein distance between two strings is the minimum number of single-character edits (character insertions, deletions, or substitutions) required to change one string into the other.
[0647] Similar to the virtual range tries discussed above, one aspect of the present invention is directed to a preferred virtual fuzzy match trie that uses a Boolean AND to intersect with a storage trie such as an index trie to return a matching key string or a document including the matching key string. As described above, the index trie can store each occurrence of a term and a document ID as two key parts (string, long).
[0648] According to this aspect of the present invention, data is retrieved from an electronic database or an information retrieval system by performing approximate string matching. First, a search string is obtained. Next, a match trie storing approximate strings including the search string and / or variants of the search string is built. The match trie is combined with the storage trie using an intersection operation, where the storage trie stores the set of strings stored in the electronic database or information retrieval system or the set of result strings queried by the electronic database or information retrieval system. In other words, at least a portion of the result trie is evaluated, where the result trie is a combination of the match trie and the storage trie storing the set of strings stored in the electronic database or information retrieval system using an intersection operation. The result trie is a trie that stores the set of keys obtained when the set of keys stored in the storage trie intersects with the set of keys stored in the match trie.
[0649] The storage trie can be an index trie, for example, storing strings included by documents and corresponding document identifiers as two key parts, such as (string, long). The strings included in the set of strings stored in the match trie are typically permanently stored in the electronic database or information retrieval system and / or represent entries in the electronic database or information retrieval system.
[0650] Similar to the intersection of tries described above, a result trie is obtained. The set of children of each node in the result trie is the intersection of the sets of children of the corresponding nodes in the match trie and the store trie, where if the same key is associated with nodes of different tries, the nodes of different tries will correspond to each other. Typically, the match trie, the store trie, and the result trie have the same structure or format. Unless otherwise specified in this section, all aspects of the trie intersection operation discussed above also apply to the intersection of tries in the context of fuzzy search.
[0651] As described above, a trie includes one or more nodes, each child node is associated with a key part, and the path from the root node in the trie to another node defines the key associated with that node, which is the concatenation of the key parts associated with the nodes on the path. The trie data structure described above can be used to implement the trie. However, unlike the example described above, the match trie is typically an undirected cycle. This means that a child node in the match trie may have more than one parent node. FIG. 48B to FIG. 48E Examples of undirected cycles are provided in, which will be described in detail below. In contrast, each child node in the store trie and the result trie typically has only one parent node.
[0652] The fuzzy search according to the present invention is particularly efficient, and the match trie is a virtual trie, which is dynamically generated during the intersection of the match trie and the store trie. Only those parts of the virtual trie that are necessary for the intersection of the match trie and the store trie are (dynamically) generated, which is sometimes also referred to as "lazy evaluation". Unless otherwise specified in this section, all aspects of the virtual trie discussed above also apply to the use of a virtual match trie in the case of fuzzy search.
[0653] As the output of the fuzzy search, strings and / or other data items are provided, such as document identifiers associated with the result set of the nodes of the result trie. Typically, the match trie includes a set of match nodes, each match node is associated with one or more keys, and the one or more keys correspond to one of the strings from the approximate string set. In this case, the result set of a node can be the set of nodes of the result trie corresponding to the set of match nodes in the match trie (if the key associated with the node of the result trie and the key associated with the node of the match trie are the same, the node of the result trie corresponds to the node of the match trie). This means that only those string data items such as document identifiers are provided as output, and these string data items are associated with the nodes of the result trie corresponding to the match nodes of the match trie.
[0654] Reference FIG. 48A to FIG. 48E , it will now be described how a (virtual) matching trie can be generated that stores an approximate string set including search strings and / or variants of search strings. The matching trie is derived from a finite automaton representing the approximate string set. First, a non-deterministic finite automaton representing the approximate string set is established. In a non-deterministic automaton, each transition between two states is typically associated with a specific character, wildcard, or empty string included in the search string. From the non-deterministic finite automaton, a deterministic finite automaton that also represents the approximate string set can be derived, where the transitions between two states of the deterministic finite automaton are typically associated with a specific character or wildcard included in the search string. Then the matching trie is derived from the deterministic finite automaton.
[0655] Fig.48A A non-deterministic finite automaton (NFA) is shown to match the search string "abc" with a maximum edit distance of 2, also known as the Levenshtein automaton. As is known to those skilled in the art, a finite automaton includes states and transitions between states. The state labeled with reference numeral 0001 is the start state, the state 0006 matches the end state with 0 edits, the state 0007 is the state with 1 edit, and the state 0008 is the state with 2 edits. For a larger edit distance, it is necessary to add additional state rows at the top. Larger strings will result in additional states on the right.
[0656] State 0002 is the state transition that matches the first character ("a"). For "abc" as the input, it reaches the final state 0006. State 0003 is the state transition for inserting a character at the start, for example, if "xabc" is provided to the automaton. The state transition 0004 reflects character replacement, for example, if "ybc" is provided to the automaton. Finally, the transition 0005 reflects character deletion, for example, if "bc" is provided to the automaton.
[0657] It can be seen that Fig.48A the automaton is non-deterministic. For example, providing "a" as the first input character will result in states 01, 11, 10, 12, 21, 22, and 32. This set of states contains a matching state (32, reference numeral 0008) because "abc" needs to have two characters deleted to get "a".
[0658] The NFA can be converted to a deterministic finite automaton (DFA) using, for example, the so-called Powerset construction method. Other methods for efficiently creating the Levenshtein automaton DFA include the method proposed by Klaus Schulz and Stoyan Mihov. Fig.48BA DFA is shown that matches "abc" with an edit distance of 1 to reduce the number of states in the example. The state labeled with reference numeral 1001 is the start state. Reference numeral 1002 refers to the transition for a specific character ("b"). Reference numeral 1003 refers to the transition for any other character (wildcard). The gray state labeled 1005 is the match state.
[0659] In a preferred embodiment, the matching trie and the parent nodes in the storage trie include bitmaps, and the value of the key part of the child nodes in the trie is determined by the value of the bit (set) in the bitmap included by the parent node of the child node, depending on which bit the child node is associated with. Such a trie data structure has been described in the above example. They allow for particularly efficient intersection operations because the intersection of the child nodes of the matching trie and the child nodes of the storage trie can be achieved by combining the bitmaps of each of the child nodes using the intersection operation.
[0660] Transitions between the states of a finite automaton can be associated by encoding specific characters or wildcards associated with the transitions to obtain an enhanced finite automaton to derive a matching trie with such a data structure from a (deterministic) finite automaton, the encoding consisting of or represented by one or more bitmaps, the length and / or format of the one or more bitmaps being equal to the bitmap included by the parent node of the matching trie. For the encoding of a specific character, exactly one bit is set in each of the bitmaps included or represented by the encoding. For the encoding of a wildcard, the bits of all valid character encodings are set in the bitmap included or represented by the encoding, thus "masking" all valid character encodings (or the bits of all valid character encodings, except for the encoding of the specific character associated with the state from which the transition departs). In other words, the encoding of a wildcard is the OR combination of the bitmaps of all valid character encodings (or the OR combination of the bitmaps of all valid character encodings, except for the encoding of the specific character associated with the state from which the transition departs).
[0661] Fig.48C Shows Fig.48B such an enhancement of the transitions of the DFA, where an encoding scheme for Unicode characters and Unicode strings as described above with reference to Fig.21 is used. For readability, only 10-bit Unicode character encodings are used. These encodings use two 64-bit bitmaps each. As described above, since exactly one bit is set in each of these bitmaps, each of them is capable of encoding 6 bits (2 6 = 64).
[0662] Encoding 2001 represents "a" encoded as two 6-bit values 1 ("000001") and 33 ("10001"), with each bit position encoded as
[0663]
[0664] and
[0665]
[0666] In hexadecimal notation, where 0000 = 0, 0001 = 1, 0010 = 2, 0011 = 3, …, 1111 = F, this corresponds to 0x0000000000000002 and 0x0000000200000000.
[0667] Encoding 2003 represents the “b” encoded as two 6-bit values 1 (“000001”) and 34 (“10010”), with each bit position encoded as
[0668]
[0669] and
[0670]
[0671] In hexadecimal notation, this corresponds to 0x0000000000000002 and 0x0000000400000000.
[0672] In the case of a wildcard, set the bits encoding all allowed characters (in this case the full Unicode alphabet). For example, the bits of encoding 2002 are set to represent all 10-bit encoded Unicode characters, i.e., the 6-bit values “000000” … “001111” and “000000” … “111111”, with each bit position encoded as
[0673]
[0674] and
[0675]
[0676] In hexadecimal notation, this corresponds to 0x000000000000FFFF and 0xFFFFFFFFFFFFFFFFFF.
[0677] Encodings 2004 and 2005 represent the same character “b”, one for a non-final case and one for a matching state, i.e., “b” is the last letter in the input string. Encodings 2006 and 2007 show a simulation of the wildcard case. 2008 and 2009 show the case of the final matching state for the character “c” and the wildcard.
[0678] The complete wildcard mask for all Unicode characters that result in a non-matching state (as explained above with reference to Fig.21 the 10-bit, 15-bit, and 21-bit encodings) is:
[0679]
[0680] The complete wildcard mask for all Unicode characters that result in a matching state is:
[0681]
[0682] Fig.48D The top of the result of the matching trie is shown, with each 10-bit Unicode character represented by two key parts. Since the boolean AND operator on the trie first requires a bitmap and then descends to the matching child nodes, a specific character is encoded as a single bit, and the wildcard is encoded as a bitmap, with bits set for all valid character encodings.
[0683] Such a matching trie can be directly derived from any of the finite automata described above, in particular the enhanced finite automaton. However, in the case where the characters stored in the matching trie, storage trie, or result trie are encoded by more than one key part, i.e., by nodes at more than one level, of the corresponding trie, the matching trie is preferably derived from a complete finite automaton representing an approximate set of strings. Preferably, M is between 2 and 4.
[0684] A complete finite automaton from a preferably deterministic finite automaton, more preferably from an enhanced finite automaton, as described above, by replacing a transition between two states of the finite automaton with one or more sequences of M - 1 levels of intermediate states and M transitions linking two states via M - 1 intermediate states, preferably for each transition, or by associating a transition between two states of the finite automaton, preferably for each transition, with one or more sequences of M - 1 levels of intermediate states and M transitions linking two states via M - 1 intermediate states. Thus, states that are not relevant to the complete string are added to the finite automaton.
[0685] For example, in the Fig.48B and Fig.48C finite automaton, there is a transition leaving state 0 and ending in state 1. In the complete finite automaton, from Fig.48D it can become apparent that this transition is replaced by the intermediate state 110, and the transition sequence includes a transition between state 1 and state 110 and another transition between state 110 and state 1. As another example, in Fig.48B and Fig.48CIn the finite automaton, there is a transition that departs from state 0 and ends at state 10. In the complete finite automaton, this transition is replaced by intermediate state 110, and the transition sequence includes a transition between state 1 and state 110 and another transition between state 110 and state 10. As a third example, in the Fig.48B and Fig.48C finite automaton, there is a transition that departs from state 0 and ends at state 14. In the complete finite automaton, this transition is replaced by intermediate states 110 and 111 and two transition sequences. The first transition sequence includes a transition between state 1 and state 110 and another transition between state 110 and state 14. The second transition sequence includes a transition between state 1 and state 111 and another transition between state 111 and state 14.
[0686] Each of the M transitions in the sequence is associated with an intermediate encoding that consists of or is represented by a bitmap, whose length and / or format is equal to the bitmap included in the parent node of the matching trie, and wherein the matching trie is derived from the complete finite automaton. This encoding is referred to here as "intermediate" because it represents only a part of the encoding of the entire character, and in the example of the complete finite automaton, from which half of the encoding of the entire character of the matching trie that can be derived Fig.48D can be derived.
[0687] For example, in the complete finite automaton from which the matching trie that can be derived Fig.48D can be derived, the transition between state 0 and state 110 is associated with an encoding that consists of or is represented by the following bitmap:
[0688]
[0689] In hexadecimal representation, it corresponds to 0x0000000000000002. The association of this encoding with the transition between state 0 and 110 is indicated by the upper dashed arrow between the encodings 2001 of the transition. The "1" that the arrow points to represents bit number 1 (the second bit) in the bitmap included in the parent node 0 of the matching trie that can be derived Fig.48D from.
[0690] As another example, in the complete finite automaton from which the matching trie that can be derived Fig.48D can be derived, the transition between state 0 and state 111 is associated with an encoding that consists of or is represented by the following bitmap:
[0691]
[0692] In hexadecimal representation, this corresponds to 0x000000000000FFFF. The association of this encoding with the transition between state 0 and 111 is indicated by the upper dashed arrow between the encodings 2002 of the transition. The "0, 2... 63" pointed to by this arrow represents the bit numbers 0, 2... 63 (the first, third... 64th bits) in the bitmap included in the parent node 0 of the matching trie of Fig.48D In the case where a transition between two states of a finite automaton is associated with a specific character, the concatenation of the bitmaps included or represented by the intermediate encodings associated with M transitions of the sequence is the encoding of the specific character, and exactly one bit is set in each of the bitmaps.
[0693] In the case where a transition between two states of a finite automaton is associated with a specific character, the concatenation of the bitmaps included or represented by the intermediate encodings associated with M transitions of the sequence is the encoding of the specific character, and exactly one bit is set in each of the bitmaps.
[0694] In Fig.48B and Fig.48C In the example of the finite automaton, the transition between state 0 and 1 is associated with the character 'a'. Therefore, in the complete finite automaton of the matching trie from which Fig.48D can be derived, the concatenation of the bitmaps included or represented by the intermediate encodings associated with the two transitions between state 0 and 1 (via state 110) is the encoding of the character 'a', where exactly one bit is set in each of the bitmaps. The encoding is as follows:
[0695]
[0696] and
[0697]
[0698] In hexadecimal representation, this corresponds to 0x0000000000000002 and 0x0000000200000000.
[0699] If a transition between two states of a finite automaton is associated with a wildcard, the concatenation of the bitmaps included or represented by the intermediate encodings associated with M transitions of the sequence includes an encoding in which the bits of all valid character encodings are set in the bitmap included or represented by the encoding, or the bits of all valid character encodings except for the encoding of the specific character associated with the state from which the transition departs.
[0700] In Fig.48B and Fig.48C In the example of the finite automaton, the transition between state 0 and 14 is associated with a wildcard. Therefore, in the complete finite automaton of the matching trie from which Fig.48D can be derived, the concatenation of the bitmaps included or represented by one of the sequences of intermediate encodings associated with the two transitions between state 0 and 14 (via state 111) is the encoding of the wildcard. The encoding is as follows:
[0701]
[0702] and
[0703]
[0704] In hexadecimal representation, this corresponds to 0x000000000000FFFF and 0xFFFFFFFFFFFFFFFFFF.
[0705] Furthermore, in the case where a transition between two states of a finite automaton is associated with a wildcard, the concatenation of the bitmaps included or represented by the intermediate encodings associated with M transitions of a sequence typically will include one or more encodings that include one or more portions of the encoding of a particular character and one or more portions of the encoding, where the bits of all valid character encodings are set in the bitmap included or represented by the encoding, or all bits of the valid character encodings except for those of the encoding of the particular character associated with the state from which the transition departs.
[0706] For example, in a complete finite automaton of a matching trie from which one can be derived Fig.48D the concatenation of the bitmaps included or represented by one of the sequences of intermediate encodings associated with two transitions (via state 110) between states 0 and 14 is an encoding that includes a first portion of the encoding of the character 'a' (which is also the first portion of the encoding of the character 'b') and a second portion of the encoding of the wildcard. The encoding is as follows:
[0707]
[0708] and
[0709]
[0710] In hexadecimal representation, it corresponds to 0x0000000000000002 and 0xFFFFFFFFFFFFFFFFFF.
[0711] An augmented finite automaton or a complete finite automaton can be represented or stored in a data structure including multiple rows, each row representing a state of the augmented finite automaton or the complete finite automaton and including a tuple for each transition departing from the state, each tuple including the encoding associated with the transition and a reference to the end state of the transition. Fig.48E Shown is such a data structure as an array of arrays, which can be used to represent the states of a complete finite automaton and from which one can be derived in a particular efficient manner Fig.48DThe matching trie. For each state at which the transformation ends, such a data structure can include information as to whether the state is a matching state, preferably encoded as a bit in each reference to the state. The data structure typically includes a row for each of the respective states of the augmented finite automaton or complete finite automaton from which the transformation departs.
[0712] The benefit of this (virtual) matching trie approach is good performance, because of the simplicity of the implementation leading to efficient execution, because there is no need to use complex state machines, etc., and also because of the way the AND operator works on the bitmaps in the preferred embodiments.
[0713] Indexing methods and performance measurements
[0714] To measure the performance of range queries according to various embodiments of the present invention, experiments were conducted, the results of which will be discussed below with reference to Fig.53 to FIG. 70. A geographical location application with spatial queries was used as an example. Sample data was exported from OpenStreetMap data for the European region, which is available from http: / / download.geofabrik.de / europe.html. Each entry in the example database contains a 10-digit ID, values for each latitude and longitude with seven positions after the decimal point, and a string representing the street name and house number.
[0715]
[0716]
[0717] A first experiment was conducted to see how the index size (the number of index records) affects the query processing performance for a constant result size. The database was queried to return the IDs of all locations within a small rectangle in the Munich area (longitude 11.581981 + / - 0.01 and latitude 48.135125 + / - 0.01, as Fig.49 shown). The database contained approximately 28.5 million records in total, and 3,143 of them matched.
[0718] For the first series of measurements, the matching records were first loaded, and then approximately 28.5 million other records were added to the database. These other records are shown as the shaded area in Fig.50 Each time 100,000 records were added, the query was made and its execution time was measured. Clearly, each query returned 3,143 matches.
[0719] For the second series of measurements, the matching records within the rectangle were also first loaded, but the remaining records were loaded without the records in the "bands" with no matching longitude or latitude. Fig.51 The remaining records loaded in the second series of measurements are shown.
[0720] The first experiment was conducted on five different methods of indexing and querying geographical locations.
[0721] In the prior art method, here the SpatialPrefixTree of the Apache Lucene 6.0.1 database engine called "prior art index" is used, which supports a spatial index that combines longitude and latitude (https: / / lucene.apache.org, package org.apache.lucene.spatial.prefix.tree). The SpatialPrefixTree is not a trie, but is used to create a reverse index optimized for spatial queries. Thus, it provides a good benchmark for the method adopted here.
[0722] To query all locations within a specified rectangle, range queries are performed in both dimensions. Each query returns a set of IDs of records representing those located in the specified longitude or latitude bands, which are collected by a cursor-based iterator to create a temporary result set. The temporary result sets are intersected to obtain the final result set. As Fig.52 shown, where the shaded areas mark the specified longitude and latitude bands (temporary result sets), and the black area marks the specified rectangle (final result set). The SpatialPrefixTree is not a trie, but is used to create a reverse index optimized for spatial queries. By creating several indexes for several resolution levels, query performance can be enhanced, similar to the variable-precision indexing of a trie, which will be explained below with reference to Figures 58 to 60 it.
[0723] Fig.53 The measurement results of the prior art method are shown. The x-axis indicates the number of records loaded into the database, and the y-axis indicates the number of queries executed per second on a logarithmic scale. It can be observed that with or without the matching bands loaded into the database, the query performance decreases as the number of records in the database increases. This performance behavior is typical for prior art databases, where the access time depends on the fill level of the index.
[0724] A first method of indexing and querying geographical locations using the trie described above is made, which is referred to herein as the "standard index". A latitude index is created by the first two-dimensional trie of the preferred embodiment, and another longitude index is created by the second two-dimensional trie of the preferred embodiment. In other words, each in the two-dimensional index trie stores two items respectively, namely (latitude, ID) or (longitude, ID). The two-dimensional index trie is stored in a non-interleaved manner, as shown in the trie above Fig.44 shown, where the latitude / longitude forms the first key part, and the ID forms the second key part. From an abstract perspective, this results in a trie containing all latitudes / longitudes, where each leaf node is the root node of a trie including the IDs of all positions with the corresponding latitude / longitude, as Figure 54 shown.
[0725] A range query on the first key part (latitude / longitude) returns the trie root node of the IDs of the positions with matching latitude / longitude. To deliver these results using the trie interface, a list of trie root nodes is combined using a multi-OR operator that provides the trie interface, as Figure 55 shown. Therefore, as Figure 56 shown, the standard index according to the present invention uses two index tries 5601, 5603, each of which has two key parts: (latitude, ID) or (longitude, ID). For the first key part of the trie, a range query is performed using the virtual latitude / longitude tries 5602, 5604. The results of each in the range query are combined using the multi-OR operators 5605, 5606. The results of the two multi-OR operators are combined using the AND operator 5607, which delivers the final result.
[0726] Figure 57 The measurement results of the standard index are shown. It can be observed that in the case where no matching records are loaded into the database, the query processing time is constant, which is due to the fact that the trie has a constant access time, i.e., independent of its filling level. Given that the SpatialPrefixTree of the Lucene database engine has been optimized for spatial queries, the query processing time is also about 10 times shorter than using the prior art method, which is very obvious.
[0727] In the case where matching records are loaded into the database, the query performance of the standard index degrades because more and more independent latitude and longitude query results must be combined. The performance of the standard index with matching records is not very good because a long list of nodes (IDs) appears. To improve the performance, it is necessary to reduce the number of nodes that must be combined by the OR operation.
[0728] In the method referred to herein as "variable precision indexing", the number of nodes that need to be ORed can be significantly reduced by maintaining multiple indexes with variable precision and creating a hierarchy. This is comparable to the concept of creating several indexes for several resolution levels in the prior art database engines mentioned above. For example, using ranges in a trie that stores 2-digit decimal numbers, the indexes at each level can represent prefixes:
[0729] -0, 1, 2,... 9
[0730] -00, 01, 02,..., 09, 10,... 99
[0731] By way of example, Figure 58 shows the index (the first key part) of the first level of the X values, and Figure 59 shows the index of the second level of the X values, where the second key part (the Y item) can be an ID. Figure 58 The entries in the first-level index shown in Figure 58 (the leaf nodes or X items of the first key part) contain all the IDs (the trie of the second key part) of the X values in the corresponding prefix range. A query with an X range from 10 to 53 will combine Figure 58 the query results of the range 1..4 on the index of Figure 58 (two hits, marked by the dashed rectangles) with Figure 59 the query results of the range 50..53 on the index of Figure 59 (one hit, marked by the dashed rectangle). This results in only three nodes needing to be combined in the subsequent multi-OR operation, instead of five nodes when only using the Figure 59 index trie of Figure 59 .
[0732] The actual experiments conducted by the inventors used 6 bits for each precision step. The precision length is stored in the first byte of the key. The trie has 11 levels, and the nodes in the trie have at most 16 child nodes at the root node level and at most 64 child nodes at the subsequent 10 levels (including the leaf nodes). This means that for the left and right parts of the trie that need to be ORed, there are at most 16 - 2 nodes at the root node level and 2*(64 - 1) nodes at the next level. This results in an upper limit of 16 - 2 + 2*(64 - 1)*10 = 1274 tries that must be ORed. Each key value that accepts range queries is stored at 11 precision levels. As described above, range queries based on the virtual range trie are used to execute the queries. A list of all the keys required to select the desired nodes is created.
[0733] Figure 60Shows the measurement results of variable precision indexing. When the matching segments are not loaded, the query execution is very fast (about 10 times faster than when using a standard index), and the time is constant. Even when the matching records are loaded into the database, the query performance is almost constant. However, as will be shown below, variable precision indexing is expensive in terms of memory requirements and index performance.
[0734] Another method of indexing and querying geographical locations (referred to herein as "two-dimensional indexing") shows that there is a solution that is faster than a standard index but without the disadvantages of variable precision indexing. In this method, as Figure 61 shown, the longitude / latitude and ID of database entries are stored in an interleaved trie having the structure discussed above, e.g., referring to Figure 32 and Figure 46 . Just like a standard index, the latitude and longitude are indexed separately in two different tries, i.e., trie 6110 (longitude, ID) and trie 6120 (latitude, ID).
[0735] To query a rectangle, a virtual range trie 6130 specifying a longitude range and a virtual range trie 6140 specifying a latitude range are created. When the longitude index is intersected with the longitude range trie and at the same time the latitude index is intersected with the latitude range trie, the intermediate results of the combined intersection are as will be explained below.
[0736] In the first step, a bitwise AND operation is performed between the root node 6111 of the longitude index trie 6110 and the bitmap of the root node of the longitude range trie 6130, as indicated by the arrow 6151 in Figure 61 . Similarly, a bitwise AND operation is performed between the root node 6121 of the latitude index trie 6120 and the bitmap of the root node of the latitude range trie 6140, as indicated by the arrow 6152 in Figure 61 .
[0737] Figure 61 The nodes 6112 and 6122 in represent the nodes at level 2 of the index tries 6110 and 6120. Each of these nodes at level 2 is the root node of the ID key part in the index. Thus, the bitwise AND operation between the root node 6111 of the longitude index trie 6110 and the bitmap of the root node of the longitude range trie 6130 produces the first key part of the IDs belonging to the locations with the specified longitude range, and the bitwise AND operation between the root node 6121 of the latitude index trie 6120 and the bitmap of the root node of the latitude range trie 6140 will produce the first key part of the IDs belonging to the locations with the specified latitude range.
[0738] In the second step, the keys in the index trie that do not belong to the positions falling within the specified longitude and latitude ranges are filtered as follows: The bitmaps of the nodes of the longitude / latitude index tries 6110 / 6120 generated by the first step are combined by a bitwise OR operation, and a bitwise AND operation is performed between the results of the bitwise OR operation, as indicated by the arrow 6153 in Figure 61 .
[0739] Figure 61 The nodes 6113 and 6123 in represent the nodes at level 3 of the index tries 6110 and 6120. Each of these nodes at level 3 stores the second key part of the longitude / latitude key parts stored in the index trie. The operations in the third step continue, and the nodes at level 3 belong to the keys that have not been filtered out in the second step. As indicated by the arrows 6154, 6155 in Figure 61 , a bitwise AND operation is performed between the bitmaps of these nodes and the corresponding nodes (at level 2) of the corresponding range tries 6130, 6140.
[0740] Figure 61 The nodes 6114 and 6124 in represent the nodes at level 4 of the index tries 6110 and 6120. Each of these nodes at level 4 stores the second key part of the ID key parts stored in the index trie. The operations in the fourth step continue, and the nodes at level 4 are generated as a result of the bitwise AND operation in the third step. Similar to the second step, the bitmaps of the nodes of the longitude / latitude index tries 6110 / 6120 having the same parent node generated by the first step are combined by a bitwise OR operation, and a bitwise AND operation is performed between the corresponding results of the bitwise OR operation, as indicated by the arrow 6156 in Figure 61 .
[0741] The operations continue in the same way until the leaf nodes of the index trie are reached. In summary, two indexes are combined using a matcher, which only returns a "view" of the alternating indexes with the first dimension (x). The second dimension is suppressed in the output of the matcher.
[0742] Figure 62 shows the measurement results of the two-dimensional index. Also, when no matching segments are loaded into the database, the query time is constant. In the case of loading matching segments into the database, the query performance is significantly improved compared to the standard index.
[0743] The strategy of matching and suppressing dimensions can in principle be applied to more than two dimensions. However, this results in a large chain of nodes: Each node in the x dimension may have 64 child nodes in the y dimension, and these child nodes may in turn have 64 child nodes, already amounting to 4096 child nodes in total.
[0744] In the last method of indexing and querying geographical locations (referred to herein as "single-index indexing"), only one multi-dimensional index is created, which stores both longitude and latitude in an interleaved trie, as discussed above, for example, with reference to Figure 32 and Figure 46 . A schematic diagram of the index trie is shown in Figure 63 , where the first key part X representing longitude and the second key part Y representing latitude are stored in an interleaved manner, and the third key part representing the ID of the location is stored in a sub-trie, which depends on the leaf nodes of the X / Y trie. The IDs stored in each of the sub-tries belong to a set of locations with the same longitude and latitude.
[0745] To query a rectangle, as described above, for example, with reference to Figure 40 , a two-dimensional range query is performed. As shown in Figure 64 , the range of longitude is specified by the first one-dimensional virtual range trie 6402, and the range of latitude is specified by the second one-dimensional virtual range trie 6403. The two one-dimensional range tries are combined by the interleaving operator 6404 to form an interleaved two-dimensional range trie. To obtain the set of ID sub-tries belonging to the matching longitude and latitude, the AND operator 6405 performs an intersection between the two-dimensional range trie and the X / Y part of the index trie 6401.
[0746] Figure 65 Shows the measurement results of single-index indexing. It can be observed that single-index indexing provides perfect scalability, and even when the matching band is loaded, the query performance does not significantly depend on the number of records in the database. Note that the peak in performance is caused by Java virtual machine garbage collection.
[0747] Figure 66 Shows the results of a second experiment, where the query performance (queries per second) was measured on increasing result sizes. The latitude and longitude of the query rectangle were increased from + / - 0.01 to + / - 0.20 degrees. First, 28.5 million records were loaded into the database. It can be observed that for larger results, the variable-precision index scales best (note that due to memory limitations, only partial data could be loaded for the variable-precision index). Single-index indexing also provides good performance. This is especially true when the ID is also stored in an interleaved manner in addition to latitude and longitude. It is believed that the performance gain of the triple-interleaved index is due to an increase in the number of common prefix paths, which is also evidenced by the smaller memory footprint (46.01 vs. 44.76 bytes / point). The performance may be better due to fewer branches in the tree, which may result in fewer recursive steps to traverse the tree.
[0748] The prior art indexing using the Lucene database delivered nearly constant results for all result sizes, but the performance level was lower than the indexing with a single index. For smaller result sizes, the performance of the standard index (the two indexes are not interleaved) and the two-dimensional index (the two indexes are interleaved) was better than the prior art indexing using the Lucene database, but for larger result sizes, the performance was worse. Figure 67 Shown with Figure 66 The same results, but with a logarithmic scale on the x-axis. In this representation, it can be seen that the query performance of the index according to the invention decreases linearly as the result size increases. This means that in particular the two single-index approaches have linear scalability, i.e. doubling the result size doubles the query time.
[0749] In the third experiment, the indexing performance (entries indexed per second) was measured. 28.5 million records were loaded and the time was measured every time 100,000 records were added. It can be observed that the state-of-the-art index and all trie-based approaches provide almost constant performance with respect to index growth. The results are summarized in Figure 68 As mentioned above, variable precision indexing performs poorly because 2 x 11 index entries must be created for each record. The prior art index (Lucene) whose approach is similar to variable precision indexing performs even worse. Both the standard trie-based index and the two-dimensional index create 2 index entries per record and therefore have similar indexing performance. The trie-based index creates only one index entry per record and performs best. Note that in the trie-based approach, the black columns represent the results of the uncompressed trie and the light columns represent the results of the trie using bitmap compression.
[0750] Figure 69 The memory space required for different indexes per index position is compared. The state-of-the-art Lucene index uses more space than a standard trie-based index and a two-dimensional index with bitmap compression (light columns) and a single index even without bitmap compression. The trie-based variable precision index requires more space than any other approach.
[0751] It can be concluded that for low- or medium-cardinality attributes, standard indexing works well enough. For example, product prices typically do not have a continuous value space, but discrete values such as 3.99, 4.49, 4.89, etc. Instead of storing something like the order date as a timestamp with millisecond precision, it can be stored with day or hour precision to meet the requirement of making the value space "more" discrete. For columns with a continuous value space, variable-precision indexing provides better performance, especially if used in multi-dimensional queries. However, due to slow indexing speed and high memory requirements, variable-precision indexing is only recommended for static applications and where sufficient memory is available. For closely related dimensions, single-index indexing is the best solution for medium expected result sizes.
[0752] Even though multi-dimensional indexing has been presented here in the context of spatial queries, trie-based range queries can be applied to many other cases, such as for graph databases. Property graph databases are based on nodes connected by edges, where both nodes and edges have properties. If each node and edge is represented by a unique ID, these three IDs can be used as dimensions to represent and query node-edge-node triples. Note that this also applies to the context of the Resource Description Framework (RDF) and its subject-predicate-object expressions, which are called triples in RDF terminology.
[0753] As mentioned above, the present invention can also be easily used for full-text search applications by storing each occurrence of a term and the document ID as two key parts (string, long). Since the present invention is based on a prefix tree, it inherits the string search functionality of the prefix tree. For example, it can be used to efficiently implement fuzzy (similarity) search.
[0754] In fact, measurements conducted by the inventors have shown competitive performance in information retrieval applications. In an experiment conducted shortly before the priority date of the present application, 500,000 English Wikipedia articles were indexed. Figure 70A Shows the average indexing performance of an embodiment of the present invention in characters per second compared to the Lucene information retrieval software library. It can be seen that when using an array of long integers (black columns), the system of the present invention only delivers slightly lower indexing performance. As expected, when using bitmap compression with an array of bytes (light columns), the indexing performance is slightly reduced.
[0755] Figure 70B Compares the index size (index size / text size, in %). Surprisingly, even the memory model based on an array of long integers (black columns) can be comparable to Lucene - despite the alignment loss. The memory model based on bitmap compression and having an array of bytes (light columns) requires less space than Lucene.
[0756] Figure 70C shows the query performance (queries / second). The term query in this experiment searches all documents that contain the words "which" or "his" with the combination "from" to provide a certain level of complexity and quantity. Using multiple concurrent threads, the performance of the system of the present invention is up to seven times higher than that of Lucene.
[0757] The fuzzy query in this experiment searches for documents that contain words similar to "chica". The similarity is defined by an edit distance (Levenshtein distance) of 1, i.e., at most one character can be deleted, inserted, or replaced. In this discipline, the invented system is proven to be four to six times faster than Lucene. It is worth noting that both Lucene and the system of the present invention deliver exactly the same number of result documents: 319,809 for the term query and 30,994 for the fuzzy query.
[0758] Shortly before the filing date of this application, approximately one year later, the experiment described above was repeated. The results of the repeated experiment can be seen in Figures 70D to 70F in.
[0759] From Figure 70D it can be seen that the indexing performance is now lower than in the previous experiment. In the system of the present invention, it is lower because memory management has now been fully implemented at the expense of some performance. In the Lucene system, it is lower because the Lucene index runs on only one thread / CPU core. This is done to obtain better comparability because the system according to the present invention is also implemented with only one thread / CPU core. Note that the index of the system of the present invention can also be parallelized, but the inventors have not yet implemented it.
[0760] Figure 70E shows that the index size actually remains unchanged compared to the earlier experiment.
[0761] Figure 70F shows that the query performance of the system of the present invention has improved compared to the earlier experiment, especially in fuzzy search. In the earlier experiment, only the DFA shown in Figure 48C was pre-generated, which only includes nodes corresponding to the states associated with the full string. Figure 48D The intermediate states (states of the finite automaton not associated with the complete string) shown in are due to the fact that character encoding requires two or more key parts, which are dynamically determined during the intersection operation. Instead, in the later experiment, a finite automaton with a complete set of states and transitions was pre-generated and stored as an array of arrays (matrix), as shown in Figure 48EAs shown in. This method doubles the query performance. Compared with the earlier experiments, the improvement of the Lucene system is similar because the Lucene system has made some optimizations between the priority date and the filing date of this application.
Claims
1. A method for retrieving data from an electronic database or information retrieval system, comprising the following steps: Obtain representations of two or more higher-order input tries; Evaluate at least a portion of a higher-order result trie, the higher-order result trie being a combination of higher-order input tries using higher-order logical operations; and Provide as output: a representation of a result trie, keys associated with nodes of the result trie, or a set of keys or values derived from keys or values associated with nodes of the result trie, wherein a representation of at least one higher-order input trie is provided as the output of a combination of lower-order logical operations that perform the following steps: Obtain representations of two or more lower-order input tries; At least partially evaluate a lower-order result trie, the lower-order result trie being a combination of lower-order input tries using lower-order logical operations; and Provide a representation of the lower-order result trie as output, and wherein during the step of evaluating at least a portion of the higher-order result trie, at least a portion of the lower-order result trie is evaluated, wherein at least some portions of the lower-order result trie are not evaluated, the at least some portions not being required for evaluating at least a portion of the higher-order result trie, and only those portions of the lower-order result trie that are required for evaluating at least a portion of the higher-order result trie are evaluated, wherein the portion of the trie is a node and / or branch of the trie, wherein evaluating a trie or a portion of a trie includes traversing the trie or the portion of the trie, wherein the higher-order logical operations and lower-order logical operations include: combination using AND Boolean operation, OR Boolean operation, ANDNOT Boolean operation, or XOR Boolean operation, wherein the input trie stores a set of keys stored in the electronic database or information retrieval system or a result key set derived from one or more sets of keys stored in the electronic database or information retrieval system.
2. The method according to claim 1, wherein Provide the keys associated with the leaf nodes of the result trie as output.
3. The method according to any one of claims 1 or 2, wherein The set of keys provided as output is provided in the trie.
4. The method according to any one of claims 1 or 2, wherein The set of keys provided as output is provided by a cursor or iterator.
5. The method according to any one of claims 1 or 2, wherein The representation of the trie is a physical trie data structure or a handle to a trie operator capable of evaluating the trie.
6. The method according to any one of claims 1 or 2, wherein The representation of the trie represents the root node of the trie and implements a trie node interface, where the trie node interface represents a trie node and provides methods or functions for querying and / or traversing the child nodes of the trie node it represents.
7. The method according to claim 1, wherein evaluating the trie or a part of the trie includes providing the following information: the traversed node exists in the trie.
8. The method according to claim 1, wherein At any given moment during the evaluation of a trie or a portion of a trie, at least those portions of the physical data structure that implement the trie or the portion of the trie that are required to be implemented at that moment for traversing the trie or the portion of the trie.
9. The method according to any one of claims 1 or 2, wherein At any given moment, less than the entire physical data structure of the trie or the portion of the trie is implemented, and only those portions of the physical data structure of the trie or the portion of the trie that are required to be implemented at that moment for traversing the trie or the portion of the trie are implemented.
10. The method according to any one of claims 1 or 2, wherein Evaluating a trie or a portion of a trie includes: implementing a physical data structure of the trie or the portion of the trie as a whole.
11. The method according to any one of claims 1 or 2, wherein At least at the moment of obtaining the representation of the low-order result trie, there is no physical data structure of the low-order result trie or the portion of the low-order result trie as a whole.
12. The method according to any one of claims 1 or 2, wherein The physical data structure of the low-order result trie or the portion of the low-order result trie as a whole is not implemented.
13. The method according to any one of claims 1 or 2, wherein After at least a portion of the high-order result trie has been evaluated, at least a portion of the low-order result trie is evaluated.
14. The method according to claim 1 or 2, wherein, The portions of the high-order result trie and the low-order result trie are evaluated in an interleaved manner.
15. The method according to claim 1 or 2, wherein, The step of providing the representation of the low-order result trie as output is performed before the step of having at least partially evaluated the low-order result trie.
16. The method according to claim 1 or 2, wherein, The low-order logical operation is the same logical operation as the high-order logical operation.
17. The method according to claim 1 or 2, wherein, The low-order logical operation is a different logical operation from the high-order logical operation.
18. The method according to claim 1 or 2, wherein, Provide at least one of the low-order input tries as the output of an even lower-order combinatorial operator, and evaluate at least a portion of the at least one of the low-order input tries during the step of evaluating at least a portion of the low-order result trie.
19. The method according to claim 1 or 2, wherein, A trie includes one or more nodes, each child node is associated with a key portion, and the path from the root node in the trie to another node defines a key associated with that node, and the key is the concatenation of the key portions associated with that node on the path.
20. The method according to claim 1 or 2, wherein, The combination of two or more input tries using a logical operation is a result trie that stores the set of keys obtained when combining the corresponding sets of keys stored by the respective input tries using the logical operation.
21. The method according to claim 1 or 2, wherein, If the logical operation is a difference set, the parent node of the result trie is the parent node of the first input trie, and the leaf nodes of the parent node of the result trie are the combination of the set of child nodes of the corresponding parent node in the first input trie and the set of child nodes of any corresponding parent node in the other input tries using the logical operation, and If the logical operation is not a difference set, the set of child nodes of each node in the result trie is the combination of the sets of child nodes of the corresponding nodes in the input tries using the logical operation; and If the keys associated with the nodes of different tries are the same, two or more nodes of different tries correspond to each other.
22. The method according to claim 1 or 2, wherein, The set of child nodes of the nodes of the result trie is determined by combining the sets of child nodes of the nodes of the input tries corresponding to the nodes of the result trie using the logical operation.
23. The method according to claim 1 or 2, wherein, The step of evaluating a result trie or a portion of a result trie includes: Performing a combination function on the root node of the result trie; Wherein, performing the combination function on the input node of the result trie includes: Determining the set of child nodes of the input node of the result trie by combining the sets of child nodes of the nodes of the input tries corresponding to the input node of the result trie using the logical operation; and Execute the combination function for each child node determined for the input node of the result trie.
24. The method according to claim 23, wherein, The step of evaluating the result trie or a part of the result trie is performed using depth - first traversal, breadth - first traversal, or a combination thereof.
25. The method according to claim 24, wherein, Performing the step of evaluating the result trie or a part of the trie in depth - first traversal includes: executing the combination function for one of the child nodes of the input node and traversing the sub - trie formed by the child node before executing the combination function for the next sibling node of the child node.
26. The method according to claim 24, wherein, Performing the step of evaluating the result trie or a part of the result trie in breadth - first traversal includes: executing the combination function for each child node determined for the input node of the result trie, and determining the set of child nodes for each child node determined for the input node of the result trie before executing the combination function for any grand - child node of the input node of the result trie.
27. The method according to claim 1 or 2, wherein, The nodes in the input trie include bitmaps.
28. The method according to claim 27, wherein, The value of the key part of the child node in the trie is determined by the value of the bit(s) in the bitmap included in the parent node of the child node, depending on which bit the child node is associated with.
29. The method according to claim 27, wherein, Combining the sets of child nodes of the nodes in the input trie corresponding to the nodes in the result trie using logical operations includes: combining the bitmaps of each node in the input trie corresponding to the nodes in the result trie bit - by - bit using logical operations to obtain a combined bitmap.
30. The method according to claim 29, wherein, Determine and / or evaluate the child nodes of the nodes in the result trie based on the combined bitmap.
31. The method according to claim 29, wherein, Perform the step of combining the bitmaps before determining and / or evaluating the child nodes of the nodes in the result trie and / or before the evaluation of the result trie traverses to the child nodes of the nodes in the result trie.
32. The method according to claim 1 or 2, wherein The logical operation is intersection, union, difference, or exclusive - or.
33. The method according to claim 1 or 2, wherein Using logical operations includes: using bit - wise AND boolean operation, bit - wise OR boolean operation, bit - wise AND NOT boolean operation, or bit - wise XOR boolean operation to combine the bitmaps of the nodes.
34. A computer program, in particular a database application program or an information retrieval system program, comprising instructions for performing the method according to any one of claims 1 to 33.
35. A computer-readable medium comprising one or more processors and a memory, the processor performing the method according to any one of claims 1 to 33 when executing instructions stored in the memory.
36. A non-transitory computer-readable medium having stored thereon the computer program according to claim 34.
Citation Information
Patent Citations
Method for increasing storage capacity in a multi-bit trie-based hardware storage engine by compressing the representation of single-length prefixes
US20040111440A1