A construction method of a superset index structure combining TRIE and LOUDS

By combining TRIE and LOUDS, a hybrid index structure is built, and using TRIE's query efficiency and LOUDS's space compression capability, the problem of TRIE's low space efficiency when storing large-scale collection data is solved, and efficient superset query and storage is achieved.

CN114185893BActive Publication Date: 2025-06-27YUNNAN NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111522608.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-14
Publication Date
2025-06-27
Estimated Expiration
2041-12-14

AI Technical Summary

Technical Problem

When storing large-scale collection data, the existing TRIE superset query structure has low spatial storage efficiency and needs to be compressed to reduce space overhead.

Method used

Combining the TRIE and LOUDS structures, the upper part uses TRIE for frequent access queries, and the lower part uses LOUDS for efficient compression storage, connecting the two parts through the joint node.

Benefits of technology

The dual optimization of the index structure in time and space efficiency is realized, providing high query efficiency and high spatial compression rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114185893B_ABST
    Figure CN114185893B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for constructing a superset index structure combining TRIE and LOUDS, belonging to the technical field of set and string processing. The present invention includes a data preprocessing stage, an index structure construction stage, and a superset query stage. In the data preprocessing stage, the sets and elements in the original set dataset are mapped and sorted. In the index structure construction stage, a hybrid index structure with TRIE at the upper part and LOUDS at the lower part is constructed. In the superset query stage, a query is given, and all sets that are subsets of the given query are retrieved on the constructed hybrid index structure. The present invention can make full use of the high query efficiency of TRIE and the high space compressibility of LOUDS, enabling the frequently accessed upper part to have a fast query speed and the less frequently accessed lower part to have high compression performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for constructing a superset index structure combining TRIE and LOUDS, and belongs to the technical field of set and string processing. Background Art

[0002] Superset queries for sets are widely used in many fields such as decision support systems, publish-subscribe systems, data mining, etc. For example, in a publish-subscribe system, users can register their keywords, and the system can deploy a superset query algorithm to perform superset queries on each incoming object to obtain the users who have subscribed to that object. In a decision support system, the system can obtain the positions that match the capabilities of a certain job seeker through a superset query.

[0003] TRIE is the most common and efficient superset query structure, but its space storage efficiency is not high. When storing large-scale set data, a large amount of storage space is often required. Therefore, it needs to be compressed, and LOUDS is the currently known TRIE structure with the highest space compression efficiency, which can effectively reduce the space overhead. Therefore, by combining TRIE and LOUDS, representing the frequently accessed upper layer and the lower layer with high storage space overhead using TRIE and LOUDS respectively, the constructed index structure can have high time and space efficiency at the same time. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a method for constructing a superset index structure combining TRIE and LOUDS, so that the constructed index structure has high time and space efficiency at the same time, thereby solving the above problems.

[0005] The technical solution of the present invention is: a method for constructing a superset index structure combining TRIE and LOUDS, including a data preprocessing stage, an index structure construction stage, and a superset query stage. In the data preprocessing stage, the sets and elements in the original set dataset are mapped and sorted. In the index structure construction stage, a hybrid index structure with TRIE in the upper part and LOUDS in the lower part is constructed. In the superset query stage, given a query, all sets that are subsets of the given query are retrieved on the constructed hybrid index structure.

[0006] The specific steps are as follows:

[0007] Step1: Map and sort the sets and elements in the initial set dataset.

[0008] Step2: Construct an index structure combining TRIE and LOUDS.

[0009] Step3: Perform a superset query in the constructed index structure.

[0010] The specific content of Step1 is as follows:

[0011] Step 1.1: Count the frequencies of each element in the dataset, and then sort the elements in descending order of frequency.

[0012] Step 1.2: Replace each element in each set with its index in the descending order of frequency, and sort the elements within the set in ascending order of value.

[0013] Step 1.3: Sort all the sets sorted in ascending order of element value in ascending order of set size.

[0014] Step 1.4: Number the sorted sets sequentially starting from 0. Denote the number of set s as s.id, denote the finally processed dataset as D, and |D| as the number of sets in the dataset.

[0015] The specific method for comparing the sizes of sets in Step 1.3 is as follows:

[0016] Step 1.3.1: For two specific sets s and t, the loop variable i ranges from 0 to min(|s|, |t|) - 1, and Steps 1.3.2 to 1.3.3 are loop-executed, where |s| and |t| respectively represent the number of elements in s and t, and min(|s|, |t|) represents the minimum value of |s| and |t|.

[0017] Step 1.3.2: If s[i] > t[i], then s > t. If s[i] < t[i], then s < t. Otherwise, it means s[i] = t[i], and continue to compare the next element. Here, s[i] and t[i] respectively represent the i-th element of s and t.

[0018] Step 1.3.3: If the first min(|s|, |t|) elements of s and t are all equal, then continue to judge whether |s| > |t| holds. If it holds, then s > t. If |s| < |t|, then s < t. Otherwise, s = t.

[0019] The specific content of Step 2 is as follows:

[0020] Step 2.1: Build a TRIE index for the first d elements of each set in the dataset. Hereinafter, d is also referred to as the segmentation depth.

[0021] Step 2.2: Build a hierarchical LOUDS index for the elements after the d-th element of each set.

[0022] Step 2.3: Merge the hierarchical LOUDS structures into a single-layer LOUDS structure. Unless otherwise specified, the LOUDS structure referred to in the present invention refers to a single-layer LOUDS structure.

[0023] The specific content of Step 2.1 is as follows:

[0024] Step 2.1.1: Create an empty root node root, let the currently visited node current = root, and record the depth of root as 0.

[0025] Step 2.1.2: For each set s ∈ D, loop through Steps 2.1.3 to 2.1.6.

[0026] Step 2.1.3: The loop variable i ranges from 0 to min(|s| - 1, d - 1), and Steps 2.1.4 to 2.1.6 are looped through.

[0027] Step 2.1.4: If a node with label s[i] cannot be found among the children of current, add a child node N with label s[i] under the current node, and let current = N. If a node N with label s[i] already exists among the children of current, directly let current = N.

[0028] Step 2.1.5: If i = |s| - 1, then N is called an end point, that is, the node corresponding to the last element of a certain set. Record N.ESets as the binary tuple <startid, count>, where N.Esets represents all the sets that end at node N, startid represents the id of the first set that is exactly the same as s, and count represents the number of sets that are exactly the same as s.

[0029] Step 2.1.6: If i = d - 1, it indicates that N is a node in the last layer of the TRIE. If |s| > d at this time, then N is called an articulation point. The articulation point is the link connecting the upper - layer TRIE and the lower - layer LOUDS. The articulation points are numbered sequentially from 0 according to their insertion order into the TRIE.

[0030] Step 2.1.7: Let C represent the total number of articulation points in the finally constructed TRIE.

[0031] The specific content of Step 2.2 is as follows:

[0032] Step 2.2.1: The hierarchical LOUDS structure constructs 1 integer array Elements, 3 bit arrays NotLeaf, StartOfChild, and EndofSet for each layer, and also includes an array ESets for storing binary tuples <startid, count>.

[0033] Step 2.2.2: For the currently inserted set s, loop through Steps 2.2.3 to 2.2.7.

[0034] Step 2.2.3: Let the loop variable i range from d to |s| - 1, and loop to execute Step 2.2.4 to Step 2.2.7.

[0035] Step 2.2.4: Insert s[i] into the Elements array of the (i - d)-th layer, and denote the position where s[i] is inserted into Elements as p.

[0036] Step 2.2.5: If i!= |s| - 1, set the position p of NotLeaf in the (i - d)-th layer to 1.

[0037] Step 2.2.6: If s[i] is the first child of its parent node, set the position p of StartOfChild in the (i - d)-th layer to 1.

[0038] Step 2.2.7: If i = |s| - 1, then set the p-th position of EndofSet in the (i - d)-th layer to 1, and construct the binary tuple <startid, count> and insert it into the ESets of the (i - d)-th layer. Similar to the ESets of the TRIE node, startid represents the id of the first set that is exactly the same as s, and count represents the number of sets that are exactly the same as s.

[0039] The specific content of Step 2.3 is as follows:

[0040] Step 2.3.1: Starting from the 0-th layer of the LOUDS hierarchy, concatenate the Elements of each layer in sequence to form the final array. For convenience, without causing ambiguity, the final array is still called Elements.

[0041] Step 2.3.2: Concatenate the arrays of NotLeaf, StartOfChild, EndofSet, and ESets of the LOUDS hierarchy in the same way to form the final arrays of NotLeaf, StartOfChild, EndofSet, and ESets of the single-layer LOUDS.

[0042] Step 2.3.3: The finally formed single-layer Elements, NotLeaf, StartOfChild, EndofSet, and ESets are the single-layer LOUDS structure used for superset query.

[0043] The specific content of Step 3 is as follows:

[0044] Step 3.1: Given a query set q, set i = 0, and the current node current = root.

[0045] Step3.2: Sequentially check whether q[j] (i ≤ j ≤ |q| - 1) exists among the children of the current node. If it exists, denote the node corresponding to q[j] as N.

[0046] Step3.3: If N.ESets is not empty, add N.ESets to the result set R.

[0047] Step3.4: If N has children, set current = N, i = j + 1, and then go to Step3.2 to continue searching in the TRIE.

[0048] Step3.5: If N is an articulation point, obtain the articulation point number pno and then transfer to search in LOUDS.

[0049] Step3.6: Obtain the position p1 of the (pno + 1)-th 1 and the position p2 of the next 1 in the StartOfChild array. If the next 1 does not exist, then p2 is the length of the array minus 1.

[0050] Step3.7: Sequentially check whether q[k] (j + 1 ≤ k ≤ |q| - 1) exists between the positions p1 and p2 - 1 in the Elements array. If it exists, obtain its position p in the Elements.

[0051] Step3.8: If the value at the EndofSet[p] position is 1, then count the number c1 of 1s from 0 to p (including p) in the EndofSet, and then insert the binary tuple corresponding to ESets[c1 - 1] into the result set R.

[0052] Step3.9: If the value at the NotLeaf[p] position is 1, then count the number c2 of 1s from 0 to p (including p) in the NotLeaf, set pno = c2 + C, and then go to Step3.6 to continue searching in LOUDS.

[0053] Step3.10: The query ends, and what is stored in R is the result of the final superset query.

[0054] The present invention can make full use of the high query efficiency of the TRIE and the high space compressibility of the LOUDS, enabling the upper part that is frequently accessed to have a fast query speed, while the lower part that is less frequently accessed has high compression performance.

[0055] The beneficial effects of the present invention are as follows: In view of the different requirements of different parts of the index structure, the present invention combines the TRIE with high query efficiency and the LOUDS with high storage efficiency, and has the following advantages:

[0056] 1. High query efficiency, and most superset queries are performed in the TRIE with high query efficiency in the upper layer.

[0057] 2. High space compression rate. Most of the storage space overhead is concentrated in the lower layer and is efficiently compressed by LOUDS.

[0058] 3. The upper and lower parts are connected by articulation points, and the transition from TRIE to LOUDS is natural and convenient. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 is a schematic diagram of the TRIE and LOUDS hybrid structure of the present invention;

[0060] Figure 2 is a schematic diagram of the hierarchical LOUDS of the present invention;

[0061] Figure 3 is a comparison chart of the query time of the present invention;

[0062] Figure 4 is a comparison chart of the index space of the present invention;

[0063] Figure 5 The present invention is a comparison chart of the comprehensive cost of query and space. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0064] The present invention will be further described below in conjunction with the drawings and specific embodiments.

[0065] A method for constructing a superset index structure combining TRIE and LOUDS specifically includes:

[0066] Step1: Preprocessing stage: Map and sort the sets and elements in the initial set dataset;

[0067] Step2: Index structure construction stage: Construct an index structure combining TRIE and LOUDS;

[0068] Step3: Superset query stage: Execute a superset query in the constructed index structure;

[0069] In this embodiment, it is assumed that the example dataset is the following 9 sets: 4 4 5 7 8 1 4 5 8 2 4 6 1 4 5 4 5 7 8 1 4 3 5 6

[0079] Specifically, Step1 is as follows:

[0080] Step1.1: Count the frequencies of each element in the dataset, and then sort the elements in descending order of frequency;

[0081] In this embodiment, the frequencies of each element in the set are 1:3, 2:1, 3:1, 4:7, 5:5, 6:2, 7:2, 8:3 respectively. The number before the colon is the element, and the number after the colon is the frequency. After sorting in descending order of frequency, it is 4:7, 5:5, 1:3, 8:3, 6:2, 7:2, 2:1, 3:1;

[0082] Step1.2: Replace each element in the set with its subscript in the descending order of frequency, and sort the elements in the set in ascending order of value;

[0083] In this embodiment, by replacing 4 with 1, 5 with 2, 1 with 3, 8 with 4, 6 with 5, 7 with 6, 2 with 7, 3 with 8, and sorting the elements in the set in ascending order of value, the dataset is converted into the following set: 1 1 2 4 6 1 2 3 4 1 5 7 1 3 2 1 2 4 6 1 3 2 5 8

[0093] Step1.3: Sort all the sets sorted in ascending order of element value in ascending order of set size;

[0094] In this embodiment, after sorting all the sets sorted in ascending order of element value in ascending order of set size, it is converted into the following set: 1 1 2 3 4 1 2 4 6 1 2 4 6 1 3 1 3 1 5 7 2 2 5 8

[0104] Step1.4: Number the sorted sets sequentially starting from 0. Denote the number of set s as s.id, denote the finally processed dataset as D, and |D| as the number of sets in the dataset;

[0105] In this embodiment, the set dataset D obtained after numbering the sets is as follows. The first number before the set is the set id; 0 1 1 1 2 3 4 2 1 2 4 6 3 1 2 4 6 4 1 3 5 1 3 6 1 5 7 7 2 8 2 5 8

[0115] The specific method for comparing the set sizes in Step 1.3 is as follows:

[0116] Step 1.3.1: For two specific sets s and t, the loop variable i ranges from 0 to min(|s|, |t|) - 1, and Steps 1.3.2 to 1.3.3 are loop-executed, where |s| and |t| respectively represent the number of elements in s and t, and min(|s|, |t|) represents the minimum value of |s| and |t|;

[0117] Step 1.3.2: If s[i] > t[i], then s > t; if s[i] < t[i], then s < t; otherwise, it means s[i] = t[i], and the next element is continued to be compared; where s[i] and t[i] respectively represent the i-th elements of s and t;

[0118] Step 1.3.3: If the first min(|s|, |t|) elements of s and t are all equal, then continue to judge whether |s| > |t| holds. If it holds, then s > t. If |s| < |t|, then s < t. Otherwise, s = t;

[0119] In this embodiment, for example, s = {1, 2, 4, 6}, t = {1, 2, 3, 4}. First, compare the first elements. Since 1 and 1 are equal, then compare the second elements. 2 and 2 are equal. Then continue to compare the third elements. Since 4 is greater than 3, so s > t;

[0120] The specific content of Step 2 is as follows:

[0121] Step 2.1: Build a TRIE index for the first d elements of each set in the dataset. Hereinafter, d is also referred to as the segmentation depth;

[0122] Step 2.2: Build a hierarchical LOUDS index for the elements after the d-th element of each set;

[0123] Step 2.3: Merge the hierarchical LOUDS structures into a single-layer LOUDS structure. Unless otherwise stated, the LOUDS structure referred to in this patent means a single-layer LOUDS structure;

[0124] In this embodiment, assume d = 2;

[0125] The specific content of Step 2.1 is as follows:

[0126] Step 2.1.1: Create an empty root node root, let the currently visited node current = root, and record the depth of root as 0;

[0127] Step 2.1.2: For each set s ∈ D, loop through Steps 2.1.3 to 2.1.6;

[0128] Step 2.1.3: The loop variable i ranges from 0 to min(|s| - 1, d - 1), and Steps 2.1.4 to 2.1.6 are looped through;

[0129] Step 2.1.4: If a node with label s[i] cannot be found among the children of current, then add a child node N with label s[i] under the current node current, and let current = N; if a node N with label s[i] already exists among the children of current, then directly let current = N;

[0130] Step 2.1.5: If i = |s| - 1, then N is called an end point, that is, the node corresponding to the last element of a certain set. Record N.ESets as the binary tuple <startid, count>, where N.Esets represents all sets that end at node N, startid represents the id of the first set that is exactly the same as s, and count represents the number of sets that are exactly the same as s;

[0131] Step 2.1.6: If i = d - 1, it indicates that N is a node in the last layer of the TRIE. If |s| > d at this time, then N is called a joint point. The joint point is the link connecting the upper - layer TRIE and the lower - layer LOUDS; number the joint points sequentially from 0 according to their insertion order into the TRIE;

[0132] In this embodiment, initially root is an empty node, and current = root;

[0133] The following takes the first two sets as an example to illustrate the index construction process;

[0134] For the set s = {1} with id = 0, when i = 0, since a node with label 1 cannot be found among the children of root, a new child node N is directly created under the root node root, and let current = N;

[0135] Since i = |s| - 1 = 0, then N is an end point, and set N.ESets = <0, 1>, indicating that the starting id of the set that is the same as s is 0 and the quantity is 1;

[0136] For the set s = {1, 2, 3, 4} with id = 0, when i = 0, since a node N with label 1 is found under root, directly set current = N;

[0137] Since i ≠ |s| - 1, then N is not a terminal node. Since i ≠ d - 1, then N is not an articulation point either. After setting i = 1, loop back to Step2.1.3 to continue execution;

[0138] At this time, a node N with label 2 cannot be found under current, so add a child node N with label 2 under current and set current = N;

[0139] Since i ≠ |s| - 1, then N is not a terminal node. Since i = d - 1 and |s| > d, then N is an articulation point. Go to LOUDS to continue inserting the remaining elements 3 and 4 of s;

[0140] Step2.1.7: Let C represent the total number of articulation points in the finally constructed TRIE;

[0141] In this embodiment, the finally created index structure contains C = 2 articulation points, numbered 0 and 1 respectively. As Figure 1 shown, the articulation points represent the nodes with a yellow background;

[0142] The specific content of Step2.2 is as follows:

[0143] Step2.2.1: The core of the hierarchical LOUDS structure lies in constructing 1 integer array Elements, 3 bit arrays NotLeaf, StartOfChild, and EndofSet for each layer. In addition, an array ESets for storing the binary tuple <startid, count> is required;

[0144] Step2.2.2: For the currently inserted set s, loop and execute Step2.2.3 to Step2.2.7;

[0145] Step2.2.3: The loop variable i ranges from d to |s| - 1, and loop and execute Step2.2.4 to Step2.2.7;

[0146] Step2.2.4: Insert s[i] into the Elements array of the (i - d)-th layer, and record the position p where s[i] is inserted into Elements;

[0147] Step2.2.5: If i!= |s| - 1, set the position p of the NotLeaf of the (i - d)-th layer to 1;

[0148] Step2.2.6: If s[i] is the first child of its parent node, set the position p of StartOfChild at the (i - d)-th layer to 1;

[0149] Step2.2.7: If i = |s| - 1, set the p-th position of EndofSet at the (i - d)-th layer to 1, and construct the binary tuple <startid, count> and insert it into ESets at the (i - d)-th layer. Similar to the ESets of the TRIE node, startid represents the id of the first set that is exactly the same as s, and count represents the number of sets that are exactly the same as s;

[0150] In this embodiment, continue to insert the remaining elements 3 and 4 of the set s;

[0151] When i = d = 2, insert s[2] = 3 into the Elements array at the 0-th layer, and obtain its position p = 0 in the Elements at the 0-th layer;

[0152] Since i ≠ |s| - 1, set NotLeaf[p] at the 0-th layer to 1;

[0153] Since s[2] is the first child of the prefix {1, 2}, set StartOfChild[p] at the 0-th layer to 1;

[0154] When i = 3, insert s[3] = 4 into the Elements array at the 1-st layer, and obtain its position p = 0 in the Elements at the 1-st layer;

[0155] Since s[3] is the first child of the prefix {1, 2, 3}, set StartOfChild[p] at the 1-st layer to 1;

[0156] Since i = |s| - 1, set the 0-th position of EndofSet at the 1-st layer to 1, and insert the binary tuple <1, 1> into ESets at the 1-st layer;

[0157] The insertion of the 2-nd set is completed. By analogy, until all sets are indexed, the final arrays of the hierarchical LOUDS structure are as Figure 2 shown. In the figure, L0 and L1 respectively correspond to the 0-th layer and the 1-st layer of the hierarchical LOUDS;

[0158] The specific content of Step2.3 is as follows:

[0159] Step2.3.1: Starting from the 0-th layer of the hierarchical LOUDS, concatenate the Elements of each layer in sequence to form the final array. For convenience, without causing ambiguity, the final array is still called Elements;

[0160] Step2.3.2: Concatenate the arrays of NotLeaf, StartOfChild, EndofSet, and ESets of the hierarchical LOUDS in the same way to form the arrays of NotLeaf, StartOfChild, EndofSet, and ESets of the final single-level LOUDS;

[0161] Step2.3.3: The finally formed single-level Elements, NotLeaf, StartOfChild, EndofSet, and ESets are the single-level LOUDS structure used for superset query;

[0162] In this embodiment, after merging the hierarchical LOUD into a single-level LOUDS, the finally constructed hybrid index structure is as Figure 1 shown. The binary tuple on the left of the node in the figure is the ESets corresponding to the node, the number on the right is the joint point number, and the arrow pointing to the right in the lower layer indicates that the lower part of the TRIE within the left dotted box is replaced by the single-level LOUDS on the right;

[0163] The specific content of Step3 is as follows:

[0164] Step3.1: Given a query set q, set i = 0, and the current node current = root;

[0165] Step3.2: Sequentially check whether q[j] (i ≤ j ≤ |q| - 1) exists among the children of the current node current. If it exists, denote the node corresponding to q[j] as N;

[0166] Step3.3: If N.ESets is not empty, add N.ESets to the result set R;

[0167] Step3.4: If N has children, set current = N, i = j + 1, and then go to Step3.2 to continue searching in the TRIE;

[0168] Step3.5: If N is a joint point, after obtaining the joint point number pno, transfer to search in the LOUDS;

[0169] Step3.6: Obtain the position p1 of the (pno + 1)-th 1 and the next position p2 of 1 in the StartOfChild array. If the next 1 does not exist, then p2 is the maximum length of the array minus 1;

[0170] Step3.7: Sequentially check whether q[k] (j + 1 ≤ k ≤ |q| - 1) exists between the positions p1 and p2 - 1 in the Elements array. If it exists, obtain its position p in the Elements;

[0171] Step 3.8: If the value at the position of EndofSet[p] is 1, then count the number of 1s, denoted as c1, from 0 to p (including p) in EndofSet. Then insert the binary tuple corresponding to ESets[c1 - 1] into the result set R.

[0172] Step 3.9: If the value at the position of NotLeaf[p] is 1, then count the number of 1s, denoted as c2, from 0 to p (including p) in NotLeaf. Set pno = c2 + C, then go to Step 3.6 and continue to search in LOUDS.

[0173] Step 3.10: The query ends, and the content stored in R is the result of the final superset query.

[0174] In this embodiment, given a query set q = {1, 2, 4, 6}, initially i = 0 and the current node current = root.

[0175] For j from i to 2, retrieve the element q[j] from the root respectively. It is found that the elements 1 (corresponding to j = 0) and 2 (corresponding to j = 1) are among the children of the root node.

[0176] For the child N of the root with label 1, since N.ESets is not empty, the corresponding ESets<0, 1> is added to the result set R.

[0177] Since N has children, set current = N, then go to Step 3.2 and continue to check that the children with labels 2 and 3 in the second layer exist among the children of current.

[0178] For the child N with label 2 in the second layer, since N.ESets is empty and it is an articulation point, obtain its articulation point number pno = 0 and then go to continue searching in LOUDS.

[0179] In the StartOfChild array of LOUDS, obtain the position p1 = 0 of the (pno + 1) = 1st 1 and the position of the (pno + 2) = 2nd 1, getting p2 = 2.

[0180] Then search whether q[2] and q[3] exist between p1 and p2 - 1 in Elements, that is, between positions 0 and 1. It is found that q[2] = 4 exists and its position p in Elements is 1.

[0181] Since NotLeaf[p] is 1, then count the number of 1s from 0 to p in NotLeaf, and the number c2 = 2. Then set pno = c2 + C = 2 + 2 = 4, and go to Step 3.6 to continue searching in LOUDS.

[0182] Obtain the position p1 = 0 of the 5th 1 in the StartOfChild array of LOUDS, where pno + 1 = 5. Since the 6th 1 does not exist, set p2 to the maximum length of the array minus 1, which is 5;

[0183] Then, between p1 and p2 - 1 in Elements, that is, between positions 5 and 5, find that q[3] = 6 exists, and its position p in Elements is 5;

[0184] Since the value at the EndofSet[p] position is 1, count the number of 1s from 0 to p (including p) in EndofSet, c1 = 4. Then insert the binary tuple corresponding to ESets[c1 - 1] = <2, 2> into the result set R;

[0185] For the child N of root with label 2, since N.ESets is not empty, add the corresponding ESets <7, 1> to the result set R;

[0186] Since N has children, after setting current = N, go to Step 3.2 and continue to detect that the children with labels 4 and 6 in the second layer do not exist in the children of current;

[0187] The query ends, and the final query result is R = {<0, 1>, <2, 2>, <7, 1>}, that is, q is a superset of sets 0, 2, 3, and 7.

[0188] The present invention can be further illustrated by the following experimental results.

[0189] Experimental environment: The CPU is Intel i7–7700CPU@3.60GHz, the memory is 16GB, the operating system is Ubuntu18.04 64-bit, and the compilation environment is Code Blocks 16.01, gcc 7.5.0.

[0190] Experimental data: The dataset used in the present invention is the DBLP dataset, which contains 781,514 set data. Each set data contains author and title information. The number of independent elements in the set is 517,326, the maximum length of the set is 219, the minimum length is 3, and the average length is 14. Randomly select 5000 sets from the dataset as queries. The query time, index storage space, and comprehensive time-space cost are respectively as Figure 3 、 Figure 4 and Figure 5As shown, the comprehensive time-space cost is the sum of the query cost and the storage cost obtained by normalizing the query time and the storage space respectively (dividing by the maximum query time and the maximum storage space). The data points at both ends of each graph represent the experimental results corresponding to LOUDS (complete LOUDS structure, corresponding to a segmentation depth of 0) and TRIE (complete TRIE, corresponding to an infinite segmentation depth).

[0191] Analysis of experimental results:

[0192] It can be seen from the experiments that the query time gradually decreases as the segmentation depth increases, while the storage space overhead gradually increases as the segmentation depth increases. TRIE corresponds to the case of the maximum storage space, while LOUDS corresponds to the case of the maximum query time. The comprehensive time-space cost reaches the optimum at a segmentation depth of about 4, thus indicating the advantage of combining TRIE and LOUDS in the present invention.

[0193] The specific embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments, and various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those of ordinary skill in the art.

Claims

1. A method for constructing a superset index structure combining TRIE and LOUDS, characterized in that: Step1: Map and sort the sets and elements in the initial set dataset; Step2: Construct an index structure combining TRIE and LOUDS; Step3: Perform a superset query in the constructed index structure; The specific content of Step2 is as follows: Step2.1: Construct a TRIE index for the first d elements of each set in the dataset. Hereinafter, d is also referred to as the splitting depth; Step2.2: Construct a hierarchical LOUDS index for the elements after the d-th element of each set; Step2.3: Merge the hierarchical LOUDS structure into a single-layer LOUDS structure; The specific content of Step2.1 is as follows: Step2.1.1: Establish an empty root node root, let the currently accessed node current = root, and record the depth of root as 0; Step2.1.2: For each set s ∈ D, loop through Steps 2.1.3 to 2.1.6; Step2.1.3: The loop variable i ranges from 0 to min(|s| - 1, d - 1), and Steps 2.1.4 to 2.1.6 are looped through; Step2.1.4: If a node with label s[i] cannot be found among the children of current, add a child node N with label s[i] under the current node, and let current = N; If a node N with label s[i] already exists among the children of current, directly let current = N; Step2.1.5: If i = |s| - 1, then N is called a terminal node, that is, the node corresponding to the last element of a certain set. Denote N.ESets as the binary tuple < starti d, count>, where N.Esets represents all the sets that terminate at node N, startid represents the id of the first set that is exactly the same as s, and count represents the number of sets that are exactly the same as s; Step2.1.6: If i = d - 1, it indicates that N is the node in the last layer of the TRIE. If |s| > d at this time, N is called a joint point. The joint point is the link connecting the upper-layer TRIE and the lower-layer LOUDS; The joint points are numbered sequentially from 0 according to their insertion order into the TRIE; Step2.1.7: Let C represent the total number of joint points in the finally constructed TRIE; The specific content of Step2.2 is as follows: Step2.2.1: For each level of the hierarchical LOUDS structure, one integer array Elements, three bit arrays NotLeaf, StartOfChild, and EndofSet are constructed, and it also includes an array ESets that stores pairs < starti d, count>; Step2.2.2: For the currently inserted set s, loop through Steps 2.2.3 to 2.2.7; Step2.2.3: The loop variable i ranges from d to |s| - 1, and Steps 2.2.4 to 2.2.7 are looped through; Step2.2.4: Insert s[i] into the Elements array of the (i - d)-th layer, and record the position p where s[i] is inserted into Elements; Step2.2.5: If i!= |s| - 1, set the position p of NotLeaf in the (i - d)-th layer to 1; Step2.2.6: If s[i] is the first child of its parent node, set the position p of StartOfChild in the (i - d)-th layer to 1; Step2.2.7: If i = |s| - 1, then set the p-th position of EndofSet in the (i - d)-th layer to 1, and construct the binary tuple < starti d, count> and insert it into the ESets in the (i - d)-th layer. Similar to the ESets of the TRIE node, startid represents the id of the first set that is exactly the same as s, and count represents the number of sets that are exactly the same as s; The specific content of Step2.3 is as follows: Step2.3.1: Starting from the 0th layer of the hierarchical LOUDS, sequentially splice the Elements of each layer to form the final array, still called Elements; Step2.3.2: Concatenate the NotLeaf, StartOfChild, EndofSet, and ESets arrays of the hierarchical LOUDS in the same way to form the NotLeaf, StartOfChild, EndofSet, and ESets arrays of the final single-layer LOUDS. Step2.3.3: The finally formed single-layer Elements, NotLeaf, StartOfChild, EndofSet, and ESets are the single-layer LOUDS structure used for superset queries.

2. The method for constructing a superset index structure combining TRIE and LOUDS according to claim 1, wherein The specific content of Step1 is as follows: Step1.1: Count the frequencies of each element in the dataset, and then arrange the elements in descending order of frequency. Step1.2: Replace the elements in each set with their subscripts in the descending order of frequency, and sort the elements in the set in ascending order of value. Step1.3: Sort all the sets sorted in ascending order of element value in ascending order of set size. Step1.4: Number the sorted sets sequentially starting from 0. Denote the number of set s as s.id, denote the finally processed dataset as D, and |D| as the number of sets in the dataset.

3. The method for constructing a superset index structure combining TRIE and LOUDS according to claim 2, characterized in that, The specific content of Step1.3 is as follows: Step1.3.1: For two specific sets s and t, the loop variable i ranges from 0 to min(|s|, |t|) - 1, and Steps 1.3.2 to 1.3.3 are loop-executed, where |s| and |t| respectively represent the number of elements in s and t, and min(|s|, |t|) represents the minimum value of |s| and |t|. Step1.3.2: If s[i] > t[i], then s > t; if s[i] < t[i], then s < t; otherwise, it means s[i] = t[i], and continue to compare the next element. Where s[i] and t[i] respectively represent the i-th element of s and t. Step1.3.3: If the first min(|s|, |t|) elements of s and t are all equal, then continue to judge whether |s| > |t| holds. If it holds, then s > t. If |s| < |t|, then s < t. Otherwise, s = t.

4. The method for constructing a superset index structure combining TRIE and LOUDS according to claim 1, characterized in that, The specific content of Step3 is as follows: Step3.1: Given a query set q, set i = 0, and the current node current = root. Step3.2: Sequentially detect whether q[j] (i ≤ j ≤ |q| - 1) exists among the children of the current node current. If it exists, denote the node corresponding to q[j] as N. Step3.3: If N.ESets is not empty, add N.ESets to the result set R. Step3.4: If N has children, then set current = N, i = j + 1, and then go to Step3.2 to continue searching in the TRIE. Step3.5: If N is an articulation point, after obtaining the articulation point number pno, transfer to search in the LOUDS. Step3.6: Obtain the position p1 of the (pno + 1)-th 1 and the position p2 of the next 1 in the StartOfChild array. If the next 1 does not exist, then p2 is the maximum length of the array minus 1; Step3.7: Sequentially detect whether q[k] exists between positions p1 and p2 - 1 in the Elements array. If it exists, obtain its position p in the Elements; Step3.8: If the value at the EndofSet[p] position is 1, then count the number c1 of 1s from 0 to p in the EndofSet, and then insert the binary tuple corresponding to ESets[c1 - 1] into the result set R; Step3.9: If the value at the NotLeaf[p] position is 1, then count the number c2 of 1s from 0 to p in the NotLeaf, set pno = c2 + C, and then go to Step3.6 to continue the search in the LOUDS; Step3.10: The query ends, and what is stored in R is the result of the final superset query.

Citation Information

Patent Citations

  • Character string dictionary indexing method and system

    CN103699647A

  • Mobile object convergent pattern mining method based on bit vector quadtree

    CN108182230A