Construction and Retrieval Methods for Index Architectures of Frequent Itemsets in Spatial Big Data

By combining R-tree and concept grid technology, a hybrid index structure is constructed, which solves the problem of low efficiency in frequent item set generation and retrieval in spatial big data, and efficient frequent item set query and index structure construction is achieved.

CN114048212BActive Publication Date: 2025-06-24HENAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111353782.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-16
Publication Date
2025-06-24
Estimated Expiration
2041-11-16

AI Technical Summary

Technical Problem

When processing spatial big data, the generation and retrieval efficiency of frequent item sets is low, resulting in large storage space and long query time.

Method used

Using concept grid technology, the spatial index structure R-tree and concept grid structure are combined to build a hybrid index structure, which is used to index frequent item sets of spatial big data to achieve efficient retrieval.

Benefits of technology

It significantly reduces the time and space cost of index structure construction, improves the efficiency of frequent item set queries, and performs better especially under multi-keyword conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114048212B_ABST
    Figure CN114048212B_ABST
Patent Text Reader

Abstract

The present invention provides a method for constructing and retrieving an index architecture for frequent itemsets in spatial big data. The construction method includes: Step A1: Take out each spatial object d from the spatial big data set, store the location information and number information of d in the R-tree, and construct a spatial index structure; Step A2: Take out the required text attributes therefrom to form a keyword set K, take out all the spatial objects in to form an object set D, and D and K are respectively used as the horizontal axis and the vertical axis of the formal context; Step A3: Traverse all the nodes of, for each concept lattice associated node, traverse the subtree of this node, take out the ids of all the data nodes of the subtree, and take out the corresponding data from the formal context formed in Step A2 to construct a new formal context F, and construct a concept lattice according to the partial order relationship in F; Step A4: Store the concept lattices corresponding to all the concept lattice associated nodes in a list, and jointly form an index structure with.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of spatial big data retrieval, and in particular to a method for constructing and retrieving an index architecture for frequent itemsets of spatial big data. Background Art

[0002] A dataset that records the spatial positions and text attribute features of a large number of spatial entity objects is spatial big data or spatial location big data. This type of data not only records the location information (latitude and longitude, spatial coordinates, etc.) of spatial entities, but also stores a large number of non-spatial attributes with complex structures. An algorithm that retrieves the nearest spatial object set using spatial location information and text attributes as retrieval keywords is called a spatial keyword query algorithm.

[0003] Due to the huge capacity and complex structure of spatial big data, implementing spatial keyword queries for spatial big data often requires the support of a specific spatial data index structure. Existing spatial big data index structures are usually hybrid index structures composed of a spatial location index and a text attribute index. Among them, common spatial index structures include R-tree, Quadtree, Grid structure, etc.; common text index structures include inverted file, bitmap index, etc. Most hybrid index structures separately construct index structures for the spatial attributes and non-spatial attributes of spatial big data to support the retrieval of corresponding attributes of spatial objects.

[0004] Frequent itemset analysis is performed on non-spatial attributes and is a common data mining application target. Currently popular frequent itemset mining algorithms include the Apriori algorithm and the Fptree algorithm, etc. Among them, the Apriori algorithm and the Fptree algorithm find all frequent items by traversing all attribute combinations and store them in a table. The retrieval of frequent items needs to be achieved by traversing the frequent item table. When the data volume is large, this method of generating and retrieving frequent items will occupy a large amount of storage space and have a low retrieval efficiency. Although combining frequent itemset mining algorithms such as the Apriori algorithm and the Fptree algorithm with a spatial index structure can achieve the purpose of frequent itemset mining for spatial big data, in order to ensure that the keywords provided by users can be indexed smoothly, the Apriori and Fptree need to set a very low minimum support degree to ensure that no part of the data is ignored. This leads to a large amount of time and space overhead for calculating frequent itemsets, and the constructed frequent itemset table is lengthy, resulting in an extremely long traversal time during querying. Summary of the Invention

[0005] In view of the problem of large time and space overheads existing in traditional frequent set mining algorithms, the present invention uses a concept lattice as a frequent item set mining tool for big spatial data. Based on concept lattice technology, a hybrid index structure for big spatial data is proposed to mine the frequent items of big spatial data, so as to obtain a set of geographical objects with the most frequent features in a region and meeting the user's index keywords, and to achieve efficient retrieval of frequent item sets of big spatial data.

[0006] On the one hand, the present invention provides a construction method for an index architecture for frequent item sets of big spatial data, including:

[0007] Step A1: Take out each spatial object d from the big spatial data set and store the position information d.p and the number information id of d into the R-tree to construct a spatial index structure

[0008] Step A2: Take out the required text attributes from the big spatial data set to form a keyword set K, which serves as the horizontal axis of the formal context; take out all the spatial objects in the big spatial data set to form an object set D, which serves as the vertical axis of the formal context;

[0009] Step A3: Traverse all the nodes of . For each concept lattice associated node, traverse the subtree of this node, take out the ids of all the data nodes in the subtree, and take out the corresponding data from the formal context formed in Step 2 to construct a new formal context F. According to the partial order relationship in F, construct the concept lattice corresponding to the current concept lattice associated node; wherein, the concept lattice associated node refers to an R-tree node whose number of data nodes in the subtree is within the range of [δ min , δ max ;

[0010] Step A4: Store the concept lattices corresponding to all the concept lattice associated nodes into the list , which, together with , forms an index structure for frequent item sets of big spatial data

[0011] Furthermore, the formula of the spatial index structure is shown in Equation (1):

[0012]

[0013] wherein, r represents the root node of , θ = [θ min , θ max is the range of the number of branches of the node, <n1, n2,..., ni > is the node set of, n i = <id, mbr, level, pn, cns, dn, ds> represents a node of, where id is the node number, mbr is the node MBR range, level is the level of the node in the tree, pn and cns are the parent node and the set of child nodes of the node respectively, and ds and dn are the set of data nodes and the number of data nodes of the subtree of the node respectively.

[0014] Furthermore, the list has the formula shown in Equation (2):

[0015]

[0016] Among them, represents a concept lattice of, nid represents the id of the concept lattice associated node in is the concept of L i , ≤ is the partial order relation, L i .F.size represents the data volume of the formal context F corresponding to L i , δ = [δ min , δ max .

[0017] On the other hand, the present invention provides a retrieval method for frequent itemsets of spatial big data, including:

[0018] Step B1: Generate a first-level frequent item query is the query coordinate, is the query keyword attribute, and satisfies k is the number of spatial objects requested by the index, is the index structure for frequent itemsets of spatial big data;

[0019] Step B2: Traverse For each in, determine the position of the MBR of each node in . If there exists and for each there exists then perform Step B3; n.mbr represents the minimum bounding rectangle of node n in the R-tree, n s represents the child node of node n, n.cns represents the set of child nodes of node n, n s .mbr represents the minimum bounding rectangle of node n s ;

[0020] Step B3: If the node n obtained in Step B2 is a concept lattice associated node, then retrieve the concept lattice associated with this node from and proceed to Step B4; if the node n obtained in Step B2 is not a concept lattice associated node, then search upward or downward for the nearest set of concept lattice associated nodes and proceed to Step B4;

[0021] Step B4: Traverse all the concept lattices obtained in Step B3 respectively. Specifically: for each concept lattice, start traversing the lattice structure from the top-level concept downward. If the intension C i of the currently traversed concept C i .intent satisfies then calculate the frequency of each direct sub-concept of this concept;

[0022] Step B5: Use the scoring formula to score and sort the extensions of each concept obtained in Step B4, and return the top k objects with higher scores as the retrieval results to the query user.

[0023] Furthermore, in Step B3, if the node n obtained in Step B2 is not a concept lattice associated node, then search upward or downward for a concept lattice associated node and proceed to Step B4, specifically including:

[0024] If the ancestor node of this node is a concept lattice associated node, then execute Step B4 for the concept lattice corresponding to this ancestor node; if the descendant node of this node is a concept lattice associated node, then retrieve the concept lattice associated node at the highest level of all sub-branches of this node and proceed to Step B4.

[0025] Furthermore, before Step B5, it also includes:

[0026] If the number of extensions of the concept obtained after Step B4 is less than k, then return to Step B2 and re-execute Steps B2 to B4 until the number of extensions is greater than k or the root node r is found.

[0027] Furthermore, the scoring formula is as shown in formula (3):

[0028]

[0029] where is the Euclidean distance, max(dist) is the maximum Euclidean distance from the query point among all objects to be sorted, and max(L.g(d i .intent).size) is the frequency of the concept to which this object belongs.

[0030] Advantageous effects of the present invention:

[0031] (1) By combining the spatial index structure R-tree with the concept lattice structure, this invention uses the R-tree structure to index the spatial attributes of spatial big data and the concept lattice structure to index the non-spatial attributes of the data in specific R-tree nodes, achieving the purpose of performing frequent item spatial keyword queries in spatial big data.

[0032] (2) Compared with the combined index structures using traditional frequent set mining algorithms such as Apriori or Fptree and spatial indexes, this invention has great advantages in terms of the time and space costs of constructing the index structure. In terms of query performance, compared with Apriori and Fptree which need to traverse long lists of frequent sets, the concept lattice only needs to traverse some concepts to complete the frequent set query, with high query performance. Especially under the condition of multiple keywords, the advantage is more obvious.

[0033] (3) The new index result scoring mechanism designed in this invention pays more attention to frequency, giving users more reference options. Brief Description of the Drawings

[0034] Figure 1 It is a schematic flowchart of the construction method of the index architecture for frequent item sets of spatial big data provided by the embodiment of this invention. Detailed Embodiments

[0035] To make the purpose, technical solutions and advantages of this invention clearer, the technical solutions in the embodiments of this invention will be clearly described below in conjunction with the drawings in the embodiments of this invention. Obviously, the described embodiments are part of the embodiments of this invention, rather than all of them. Based on the embodiments of this invention, all other embodiments obtained by those of ordinary skill in the art without creative work belong to the scope of protection of this invention.

[0036] Embodiment 1

[0037] The embodiment of this invention provides a construction method for an index architecture for frequent item sets of spatial big data, including:

[0038] S101: Take out each spatial object d from the spatial big data set and store the location information d.p and the number information id of d into the R-tree, constructing the spatial index structure

[0039] It can be understood that each spatial object stored in the R-tree is called a leaf node of the R-tree, also called a data node. According to the classic R-tree construction algorithm, with the insertion of data nodes, some nodes in the tree split and the R-tree grows upward. As the data set All the spatial objects in have been inserted, and the R-tree construction is completed.

[0040] In this embodiment, the symbol is used to represent the R-tree;

[0041] As an implementable manner, the formula of the spatial index structure is shown in Equation (1):

[0042]

[0043] where r represents the root node, θ = [θ min , θ max is the range of the number of branches of the node, <n1, n2,..., n i > is the node set, n i = <id, mbr, level, pn, cns, dn, ds> represents a node of, id is the node number, mbr is the node MBR range, level is the level of the node in the tree, pn and cns are the parent node and the child node set of the node respectively, ds and dn are the data node set and the number of data nodes of the subtree of the node respectively.

[0044] S102: Extract the required text attributes from the spatial big data set to form a keyword set K, which serves as the horizontal axis of the formal context; Extract all the spatial objects from the spatial big data set to form an object set D, which serves as the vertical axis of the formal context;

[0045] For whether each object in the object set D satisfies the attributes of the ordinate, it is 1 or 0 at the corresponding position. In this embodiment, the formal context in step S102 can be saved as a DataFrame structure in Python and stored in memory, or saved as a csv file and stored on disk.

[0046] S103: Traverse all the nodes of , for each concept lattice associated node, traverse the subtree of the node, extract the ids of all the data nodes of the subtree, and extract the corresponding data from the formal context formed in step S102 to construct a new formal context F. According to the partial order relationship in F, construct the concept lattice corresponding to the current concept lattice associated node; where, the concept lattice associated node refers to the R-tree node whose number of data nodes in the subtree is in the range of [δ min , δ max ;

[0047] The object of the present invention is to index spatial objects representing regional features. It can be understood that if the formal context is too large, the area covered by a formal context is also too large, and the regional features of the indexed spatial objects are not obvious, and the construction and indexing loss of the concept lattice will also cause waste. On the contrary, if the formal context is too small, the regional characteristics cannot be reflected, and the advantages of the concept lattice cannot be reflected. Therefore, in the embodiments of the present invention, a data volume range parameter δ = [δ min , δ max is set as the basis for the concept lattice associated with the R-tree node, and the R-tree node with the number of data nodes in the subtree within the range of [δ min , δ max is defined as a "concept lattice associated node".

[0048] As an implementable manner, the formula for the concept lattice L constructed in this step is: where id is the id of the concept lattice associated node, F is the formal context of L, is the concept of L i , and ≤ is the partial order relation.

[0049] S104: Store the concept lattices corresponding to all concept lattice associated nodes into the list , and together form an index structure for frequent itemsets of spatial big data

[0050] As an implementable manner, the formula of the list is shown in formula (2):

[0051]

[0052] where represents a concept lattice of , nid represents the id of the concept lattice associated node in , is the concept of L i , ≤ is the partial order relation, and L i .F.size represents the data volume of the formal context F corresponding to L i , and δ = [δ min , δ max .[[]]

[0053] A construction method for an index architecture for frequent item sets in spatial big data provided by an embodiment of the present invention combines a spatial index structure R-tree with a concept lattice structure, achieving the purpose of performing frequent item spatial keyword queries in spatial big data. Compared with using traditional frequent set mining algorithms such as Apriori or Fptree combined with a spatial index structure, it greatly saves the time and space overhead of constructing the index structure.

[0054] Embodiment 2

[0055] Based on the above embodiment, an embodiment of the present invention provides a retrieval method for frequent item sets in spatial big data, including the following steps:

[0056] S201: Generate a first-level frequent item query is the query coordinate, is the query keyword attribute, and satisfies k is the number of spatial objects for the index request, is the index structure for frequent item sets in spatial big data;

[0057] S202: Traverse Compare each node MBR in with to determine the position. If there exists and for each there exists then proceed to step S203; n.mbr represents the minimum bounding rectangle of node n in the R-tree, n s represents the child node of node n, n.cns represents the set of child nodes of node n, n s .mbr represents the minimum bounding rectangle of node n s ;

[0058] S203: If the node n obtained in step S202 is a concept lattice association node, then retrieve the concept lattice associated with this node from and proceed to step S204;

[0059] If the node n obtained in step S202 is not a concept lattice association node, then search up or down for the nearest set of concept lattice association nodes. If the ancestor node of this node is a concept lattice association node, then execute step S204 for the concept lattice corresponding to this ancestor node; if the descendant node of this node is a concept lattice association node, then retrieve the highest-level concept lattice association nodes of all sub-branches of this node and proceed to step S204;

[0060] For example, if is in For the fifth - layer nodes, if the third layer of the sub - trees of these nodes are all "concept lattice nodes", then the corresponding concept lattices of these nodes are entered into step S204 for operation, and the results are integrated.

[0061] S204: Traverse all the concept lattices obtained in step S203 respectively. Specifically: For each concept lattice, start traversing the lattice structure from the top - level concept downwards. If the current concept C i of the intension C i .intent satisfies then calculate the frequency of each direct sub - concept of this concept;

[0062] Specifically, in order to make the index results more region - characteristic, this step deeply mines the hierarchical association of the concept lattice, calculates the frequency of each direct sub - concept of this concept, which is equivalent to further refining the object set that has already satisfied the keyword . Under such refinement, the regional characteristics of these objects can be better reflected. In this step, the number of extensions of these concepts is respectively counted as the frequency of the concept extension. The extension and its frequency of each concept enter step S205 as the result of step S204.

[0063] It should be noted that if in the concept lattice, points are used to represent concepts and lines are used to represent associations, "the direct sub - concept of a concept" can be understood as the sub - concept directly associated with this concept.

[0064] As an implementable way, before performing step S205, it also includes:

[0065] If the number of extensions of the concept obtained after step S204 is less than k, then return to step S202, and re - execute steps S202 to S204 until the number of extensions is greater than k or the root node r is found.

[0066] S205: Use the scoring formula to score and sort the extensions of each concept obtained in step S204, and return the top k objects with higher scores as the retrieval results to the query user.

[0067] As an implementable way, the scoring formula is as shown in formula (3):

[0068]

[0069] Among them, is the Euclidean distance, max(dist) is the maximum Euclidean distance from the query point among all the objects to be sorted, and max(L.g(d i .intent).size) is the frequency of the concept to which this object belongs.

[0070] Specifically, is the Euclidean distance evaluation score, with a value range of (0, 1). The closer the distance, the higher the score. And is the frequency score, with an integer value. It can be seen that the frequent item space keyword query in this embodiment takes the object frequency as a higher priority, and the final index result can better reflect the regional characteristics. By sorting the result of step S204 using the above scoring formula, the top k objects are returned to the query user as the result. Thus, a frequent set space keyword query is completed.

[0071] The retrieval method provided by the embodiment of the present invention further refines the concepts in the concept lattice that meet the requirements of the index keywords, so that the result can better reflect the regional characteristics. Moreover, a new scoring mechanism is used to sort the result obtained after the spatial index and the frequent set operation. This scoring mechanism can further highlight the regional characteristics of the index result.

[0072] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A construction method for an index architecture of frequent itemsets in spatial big data, characterized in that Including: Step A1: Take out each spatial object d from the spatial big data set and store the location information d.p and the number information id of d into the R-tree. Each spatial object stored in the R-tree is called a leaf node of the R-tree, and the spatial index structure is constructed The spatial index structure is shown by formula (1) as follows: where r represents the root node of min , θ = [θ max is the range of the number of branches of the node, <n1, n2,..., n i > is the node set of i n = <id, mbr, level, pn, cns, dn, ds> represents a node of , where id is the node number, mbr is the minimum bounding rectangle range of the node, level is the level of the node in the tree, pn and cns are the parent node and the child node set of the node respectively, and ds and dn are the data node set and the number of data nodes of the subtree of the node; Step A2: Extract the required text attributes from the spatial big dataset to form a keyword set K, which serves as the horizontal axis of the formal context; extract all spatial objects from the spatial big dataset to form an object set D, which serves as the vertical axis of the formal context. Step A3: Traverse all nodes of min . For each concept lattice associated node, traverse the subtree of this node, take out the ids of all data nodes in the subtree, and take out the corresponding data from the formal context formed in Step A2 to construct a new formal context F. According to the partial order relationship in F, construct the concept lattice corresponding to the current concept lattice associated node; wherein, the concept lattice associated node refers to an R-tree node whose number of data nodes in the subtree is within the range of [δ max . Step A4: Store the concept lattices corresponding to all the associated nodes of the concept lattices into a list in which, together with forms an index structure for frequent itemsets of spatial big data 2. The construction method of an index architecture for frequent itemsets of spatial big data according to claim 1, wherein The said list has a formula as shown in Equation (2): Among them, denotes a concept lattice, where nid denotes the id of the associated node of the concept lattice in is a concept of L i , ≤ is a partial order relation, and L i .F.size represents the data volume of the formal context F corresponding to L i , and δ = [δ min , δ max .

3. A retrieval method for frequent itemsets of spatial big data, characterized in that, Applied to the index architecture constructed by the construction method according to claim 1 or 2, the retrieval method includes: Step B1: Generate a first frequent item query is the query coordinate, is the query keyword attribute, and satisfies k is the number of spatial objects for the index request, is the index structure for the frequent item set of spatial big data; Step B2: Traverse each node MBR in and perform a position judgment with If there exists and for each there exists s then go to Step B3; n.mbr represents the minimum bounding rectangle of node n in the R-tree, n s represents the child node of node n, n.cns represents the set of child nodes of node n, and n s .mbr represents the minimum bounding rectangle of node n Step B3: If the node n obtained in Step B2 is a concept lattice associated node, then retrieve the concept lattice associated with this node from and perform Step B4; if the node n obtained in Step B2 is not a concept lattice associated node, then search upward or downward for the nearest set of concept lattice associated nodes and perform Step B4; Step B4: Traverse all the concept lattices obtained in Step B3 respectively, specifically: for each concept lattice, start traversing the lattice structure from the top-level concept downwards. If the intension C i of the current concept C i .intent satisfies then calculate the frequency of each direct sub-concept of this concept; among them, in the concept lattice, a concept is represented by a point and an association is represented by a line, then the direct sub-concept of a concept represents the sub-concept directly associated with the concept; the number of extensions of the concept is used as the frequency of the concept. Step B5: Use the scoring formula to score and sort the extensions of each concept obtained in step B4, and return the top k objects with higher scores as the retrieval results to the query user; the scoring formula is as shown in formula (3): Among them, represents a certain extension of the current concept, where 0 ≤ i ≤ k, is the Euclidean distance, and max(dist) is the maximum Euclidean distance from the query point among all objects to be sorted, is the frequency of the concept to which the object belongs.

4. A retrieval method for frequent itemsets of spatial big data according to claim 3, characterized in that In step B3, if the node n obtained in step B2 is not a concept lattice associated node, then search up or down for a concept lattice associated node and perform step B4, specifically including: If the ancestor node of this node is a concept lattice associated node, then perform step B4 for the concept lattice corresponding to this ancestor node; if the descendant node of this node is a concept lattice associated node, then take out the concept lattice associated nodes at the highest level of all sub-branches of this node and perform step B4.

5. A retrieval method for frequent itemsets of spatial big data according to claim 3, characterized in that, Before step B5, it further includes: If the number of extensions of the concept obtained after step B4 is less than k, then return to step B2 and re-execute steps B2 to B4 until the number of extensions is greater than k or the root node r is found.

Citation Information

Patent Citations

  • Formal concept lattice-based faceted search method and system

    CN107391584A

  • Concept lattice-based association rule optimization method and visual display method

    CN112597236A