A Nearest Neighbor Spatial Set Keyword Query Method Based on MapReduce

The MapReduce-based Hilbert R-tree indexing with leaf node pruning and circular scanning pruning optimizes spatial collection keyword queries, addressing inefficiencies in large data environments by reducing index size and query response time.

CN116303521BActive Publication Date: 2025-07-15XIAN UNIV OF TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211325865.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-27
Publication Date
2025-07-15
Estimated Expiration
2042-10-27

AI Technical Summary

Technical Problem

The existing technology cannot effectively implement multi-keyword spatial set query in large data environments, especially when data updates are fast and data volumes are large, real-time performance is insufficient. The existing algorithm assumes that spatial index and inverted index have been completed, and there are unripe applications.

Method used

Using a MapReduce-based method, combined with distributed Hilbert R-tree index, two-point pairing algorithm and leaf node diagonal pruning method, fast spatial collection keyword query is performed through the circular scanning pruning method, and real-time index establishment of map applications is used using the Hadoop framework.

Benefits of technology

Effectively reduce the index volume, shorten the search time, improve the system response speed, promote the practicality of keyword query in the collection space, and adapt to heterogeneous data sources and fast update data environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116303521B_ABST
    Figure CN116303521B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for nearest neighbor spatial set keyword query based on MapReduce. Based on the Hadoop framework, by using the way of establishing a distributed Hilbert R-tree index in real time, the method adopts a two-point pairing algorithm combined with a leaf node diagonal pruning method and a circular scanning pruning method to complete the fast query of spatial set keywords, and experiments are carried out in map applications. The experimental results show that the design of the distributed Hilbert R-tree index can effectively reduce the index volume and cope with the situation of heterogeneous data sources and rapid data updates; the two-point pairing algorithm can greatly reduce the retrieval space, speed up the response speed of the system, and promote the practical application of set spatial keyword query.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data query, and particularly relates to a nearest neighbor spatial set keyword query method. Background Art

[0002] With the emergence of positioning technologies and the popularity of mobile devices, many web pages, social networks, and other data have added "location" tags. Spatial keyword queries based on keyword and location matching have grown rapidly on the Internet, greatly enriching the resources of map-based location services, changing the way people query locations in daily life, and better serving people's lives. Among various location information queries based on keywords, single-keyword queries are relatively common, such as searching for "gas station" or "convenience store" in Baidu Map. As requests become increasingly complex, it becomes more and more difficult for multiple keywords to be contained by a single spatial object, and the query target gradually changes from a single spatial object to a set of multiple spatial objects. For example, querying {"gas station", "restaurant"} can find nearby location areas.

[0003] This method of searching from a single object to the area where the object is located is also called spatial set query, which has attracted the attention of researchers in recent years. Due to the limited storage and computing capabilities of traditional regional databases, spatial set queries in a large data environment are still in theoretical research, and there is no mature map-based spatial set query application.

[0004] The methods for spatial object queries facing big data mainly include the following two: 1) directly applying the existing environment for parallel computing of big data to solve spatial query problems, such as using MapReduce to solve K-NN queries; 2) extending the existing big data systems to solve spatial range query problems, such as using TAREEG based on Hadoop to solve spatial query problems and MoDisSENSE to solve time and space level queries. Currently, the set spatial keyword query facing big data is mainly in the research stage and there is no relevant practical application. The reason is that the query algorithms mostly use a direct combination method to implement algorithms in a big data environment; and all existing algorithms assume that the spatial index and inverted index have been completed, and there are still relatively large defects in terms of fast data update, large data volume, and real-time performance.

[0005] Hilbert R-tree is an effective index structure based on the basic structure of R-tree. Using the excellent clustering property of Hilbert curve, it maps high-dimensional data to one dimension, preserves most of the spatial information, and realizes the effective organization of R-tree data. This method has the characteristics of a fast tree structure, few external memory accesses, and easy parallel processing. Hilbert R-tree is suitable for dealing with the situation of unevenly scattered point data sets. For actual point data, the construction of Hilbert R-tree can be divided into static and dynamic. Summary of the Invention

[0006] In order to overcome the deficiencies of the prior art, the present invention provides a method for nearest neighbor spatial set keyword query based on MapReduce. Based on the Hadoop framework, by using the way of establishing a distributed Hilbert R-tree index in real time, the method adopts a two-point pairing algorithm combined with a leaf node diagonal pruning method and a circular scanning pruning method to complete the rapid query of spatial set keywords, and experiments are carried out in map applications. The experimental results show that the design of the distributed Hilbert R-tree index can effectively reduce the index volume and cope with heterogeneous data sources and rapid data updates; the two-point pairing algorithm can greatly reduce the retrieval space, speed up the response speed of the system, and promote the practical application of set space keyword query.

[0007] The technical solutions adopted by the present invention to solve its technical problems include the following steps:

[0008] Step 1: The format of map location object data includes: 1) keywords contained in the data; 2) longitude and latitude coordinates of the data; all data is stored in HDFS;

[0009] Step 2: Set the query keyword, read the object data from the block of HDFS. During the process of reading the object data, according to the keywords contained in the object data, place the object data that meets the query conditions in the memory, and generate the keyBit of the object data according to the query keyword. The keyBit is the binary code representing the keywords contained in the object data; then serialize the HilbertCurve object, calculate its Hilbert value according to the longitude and latitude coordinates of the HilbertCurve object, establish a Hilbert R-tree, and pass the Hilbert R-tree to all computing nodes;

[0010] Step 3: Leaf node diagonal pruning method;

[0011] Step 3-1: Generate a leaf node, and judge the query keywords contained in the leaf node. If the leaf node contains all the query keywords, then set the diagonal distance of the leaf node as the threshold;

[0012] Step 3-2: Repeat Step 3-1. If the newly generated leaf node contains all the query keywords and the diagonal distance is less than the threshold, then update the threshold with the diagonal distance of the newly generated leaf node until all the leaf nodes are generated; the finally obtained threshold is denoted as S;

[0013] Step 4: Circular scanning pruning method;

[0014] Traverse all object data. Taking the data object as the center of the circle, construct a circle with √3*S as the diameter, and determine whether all query keywords are included within the area of this circle. If not all query keywords are included, delete the object data from the object set;

[0015] Step 5: Traverse the remaining object data, and pair the object data pairwise using the two-point pairing method to form object pairs;

[0016] Step 6: Object pairing method;

[0017] Step 6-1: Pair the root node of the Hilbert R-tree with itself;

[0018] Step 6-2: For non-leaf node pairs, list all child node pairs. If the minimum distance between the rectangles corresponding to the two nodes is greater than S, ignore the node pair; otherwise, continue to pair the child nodes; for leaf node pairs, list all object pairs;

[0019] Step 6-3: Repeat 6-2 until all object pairs are paired;

[0020] Step 7: If the distance between the object pair is less than the threshold S, set the distance of the object pair as the key, and the two object data as the value and output to reduce;

[0021] Step 8: In the reduce stage, scan the area where each object pair is located. If an object set with the key as the diameter can be found, it is used as the candidate output result; select the top ten candidates with the smallest distance for serialization and output to the corresponding file.

[0022] Furthermore, the specific two-point pairing method is as follows:

[0023] The goal of the collective space keyword query is to find a set of several objects with the closest positions, where the measure of closeness is described by the "diameter" of the object set, that is, the maximum value of the pairwise distances in the object set is used as the diameter; the collective keyword query is to find an object set that can contain all keywords and has the smallest diameter;

[0024] First, combine the objects pairwise to form object pairs, then sort the object pairs in ascending order according to the distance, and finally expand each object pair into an object set one by one. If an object pair can be expanded into an object set, then this object set is the desired result.

[0025] Furthermore, the specific leaf node diagonal pruning method is as follows:

[0026] The construction of the Hilbert R-tree adopts a bottom-up approach. First, leaf nodes are generated. The leaf nodes use the minimum bounding rectangle as their spatial attribute. If a leaf node contains all the query keywords, then a set of objects that satisfy all the query keywords can be found in this leaf node, and its diameter is less than the diagonal distance of this leaf node. Therefore, the diagonal distance of this leaf node is the upper limit of the result of the set space keyword query.

[0027] When generating the Hilbert R-tree, the generated leaf nodes are judged. If all the query keywords are included, then the diagonal distance of this leaf node is the upper limit of the result. After that, as the leaf nodes are generated, this upper limit value is updated. When all the leaf nodes are generated, the diagonal distance of the smallest leaf node that meets this condition is the minimum upper limit. This minimum upper limit is used as the threshold.

[0028] The beneficial effects of the present invention are as follows:

[0029] 1. A query strategy based on spatial big data is proposed for the practical problem of set space keyword query, enriching the location service based on the map.

[0030] 2. A query algorithm using Hilbert R-tree spatial search is designed: based on the two-point pairing method, the diagonal pruning of leaf nodes and the circular scanning pruning process are added, effectively reducing the traversal of the retrieval space.

[0031] 3. The specific implementation methods and steps of the algorithm in the MapReduce parallel environment are completed. Experiments prove that this algorithm can read query-related data into memory and complete the set keyword query ability in the case of large-scale data.

[0032] 4. The experimental results show that the proposed query strategy can effectively reduce the time overhead: it proves the significant improvement of the Hilbert R-tree spatial search algorithm for keyword search efficiency in the MapReduce environment. Description of the Drawings

[0033] Figure 1 It is a comparison of the retrieval times of the query algorithm proposed in the present invention in the MapReduce environment for different numbers of machines and different numbers of keywords.

[0034] Figure 2 It is a schematic diagram of the scanning area of the two-point pairing algorithm of the present invention.

[0035] Figure 3 It is a schematic diagram of the diagonal pruning method of the leaf nodes of the present invention.

[0036] Figure 4It is a schematic diagram of the circular scanning pruning method of the present invention.

[0037] Figure 5 It is the flowchart of the query algorithm proposed by the present invention;

[0038] Figure 6 It is the comparison of the effectiveness verification results of the pruning method proposed by the present invention. Detailed implementation manners

[0039] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0040] Aiming at the problem of set spatial keyword query in map applications, based on the MapReduce computing framework, the present invention proposes a set keyword query strategy in the big data environment. Using the Hilbert R-tree as an index and adopting a combination method based on two-point pairing can effectively improve the response speed of the query system. Use MapReduce to complete the design and implementation of the spatial set query algorithm based on the big data environment, and query applications based on actual map positions. In addition, the algorithms involved in the present invention can be applied to problems related to location, social media data analysis, and location-sensitive business operations. It plays an important role in personal activities, company activities, and urban operations and has good application prospects.

[0041] The core idea of the present invention is divided into two points: First, in the case of a large amount of data for map applications, due to limited memory, it is impossible to directly read all the data into memory. However, the amount of data related to the keywords actually retrieved is only a very small part of the overall data. It is possible to establish a real-time index only for this small part of the data in memory and complete the query work using this index and an efficient query algorithm. Second, a two-point pairing algorithm in the parallel framework Hadoop environment is proposed, and corresponding pruning rules are designed, which can complete the compression and pruning of the retrieval space under cluster machines and improve the query speed.

[0042] The present invention designs a spatial data query algorithm based on the Hilbert R-tree: based on the two-point pairing method, the diagonal pruning of leaf nodes and the circular scanning pruning process are added.

[0043] 1. Two-point pairing method

[0044] Basic idea of the two-point pairing method: The goal of keyword query in the set space is to find a set of several objects that are the most "similar" in location. The measure of "similarity" is described by the "diameter" of the object set, that is, the maximum value of the pairwise distances in the object set is used as the diameter. The larger the diameter of an object set, the farther the distances between the objects in the set. The set keyword query is to find an object set that can contain all keywords and has the smallest diameter. The essence of solving the set keyword query is to compare the diameters of different object sets. Since the number of object sets grows exponentially, the present invention uses a two-point pairing algorithm to complete the search process of the object set. First, the objects are pairwise combined to form "object pairs", then the "object pairs" are sorted in ascending order, and finally, the "object pairs" are gradually expanded into "object sets" one by one. If an object pair can be expanded into an object set, then this object set is the desired result. With the help of a tree-shaped space index, this method can change the node set that originally needed to be enumerated exponentially into a polynomial level, avoiding the query delay caused by the explosive growth during the node enumeration process.

[0045] When expanding an "object pair" into an "object set", the supplementary object points will be concentrated in the area near the object pair, as Figure 2 shown in the shaded part, that is, the overlapping area of two circles with A and B as the centers and |AB| (the distance between A and B) as the radius. In this way, it can be ensured that the distance between the objects in the area and A or B is not greater than |AB|. In the algorithm, first, it is determined whether the area contains all the query keywords, and then the combination of objects is carried out. Here, only the objects in the area need to be combined, which can greatly reduce the overhead of object combination compared with the overall query objects.

[0046] 2. Pruning of leaf node diagonals and circular scanning pruning

[0047] Basic idea of pruning leaf node diagonals: The construction of the Hilbert R-tree is carried out in a bottom-up manner. First, leaf nodes are generated, and the nodes use the minimum bounding rectangle as their spatial attribute. If a node contains all the query keywords, then it is certain that an object set that satisfies all the keywords can be found in this node, and its diameter is less than the diagonal distance of this node. Therefore, the diagonal distance of this node is the upper limit of the result of the keyword query in the set space. As Figure 3 shown, the rectangle of a leaf node contains four objects A, B, C, and D, and A, B, C, and D contain all the query keywords. Then dis(A, C), which is the diagonal distance, is the diameter of the object set {A, B, C, D}, and it is less than or equal to the diagonal distance of the rectangle.

[0048] In this way, when generating the Hilbert R-tree, the generated leaf nodes can be judged. If all keywords are included, the diagonal distance of the leaf node is the upper limit of the result. Then, as the leaf nodes are generated, the upper limit value is updated. When all leaf nodes are generated, the diagonal distance of the smallest leaf node that meets this condition is the minimum upper limit. Using this minimum upper limit as a threshold, when generating object pairs, all object pairs must be smaller than this threshold, which can effectively reduce the generation of object pairs and save memory overhead.

[0049] The basic idea of circular scan pruning is to further judge a single object on the basis of the diagonal pruning rule of leaf nodes. When there is an object point that does not meet the condition, this object is directly deleted from the memory.

[0050] In Figure 4 the intersecting part is the area searched for generating the "object set" from the "object pairs" mentioned in the previous text. Suppose MN is the "object pair" to be judged, and the smallest object set exists in the shaded part. Then it can be determined that the diameter of the object set is less than 3*|MN| (the distance between M and N), that is, the green circle in the figure. In addition, in the two-point pairing method of diagonal pruning of leaf nodes, the diagonal distance S of the smallest leaf node can be obtained as a threshold by judging the leaf nodes. It can be known that the diameter of the object set is less than S, and |MN| is less than or equal to S. Using a circle with a diameter of √3*S, with each object as the center point, a region is rotated. If it is found that after a full rotation, all query keywords are not found in this circular region, then this object cannot be a point in the diameter of the object set and can be deleted from the dataset.

[0051] 3. Specific steps of the algorithm based on the MapReduce framework Figure 5 as shown.

[0052] Under the Hadoop distributed framework, the specific query steps of the present invention are designed using the MapReduce framework. The format of the object data includes: 1) the keywords included in the data; 2) the longitude and latitude coordinates of the data. All data is stored in HDFS.

[0053] 1): In the Mapper class, data is read from the blocks of HDFS. During the data reading process, according to the keywords of the object, the points that meet the query conditions are placed in the memory, and the keyBit of the data is generated according to the query keywords. The keyBit is the binary code indicating that the data contains keywords. Then the HilbertCurve object is serialized, the hilbert value is obtained according to the longitude and latitude coordinates of the object, a Hilbert R-tree is established, and the Hilbert R-tree is passed to all computing nodes.

[0054] 2): Generate leaf nodes, and judge the keywords included in the leaf nodes. If all keywords are included, then the diagonal distance of this node is the threshold value.

[0055] 3): Repeat step 2). If a certain leaf node satisfies all keywords and the diagonal distance is less than the threshold value, then update the threshold value to this diagonal distance until all leaf nodes are generated.

[0056] 4): Traverse all objects o, rotate a circle with a diameter of √3*S around o, and judge whether all the keywords to be queried are included in the scanned circular area of this point. If all keywords are not satisfied after one week, then delete this point from the object set.

[0057] 5): Traverse the points and pair the remaining objects in pairs.

[0058] 6): Use the Hilbert R-tree to judge whether all the keywords to be queried are included in the area of the point pair in the Hilbert R-tree.

[0059] 7): Use the point pair distance as the key and the two point objects as the value and output to reduce.

[0060] 8): In the reduce stage, scan the area where each object pair <A, B> is located. If an object set with <A, B> as the diameter can be found, then use it as the candidate output result. Select the top ten candidate objects for serialization and output them to the corresponding file if applicable. Specific embodiments:

[0062] To prove the effectiveness of the present invention, a comparison of experimental results was carried out: The experimental data is a text file named with keywords. Each line contains a data point, and this data point has three attributes, id, x coordinate, and y coordinate. 195,678 groups of data were set, and 376 keywords were set.

[0063] Such as Figure 1 . Applying the algorithm proposed by the present invention, when the number of keywords is 3, 4, and 5 respectively, the running times under different numbers of machines are compared. Obviously, it can be seen that as the number of machines increases, the running time becomes smaller and smaller, showing a linear decrease. This shows that the search efficiency in a distributed environment is much higher than that in a single machine. Therefore, putting the Hilbert R-tree spatial search algorithm in the MapReduce environment will significantly improve its search efficiency.

[0064] In addition, to verify the effectiveness of the two pruning methods, the running results of diagonal pruning of leaf nodes and circular scanning pruning will be compared with those without using these two pruning methods. The experiment is run on three machines in a MapReduce environment, and the effectiveness of the two pruning algorithms is verified by comparing the relationship between different keyword quantities and running times. As Figure 6 it can be clearly seen that the method with the pruning method has a certain improvement in running time when searching for the same number of keywords.

Claims

1. A method for nearest neighbor spatial set keyword query based on MapReduce, characterized in that, It includes the following steps: Step 1: The format of the map location object data includes: 1) the keywords contained in the data; 2) the longitude and latitude coordinates of the data; all data is stored in HDFS; Step 2: Set the query keyword, read the object data from the block of HDFS. During the reading process of the object data, according to the keywords contained in the object data, the object data that meets the query conditions is placed in the memory, and the keyBit of the object data is generated according to the query keyword. The keyBit is a binary code representing the keywords contained in the object data; then serialize the HilbertCurve object, calculate its Hilbert value according to the longitude and latitude coordinates of the HilbertCurve object, build a Hilbert R-tree, and pass the Hilbert R-tree to all computing nodes; Step 3: The diagonal pruning method for leaf nodes; Step 3-1: Generate a leaf node, and judge the query keywords contained in this leaf node. If this leaf node contains all query keywords, then set the diagonal distance of this leaf node as the threshold; Step 3-2: Repeat Step 3-1. If the newly generated leaf node contains all the query keywords and the diagonal distance is less than the threshold, update the threshold with the diagonal distance of the newly generated leaf node until all leaf nodes are generated; the finally obtained threshold is denoted as S ; Step 4: The circular scanning pruning method; Traverse all object data. With the data object as the center, * S construct a circle with the diameter. Determine whether all query keywords are included in the area of this circle. If not all query keywords are included, delete the object data from the object set; Step 5: Traverse the remaining object data, and use the two-point pairing method to pair the object data pairwise to form object pairs; the specific two-point pairing method is as follows: The goal of the set space keyword query is to find a set of several objects with the closest positions. Among them, the measurement method of proximity is described by the "diameter" of the object set, that is, the maximum value of the pairwise distances in the object set is used as the diameter; the set keyword query is to find an object set that can contain all keywords and has the smallest diameter; First, combine the objects pairwise to form object pairs, then sort the object pairs in ascending order according to the distance, and finally expand each object pair into an object set one by one. If an object pair is expanded into an object set, then this object set is the required result; Step 6: Use the Hilbert R-tree, the object pairing method; Step 6-1: Pair the root node of the Hilbert R-tree with itself; Step 6-2: For non-leaf node pairs, list all child node pairs. If the minimum distance between the rectangles corresponding to the two nodes is greater than S , then ignore this node pair; otherwise, continue to pair the child nodes. For leaf node pairs, list all object pairs. Step 6-3: Repeat 6-2 until all object pairs are paired; Step 7: If the distance of the object pair is less than the threshold S , then set the distance of the object pair as the key, and output the two object data as the value to the reduce; Step 8: In the reduce stage, scan within the area where each object pair is located. If an object set with the key as the diameter can be found, it is used as a candidate output result; serialize the top ten candidate objects with the smallest distance and output them to the corresponding file.

2. The method for querying the nearest neighbor spatial set keyword based on MapReduce according to claim 1, wherein The specific diagonal pruning method for leaf nodes is as follows: The Hilbert R-tree is constructed in a bottom-up manner. First, leaf nodes are generated. The leaf nodes use the minimum bounding rectangle as their spatial attribute; if a leaf node contains all query keywords, then an object set that meets all query keywords can be found in this leaf node, and its diameter is less than the diagonal distance of this leaf node. Therefore, the diagonal distance of this leaf node is the upper limit of the result of the set space keyword query; When generating the Hilbert R-tree, determine the generated leaf nodes. If all query keywords are included, the diagonal distance of the leaf node is the upper limit of the result. Then, as leaf nodes are generated, update this upper limit value. When all leaf nodes are generated, the diagonal distance of the smallest leaf node that meets this condition is the minimum upper limit; use this minimum upper limit as the threshold.

Citation Information

Patent Citations

  • Road network space keyword search method

    CN104376112A

  • Reverse rearest neighbor query processing method for road network geographical social keywords

    CN107145526A