Index tree construction method, data query method and electronic device
By building an index tree, using spatial and layer structure indexing methods, combining the unstructured layer structure of the small world network and the Delaunay Triangle Network, the unstructured data query problem with spatial scope constraints is solved, and efficient approximate nearest neighbor query is achieved, which is suitable for various terminal devices and servers.
Patent Information
- Application Number
- CN202310085421.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-11
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2043-01-11
AI Technical Summary
The prior art lacks a fast and efficient solution to deal with approximate nearest neighbor queries of unstructured data with spatial scope constraints. Especially in hybrid query analysis, it is difficult to efficiently find query targets that conform to spatial scope and unstructured feature constraints.
Build an index tree, index the root node and child nodes through spatial indexing, and index leaf nodes using layer structure indexing. Combined with the hierarchy, the union layer structure of the small world network and the Delaunay Triangle Network can be used to achieve efficient query of unstructured data.
It realizes efficient approximate nearest neighbor query for unstructured data under spatial scope constraints, improves query speed and accuracy, and is suitable for data queries of various terminal devices and servers.
Smart Images

Figure CN116186184B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of cloud computing technology, and in particular to an index tree construction method, a data query method, and an electronic device. Background Art
[0002] With the advent of the big data era and the widespread use of sensors such as positioning technology, cameras, and audio equipment, mobile device applications can collect massive amounts of geospatial and unstructured information (e.g., images, audio, and historical reviews) about target objects. Location-based electronic maps store not only the location information of target objects or points of interest (POIs), but also unstructured information such as images and reviews. This poses challenges for hybrid query analysis of spatial and unstructured data. For example, find the top k people with the highest similarity to a given photo within 3 kilometers of a central square. This query has the spatial range constraint of "within 3 kilometers of the central square" and the unstructured feature constraint of "similar to the given photo." This query can be summarized as a k-approximate nearest neighbor query on unstructured data with a spatial range constraint.
[0003] Currently, there is a lack of a fast and efficient query solution for approximate nearest neighbor queries on unstructured data with spatial range constraints. Summary of the Invention
[0004] Embodiments of the present application provide an index tree construction method, a data query method, and an electronic device to implement approximate nearest neighbor query of unstructured data with spatial range constraints.
[0005] In a first aspect, an embodiment of the present application provides a method for constructing an index tree, comprising:
[0006] An object data set is obtained, wherein the objects in the object data set include unstructured data and location information corresponding to the unstructured data; an index tree is constructed based on the storage locations of the objects in a data table; the root node and child nodes of the index tree respectively index the multiple storage locations using a spatial index method; and the leaf nodes of the index tree index the multiple storage locations using a layered structure index method.
[0007] In a second aspect, an embodiment of the present application provides a data query method, including:
[0008] receiving a data query request, wherein the data query request includes unstructured data and a spatial constraint range;
[0009] Based on a pre-built index tree, the query targets that meet the spatial constraint range and match the unstructured data are queried; the root node and child nodes of the index tree index multiple storage locations through spatial indexing respectively; the leaf nodes of the index tree index multiple storage locations through hierarchical indexing.
[0010] In a third aspect, an embodiment of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory, wherein the processor implements any of the above-described methods when executing the computer program.
[0011] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements any of the methods described above.
[0012] Compared with the prior art, this application has the following advantages:
[0013] The present application provides an index tree construction method, a data query method, and an electronic device. First, an object data set is obtained, wherein the objects in the object data set include unstructured data and location information corresponding to the unstructured data; then, an index tree is constructed based on the storage location of the objects in the data table; the root node and child nodes of the index tree respectively index multiple storage locations using a spatial index method; and the leaf nodes of the index tree index multiple storage locations using a layered index method. Upon receiving a data query request, the index tree is used to perform a data query. The spatial index method of the index tree can be used to make the query target meet the spatial constraint range. The layered index method can be used to query the query target that matches the unstructured data carried in the data query request, thereby realizing an approximate nearest neighbor query of unstructured data with spatial range constraints.
[0014] The above description is only an overview of the technical solution of this application. In order to more clearly understand the technical means of this application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of this application more obvious and easy to understand, the specific implementation methods of this application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the multiple drawings represent the same or similar components or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings only depict some embodiments according to the present application and should not be regarded as limiting the scope of the present application.
[0016] Figure 1 A schematic diagram of an application scenario of the technical solution of this application;
[0017] Figure 2This is a flowchart of an index tree construction method according to an embodiment of the present application;
[0018] Figure 3 is a schematic diagram of an RH index tree according to an embodiment of the present application;
[0019] Figure 4 is a schematic diagram of an RHDN index tree according to an embodiment of the present application;
[0020] Figure 5 Schematic diagram of a HDN structure of a union layer according to an embodiment of the present application;
[0021] Figure 6 This is a flow chart of a data query method according to an embodiment of the present application;
[0022] Figure 7 is a schematic diagram showing how query time varies with the size of a spatial query area in two different data sets according to an embodiment of the present application;
[0023] Figure 8 is a schematic diagram of a measurement of query time and accuracy in two different data sets according to an embodiment of the present application;
[0024] Figure 9 This is a structural block diagram of an index tree construction device according to an embodiment of the present application;
[0025] Figure 10 is a structural block diagram of a data query device according to an embodiment of the present application; and
[0026] Figure 11 A block diagram of an electronic device used to implement an embodiment of the present application. DETAILED DESCRIPTION
[0027] Hereinafter, only certain exemplary embodiments are briefly described. As will be appreciated by those skilled in the art, the described embodiments may be modified in various ways without departing from the spirit or scope of the present application. Therefore, the drawings and description are to be regarded as illustrative in nature and not restrictive.
[0028] To facilitate understanding of the technical solutions of the embodiments of the present application, the following describes the related technologies of the embodiments of the present application. The following related technologies can be combined with the technical solutions of the embodiments of the present application as optional solutions, and all of them fall within the scope of protection of the embodiments of the present application.
[0029] Figure 1 This is a schematic diagram of an application scenario of the technical solution of this application. Figure 1 As shown, the server can receive a query request sent by a terminal device, and the terminal device may include: a mobile terminal, augmented reality (AR) glasses, and a fixed terminal, etc.
[0030] For example, a user wants to find a nearby Japanese restaurant. Using a mobile device like a cell phone, they can send a query request to the server. The query includes a photo of a Japanese restaurant, and the query range is set to be within one kilometer of the user. After receiving the query request, the server uses a pre-built index tree to search the database for the location information of multiple Japanese restaurants within one kilometer of the user that are similar to the photo. The server then returns the query results to the user, allowing them to select and visit the restaurant they most desire.
[0031] For another example, this query can be applied to augmented reality applications. When a user puts on AR glasses, the AR glasses send a data query request to the server. The query request includes images of the user's surrounding objects, obtained through the AR glasses. The default spatial constraint range is pre-configured to be 1.5 kilometers. The server searches the database for objects within 1.5 kilometers that are similar to the images of the surrounding objects and returns the information to the AR glasses. The AR glasses then display information about objects similar to the surrounding objects seen by the user. Unstructured information such as images can be converted into high-dimensional vectors for querying.
[0032] In addition, the technical solution of the present application may also not send a query request to the server, but build an index tree through the terminal device, and after receiving the user's query instruction, directly query the database of the terminal device and display the query results to the user.
[0033] The present application embodiment provides a method for constructing an index tree. The method in this embodiment can be applied to a computing device, which may include: a server, a mobile terminal, a fixed terminal, etc. Figure 2 FIG. 1 is a flowchart of a method for constructing an index tree according to an embodiment of the present application, including:
[0034] Step S201: Acquire an object data set, where the objects in the object data set include unstructured data and location information corresponding to the unstructured data.
[0035] Unstructured data can include at least one of images, videos, audio, and historical reviews. In practical applications, unstructured data can be converted into high-dimensional vectors for computation. The location information corresponding to the unstructured data can include a geographic location associated with the unstructured data. For example, if the unstructured data is a photo of a store, the corresponding location information would be the store's geographic location.
[0036] Step S202: construct an index tree according to the storage location of the object in the data table; the root node and child nodes of the index tree index multiple storage locations using spatial indexing; and the leaf nodes of the index tree index multiple storage locations using hierarchical indexing.
[0037] By storing each object in the object dataset into a database table, we can obtain the storage location of each object in the table. Using the object storage location, we can build an index tree. The index tree consists of a root node, child nodes, and leaf nodes, and each node indexes the storage location of multiple objects.
[0038] The root node and child nodes of the index tree index multiple storage locations respectively through spatial indexing. The spatial indexing method can include a tree structure for spatial range query, such as an R (rectangle) tree, etc. Each node in the index tree corresponds to a minimum bounding rectangle (MBR).
[0039] The leaf nodes of the index tree index multiple storage locations using a hierarchical indexing approach. The hierarchical indexing approach can include multiple layers, each of which indexes multiple storage locations using a graph-based indexing approach. Graph-based indexing can be used to calculate similarity between high-dimensional vectors.
[0040] The index tree construction method provided by the embodiment of the present application first obtains an object data set, wherein the objects in the object data set include unstructured data and location information corresponding to the unstructured data; then, an index tree is constructed based on the storage location of the objects in the data table; the root node and child nodes of the index tree respectively index multiple storage locations through a spatial index method; and the leaf nodes of the index tree index multiple storage locations through a layered index method. Upon receiving a data query request, the index tree is used to perform a data query. The spatial index method of the index tree can be used to make the query target meet the spatial constraint range, and the layered index method can be used to query the query target that matches the unstructured data carried in the data query request, thereby realizing an approximate nearest neighbor query of unstructured data with spatial range constraints.
[0041] In one implementation, leaf nodes of the index tree index multiple storage locations at least through a layer structure obtained based on a hierarchical navigable small world (HNSW) network.
[0042] In one example, an RH index tree is obtained based on the R-tree and HNSW, such as Figure 3As shown in Figure 1, the RH index tree includes a root node mbr1, child nodes mbr2, mbr3, and mbr4, and leaf nodes mbr5, mbr6, and mbr7. Leaf nodes mbr5, mbr6, and mbr7 index multiple storage locations using the graph indexing methods HNSW1, HNSW2, and HNSW3, respectively. Each layer of the layer structure obtained based on HNSW includes multiple object nodes, each of which is a storage location for an object. The edges between object nodes are paths between object nodes, which can be set according to specific needs. Figure 3 As shown, the bottom layer of HNSW1 includes 7 object nodes, the middle layer includes 4 object nodes, and the top layer includes 2 object nodes. The object nodes in the middle layer are a subset of the object nodes in the bottom layer, and the object nodes in the top layer are a subset of the object nodes in the middle layer. The edges between the object nodes in the middle and top layers are determined by the object nodes in the bottom layer. That is, if there is an edge between two object nodes in the bottom layer, there is also an edge between the object nodes in the middle and top layers. If there is no edge between two object nodes in the bottom layer, there is no edge between the object nodes in the middle and top layers.
[0043] In addition, other forms of index trees can also be obtained based on HNSW, as shown in the following embodiments:
[0044] In one implementation, leaf nodes of the index tree index multiple storage locations through a union layer structure of a layer structure obtained based on a hierarchical navigable small-world network and a layer structure obtained based on a Delaunay triangulation network.
[0045] In one example, the RHDN index tree is obtained based on the R-tree, HNSW and Delaunay triangulation network, such as Figure 4 As shown, the RHDN index tree includes a root node mbr1, child nodes: mbr2, mbr3 and mbr4, and leaf nodes: mbr5, mbr6 and mbr7. Leaf nodes mbr5, mbr6 and mbr7 index multiple storage locations through layer structures HDN1, HDN2 and HDN3 respectively. The layer structure HDN is a union layer structure of the layer structure obtained based on HNSW and the layer structure HD obtained based on Delaunay triangulation. Each layer of HDN includes multiple object nodes, each object node is the storage location of an object, and the edges between object nodes are paths between object nodes, which can be set according to specific needs. Figure 4As shown, the bottom layer (Layer = 0) of HDN1 includes 7 object nodes, the middle layer (Layer = 1) includes 4 object nodes, and the top layer (Layer = 2) includes 2 object nodes. The object nodes in the middle layer are a subset of the object nodes in the bottom layer, and the object nodes in the top layer are a subset of the object nodes in the middle layer. The edges between the object nodes in the middle and top layers are determined by the object nodes in the bottom layer. That is, if there is an edge between two object nodes in the bottom layer, there is also an edge between the object nodes in the middle and top layers. If there is no edge between two object nodes in the bottom layer, there is no edge between the object nodes in the middle and top layers.
[0046] In one implementation, the object nodes in the layers of the union layer structure are the same as the object nodes of the layer structure obtained based on the hierarchical navigable small-world network or the layer structure obtained based on the Delaunay triangulation network; the paths between the object nodes of the union layer structure are obtained by the union of the paths between the object nodes of the layer structure obtained based on the hierarchical navigable small-world network and the layer structure obtained based on the Delaunay triangulation network.
[0047] In one example, the layer structure obtained based on HNSW and the layer structure HD obtained based on Delaunay triangulation determine the union layer structure HDN, such as Figure 5 As shown, the object nodes in the layers of the HDN are the same as the object nodes in the layer structure obtained based on HNSW or the layer structure HD obtained based on the Delaunay triangulation network. That is, Layer = 0 of HNSW contains object nodes 1, 2, 3, 4, 5, 6, and 7, and Layer = 0 of HD contains object nodes 1, 2, 3, 4, 5, 6, and 7. Therefore, Layer = 0 of the HDN contains object nodes 1, 2, 3, 4, 5, 6, and 7. The highest layer that the object nodes in the layers of the HD can be placed in is set according to the method of object nodes in the layers of the layer structure obtained based on HNSW. The paths between the object nodes in the layers of the HDN are represented by edges. The edges between the object nodes of the HDN are the union of the edges between the object nodes of the layer structure obtained based on HNSW and the layer structure HD obtained based on the Delaunay triangulation network. That is, if there is an edge between two object nodes in the layer structure obtained by HNSW, but no edge between the object nodes of the HD, the union of the two will result in an edge between the object nodes of the HDN. There is an edge between the object nodes 1 and 6 of Layer=0 of HNSW, but not between the object nodes 1 and 6 of Layer=0 of HD. However, there is an edge between the object nodes 1 and 6 of Layer=0 of HDN.
[0048] After the index tree is built, you can use the index tree to query data, as shown in the following example:
[0049] The present application also provides a data query method. The method in this embodiment can be applied to a computing device, which may include a server, a mobile terminal, a fixed terminal, etc. Figure 6 The flowchart of the data query method according to one embodiment of the present application is shown, including:
[0050] Step S601: receiving a data query request, where the data query request includes unstructured data and a spatial constraint range.
[0051] Step S602: Based on the pre-built index tree, query targets that meet the spatial constraint range and match the unstructured data are searched; the root node and child nodes of the index tree index multiple storage locations through spatial indexing respectively; and the leaf nodes of the index tree index multiple storage locations through hierarchical indexing.
[0052] The data query request includes unstructured data and a spatial constraint range. The purpose is to search for query targets that match the unstructured data within the spatial constraint range. Unstructured data can include at least one of images, videos, audio, and historical reviews. In practical applications, unstructured data can be converted into high-dimensional vectors for calculation.
[0053] Optionally, the data query request may further include the number of query targets. For example, the k stores that are most similar to the store photos included in the request and are within one kilometer of the user's current location are queried as matching query targets.
[0054] It is understandable that, when the query request does not include the number of query targets, matching query targets may be returned according to a pre-configured number.
[0055] Based on the pre-built index tree, after searching for query targets that meet the spatial constraints and match the unstructured data, information related to the query target is returned to the requester. This information includes the geographic location of similar stores and historical evaluation information.
[0056] The data query method provided by the embodiment of the present application first obtains an object data set, wherein the objects in the object data set include unstructured data and location information corresponding to the unstructured data; then, an index tree is constructed based on the storage location of the objects in the data table; the root node and child nodes of the index tree respectively index multiple storage locations using a spatial index method; and the leaf nodes of the index tree index multiple storage locations using a layered index method. Upon receiving a data query request, the index tree is used to perform a data query. The spatial index method of the index tree can be used to make the query target meet the spatial constraint range, and the layered index method can be used to query the query target that matches the unstructured data carried in the data query request, thereby realizing an approximate nearest neighbor query of unstructured data with spatial range constraints.
[0057] In one implementation, the leaf nodes of the index tree index multiple storage locations using at least a layer structure derived from a hierarchical navigable small-world network. Based on the pre-built index tree, query targets that meet spatial constraints and match unstructured data include:
[0058] A query is performed from the root node to the leaf node in the index tree to determine the target leaf node that intersects with the spatial constraint range. If the spatial constraint range completely covers the minimum bounding rectangle corresponding to the target leaf node, the layer structure obtained based on the hierarchical navigable small-world network is used to query the query target that meets the spatial constraint range and matches the unstructured data.
[0059] The target leaf node can be one or more, depending on the actual situation of the specific query. Each layer of the target leaf node includes multiple object nodes, each object node corresponds to a storage location, and there are paths between the object nodes.
[0060] The query targets that match the unstructured data can be multiple data items that are similar to the unstructured data. The multiple data items that are similar to the unstructured data are sorted from highest to lowest in terms of similarity, and the top-ranked data items are used as the query targets that match the unstructured data. The number of query targets can be determined based on the number of query targets included in the data query request. If the data query request does not include the number of query targets, a pre-configured number of query targets can be returned.
[0061] In one implementation, leaf nodes of the index tree index multiple storage locations using at least a layer structure derived from a hierarchical navigable small-world network. Based on the pre-built index tree, searching for a query target that satisfies a spatial constraint range and matches the unstructured data includes: querying the index tree from the root node to the leaf nodes, determining a target leaf node that intersects the spatial constraint range; and if the spatial constraint range does not completely cover the minimum bounding rectangle corresponding to the target leaf node, performing a brute force search within the target leaf node to obtain a query target that satisfies the spatial constraint range and matches the unstructured data.
[0062] Among them, performing a brute force search in the target leaf node means obtaining the unstructured data corresponding to each storage location according to the storage location in each layer of the target leaf node, and calculating the similarity with the unstructured data in the query request in turn to obtain the query target.
[0063] In one example, given a dataset D and a query q, each object in D contains an unstructured data (such as an image, audio, and historical record, etc.) and a geographic coordinate information, and the query q contains a spatial constraint range R, an unstructured data v (a high-dimensional vector), and an integer k. The goal of the query is to return the k objects in the dataset D that are within the spatial constraint range R and are closest to v.
[0064] The process of performing k-approximate nearest neighbor query q = (k, R, v) on unstructured data with spatial range constraints on the RH tree using the pre-built RH index tree for data query consists of two stages:
[0065] Phase 1: Find the set S of leaf nodes that intersect with R. The leaf nodes in set S are the target leaf nodes.
[0066] Phase 2: For each target leaf node s∈S, if the MBR of s is completely covered by R, HNSW is used to search for the unstructured data of the k nearest neighbors of vector v and return it to the requester as the query target. Otherwise, a brute force search is performed by iterating over each object node in the target leaf node s, or HNSW is used to calculate the query results for the object nodes within R in the intermediate result list during the search process. The unstructured data of the k nearest neighbors is obtained and returned to the requester as the query target.
[0067] In one implementation, leaf nodes of an index tree index multiple storage locations using a union hierarchical structure derived from a hierarchically navigable small-world network and a hierarchical structure derived from a Delaunay triangulation network. Based on the pre-built index tree, searching for a query target that satisfies a spatial constraint range and matches unstructured data involves performing a query from the root node to the leaf nodes in the index tree, determining a target leaf node that intersects the spatial constraint range, and if the spatial constraint range completely covers the minimum bounding rectangle corresponding to the target leaf node, then using the hierarchical structure derived from the hierarchical navigable small-world network in the union hierarchical structure to search for a query target that satisfies the spatial constraint range and matches the unstructured data.
[0068] The target leaf node can be one or more, depending on the actual situation of the specific query. Each layer of the target leaf node includes multiple object nodes, each object node corresponds to a storage location, and there are paths between the object nodes.
[0069] The query targets that match the unstructured data can be multiple data items that are similar to the unstructured data. The multiple data items that are similar to the unstructured data are sorted from highest to lowest in terms of similarity, and the top-ranked data items are used as the query targets that match the unstructured data. The number of query targets can be determined based on the number of query targets included in the data query request. If the data query request does not include the number of query targets, a pre-configured number of query targets can be returned.
[0070] In one implementation, the leaf nodes of the index tree index multiple storage locations using a layer structure derived from a hierarchical navigable small-world network and a layer structure derived from a Delaunay triangulation network. The leaf nodes of the index tree index multiple storage locations using a layer structure derived from a hierarchical navigable small-world network and a layer structure derived from a Delaunay triangulation network. Based on the pre-built index tree, queries that meet spatial constraints and match unstructured data include:
[0071] A query is performed from the root node to the leaf node in the index tree to determine the target leaf node that intersects with the spatial constraint range. If the spatial constraint range does not completely cover the minimum bounding rectangle corresponding to the target leaf node, a sample set of object nodes is obtained from the target leaf node to determine the selectivity of the object node sample set in the spatial constraint range. Based on the comparison result of the selectivity and the selectivity threshold, the layer structure obtained by using the hierarchical navigable small-world network in the union layer structure or the layer structure obtained by using the Delaunay triangulation network is determined to query the query target that meets the spatial constraint range and matches the unstructured data.
[0072] The calculation of the selection rate specifically includes: selecting multiple object nodes in the target leaf node according to a preset selection method to obtain an object node sample set, finding out how many object nodes in the object node sample set are within the spatial constraint range, and the number of object nodes within the spatial constraint range as a percentage of the total number of object nodes in the object node sample set, which is the selection rate of the object node sample set within the spatial constraint range.
[0073] Optionally, if the selectivity is greater than the selectivity threshold, the layer structure obtained based on the hierarchical navigable small-world network in the union layer structure is used to query a query target that meets the spatial constraint range and matches the unstructured data.
[0074] If the selection rate is less than the selection rate threshold, the layer structure obtained based on the Delaunay triangulation in the union layer structure is used to query the query target that meets the spatial constraint range and matches the unstructured data.
[0075] If the selection rate is equal to the selection rate threshold, the layer structure obtained based on the hierarchical navigable small-world network or the layer structure obtained based on the Delaunay triangulation network is used to query the query target that meets the spatial constraint range and matches the unstructured data.
[0076] When constructing the index tree, the query rate threshold is calculated and stored in the leaf node of the index tree. When performing data query, the selectivity threshold is compared with the selectivity.
[0077] In one example, given a dataset D and a query q = (k, R, v), the query aims to return the k objects in dataset D that are within the spatial constraint range R and closest to v. Using the RHDN index tree for data query, the selectivity threshold se* is calculated as follows:
[0078] The query efficiency of HD and HNSW is expressed by the number of visited object nodes. The query efficiency will decrease as the number of visited object nodes increases. For HD and HNSW, the first step of the query is to traverse the object nodes from the top graph to the bottom graph. At this stage, the number of visited object nodes in each layer is constrained by a constant. Among them, G is a geometric series that can be obtained based on a logarithmic function and is used to calculate the probability of a node object being placed in each layer. The higher the number of layers, the smaller the value of G. Assuming that the selectivity of query q is se, there are N leaf nodes i. i objects, the average degree of HD (the average number of edges connected to each object node) is m d , the maximum number of layers is maxL (starting from Layer=0). For HD, after the first step of the query, all object nodes in R are visited, and the number of object nodes visited by HD is N d Calculate according to the following formula (1):
[0079] N d =(maxL+1)*m d *G+se*N i (1)
[0080] For HNSW, we start from an object node close to v and iterate on the neighboring nodes of the current nearest object node until we find the e object nodes closest to v in R. Assuming that the spatial data and high-dimensional vectors are independent of each other, in order to obtain the e object nodes in R, we have an average of object nodes and their neighbor nodes need to be visited. However, some neighbor nodes may be visited repeatedly. Assume that the repetition rate is θ h , represents the ratio of neighbor nodes that have been visited. In general, Among them, m h represents the average degree of the layers except the bottom layer in HNSW, m h0Indicates the average degree of the bottom layer. Take θ h Therefore, the number of points that HNSW needs to visit is N h According to the following formula (2), it is calculated by the total amount of data N in leaf node i. i Restricted.
[0081]
[0082] As the selectivity increases, the query efficiency of HD gradually decreases, while the opposite is true for HNSW. Therefore, the two formulas are equivalent to obtain the selectivity se when the query efficiency of HD and HNSW is the same * :
[0083]
[0084] Where P = maxL·G·(m h -m d )+G·(m h0 -m d ). When the selection rate of q for leaf node i is less than se * When , HD is used to execute the query; otherwise, HNSW is used.
[0085] In one example, a data query using a pre-built RHDN index tree is performed. The process of performing a k-approximate nearest neighbor query q = (k, R, v) on unstructured data with spatial range constraints on RHDN includes two stages:
[0086] Phase 1: Find the set S of leaf nodes that intersect with R. The leaf nodes in set S are the target leaf nodes.
[0087] Phase 2: For each target leaf node s∈S, if s's MBR is fully covered by R, the HNSW portion of the HDN (union-level structure) is used to search for the k vectors closest to v. Otherwise, the selectivity of R for the target leaf node is calculated. Based on this selectivity, the HDN or HNSW portion is used to determine the query result. If the selectivity is greater than the selectivity threshold, the HNSW portion of the HDN is used to calculate the query result. Otherwise, the HDN portion is used to find the target node in R. Among these target nodes, the k vectors closest to v are calculated to obtain the query target.
[0088] Among them, the method of using HD to query object nodes within R includes: starting from an arbitrary object node n in the top-level Delaunay graph of HD, moving toward R by continuously visiting the path of the n neighbor nodes closest to R. If all neighbor nodes are farther away from R than the current object node n, then the neighbor nodes of the object node n in the lower Delaunay graph will continue to be visited. In this process, if a node in R is encountered, all object nodes in R in the lower Delaunay graph are directly checked. Compared with the single-layer Delaunay graph method, the multi-layer HD is more efficient in range queries. This is because the hierarchical structure allows HD to jump out of the local optimal solution and move towards the result faster than the single-layer Delaunay graph.
[0089] The technical solution of this application fills the gap in the query of unstructured data containing geographic coordinates. It is the first to design an index and query method for k-nearest neighbor query of unstructured data with spatial range constraints, which has a good trade-off between query speed and accuracy. Figure 7 and Figure 8 shown.
[0090] Comparison method: Compare the data query method based on the RHDN index tree with the following three methods: the first is a data query method based entirely on the spatial index R-tree; the second is a data query method based entirely on HNSW; and the third is a data query method based on the RH index tree.
[0091] Experimental configuration:
[0092] Server configuration: 2.7GHz, 64-core CPU, 256G memory.
[0093] Programming language: Java.
[0094] Among them, Figure 7 The diagram below shows how query time changes with the size of the spatial query region in two different datasets (dataset 1 on the left and dataset 2 on the right). The horizontal axis represents the size of the query region, and the vertical axis represents the query time. Figure 8 The figure below shows the relationship between query time and accuracy for two different datasets (Dataset 1 on the left and Dataset 2 on the right). The horizontal axis represents query time, and the vertical axis represents accuracy. Accuracy Recall@k is calculated based on the result O of an exact query of k approximate nearest neighbors on unstructured data with spatial range constraints and the result O' of an approximate query: A larger Recall@k indicates higher accuracy in the query results.
[0095] Corresponding to the application scenario and index tree construction method provided in the embodiment of the present application, the embodiment of the present application also provides an index tree construction device. Figure 9 FIG2 is a block diagram of an index tree construction device according to an embodiment of the present application, wherein the device includes:
[0096] An acquisition module 901 is configured to acquire an object data set, where objects in the object data set include unstructured data and location information corresponding to the unstructured data;
[0097] Construction module 902 is used to construct an index tree based on the storage location of the object in the data table; the root node and child nodes of the index tree index multiple storage locations through spatial indexing respectively; the leaf nodes of the index tree index multiple storage locations through layer structure indexing.
[0098] The index tree construction device provided by the embodiment of the present application first obtains an object data set, wherein the objects in the object data set include unstructured data and location information corresponding to the unstructured data; then, an index tree is constructed based on the storage location of the objects in the data table; the root node and child nodes of the index tree respectively index multiple storage locations through a spatial index method; and the leaf nodes of the index tree index multiple storage locations through a hierarchical index method. Upon receiving a data query request, the index tree is used to perform a data query. The spatial index method of the index tree can be used to make the query target meet the spatial constraint range, and the hierarchical index method can be used to query the query target that matches the unstructured data carried in the data query request, thereby realizing an approximate nearest neighbor query of unstructured data with spatial range constraints.
[0099] In one implementation, leaf nodes of the index tree index a plurality of storage locations at least through a layer structure obtained based on a hierarchical navigable small-world network.
[0100] In one implementation, leaf nodes of the index tree index multiple storage locations through a union layer structure of a layer structure obtained based on a hierarchical navigable small-world network and a layer structure obtained based on a Delaunay triangulation network.
[0101] In one implementation, the object nodes in the layers of the union layer structure are the same as the object nodes of the layer structure obtained based on the hierarchical navigable small-world network or the layer structure obtained based on the Delaunay triangulation network; the paths between the object nodes of the union layer structure are obtained by the union of the paths between the object nodes of the layer structure obtained based on the hierarchical navigable small-world network and the layer structure obtained based on the Delaunay triangulation network.
[0102] The functions of each module in each device in the embodiments of the present application can be found in the corresponding description in the above method, and have corresponding beneficial effects, which will not be repeated here.
[0103] Corresponding to the application scenario and data query method provided by the embodiment of the present application, the embodiment of the present application also provides a data query device. Figure 10 FIG2 is a block diagram of a data query device according to an embodiment of the present application, which includes:
[0104] The receiving module 1001 is configured to receive a data query request, wherein the data query request includes unstructured data and a spatial constraint range;
[0105] The query module 1002 is used to query the query target that meets the spatial constraint range and matches the unstructured data based on a pre-built index tree; the root node and child nodes of the index tree respectively index multiple storage locations through spatial indexing; the leaf nodes of the index tree index multiple storage locations through hierarchical indexing.
[0106] The data query device provided by the embodiment of the present application first obtains an object data set, wherein the objects in the object data set include unstructured data and location information corresponding to the unstructured data; then, an index tree is constructed based on the storage location of the objects in the data table; the root node and child nodes of the index tree respectively index multiple storage locations using a spatial index method; and the leaf nodes of the index tree index multiple storage locations using a layered index method. Upon receiving a data query request, the index tree is used to perform a data query. The spatial index method of the index tree can be used to make the query target meet the spatial constraint range, and the layered index method can be used to query the query target that matches the unstructured data carried in the data query request, thereby realizing an approximate nearest neighbor query of unstructured data with spatial range constraints.
[0107] In one implementation, leaf nodes of the index tree index a plurality of storage locations at least through a layer structure obtained based on a hierarchical navigable small-world network.
[0108] In one implementation, the query module 1002 is used to: perform a query from the root node to the leaf node in the index tree to determine the target leaf node that intersects with the spatial constraint range; if the spatial constraint range completely covers the minimum bounding rectangle corresponding to the target leaf node, then use the layer structure obtained based on the hierarchical navigable small-world network to query the query target that meets the spatial constraint range and matches the unstructured data.
[0109] In one implementation, the query module 1002 is used to: query from the root node to the leaf node in the index tree to determine the target leaf node that intersects with the spatial constraint range; if the spatial constraint range does not completely cover the minimum circumscribed rectangle corresponding to the target leaf node, perform a brute force search in the target leaf node to obtain a query target that meets the spatial constraint range and matches the unstructured data.
[0110] In one implementation, leaf nodes of the index tree index multiple storage locations through a union layer structure of a layer structure obtained based on a hierarchical navigable small-world network and a layer structure obtained based on a Delaunay triangulation network.
[0111] In one implementation, the query module 1002 is used to: perform a query from the root node to the leaf node in the index tree to determine the target leaf node that intersects with the spatial constraint range; if the spatial constraint range completely covers the minimum bounding rectangle corresponding to the target leaf node, then use the layer structure obtained based on the hierarchical navigable small-world network in the union layer structure to query the query target that meets the spatial constraint range and matches the unstructured data.
[0112] In one implementation, the query module 1002 is configured to: perform a query from the root node to the leaf nodes in the index tree to determine a target leaf node that intersects with the spatial constraint range; if the spatial constraint range does not completely cover the minimum bounding rectangle corresponding to the target leaf node, obtain a sample set of object nodes from the target leaf node and determine a selectivity of the sample set of object nodes in the spatial constraint range;
[0113] According to the comparison result of the selectivity and the selectivity threshold, it is determined to use the layer structure obtained by the hierarchical navigable small-world network in the union layer structure or the layer structure obtained by the Delaunay triangulation network to query the query target that meets the spatial constraint range and matches the unstructured data.
[0114] The functions of each module in each device in the embodiments of the present application can be found in the corresponding description in the above method, and have corresponding beneficial effects, which will not be repeated here.
[0115] Figure 11 FIG. 1 is a block diagram of an electronic device for implementing an embodiment of the present application. Figure 11 As shown, the electronic device includes: a memory 1110 and a processor 1120. The memory 1110 stores a computer program that can be run on the processor 1120. When the processor 1120 executes the computer program, the method in the above embodiment is implemented. The number of the memory 1110 and the processor 1120 can be one or more.
[0116] The electronic device also includes:
[0117] The communication interface 1130 is used to communicate with external devices and perform data exchange transmission.
[0118] If the memory 1110, the processor 1120, and the communication interface 1130 are implemented independently, the memory 1110, the processor 1120, and the communication interface 1130 can be connected to each other via a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 11 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0119] Optionally, in a specific implementation, if the memory 1110, the processor 1120 and the communication interface 1130 are integrated on a chip, the memory 1110, the processor 1120 and the communication interface 1130 can communicate with each other through an internal interface.
[0120] An embodiment of the present application provides a computer-readable storage medium storing a computer program, which implements the method provided in the embodiment of the present application when the program is executed by a processor.
[0121] An embodiment of the present application also provides a chip, which includes a processor for calling and executing instructions stored in the memory from the memory, so that a communication device equipped with the chip executes the method provided in the embodiment of the present application.
[0122] An embodiment of the present application also provides a chip, including: an input interface, an output interface, a processor and a memory. The input interface, the output interface, the processor and the memory are connected through an internal connection path. The processor is used to execute the code in the memory. When the code is executed, the processor is used to execute the method provided in the embodiment of the application.
[0123] It should be understood that the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. It is worth noting that the processor may be a processor that supports the Advanced RISC Machines (ARM) architecture.
[0124] Furthermore, optionally, the above-mentioned memory may include a read-only memory and a random access memory. The memory may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may include a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may include a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available. For example, static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM) and direct memory bus random access memory (DR RAM).
[0125] In the above embodiments, all or part of the embodiments may be implemented using software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments may be implemented in the form of a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another.
[0126] In the description of this specification, the reference terms "one embodiment," "some embodiments," "example," "specific example," or "some examples" mean that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials, or characteristics described may be combined in any appropriate manner in any one or more embodiments or examples. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification, as well as features of different embodiments or examples, unless they are mutually inconsistent.
[0127] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. Throughout the description of this application, "plurality" means two or more, unless otherwise specifically defined.
[0128] Any process or method described in the flowchart or otherwise described herein can be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a specific logical function or process. The scope of the preferred embodiments of the present application includes other implementations in which the functions may be performed in a different order than shown or discussed, including performing the functions substantially simultaneously or in reverse order depending on the functions involved.
[0129] The logic and / or steps described in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, apparatus or device (such as a computer-based system, a system including a processor or other system that can fetch instructions from an instruction execution system, apparatus or device and execute instructions), or used in combination with such instruction execution systems, apparatuses or devices.
[0130] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. All or part of the steps of the above embodiment method can be completed by instructing the relevant hardware through a program, which can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0131] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing module, or each unit may exist physically separately, or two or more units may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or in the form of software functional modules. If the aforementioned integrated modules are implemented in the form of software functional modules and sold or used as independent products, they may also be stored in a computer-readable storage medium. The storage medium may be a read-only memory, a magnetic disk, or an optical disk, etc.
[0132] The above is merely an exemplary embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various modifications or substitutions within the technical scope described in this application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A method for constructing an index tree, the method comprising: Acquire an object data set, where objects in the object data set include unstructured data and location information corresponding to the unstructured data; An index tree is constructed according to the storage location of the object in the data table; the root node and child nodes of the index tree respectively index the plurality of storage locations by spatial indexing; and the leaf nodes of the index tree index the plurality of storage locations by layer structure indexing; The leaf nodes of the index tree index the plurality of storage locations through a union layer structure of a layer structure obtained based on a hierarchical navigable small-world network and a layer structure obtained based on a Delaunay triangulation network; The object nodes in the layers of the union layer structure are the same as the object nodes of the layer structure obtained based on the hierarchical navigable small-world network or the layer structure obtained based on the Delaunay triangulation network; the paths between the object nodes of the union layer structure are obtained by the union of the paths between the object nodes of the layer structure obtained based on the hierarchical navigable small-world network and the layer structure obtained based on the Delaunay triangulation network; wherein the object nodes of the middle layer are a subset of the object nodes of the bottom layer, and the object nodes of the top layer are a subset of the object nodes of the middle layer, and the edges between the object nodes of the middle layer and the top layer are determined based on the object nodes of the bottom layer, that is, if there is an edge between two object nodes of the bottom layer, then there is also an edge between the object nodes of the middle layer and the top layer; If there is no edge between the two object nodes at the bottom layer, then there is no edge between the object nodes at the middle layer and the top layer.
2. The method according to claim 1, wherein the leaf nodes of the index tree index the plurality of storage locations at least through a layer structure obtained based on a hierarchical navigable small-world network.
3. A data query method, comprising: receiving a data query request, wherein the data query request includes unstructured data and a spatial constraint range; Based on a pre-built index tree, query targets that meet the spatial constraint range and match the unstructured data are searched; the root node and child nodes of the index tree respectively index multiple storage locations using a spatial index method; and the leaf nodes of the index tree index the multiple storage locations using a layered index method. The leaf nodes of the index tree index the plurality of storage locations through a union layer structure of a layer structure obtained based on a hierarchical navigable small-world network and a layer structure obtained based on a Delaunay triangulation network; The query target that satisfies the spatial constraint range and matches the unstructured data based on the pre-built index tree includes: Performing a query from a root node to a leaf node in the index tree to determine a target leaf node that intersects the spatial constraint range; if the spatial constraint range does not completely cover a minimum bounding rectangle corresponding to the target leaf node, obtaining a sample set of object nodes from the target leaf node and determining a selectivity of the sample set of object nodes in the spatial constraint range; Determining, based on a comparison result between the selectivity and the selectivity threshold, to use the layer structure obtained based on the hierarchical navigable small-world network or the layer structure obtained based on the Delaunay triangulation network in the union layer structure to query a query target that satisfies the spatial constraint range and matches the unstructured data; A query is performed from the root node to the leaf node in the index tree to determine a target leaf node that intersects with the spatial constraint range. If the spatial constraint range completely covers the minimum bounding rectangle corresponding to the target leaf node, a layer structure obtained based on a hierarchical navigable small-world network is used to query a query target that meets the spatial constraint range and matches the unstructured data. 4 . The method according to claim 3 , wherein the leaf nodes of the index tree index the plurality of storage locations at least through a layer structure obtained based on a hierarchical navigable small-world network.
5. The method according to claim 4, wherein the querying, based on the pre-built index tree, for a query target that satisfies the spatial constraint range and matches the unstructured data comprises: A query is performed from the root node to the leaf node in the index tree to determine a target leaf node that intersects with the spatial constraint range. If the spatial constraint range does not completely cover the minimum circumscribed rectangle corresponding to the target leaf node, a brute force search is performed in the target leaf node to obtain a query target that meets the spatial constraint range and matches the unstructured data.
6. The method according to claim 3, wherein the querying, based on the pre-built index tree, for a query target that satisfies the spatial constraint range and matches the unstructured data comprises: A query is performed from the root node to the leaf node in the index tree to determine a target leaf node that intersects with the spatial constraint range. If the spatial constraint range completely covers the minimum circumscribed rectangle corresponding to the target leaf node, the layer structure obtained based on the hierarchical navigable small-world network in the union layer structure is used to query a query target that meets the spatial constraint range and matches the unstructured data.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory, wherein the processor implements the method according to any one of claims 1 to 6 when executing the computer program.
8. A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Spatial index structure tree based method and device for providing results of searching spatial objects
CN103714080A
Spatio-temporal data index building and searching methods, a spatio-temporal data index building and searching device and spatio-temporal data index building and searching equipment
CN104750708A