Skyline query method and device based on spatial index, equipment and storage medium

Through the QSTree index structure and dominant group concept, Skyline query is optimized, which solves the problems of low query efficiency and insufficient adaptability in massive data and high-dimensional environments, and realizes efficient Skyline query.

CN120336439APending Publication Date: 2025-07-18HUBEI UNIV OF ARTS & SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510330545.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing Skyline query algorithm is inefficient in massive data and high-dimensional environments, and is not adaptable to data diversity, resulting in reduced query complexity and performance.

Method used

The QSTree structure based on spatial index is adopted, and the index is constructed through the Z-order method, combining the concept of dominant groups and pruning strategies of the subspace to optimize the Skyline query process.

Benefits of technology

It significantly improves the efficiency and scope of application of Skyline query, reduces the complexity of data query, reduces expensive dominance tests, adapts to different types of data sets, and meets the query needs of complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336439A_ABST
    Figure CN120336439A_ABST
Patent Text Reader

Abstract

The invention provides a Skyline query method and device based on a spatial index, equipment and a storage medium. Comprising the steps of obtaining a to-be-queried data set, and constructing a QSTree index structure according to the to-be-queried data set; wherein the QSTree index structure comprises a root node, the root node is spatially divided into a plurality of child nodes, the child node at the tail end is a leaf node, and the to-be-queried data set is distributed in the leaf node according to coordinate values of the to-be-queried data set; determining a dominating group of a query object according to the QSTree index structure; wherein the query object is any object in the nodes, and the dominating group of the query object comprises leaf nodes and non-leaf nodes; and performing pruning operation on the QSTree index structure according to the dominating group of the query object, and determining a Skyline query result set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of data query, and particularly relates to a Skyline query method, device, equipment and storage medium based on a spatial index. Background Art

[0002] Skyline query is a typical multi-objective optimization problem and has been widely applied in fields such as intelligent recommendation, medical health, resource scheduling and decision support. For example, Skyline query can help users find the option with the best cost performance from a list of hotels, or screen out products that meet multi-dimensional preferences in a recommendation system.

[0003] In recent years, the research on Skyline query has mainly focused on optimizing the efficiency of query algorithms. Lee et al. proposed a Skyline query algorithm ZSearch based on the ZBTree index structure (Ken C.K. Lee, Wang Chien Lee, Baihua Zheng, Huajing Li, and Yuan Tian. 2010. Z-SKY: An efficient skyline query processing framework based on Z-order. VLDB Journal 19, 3 (2010), 333-362.), which is the first algorithm to alleviate the curse of dimensionality problem. ZSearch avoids performing dominance tests on all Skyline objects and determines whether an object belongs to a Skyline object through an efficient indexing mechanism.

[0004] Subsequently, Lee and Hwang et al. (Jongwuk Lee and Seung-won Hwang. 2010. BSkyTree: scalable skyline computation using a balanced pivot selection (EDBT’10). Association for Computing Machinery, New York, NY, USA, 195-206.) introduced the concept of pivot object partitioning and reduced the number of dominance tests through the strategy of incomparability between regions, thereby further improving the query efficiency.

[0005] In addition, Zhang et al. (Ji Zhang, Wenlu Wang, Xunfei Jiang, Wei-Shinn Ku, and Hua Lu. 2019. An MBR-Oriented Approach for Efficient Skyline Query Processing. In 2019 IEEE 35th International Conference on Data Engineering (ICDE). 806–817.) proposed a Skyline query method based on the minimum bounding rectangle (MBR), which makes the objects inside the MBR only need to perform dominance tests with the objects in its dependency group by introducing the dependency group of the MBR.

[0006] However, the existing technologies have the following defects: First, the processing bottleneck of massive data: With the sharp increase in the amount of data, the query efficiency of traditional algorithms has decreased significantly, making it difficult to meet the real-time requirements of today's massive data scenarios. Second, the curse of dimensionality problem: In a high-dimensional environment, the number of Skyline objects increases exponentially. This "curse of dimensionality" phenomenon greatly increases the complexity of the query, resulting in reduced pruning and query efficiency. Third, insufficient adaptability to data diversity: Traditional algorithms usually perform well only on some data sets. On other data sets, due to the limitations of their algorithm optimization strategies, they often cannot fully consider the complex relationships and differences between data, resulting in reduced computational efficiency and ineffective pruning strategies, thus affecting the overall performance. Summary of the Invention

[0007] The present invention provides a Skyline query method, device, equipment, and storage medium based on a spatial index, which can solve the curse of dimensionality problem, and can also improve the adaptability of the query method and the query efficiency in the face of massive data.

[0008] On the one hand, a Skyline query method based on a spatial index is provided, including:

[0009] Obtain a data set to be queried, and construct a QSTree index structure according to the data set to be queried; wherein, the QSTree index structure includes a root node, the root node is spatially divided into multiple child nodes, and the outermost child nodes are leaf nodes, and the data set to be queried is distributed in the leaf nodes according to its coordinate values;

[0010] Determine the dominance group of the query object according to the QSTree index structure; wherein, the query object is any object in the node, and the dominance group of the query object includes leaf nodes and non-leaf nodes;

[0011] Perform pruning operations on the QSTree index structure according to the dominance group of the query object, and determine the Skyline query result set.

[0012] Optionally, construct a QSTree index structure based on the dataset to be queried, including:

[0013] Convert each data in the dataset to be queried into a binary sequence, and arrange the binary sequences in ascending order;

[0014] Load the binary sequences into the root node, and recursively divide the root node in space to obtain child nodes and leaf nodes, forming a QSTree index structure.

[0015] Optionally, the step of recursively dividing the root node in space to obtain child nodes and leaf nodes includes:

[0016] Perform a spatial division on the root node to obtain leaf nodes;

[0017] Determine the size of the data in the leaf nodes. When the data accommodated in the leaf nodes exceeds the first threshold, divide the leaf nodes to form new leaf nodes; among them, the leaf nodes before division are used as child nodes;

[0018] Determine the size of the data in all leaf nodes. When the data volume in any one leaf node exceeds the first threshold, continue to divide the leaf nodes until the data volume in all leaf nodes is less than the first threshold;

[0019] Determine the size of the data in all leaf nodes. When the sum of the data volumes of any two adjacent leaf nodes is less than the second threshold, merge the two adjacent leaf nodes.

[0020] Optionally, the determination method of the dominance group of the query object is as follows:

[0021] Determine the object information minpts in the node, where minpts is composed of data points with the minimum coordinate values in each coordinate dimension of the current node;

[0022] Assume that there is an object p and the minpts of any node. If the corresponding dimension space coordinate of the minpts of the node in one coordinate dimension is less than the object p, and the corresponding dimension space coordinates in other dimensions are less than or equal to the object p, then determine that the node dominates the object p, and this node belongs to the dominance group of the object p.

[0023] Optionally, the steps of pruning the QSTree according to the dominance group of the query object and determining the result set include:

[0024] Use the domination group of nodes to perform pruning operations on the QSTree to remove the dominated child nodes; traverse the data points in the remaining nodes and perform domination tests on the data points in the remaining nodes to determine the result set.

[0025] Optionally, the steps of using the domination group of nodes to perform pruning operations on the QSTree to remove the dominated child nodes, traversing the data points in the remaining nodes, and performing domination tests on the data points in the remaining nodes to determine the result set include:

[0026] Take the first child node of the root node as the current node;

[0027] Use the domination test algorithm to determine whether the minpts of the current node is dominated. If it is not dominated, traverse the child nodes of the current node from left to right to determine the undominated child node as the new current node; if the minpts of the current node is dominated, it means the current node is dominated and prune the current node, and reselect the sibling node that dominates the current node as the new current node, and determine whether the current node is dominated;

[0028] Determine whether the current node is a leaf node. If it is a leaf node, traverse all the data points stored in the leaf node and use the domination test algorithm to determine whether these data points are dominated. If the data point is dominated, discard it; if the data point is not dominated, add it to the result set; if the current node is not a leaf node, continue to traverse the child nodes of the current node to determine the undominated child nodes in the current node until the current node is a leaf node;

[0029] Determine whether the traversal of the child nodes of the root node is completed. If the traversal is completed, output the result set; if the traversal is not completed, continue to select the child node of the root node as the current node until the traversal of the child nodes of the root node is completed, and output the result set.

[0030] Optionally, the steps of using the domination test algorithm to determine whether p is dominated are as follows:

[0031] Select an object p in a node;

[0032] Take the node where p is located as the current node;

[0033] If the current node has sibling nodes, traverse the sibling nodes of the current node to determine whether the sibling nodes of the current node dominate p; if the sibling nodes do not dominate p, it is proved that all the data under the current node has no possibility of dominating p; if the sibling nodes dominate p, there is a possibility that the data points in the node dominate p. Find all the child nodes belonging to the domination group p in the node to replace the node; or directly use the pointers of the node to access the Skyline result set under the node to determine whether there are truly dominating data points. If so, p is dominated.

[0034] Visit the parent node of the current node and repeat the above process until the Skyline result sets pointed to by the pointers of all nodes do not dominate p, which proves that p is not dominated.

[0035] In a second aspect, a Skyline query device based on a spatial index is provided, including:

[0036] An acquisition module for acquiring a data set to be queried and constructing a QSTree index structure according to the data set to be queried; wherein, the QSTree index structure includes a root node, the root node is spatially divided into multiple child nodes, and the outermost child nodes are leaf nodes. The data set to be queried is distributed in the leaf nodes according to its coordinate values.

[0037] A first determination module for determining the domination group of a query object according to the QSTree index structure; wherein, the query object is any object in the node, and the domination group of the query object includes leaf nodes and non-leaf nodes.

[0038] A second determination module for pruning the QSTree index structure according to the domination group of the query object to determine the Skyline query result set.

[0039] In a third aspect, an electronic device is provided, including the Skyline query device based on a spatial index as described above.

[0040] In a fourth aspect, a computer-readable storage medium is provided, in which at least one program code is stored, and the program code is executed by a processor to implement the Skyline query method based on a spatial index as described in any one of the above.

[0041] The beneficial effects of the technical solution provided by the present invention are:

[0042] First, regarding the processing bottleneck of massive data: The spatial index structure (QSTree) based on the PR quadtree designed in the present invention effectively reduces the complexity of data query through a hierarchical structure of spatial partitioning. Compared with traditional R-Tree and ZBTree index structures, there is no spatial overlap between sibling nodes. The present invention introduces the concept of the dominance group of subspaces, which can minimize expensive dominance tests. On the other hand, combined with the pruning optimization strategy, the real-time query efficiency in the massive data scenario is significantly improved. Therefore, this technology can complete Skyline queries with a lower time complexity in large-scale data scenarios.

[0043] Second, regarding the curse of dimensionality problem: The present invention reduces the number of incomparable cases significantly through the method of spatial partitioning and by introducing the concept of the dominance group of subspaces. On the other hand, as the dimension gradually increases, the number of incomparable spaces in high-dimensional subspaces gradually increases, significantly reducing the curse of dimensionality problem. Therefore, for the curse of dimensionality problem, this technology can effectively improve the query efficiency.

[0044] Third, regarding the insufficient adaptability to data diversity: The performance of traditional algorithms on heterogeneous data sets is limited by the characteristics of data distribution, while the QSTree index structure designed in the present invention does not overly focus on data but is committed to space. For different types of data sets, only the subspace to which the data belongs needs to be considered. Therefore, whether it is a uniformly distributed or anti-correlated data set, the present invention can efficiently perform query processing, thus ensuring that the query efficiency is not reduced due to data heterogeneity.

[0045] Through the improvement of the above technical solutions, the present invention overcomes the problems of the prior art in massive data processing, complexity in high-dimensional environments, and insufficient adaptability to heterogeneous data, significantly improving the efficiency and application scope of Skyline queries and meeting the Skyline query requirements in today's complex scenarios. Brief Description of the Drawings

[0046] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0047] Figure 1 It is a flowchart of a Skyline query method based on a spatial index provided by the present invention;

[0048] Figure 2 It is a flowchart of a construction method of a QSTree index structure provided by the present invention;

[0049] Figure 3 Structural schematic diagram of a QSTree provided by the present invention;

[0050] Figure 4 Another structural schematic diagram of a QSTree provided by the present invention;

[0051] Figure 5 Structural schematic diagram of incomparability in two-dimensional space provided by the present invention;

[0052] Figure 6 Structural schematic diagram of incomparability in three-dimensional space provided by the present invention;

[0053] Figure 7 Flowchart of a dominance test algorithm provided by the present invention;

[0054] Figure 8 Flowchart of an ISSQ algorithm provided by the present invention;

[0055] Figure 9 Schematic diagram of an ISSQ execution process provided by the present invention;

[0056] Figure 10 Block diagram of the structure of a Skyline query device based on a spatial index provided by the present invention;

[0057] Figure 11 Block diagram of the structure of an electronic device provided by the present invention.

[0058] The reference numerals are as follows:

[0059] 11: Acquisition module; 12: First processing module; 13: Second processing module;

[0060] 21: Processor; 22: Memory. Detailed implementation manners

[0061] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.

[0062] Given a multi-dimensional dataset, the Skyline query aims to find all objects that are not dominated by any other data objects in the dataset. If object p is not worse than p' in any dimension and there is at least one dimension in p that is better than p', then p dominates p'. Without loss of generality, this description assumes that the smaller the dimension value, the better, and conversely, the larger the value, the worse.

[0063] Figure 1 The flowchart of a Skyline query method based on spatial index provided by the present invention. Refer to Figure 1 , including:

[0064] S101. Obtain the dataset to be queried, and construct a QSTree index structure according to the dataset to be queried; wherein, the QSTree index structure includes a root node, the root node is spatially divided into multiple child nodes, and the outermost child nodes are leaf nodes, and the dataset to be queried is distributed in the leaf nodes according to its coordinate values.

[0065] In an example, step S101 includes:

[0066] Step 1. Convert the data to be queried into a binary sequence, and arrange the binary sequence in ascending order.

[0067] Among them, the data to be queried is converted into a binary sequence by the Z-order method. The Z-order method is a method of mapping multi-dimensional data to a one-dimensional space, which is often used in fields such as dimension reduction, storage reduction, and space partitioning. Z-order interleaves the multi-dimensional values bit by bit into one value to achieve the purpose of converting a high-dimensional space into one dimension.

[0068] Specifically, mapping a three-dimensional space data (3, 5, 1) to a one-dimensional space by the Z-order method, first, the values of each dimension need to be converted into binary (011, 101, 001) representation. Then, according to the Z-order method, the binary bits of each dimension are alternately combined to form a new binary sequence "010100111". In addition, "010100111" is also composed of three groups of d-bits ("010", "100", "111"), and each group from left to right corresponds to a more fine-grained space partition.

[0069] Specifically, d-bit is the binary bit used to represent space partitioning. By different combinations of d-bit binary, the partitioning of space can be achieved. Moreover, each individual bit within d-bit represents a single partition along one dimension. Using the Z-order method can not only improve the efficiency of index construction, but also effectively avoid the problem of space overlap. On the other hand, sorting the binary sequence formed by the Z-order method can ensure the progressive nature of Skyline queries. In addition, objects with smaller binary sequences contribute significantly to eliminating non-Skyline objects. By utilizing the Z-order method, Skyline objects can be effectively identified while pruning a large number of irrelevant objects.

[0070] Step 2: Load the binary sequence into the root node, and recursively partition the root node in space to obtain child nodes and leaf nodes, forming a QSTree index structure.

[0071] In this embodiment, the root node is the node including all data to be queried. In space, the root node can be divided into multiple child nodes, and the child nodes at the outermost end are identified as leaf nodes. Among them, the child nodes are the nodes between the root node and the leaf nodes. For example, the root node can be divided into multiple first child nodes in space, and the first child nodes can be divided into multiple second child nodes, and the second child nodes can be further divided into multiple third child nodes in space. If the third child node is the outermost end node, then the third child node is the leaf node.

[0072] In one example, Step 2 includes:

[0073] The first step: Partition the root node in space to obtain leaf nodes.

[0074] The second step: Determine the size of the data in the leaf nodes. When the data accommodated in the leaf nodes exceeds the first threshold, divide the leaf nodes to form new leaf nodes; among them, the leaf nodes before division serve as child nodes.

[0075] The third step: Determine the size of the data in all leaf nodes. When the data volume in any one leaf node exceeds the first threshold, continue to divide the leaf nodes until the data volume in all leaf nodes is less than the first threshold.

[0076] The fourth step: Determine the size of the data in all leaf nodes. When the sum of the data volumes of any two adjacent leaf nodes is less than the second threshold, merge the two adjacent leaf nodes.

[0077] In this embodiment, ensure that the data volume in the leaf nodes is small, so that by using the dominating set and incomparable nodes of the nodes, a large number of nodes can be removed, and only the data points in a small number of nodes need to be screened, thus saving a large amount of screening time.

[0078] Figure 2 This is a flowchart of a method for constructing a QSTree index structure provided by the present invention. Refer to Figure 2 , first, map all data into binary sequences by the Z-order method, and sort these binary sequences in descending order to facilitate partitioning using the d-bit of the binary sequences.

[0079] Subsequently, load all data into a root node and recursively partition the root node. If the number of data in any leaf node (pop) exceeds the capacity upper limit C, then the node needs to be divided.

[0080] Specifically, extract the largest binary sequence from the data in pop, determine the position pos of the d-bit required for the next division. The pos is used to locate the position of extracting the d-bit from all binary sequences in the current node. And determine whether the data can be divided into d spaces through the d-bit. If the number of partitions is 1, it proves that the partition is unsuccessful, and continue to recursively select the next group of d-bits for division (that is, shift the position of pos to the right by d bits).

[0081] When the partition is successful, it is necessary to further determine whether the data volume of each child node after division is less than the threshold G. If it is less than G, check whether adjacent child nodes can be merged. If the data volume after merging does not exceed C, perform the merge operation. This process is repeated until the data volume of all leaf nodes does not exceed C.

[0082] In addition, in order to enable the QSTree to effectively support Skyline queries, additional fields need to be attached to each node of the QSTree. The additional fields can be abstracted as <minpts, pointers>, where minpts represents the minimum values of each dimension of the data objects in all leaf nodes under the current node. The pointers field stores the start (pointers.pre) and end (pointers.last) positions of the Skyline result sets corresponding to all leaf nodes under the current node. At the same time, the pointers field can track the number of Skyline objects calculated in the current node, so as to use sequential I / O to reduce unnecessary random I / O operations. During the process of initializing the node, both of these pointers point to null, and as the Skyline objects are calculated, these two pointers will be dynamically updated.

[0083] Figure 3 and Figure 4 is the structural schematic diagram of the QSTree provided by the present invention. Among them, the root node [0, 63] is divided into child nodes [0, 15], [16, 31], [32, 47], [48, 63]. p1 to p 11 represent data points.

[0084] Figure 5 and Figure 6 is another structural schematic diagram of the QSTree provided by the present invention. Among them, Figure 5 represents the schematic diagram of spatial incomparability in the two-dimensional space, Figure 6 represents the schematic diagram of spatial incomparability in the three-dimensional space. Sub0 to Sub7 represent subspaces (that is, child nodes).

[0085] Of course, Figure 5 and Figure 6This is only an example provided by the present invention. In actual situations, Sub0 can also be divided into more subspaces (i.e., child nodes), and the present invention does not limit this.

[0086] S102. Determine the dominance group of the query object according to the QSTree index structure; wherein, the query object is any object in the node, and the dominance group of the query object includes leaf nodes and non-leaf nodes.

[0087] Among them, the determination method of the dominance group of the node is as follows:

[0088] Determine the object information minpts in the node, where minpts is composed of data points with the minimum coordinate values in each coordinate dimension of the current node.

[0089] Suppose there is an object p and the minpts of any node. If the corresponding dimension space coordinates of the minpts of the node in one coordinate dimension are less than the object p and the corresponding dimension space coordinates in other dimensions are less than or equal to the object p, then determine that the node dominates the object, and this node belongs to the dominance group of the object p.

[0090] S103. Perform a pruning operation on the QSTree index structure according to the dominance group of the query object to determine the result set.

[0091] In one example, step S103 includes:

[0092] Use the dominance group of the node to perform a pruning operation on the QSTree to remove the dominated child nodes; traverse the data points in the remaining nodes, perform a dominance test on the data points in the remaining nodes, and determine the result set.

[0093] Exemplarily, step S103 includes:

[0094] The first step: Take the first child node of the root node as the current node.

[0095] The second step: Use the dominance test algorithm to determine whether the minpts of the current node is dominated. If it is not dominated, traverse the child nodes of the current node from left to right, and determine the undominated child node as the new current node; if the minpts of the current node is dominated, it means that the current node is dominated and prune the current node, and re-select the sibling node that dominates the current node as the new current node, and determine whether the current node is dominated.

[0096] Among them, taking the root node as an example, the root node is divided into multiple first child nodes, and the first child nodes are divided into multiple second child nodes. Then, the first child nodes are siblings of each other, the root node is the father node of the first child nodes, and the second child nodes are the child nodes of the first child nodes.

[0097] Step 3: Determine whether the current node is a leaf node. If it is a leaf node, traverse all data points under the leaf node, and use the dominance test algorithm to determine whether it is dominated. If it is dominated, discard it; if it is not dominated, add it to the result set. If the current node is not a leaf node, continue to traverse the child nodes of the current node to determine the non-dominated child nodes in the current node until the current node becomes a leaf node.

[0098] Step 4: Determine whether the traversal of the child nodes of the root node is completed. If the traversal is completed, output the result set; if the traversal is not completed, continue to select the child nodes of the root node as the current node until the traversal of the child nodes of the root node is completed, and then output the result set.

[0099] In one example, the steps of using the dominance test algorithm to determine whether p is dominated are as follows:

[0100] Step 1: Select an object p in a node, where p is the minpts in the node or a data object in the leaf node. If p is the minpts in the node, determine whether the node is pruned. If p is a data object in the leaf node, determine whether p is a Skyline object.

[0101] Step 2: Take the node where the object p is located as the current node, and find the dominance group of the object p.

[0102] Step 3: If the current node has sibling nodes, traverse the sibling nodes of the current node to determine whether the sibling nodes of the current node dominate p. If the sibling nodes do not dominate p, it proves that the current node does not belong to the dominance group of the object p, that is, all data under the current node has no possibility of dominating p. If the sibling nodes dominate p, it proves that the current node belongs to the dominance group of the object p, that is, there is a possibility that the data points in the node dominate p.

[0103] Step 4: Visit the parent node of the current node and repeat the above process until a dominance group of the object p is found.

[0104] Step 5: Traverse the dominance group of the object p. Determine whether to continue to find all child nodes in the node that belong to the dominance group p to replace the node through a threshold; or directly use the pointers of the node to access the Skyline result set under the node to determine whether there are data points that truly dominate p.

[0105] Step 6: Repeat the above process until the traversal of the dominance group of the object p is completed, and there is no Skyline object in the Skyline result set of any node during the traversal process, which proves that p is not dominated; otherwise, p is dominated.

[0106] Figure 7The flowchart of a domination test algorithm provided by the present invention. Refer to Figure 7 , to determine whether an object p is dominated, the following steps can be adopted:

[0107] Step 1: If the object p is the minimum value minpts of the node, add the node cur where the object p is located to the stack.

[0108] If the object p is the minimum value minpts of the node, it is necessary to determine whether the node cur is pruned; if the object p is not the minimum value minpts of the node, then p is a data point, and then it is necessary to determine whether the object p is a skyline object.

[0109] Step 2: Visit the parent node pcur of cur.

[0110] Step 3: Traverse the child nodes child of pcur from left to right until cur; if child dominates the object p, add child to the stack.

[0111] The child node must be on the left side of cur and does not include cur, and the right side of cur has not been traversed; if child dominates p, it proves that there is a possibility that the Skyline object of the already calculated child dominates the object p.

[0112] Step 4: Let cur be pcur, and repeat Step 3 until cur is empty.

[0113] Step 5: Pop the node pop from the stack, visit the pointers field of the node pop, if pointers.last - pointers.pre <= Q (at this time, there is no need to traverse the child nodes of pop), then directly traverse the skyline objects in SL through pointers. At this time, if there is an object in SL that dominates p, end the algorithm and return the result true (the object p is dominated); otherwise, traverse the child node pchild of the pop node. If pchild dominates the object p, add pchild to the stack.

[0114] Step 6: Repeat Step 5 until the stack is empty. If it does not end in Step 5, return the result false (the object p is not dominated).

[0115] Figure 8 The flowchart of an ISSQ algorithm provided by the present invention. Refer to Figure 8, where the ISSQ algorithm (i.e., the Skyline result set query algorithm) first pushes the root node root of the QSTree onto the stack stack, and then continuously pops nodes from the stack, which is called pop. During the popping process, the pointers.pre and pointers.last of the pop node are pointed to the position of the current Skyline result set. Then, it is determined whether pop is dominated through the dominance test algorithm. If the dominance test algorithm returns TRUE, pop will be pruned. On the other hand, if pop is a non-leaf node, its child nodes are added to the stack. Otherwise, the dominance test algorithm is used to check whether each object in pop is dominated. If there is at least one Skyline object in pop, the pointers.last pointer of pop is updated. After accessing all the objects in pop, the pointers.last pointers of the ancestor nodes of pop are dynamically updated. This process is repeated until the stack is empty.

[0116] Next, the dataset in Figure 3 and Figure 4 will be used to illustrate the ISSQ algorithm. First, assume that the threshold Q = 2, and use Figure 9 to record the execution process of the algorithm, which details the state changes of the stack Stack, the currently processed node cur, the SDG (i.e., the dominance group) of the subspace, and the Skyline result set SL. First, the root node [0, 63] of the QSTree is added to the stack. The node [0, 63] is popped and not dominated because the parent node of the current root node is NULL, and its child nodes are added. In particular, the leaf nodes {p10, p11} are discarded due to the properties of the subspace.

[0117] The nodes in the stack are {[32, 47], [26, 31], [0, 15]}. At this time, the top of the stack is [0, 15]. Pop this node and it is not dominated. Then push the child nodes {p0, p1} and {p2} of the node [0, 15] into the stack. At this time, the top node of the stack is {p0, p1} and pop the leaf node {p0, p1}. Perform a dominance test on the leaf node {p0, p1}. Since the dominance group of the leaf node {p0, p1} is empty, it is not dominated. Perform a dominance test on its data {p0, p1}. The dominance groups of objects p0 and p1 are the same, which is the node {p0, p1}. At this time, add object p0 and object p1 to the Skyline result set, and dynamically update the pointers.last pointers of the nodes {p0, p1}, [0, 15] and [0, 63]. Subsequently, pop the top node {p2} of the stack. The dominance group of p2 is {p2}, so p2 is not dominated and is added to the Skyline result set. At this time, the Skyline result set is {p0, p1, p2}. In the subsequent dominance test process, when accessing the node [0, 15], there is no need to access its child nodes because the size of the Skyline result set of the node [0, 15] is 3, which is greater than the threshold Q. Then, pop the node [16, 31]. Its dominance group is {[0, 15]}. Since the Skyline set of the node [0, 15] is {p0, p1, p2} and p0 dominates the minpts = [1, 5] of the node [16, 31], at this time, the node [16, 31] is pruned. Finally, the leaf nodes {p6, p7} are dominated by the Skyline object p2 in the node [0, 15] and are discarded. p8 is not dominated by the Skyline result set in the node [0, 15] and is added to the Skyline result set. At this time, the stack is NULL, the algorithm ends, and the Skyline result set is {p0, p1, p2, p8}.

[0118] The beneficial effects brought by the technical solution provided by the present invention are:

[0119] On the one hand, by introducing the Z-order method to construct the QSTree index structure, this application can not only ensure the incrementality of skyline queries, but also facilitate pruning strategies. Meanwhile, this application introduces the concept of SDG (Spatial Dominance Group). Firstly, potential incomparable spaces can be filtered through bit operations (in QSTree, d-bits are used to represent the positions of child nodes, and there are 2^d child nodes in total. Assume that nodes n1 and n2 are sibling nodes, where the d-bit value of n1 is represented as i, the d-bit value of n2 is represented as j, and i > j. If i & j ≠ j, then the objects in n1 cannot dominate any objects in n2). Then, it is judged whether it is dominated through the field minpts, minimizing the expensive dominance test. Then, two pointers in pointers are used to track the number of Skyline objects of the current node, thus avoiding unnecessary I / O operations. On the other hand, using the Z-order method to construct the QSTree index structure can simplify the construction process and reduce the storage consumption of the index. Meanwhile, using the QSTree index can filter out potential incomparable spaces through bit operations, reducing the expensive dominance test. Moreover, the dominance test algorithm only needs to perform dominance tests in the dominance groups of the current object's node and its ancestor nodes, avoiding incomparability of nodes and being able to efficiently process large Skyline result sets. The ISSQ algorithm performs efficient pruning through the depth-first traversal strategy and minpts, effectively avoiding the situation where algorithm performance is reduced due to data diversity.

[0120] Regarding the processing bottleneck of massive data: The spatial index structure (QSTree) based on the PR quadtree designed by the present invention effectively reduces the complexity of data queries through a hierarchical structure of spatial partitioning. Compared with traditional R-Tree and ZBTree index structures, there is no spatial overlap between sibling nodes. The present invention introduces the concept of the dominance group of subspaces, which can minimize the expensive dominance test. On the other hand, combined with pruning optimization strategies, the real-time query efficiency in massive data scenarios is significantly improved. Therefore, this technology can complete skyline queries with a lower time complexity in large-scale data scenarios.

[0121] Regarding the curse of dimensionality problem: The present invention greatly reduces the number of incomparabilities through the method of spatial partitioning and by introducing the concept of the dominance group of subspaces. On the other hand, as the dimension gradually increases, the number of incomparable spaces in high-dimensional subspaces gradually increases, significantly reducing the curse of dimensionality problem. Therefore, this technology can effectively improve the query efficiency for the curse of dimensionality problem.

[0122] Insufficient adaptability to data diversity: The performance of traditional algorithms on heterogeneous data sets is limited by the characteristics of data distribution. However, the QSTree index structure designed in the present invention does not overly focus on data but is committed to space. For different types of data sets, only the subspace to which the data belongs needs to be considered. Therefore, whether it is a uniformly distributed or anti-correlated data set, the present invention can efficiently perform query processing, thereby ensuring that the query efficiency is not reduced due to data heterogeneity.

[0123] Through the improvement of the above technical solutions, the present invention overcomes the problems of the prior art in mass data processing, complexity in high-dimensional environments, and insufficient adaptability to heterogeneous data, significantly improves the efficiency and application scope of Skyline queries, and can meet the Skyline query requirements in today's complex scenarios.

[0124] Figure 10 It is a structural block diagram of a Skyline query device based on a spatial index provided by the present invention.

[0125] See Figure 10 , including:

[0126] An acquisition module 11, configured to acquire a data set to be queried and construct a QSTree index structure according to the data set to be queried; wherein, the QSTree index structure includes a root node, the root node is spatially divided into multiple sub-nodes, and the outermost sub-nodes are leaf nodes, and the data set to be queried is distributed in the leaf nodes according to its coordinate values;

[0127] A first determination module 12, configured to determine a domination group of a query object according to the QSTree index structure; wherein, the query object is any object in the node, and the domination group of the query object includes leaf nodes and non-leaf nodes;

[0128] A second determination module 13, configured to perform a pruning operation on the QSTree index structure according to the domination group of the query object to determine a Skyline query result set.

[0129] Figure 11 It is a structural block diagram of an electronic device provided by the present invention. See Figure 11 , including: The electronic device may include Figure 10 The Skyline query device based on a spatial index described above. Generally, the electronic device includes: a processor 21 and a memory 22.

[0130] The processor 21 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 21 may be implemented in at least one of the hardware forms of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 21 may also include a main processor and a co-processor. The main processor is a processor used to process data in the wake state, also known as the CPU (Central Processing Unit); the co-processor is a low-power processor used to process data in the standby state. The memory 22 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 22 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 22 is used to store at least one instruction, and the at least one instruction is used to be executed by the processor 21 to implement the spatial index-based Skyline query method executed by the electronic device provided in the method embodiments of the present application.

[0131] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A Skyline query method based on spatial indexing, characterized in that, Including: Obtain the dataset to be queried, and construct a QSTree index structure based on the dataset to be queried; wherein, the QSTree index structure includes a root node, the root node is spatially divided into multiple child nodes, and the outermost child nodes are leaf nodes, and the dataset to be queried is distributed in the leaf nodes according to its coordinate values; Determine the dominance group of the query object according to the QSTree index structure; wherein, the query object is any object in the node, and the dominance group of the query object includes leaf nodes and non-leaf nodes; Perform a pruning operation on the QSTree index structure according to the dominance group of the query object to determine the Skyline query result set.

2. The Skyline query method based on spatial index according to claim 1, wherein ,Constructing a QSTree index structure according to the dataset to be queried includes: Convert each data in the dataset to be queried into a binary sequence, and arrange the binary sequences in ascending order; Load the binary sequences into the root node, and recursively divide the root node spatially to obtain child nodes and leaf nodes, forming a QSTree index structure.

3. The Skyline query method based on spatial index according to claim 2, wherein, The step of recursively dividing the root node spatially to obtain child nodes and leaf nodes includes: Spatially divide the root node to obtain leaf nodes; Determine the size of the data in the leaf nodes. When the data accommodated in the leaf nodes exceeds the first threshold, divide the leaf nodes to form new leaf nodes; wherein, the leaf nodes before division are used as child nodes; Determine the size of the data in all leaf nodes. When the data volume in any one leaf node exceeds the first threshold, continue to divide the leaf nodes until the data volume in all leaf nodes is less than the first threshold; Determine the size of the data in all leaf nodes. When the sum of the data volumes of any two adjacent leaf nodes is less than the second threshold, merge the two adjacent leaf nodes.

4. The method for Skyline query based on spatial index according to any one of claims 1 to 3, characterized in that, The determination method of the dominance group of the query object is as follows: Determine the object information minpts in the node, where minpts is composed of data points with the minimum coordinate values in each coordinate dimension of the current node; Assume that there is an object p and the minpts of any node. If the corresponding dimensional space coordinate of the minpts of the node in one coordinate dimension is less than the object p, and the corresponding dimensional space coordinates in other dimensions are less than or equal to the object p, then it is determined that the node dominates the object p, and the node belongs to the dominance group of the object p.

5. The Skyline query method based on spatial index according to any one of claims 1 to 3, characterized in that, The steps of performing a pruning operation on the QSTree according to the dominance group of the query object to determine the result set include: Use the dominance group of the node to perform a pruning operation on the QSTree to remove the dominated child nodes; traverse the data points in the remaining nodes, and perform a dominance test on the data points in the remaining nodes to determine the result set.

6. The Skyline query method based on spatial index according to claim 5, wherein Use the dominance group of the node to perform a pruning operation on the QSTree to remove the dominated child nodes; The steps of traversing the data points in the remaining nodes and performing a dominance test on the data points in the remaining nodes to determine the result set include: Take the first child node of the root node as the current node; Use the dominance test algorithm to determine whether the minpts of the current node is dominated. If it is not dominated, traverse the child nodes of the current node from left to right to determine the undominated child node as the new current node. If the minpts of the current node is dominated, it indicates that the current node is dominated and prune the current node, and re-select the sibling node that dominates the current node as the new current node, and determine whether the current node is dominated. Determine whether the current node is a leaf node. If it is a leaf node, traverse all the data points stored in the leaf node, and use the dominance test algorithm to determine whether these data points are dominated. If the data point is dominated, discard it. If the data point is not dominated, add it to the result set. If the current node is not a leaf node, continue to traverse the child nodes of the current node to determine the undominated child nodes in the current node until the current node is a leaf node. Determine whether the traversal of the child nodes of the root node is completed. If the traversal is completed, output the result set. If the traversal is not completed, continue to select the child nodes of the root node as the current node until the traversal of the child nodes of the root node is completed, and output the result set.

7. The Skyline query method based on spatial index according to claim 6, characterized in that The steps of using the dominance test algorithm to determine whether p is dominated are as follows: Select an object p in a node. Take the node where p is located as the current node. If the current node has sibling nodes, traverse the sibling nodes of the current node to determine whether the sibling nodes of the current node dominate p. If the sibling node does not dominate p, it proves that there is no possibility that the data under the current node dominates p. If the sibling node dominates p, there is a possibility that the data points in the node dominate p. Find all the child nodes belonging to the dominance group of p in the node to replace the node; or directly use the pointers pointer of the node to access the Skyline result set under the node to determine whether there is a data point that truly dominates p. If so, p is dominated. Access the parent node of the current node and repeat the above process until the Skyline result sets pointed to by the pointers of all nodes do not dominate p, which proves that p is not dominated.

8. A Skyline query device based on spatial indexing, characterized in that It includes: An acquisition module for acquiring a dataset to be queried and constructing a QSTree index structure according to the dataset to be queried; wherein, the QSTree index structure includes a root node, the root node is spatially divided into multiple child nodes, and the outermost child nodes are leaf nodes, and the dataset to be queried is distributed in the leaf nodes according to its coordinate values. A first determination module for determining the dominance group of the query object according to the QSTree index structure; wherein, the query object is any object in the node, and the dominance group of the query object includes leaf nodes and non-leaf nodes. A second determination module for pruning the QSTree index structure according to the dominance group of the query object to determine the Skyline query result set.

9. An electronic device, characterized in that, It includes the Skyline query device based on spatial index as described in claim 8.

10. A computer-readable storage medium, characterized in that, At least one program code is stored in the computer-readable storage medium, and the program code is executed by a processor to implement the spatial index-based Skyline query method according to any one of claims 1 to 7.