A generalized approximate clustering skyline query processing method in road network environment
By adopting the DBSCAN algorithm and SSR-tree index in a road network environment, the problems of too few skyline query result sets and repeated judgment of similar points are solved, achieving more efficient query result enrichment and accuracy.
Patent Information
- Application Number
- CN202310942112.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-29
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2043-07-29
AI Technical Summary
During the processing of skyline queries in road networks, there are problems such as too few result sets and repeated judgment of a large number of similar points. In addition, existing algorithms cannot effectively handle objects with similar distances in road networks.
A variant of the DBSCAN algorithm is used for approximate clustering processing to construct an SSR-tree index. The judgment is dominated by generalized clustering to enrich the skyline query result set and reduce the judgment between approximate objects.
It improves the efficiency of skyline queries, enriches the query result set, reduces the judgment between similar objects, and meets the query needs of users.
Smart Images

Figure CN116860834B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data query, and in particular to a generalized approximate aggregation skyline query processing method in a road network environment. Background Art
[0002] Skyline queries, specifically targeting recommender systems, are used to solve multi-objective decision-making problems. Skyline queries are a query method used to find optimal solutions in multi-attribute datasets. Their goal is to identify a set of data objects (also called records or data points) that outperform others on multiple attributes, meaning they possess the best performance or value for certain attributes. In practical applications, skyline queries can be used in a variety of fields, such as data mining, decision support, and multi-objective optimization. They can help users identify the most valuable or promising options from large amounts of data, providing reference and insight for decision-making.
[0003] Skyline queries in road networks not only consider existing non-spatial attributes but also incorporate the network's distance attribute, making them more complex. Existing network skyline algorithms only consider distance as an attribute. If two objects are close to each other, their distance attribute values can be roughly assumed to be the same. Furthermore, the network may contain a large number of similar points that require repeated judgment, and the skyline result set may be too small. Summary of the Invention
[0004] In order to solve the above problems, the purpose of the present invention is to provide a generalized approximate clustering skyline query processing method in a road network environment, enrich the skyline query result set, reduce the judgment between approximate objects, improve efficiency, and solve the problem of too few traditional skyline query result sets.
[0005] The technical solution of the present invention is: a generalized approximate clustering skyline query processing method in a road network environment, comprising the following steps:
[0006] Step 1: For a given road network G and query dataset D, perform approximate clustering on the query dataset D using a variant of the DBSCAN algorithm to form a dataset D' with an approximate point set;
[0007] Step 2: Divide the road network based on the dataset D' and construct an SSR-tree index;
[0008] Step 3: Based on the SSR-tree index structure, perform generalized clustering dominance judgment of the road network between the approximate set and the independent points in the dataset, and output a result set with a two-layer approximate index structure.
[0009] The step of performing approximate clustering processing on the query data set D using the variant DBSCAN algorithm specifically includes:
[0010] Set the road network approximate distance d ε , if the road network distance between two objects is less than d ε , then the road network distances between the two objects are approximate; define a set of non-spatial dimension thresholds ε1, ε2, …, ε n If the difference between two objects in each non-spatial dimension is less than the non-spatial dimension threshold, the non-spatial dimensions of the two objects are similar; perform DBSCAN calculation on the midpoints of the dataset D, and the cluster radius is d ε , gather those d ε Objects that are approximated in non-spatial dimensions form an approximation set, and points that do not belong to any approximation point set are called independent points.
[0011] The SSR-tree indexing step specifically includes:
[0012] The traditional R-tree tree child nodes are connected to a two-dimensional list, and the approximate set in the dataset D' is regarded as an object in the space and stored in the two-dimensional list to construct the SSR-tree index, ensuring that the approximate set can be judged identically in the index.
[0013] The road network generalized aggregation control processing step specifically includes:
[0014] Given a dataset D' containing approximate sets and independent points, determine the dominance relationship between approximate sets, between approximate sets and independent points, and between independent points. Dominance judgment between approximate sets: If there is a point in approximate set A that dominates all points in approximate set B in all dimensions (including the road network distance dimension), then the approximate set A road network dominates approximate set B; dominance judgment between approximate sets and independent points: If an independent point p dominates all points in approximate set A in all dimensions (including the road network distance dimension), then the point p road network dominates approximate set A. Conversely, if there is a point in approximate set A that dominates independent point p in all dimensions (including the road network distance dimension), then the approximate set A dominates independent point p. The specific steps of the generalized approximate clustering skyline query processing method in the road network environment include:
[0015] Step 1: Create an empty result set to store the found skyline data points.
[0016] Step 2: Process data points one by one: Use the SSR-tree index to quickly determine the dominance relationship between objects (approximate sets and independent points) in the distance dimension.
[0017] Step 3: Domination judgment: Determine whether the current object (approximate set and independent point) is dominated by any object in the result set. If the current object is dominated, it will not be added to the result set.
[0018] Step 4: Non-domination judgment: If the current object is not dominated by any object in the result set, execute the next step.
[0019] Step 5: Dominate objects in the result set: Determine whether an object in the result set is dominated by the current object. If an object in the result set is dominated by the current object, remove the object from the result set.
[0020] Step 6: Add to result set: Add the current object to the result set.
[0021] The method of the present invention has the following beneficial effects: By finding and merging similar objects in the road network and establishing an effective SSR-tree index structure, the computational efficiency of the approximate set is improved. The proposed generalized clustering dominance method reduces the judgment between similar data, enriches the skyline query result set, and increases skyline query results, making the query results more in line with user needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 This is a flowchart of the steps of a generalized approximate clustered skyline query processing method in a road network environment of the present invention.
[0023] Figure 2 This is a flow chart of a variant DBSCAN approximate clustering algorithm according to a specific embodiment of the present invention.
[0024] Figure 3 This is an example diagram of an SSR-tree index according to a specific embodiment of the present invention. DETAILED DESCRIPTION
[0025] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. The step numbers in the following embodiments are provided for ease of description only and do not limit the order of the steps. The order of execution of the steps in the embodiments can be adaptively adjusted based on the understanding of those skilled in the art.
[0026] Reference Figure 1 and Figure 2 The present invention provides a generalized approximate clustering skyline query processing method in a road network environment, the method comprising the following steps:
[0027] Step 1: For a given road network G and query dataset D, perform approximate clustering on the query dataset D using a variant of the DBSCAN algorithm to form a dataset D' with an approximate point set; set the road network approximate distance d ε, if the road network distance between two objects is less than d ε , then the road network distances between the two objects are approximate; define a set of non-spatial dimension thresholds ε1, ε2, …, ε n If the difference between two objects in each non-spatial dimension is less than the non-spatial dimension threshold, the non-spatial dimensions of the two objects are similar; traverse the midpoints of the dataset D to perform variant DBSCAN calculations, and the cluster radius is d ε , gathering those objects that are similar in non-spatial dimensions to form an approximate set, and the points that do not belong to any approximate point set are called independent points.
[0028] The specific steps of the variant DBSCAN algorithm are as follows Figure 2 As shown:
[0029] Input: road network G and query dataset D, density radius d ε .
[0030] Output: Dataset D' with approximation set.
[0031] 1. Given a road network G and a query dataset D, all points in D are marked as unvisited, and the input radius d ε , density threshold ≡1.
[0032] 2. Randomly select an unvisited point p and mark it as visited.
[0033] 3. If the road network radius d at point p ε If there is at least one point in the domain, create a new cluster C and add p to C, otherwise mark p as an independent point.
[0034] 4. If the road network radius d at point p ε There is at least one point in the domain. Add all points in the domain of p to the set N, and loop to judge each unvisited point p' in N.
[0035] 5. Mark p' as visited and determine whether p' is non-spatially similar to p. If so, and p' has not yet been added to a cluster and marked as an independent point, add all points in p''s area to set N and add p' to C. If p' is not non-spatially similar to p, mark p' as an independent point.
[0036] 6. When the set N is empty, output cluster C.
[0037] 7. When there are no unvisited points in the dataset D, the operation ends and an approximate dataset D' is generated.
[0038] Step 2: SSR-tree index Figure 3As shown in the figure, the traditional R-tree tree child nodes are connected to the two-dimensional list, and the approximate set in the dataset D' is regarded as an object in the space and stored in the two-dimensional list. In this way, the SSR-tree index is constructed to ensure that the approximate set can be judged identically in the index.
[0039] 1. We abstract the approximate set into an object in the dataset to construct the R-tree index. We replace the root node of the R-tree index with a two-dimensional List list.
[0040] 2. The two-dimensional List is used to store approximate sets or independent points under the R-tree branch.
[0041] 3. Rapidly prune the dataset through indexing, and perform dominance judgment between approximate sets and independent points through our proposed generalized clustered skyline query.
[0042] Step 3: The generalized clustering dominance processing of the road network specifically includes: determining the dominance relationship between approximate sets, between approximate sets and independent points, and between independent points. Dominance judgment between approximate sets: If there is a point in approximate set A that dominates all points in approximate set B in all dimensions (including the road network distance dimension), then the approximate set A road network dominates approximate set B; Dominance judgment between approximate sets and independent points: If an independent point p dominates all points in approximate set A in all dimensions (including the road network distance dimension), then the point p road network dominates approximate set A. Conversely, if there is a point in approximate set A that dominates independent point p in all dimensions (including the road network distance dimension), then the approximate set A dominates independent point p; Dominance between independent points is the same as traditional skyline dominance.
[0043] The specific steps of the road network generalized clustering skyline algorithm are as follows:
[0044] Input: road network G and query dataset D', SSR-tree index.
[0045] Output: The result of the generalized clustered skyline query on the road network, which includes a two-layer approximate index structure of approximate sets and independent points.
[0046] 1. Create an empty result set to store the found skyline data points.
[0047] 2. Process data points one by one: Use the SSR-tree index to quickly determine the dominance relationship between various objects (approximate sets and independent points) in the distance dimension.
[0048] 2.a. Domination check: Check whether the current object (approximate set and independent point) is dominated by any object in the result set. If the current object is dominated, it will not be added to the result set.
[0049] 2.b. Non-domination judgment: If the current object is not dominated by any object in the result set, execute the next step.
[0050] 2.c. Dominate objects in the result set: Determine whether any object in the result set is dominated by the current object. If an object in the result set is dominated by the current object, remove the object from the result set.
[0051] 3. Add to result set: add the current object to the result set.
[0052] The above is a specific description of the preferred implementation of the present invention, but the invention is not limited to the embodiments. Those skilled in the art can make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of this application.
Claims
1. A generalized approximate clustering skyline query processing method in a road network environment, characterized by: Set the road network approximate distance d ε , if the road network distance between two objects is less than d ε , then the road network distances between the two objects are approximate; define a set of non-spatial dimension thresholds ε1, ε2, …, ε n If the difference between two objects in each non-spatial dimension is less than the non-spatial dimension threshold, the two objects are approximate in non-spatial dimensions; points that do not belong to any approximate point set are called independent points; for a given road network G and query dataset D, the specific operation steps include: Step 1: Perform approximate clustering on the query dataset D using a variant of the DBSCAN algorithm, clustering those d ε objects within the non-spatial dimension, clustering those d ε For objects with similar inner non-spatial dimensions, the objects in the dataset D are divided into approximate sets and independent points, and the objects in the approximate sets are integrated into a whole for processing, forming a dataset D' with approximate point sets and independent points; Step 2: Divide the road network based on the approximate set and independent points in the dataset D' generated in step 1, build an SSR-tree index, and add the approximate set as a whole to the SSR-tree index for the same processing; Step 3: Based on the SSR-tree index structure, the approximate set is put into the Skyline query to perform generalized clustering and domination of the road network of the approximate set and independent points in the dataset, and the result set with a two-layer approximate index structure is output; The specific steps of the variant DBSCAN algorithm are: Step 1: Given a road network G and a query dataset D, all points in D are marked as unvisited, and the input radius d ε , density threshold ≡ 1; Step 2: Randomly select an unvisited point p and mark it as visited; Step 3: If the road network radius d at point p ε If there is at least one point in the domain, create a new cluster C and add p to C, otherwise mark p as an independent point; Step 4: If the road network radius d at point p ε There is at least one point in the domain. Add all points in the domain of p to the set N, and loop to judge each unvisited point p' in N; Step 5: Mark p' as visited and determine whether p' is non-spatially similar to p. If so, and p' has not yet been added to a cluster and marked as an independent point, add all points in p''s area to set N and add p' to C. If p' is not non-spatially similar to p, mark p' as an independent point. Step 6: When the set N is empty, output cluster C; Step 7: When there are no unvisited points in the dataset D, the operation ends and an approximate dataset D' is generated; The step of generalized clustering domination processing of the road network specifically includes: for the generated data set D' and SSR-tree index, quickly querying the dominating set in D'. The generalized clustering domination processing of the road network is to process the dominating relationship between approximate sets and approximate sets, approximate sets and independent points, and independent points and independent points, and finally generate a query result set with a two-layer index structure with approximate sets. The dominance judgment between approximate sets and approximate sets is: if there is a point in the approximate set A that dominates all points in the approximate set B in all dimensions, then the approximate set A road network dominates the approximate set B; the dominance judgment between approximate sets and independent points is: if the independent point p dominates all points in the approximate set A in all dimensions, then the point p road network dominates the approximate set A; conversely, if there is a point in the approximate set A that dominates the independent point p in all dimensions, then the approximate set A dominates the independent point p.
2. The method for processing generalized approximate clustered skyline queries in a road network environment according to claim 1, characterized in that: The SSR-tree indexing step specifically includes: connecting the traditional R-tree tree child nodes to a two-dimensional list, treating the approximate set in the data set D' as an object in space and storing it in the two-dimensional list, thereby constructing an SSR-tree index to ensure that the approximate set is judged identically in the index.
Citation Information
Patent Citations
Position-based static skyline query method in road network environment
CN114064995A
Multi-source skyline query method and system based on minimum aggregation distance
CN115146020A