An optimal location query method and system considering the popularity and arrival distance of points of interest
By modeling the road network as a weighted undirected graph and dividing it into sub-graphs, combining the scoring mechanism of interest hot spots and arrival distance, the problem of ignoring the shortest travel route in the existing technology is solved, and efficient and practical travel route planning is achieved.
Patent Information
- Application Number
- CN202411929798.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-12-26
AI Technical Summary
When planning tourism routes, the prior art focuses on maximizing the popularity of attractions, while ignoring the user's preference for the shortest travel route, resulting in inefficient travel routes in the planning results, which is difficult to meet the actual needs of users.
By modeling the road network as a weighted undirected graph, dividing it into sub-graphs, and combining the scoring mechanism of point of interest heat and arrival distance in the filtering and refinement stage, potential edge segments covering all target categories are identified, and the upper bound of the score is preferred to quickly filter edge segments that cannot exceed the current highest score.
It realizes the optimal location query that ensures the shortest possible total travel distance while ensuring coverage of high-hot attractions, improving the query efficiency and practicality of the results.
Smart Images

Figure CN119357304B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of location query, and in particular to an optimal location query method and system that considers the popularity and arrival distance of a point of interest. Background Art
[0002] When tourists from other places visit a city, they usually want to book the most representative hotel in the city to maximize the total popularity (such as ratings and visits) of tourist attractions near the hotel due to the limited reach of their daily activities. At the same time, they also hope to find the shortest travel route to minimize the travel distance between these attractions. Existing technical methods often focus on maximizing the popularity of attractions, while ignoring the user's preference for the shortest travel route, resulting in inefficient travel paths in the planning results, which are difficult to meet the actual needs of users.
[0003] In response to this demand, the ideal query process is that users only need to enter their activity radius and the various places they want to reach, and the system can recommend an optimal hotel location and plan a shortest travel route covering these target locations. The "optimal" evaluation principle here is that the hotel should not only cover as many high-heat attractions as possible, but also ensure that the total travel distance to various attractions that users care about is as short as possible. The present invention refers to this type of optimal location query as an optimal location query considering the heat of points of interest and the arrival distance. The query involves a given road network, a set of points of interest on the road network, and a query input containing a query radius and a set of points of interest of a target category. The goal of the query is to find the optimal location (i.e., the maximum scoring location) in the road network, which can contain all the points of interest categories specified by the user within the query radius and has a better score than other candidate locations. The scoring mechanism comprehensively considers two factors: one is the sum of the heat values of the covered points of interest; the other is the shortest route length connecting the points of interest of the target category.
[0004] Current research mainly focuses on maximizing the sum of the heat values of points of interest on the road network, without considering the impact of the route length to the points of interest on the query results. This method cannot meet the actual needs of users for travel efficiency and rationality, and has significant limitations. Therefore, it is urgent to develop an efficient query algorithm that integrates the weight of points of interest and the shortest route length, so as to provide users with more practical travel planning results and better adapt to complex practical application scenarios. Summary of the invention
[0005] In order to solve the above-mentioned problems, the present invention provides an optimal location query method and system that considers the popularity and arrival distance of a point of interest.
[0006] In a first aspect, the present invention provides an optimal location query method considering the popularity and arrival distance of a point of interest, which adopts the following technical solution:
[0007] An optimal location query method considering the popularity and arrival distance of a point of interest, comprising:
[0008] Obtaining road network data and user query parameters; wherein the user query parameters include query radius and point of interest target category;
[0009] The road network is modeled as a weighted undirected graph, where vertices represent intersections, edges represent road segments, edge weights represent the length of the road segments, and points of interest are mapped to vertices or edges of the graph;
[0010] Divide the entire weighted undirected graph into several subgraphs, and construct a list of points of interest and a shortest distance table within each subgraph for each subgraph;
[0011] In the filtering stage, the categories of POIs are counted according to the user query parameters, and the subgraphs that fail to fully cover the target category are marked as low-potential subgraphs; the search is expanded from the low-potential subgraphs to the adjacent subgraphs until the query radius is covered, and the POIs covered in the subgraphs and the expanded area are counted, and the subgraphs that still do not contain all the POIs of the target category after the expansion are removed;
[0012] In the refinement phase, potential edge segments covering all target category interest points are identified from the candidate subgraphs, and the upper bound of the score is estimated first to quickly filter out the edge segments that are unlikely to exceed the current highest score. For the edge segments whose upper bound may exceed the current highest score, their precise scores are further calculated and the optimal result is updated.
[0013] After traversing all candidate subgraphs, the edge segment with the highest score is returned to the user as the global final query result.
[0014] Furthermore, the road network is modeled as a weighted undirected graph, including modeling the road network as a weighted undirected graph , where the vertex Represents an intersection in a road network. express and The weight of the road segment between wi,j∈W represents the road segment The interest points in the road network are mapped to the edges or vertices of the graph G.
[0015] Furthermore, the entire weighted undirected graph is divided into several subgraphs, and a list of interest points and a shortest distance table within the subgraph are constructed for each subgraph, including for the weighted undirected graph , starting from any vertex, traverse the graph using a breadth-first strategy Generate several subgraphs, and ensure that the number of vertices in each subgraph is at most , different subgraphs do not share vertices but share edges. The set of subgraphs after partitioning is expressed as , where n is the number of subgraphs. If the subgraph A vertex in At least one adjacent vertex in the graph belongs to a different subgraph , then the vertex is a boundary vertex.
[0016] Furthermore, the method of counting interest point categories according to user query parameters and marking sub-graphs that fail to completely cover the target category as low potential sub-graphs includes: , check the categories covered by each sub-graph's interest point list, if the sub-graph The category tags cannot be fully covered , it is marked as a low potential subgraph, where the scan subgraph List of points of interest , count the number of categories of interest points of the target category ,Compare and the size of the target category set ,like , then mark is a low potential subgraph, otherwise, it is marked as a candidate subgraph.
[0017] Furthermore, the search is expanded from the low potential subgraph to the adjacent subgraph until the range of the query radius is covered, including abstracting the subgraph marked as low potential in the preliminary screening as a virtual vertex, connecting it to the boundary vertex of the adjacent subgraph through an external edge, and using the Dijkstra algorithm to expand the coverage area to the adjacent subgraph with each boundary vertex as the source point until the expansion range reaches the query radius. , further count the categories and numbers of POIs within its potential coverage. If it still cannot meet the user’s query conditions, the subgraph will be cut off; otherwise, it will be marked as a candidate subgraph for further calculation.
[0018] Further, the step of identifying potential edge segments covering all interest points of the target category from the candidate subgraph includes: Starting from, Dijkstra algorithm is used to calculate the radius Range query to record the coverage of the point of interest , according to the coverage, mark the relevant edges that meet the conditions as After generating the relevant edge segments of all interest points, the candidate edge segments covering all target category interest points are identified through a scanning operation.
[0019] Furthermore, the prioritization of estimating the upper bound of the score to quickly screen out the edge segments that are unlikely to surpass the current highest score includes, after determining the candidate edge segments, calculating the scores thereof to identify the optimal edge segments, wherein, for each candidate edge segment, first calculating the upper bound of its score , avoiding the need to directly calculate the complex exact score, if Less than the maximum score found so far , then prune and discard the edge segment in advance, otherwise continue to calculate its exact score. Finally, calculate the exact score of the candidate edge segment based on the scoring formula , and the current maximum score In contrast, if ,renew And record the edge segment as the current optimal position; otherwise, discard the edge segment.
[0020] In a second aspect, an optimal location query system considering the popularity and arrival distance of a point of interest includes:
[0021] The data acquisition module is configured to acquire road network data and user query parameters; wherein the user query parameters include query radius and point of interest target category;
[0022] The modeling module is configured to model the road network as a weighted undirected graph, wherein vertices represent intersections, edges represent road segments, weights of edges represent lengths of road segments, and points of interest are mapped to positions on vertices or edges of the graph;
[0023] The partitioning module is configured to partition the entire weighted undirected graph into a plurality of subgraphs, and construct a list of interest points and a shortest distance table within the subgraph for each subgraph;
[0024] The filtering module is configured to, in the filtering stage, count the categories of interest points according to the user query parameters, mark the sub-graphs that fail to completely cover the target category as low-potential sub-graphs; expand the search from the low-potential sub-graphs to the adjacent sub-graphs until the range of the query radius is covered, count the interest points covered in the sub-graphs and the extended area, and remove the sub-graphs that still do not contain all the interest points of the target category after the extended range;
[0025] A refinement module is configured to, in the refinement phase, identify potential edge segments covering all target category interest points from the candidate subgraph, prioritize estimating the upper bound of the score to quickly filter out the edge segments that are unlikely to exceed the current highest score, and further calculate the precise score of the edge segments whose upper bound may exceed the current highest score and update the optimal result;
[0026] The feedback module is configured to, after traversing all candidate subgraphs, return the edge segment with the highest score as the global final query result to the user.
[0027] In a third aspect, the present invention provides a computer-readable storage medium storing a plurality of instructions, wherein the instructions are suitable for being loaded by a processor of a terminal device and executing the optimal location query method considering the popularity and arrival distance of a point of interest.
[0028] In a fourth aspect, the present invention provides a terminal device including a processor and a computer-readable storage medium, the processor being used to implement various instructions; the computer-readable storage medium being used to store multiple instructions, the instructions being suitable for being loaded by the processor and executing the optimal location query method that considers the popularity and arrival distance of a point of interest.
[0029] In summary, the present invention has the following beneficial technical effects:
[0030] The present invention proposes an optimal location query method and system that considers the popularity and arrival distance of points of interest in a road network environment. The search range is optimized by subgraph division, and the query efficiency is improved by combining filtering and refinement strategies. For subgraphs that do not contain any group of points of interest that meet the user-specified category within the query range, they are directly excluded from participating in subsequent calculations, thereby reducing invalid exploration. Then, for the screened candidate subgraphs, the algorithm identifies any candidate edge segments within them that include all points of interest of the specified category within the coverage range, and through the upper bound estimation of the score, the edge segments with lower score potential are screened out, and only the candidate edge segments with higher score potential are performed The shortest route calculation and their scores are evaluated. By comparing the calculated score and the current maximum score, the maximum score position and the corresponding score are updated. After all candidate subgraphs are processed, the current maximum score position and its corresponding score are the global optimal results.
[0031] The present invention adopts a pruning strategy based on Euclidean distance, and by calculating the lower bound of the distance of the point of interest access sequence, quickly eliminates the sequence that cannot contain the shortest route, thereby significantly improving the efficiency of the shortest route calculation. Compared with the prior art, the present invention adopts efficient algorithm optimization and pruning strategy, which improves the query efficiency while ensuring the query accuracy, has good scalability, and can adapt to the optimal location query needs of large-scale road networks. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 is a schematic diagram of an undirected graph G constructed in Example 1 of the present invention;
[0033] Figure 2 is a schematic diagram of subgraph division according to Embodiment 1 of the present invention;
[0034] Figure 3 is an example diagram of a sub-graph after division according to Embodiment 1 of the present invention;
[0035] Figure 4The following are comparison diagrams of the effects of different parameters on the query time of the algorithm; (a) is a comparison diagram of the effects of changes in the query radius on the query time in the OL road network; (b) is a comparison diagram of the effects of changes in the query radius on the query time in the CA road network; (c) is a comparison diagram of the effects of changes in the number of interest points on the query time in the OL road network; (d) is a comparison diagram of the effects of changes in the number of interest points on the query time in the CA road network;
[0036] Figure 5 : The following are comparison diagrams of the effects of different parameters on pruning efficiency, where (a) is a comparison diagram of the effects of changes in query radius on pruning efficiency in an OL network; (b) is a comparison diagram of the effects of changes in query radius on pruning efficiency in a CA network; (c) is a comparison diagram of the effects of changes in the number of interest points on pruning efficiency in an OL network; (d) is a comparison diagram of the effects of changes in the number of interest points on pruning efficiency in a CA network;
[0037] Figure 6 It is a flow chart of an optimal location query method considering the popularity and arrival distance of a point of interest of the present invention. DETAILED DESCRIPTION
[0038] The present invention is further described in detail below in conjunction with the accompanying drawings.
[0039] Terminology explanation:
[0040] (1) Undirected graph: The present invention models the road network as a weighted undirected graph , where the vertex Represents an intersection in a road network. Indicates connection and The road section, weight Indicates road segment The length of the vertex The position of Given.
[0041] (2) Points of interest: Points of interest It is a picture Objects located at vertices or edges in the , including the following attributes: 1) Category : The category to which the POI belongs, used to distinguish different types of POIs, such as restaurants, shops, or attractions; 2) Heat value : Indicates the popularity of a point of interest, such as a score based on user reviews or recommendations; 3) Geographic location coordinates : The spatial position of the point of interest, including its longitude and latitude. The present invention maps the point of interest to the vertex or edge of the undirected graph. If the point of interest is located on the edge, it can be located to the adjacent vertex according to its distance offset.
[0042] (3) Subgraph: Subgraph It is a picture A subset of ,and And meet: If ,but In addition, the weight set of the subgraph Corresponding to the set The edge weights in .
[0043] (4) Shortest route: For any two vertices and , whose shortest route Yes Connect and The route with the shortest route distance is recorded as .
[0044] (5) Coverage: given distance threshold ,point Coverage The road network is is the center and the radius is This range includes Departure, distance not exceeding All complete edges and their vertices, including those that partially fall within the radius , even if the entire edge segment is not completely covered.
[0045] (6) Point rating: points Rating Based on the collection of points of interest within its coverage area ,if The categories include , the calculation formula is as follows:
[0046]
[0047] in, For point Coverage The collection of points of interest contained in Indicates point of interest The heat value, It is the balance factor between the total heat value and the route length. express The sum of the heat values of the points of interest within the coverage area, Indicates connection The shortest route length required for the point of interest with the highest heat value in each category, Yes Category The point of interest with the highest medium heat value, is the shortest distance between two points of interest.
[0048] (7) Optimal location query considering the popularity and arrival distance of POI: Given a graph and the set of points of interest mapped on it , a query The goal is to find an optimal set of locations The following conditions must be met: 1) Each point in With this point as the center and the radius Covering the collection within the range All POI categories in; 2) Any point in Rating maximize.
[0049] (8) Segment score: An edge segment is an edge or part of an edge in a road network, where each point covers the same set of points of interest, and the categories of the set include all specified categories. In this case, the score of the edge segment is equal to the score of any point on it, denoted as .
[0050] Example 1
[0051] An optimal location query method considering the popularity and arrival distance of a point of interest in this embodiment includes the following steps:
[0052] Step (1): Model the road network as a weighted undirected graph, where vertices represent intersections, edges represent road segments, edge weights represent the length of the road segments, and points of interest are mapped to vertices or edges of the graph;
[0053] Step (2): Divide the entire weighted undirected graph into several subgraphs, and construct a list of interest points and a shortest distance table within each subgraph for each subgraph;
[0054] Step (3): In the filtering stage, according to the user query parameters (query radius and the set of interest point target categories ), count the interest point categories, and mark the subgraphs that fail to fully cover the target category as lower potential subgraphs. Then, expand the search from the low potential subgraph to the adjacent subgraph until the coverage radius Range, count the sub-graphs and the points of interest covered in the extended area, and remove the sub-graphs that still do not contain all the points of interest of the target category after the extended range;
[0055] Step (4): In the refinement phase, potential edge segments covering all target category interest points are identified from the candidate subgraphs, and the upper bound of the score is estimated first to quickly filter out the edge segments that are unlikely to exceed the current highest score. For the edge segments whose upper bound may exceed the current highest score, their precise scores are further calculated and the optimal result is updated;
[0056] Step (5): After traversing all candidate subgraphs, the edge segment with the highest score is returned to the user as the global final query result.
[0057] As one or more embodiments, the step (1): constructing the road network as a weighted undirected graph In this embodiment, the road network model is represented as a weighted (positive value) undirected graph , where the vertex Represents an intersection in a road network. express and The weight of the road segment between wi,j∈W represents the road segment The length of the road network, the points of interest in the road network are mapped to the edges or vertices of the graph 𝐺. Figure 1 An undirected graph is shown in Example.
[0058] As one or more embodiments, step (2): divide the graph G into multiple subgraphs, and construct a list of interest points and a shortest distance table within the subgraph for each subgraph. The specific division process is as follows:
[0059] For the graph , starting from any vertex, traverse the graph using a breadth-first strategy Generate multiple subgraphs, and ensure that the number of vertices in each subgraph is at most , different subgraphs cannot share vertices but can share edges. The set of subgraphs after partitioning is represented as , where n is the number of subgraphs. A vertex in At least one adjacent vertex in the graph belongs to a different subgraph , then the vertex is a boundary vertex.
[0060] Figure 2 Given Figure 1 An example of sub-graph partitioning. Figure 1 middle As the starting point, perform breadth-first traversal, and let , the figure Divide into four subgraphs and color the boundary vertices in each subgraph. Figure 2 As shown, the graph G is divided into .
[0061] After the sub-graph is divided, in order to ensure that the search can be carried out in the wrong The interest points are calculated one by one. Is it possible to contain locations that meet the query conditions? For each subgraph The following information is stored:
[0062] 1): Subgraph List of points of interest in : Store subgraph Detailed information of all points of interest in the graph, including their category, heat value, and their distance offset to the nearest vertex in the subgraph.
[0063] 2): Distance matrix : Store subgraph The shortest distance between any two vertices in .
[0064] Specifically: Figure 2 middle For example, Table 1 gives the subgraph The shortest distance information between subgraph vertices is recorded.
[0065] Table 1 Subgraphs The shortest distance information between vertices
[0066] <![CDATA[SG4]]> <![CDATA[v 13 ]]> <![CDATA[v 14 ]]> <![CDATA[v 15 ]]> <![CDATA[v 16 ]]> <![CDATA[v 13 ]]> 0 6 4 10 <![CDATA[v 14 ]]> 6 0 7 4 <![CDATA[v 15 ]]> 4 7 0 6 <![CDATA[v 16 ]]> 10 4 6 0
[0067] In addition, the sub-graph The list of points of interest is:
[0068] , , , , , , .
[0069] As one or more embodiments, the step (3): in the filtering stage, according to the user query parameter (query radius and the set of interest point target categories ), count the interest point categories, and mark the subgraphs that fail to fully cover the target category as lower potential subgraphs. Then, expand the search from the low potential subgraph to the adjacent subgraph until the coverage radius The range is calculated, and the points of interest covered in the sub-graph and the extended area are counted, and the sub-graphs that still do not contain all the points of interest of the target category after the extended range are removed. Specifically, step (3) includes:
[0070] Step (301): Preliminary screening of subgraphs. According to the target category set specified by the user , check the categories covered by each sub-graph's interest point list. The category tags cannot be fully covered , it is marked as a low potential subgraph. Specifically including: scanning subgraph List of points of interest , count the number of categories of interest points of the target category .Compare and the size of the target category set .like , then mark is a low potential subgraph, otherwise, it is marked as a candidate subgraph.
[0071] Step (302): Expand the search. For the subgraphs marked as having low potential in the initial screening, abstract them as virtual vertices and connect them to the boundary vertices of the adjacent subgraphs through external edges. Use the Dijkstra algorithm to expand the coverage area (to the adjacent subgraphs) with each boundary vertex as the source point until the expansion range reaches the radius , further counting the types and numbers of points of interest within its potential coverage. If it still cannot meet the user's query conditions, the subgraph is pruned; otherwise, it is marked as a candidate subgraph for further calculation. The specific steps of subgraph pruning include:
[0072] Step (3021): Expand the search range. Each border vertex of Start by searching the radius The search range is gradually expanded to the adjacent subgraph. During the expansion process, the identified points of interest are recorded into a temporary category set according to their categories. middle.
[0073] Step (3022): Target category check. Check the list of points of interest and Whether all target category interest points are included. If the target category is still missing in , the subgraph cannot satisfy the query The coverage requirement is marked as an invalid subgraph and directly excluded. Covering all the points of interest of the target category, the sub-graph Mark as a candidate subgraph for subsequent processing.
[0074] by Figure 2 Take this example to introduce the process of filtering subgraphs to obtain a set of candidate subgraphs. Assume that the POI category selected by the user is and , the query radius is First, check the coverage of each sub-graph's POI category, and then Count the number of categories it covers In the initial screening stage, it is assumed that the subgraph , , and The categories included are: : , : , : , : . According to the category coverage, sub-graph are marked as candidate subgraphs, and the remaining subgraphs , , The subgraphs are marked as low potential. Next, the search is extended for the subgraphs marked as low potential. For each low potential subgraph, the algorithm expands the search range from the boundary vertex until the query radius is reached. Specifically, from the subgraph Starting from the boundary vertex of , and After that, the points of interest identified during the expansion process are recorded in a temporary collection Then, the algorithm checks the expanded set of categories and , if all target categories are covered, then the sub-graph Mark as a candidate subgraph, otherwise, discard .
[0075] As one or more embodiments, step (4): in the refinement stage, potential edge segments covering all target category interest points are identified from the candidate subgraphs, and the upper bound of the score is estimated first to quickly filter out the edge segments that are unlikely to exceed the current highest score. For the edge segments whose upper bound may exceed the current highest score, their precise scores are further calculated and the optimal result is updated. Specifically, step (4) includes:
[0076] Step (401): Generate relevant edge segments of interest points. From each interest point in the candidate subgraph coverage Starting from, Dijkstra algorithm is used to calculate the radius Range query to record the coverage of the point of interest According to the coverage, mark the relevant edges that meet the following conditions as The relevant edge segments are: 1) the edge completely included in the coverage range; 2) the edge segment intersecting with the coverage range. For each relevant edge segment, record its starting point and end point, and use the triple Points of interest As the key and the list of related edge segments as the value, an inverted index table of interest points and related edge segments is constructed.
[0077] Step (402): Identify candidate edge segments. After generating the relevant edge segments of all interest points, a scanning operation is performed to identify candidate edge segments that cover all interest points of the target category. First, a target interest point is randomly selected. , process each unprocessed related edge segment in turn: for each related edge segment , perform a linear scan from the starting point to the end point, query the overlapping part of the edge segment with other interest point related edge segments through the inverted index, and mark all related edge segments forming the overlapping part as processed. Then, for each edge segment formed by the overlapping part of the related edge segments, check whether the interest point category it covers meets the user query condition , if the interest point list corresponding to the edge segment covers all target categories, it is marked as a candidate edge segment and retained, otherwise, the edge segment is discarded. After processing the current interest point Finally, the next interest point is selected from the unprocessed interest points, and the above operation is repeated until the relevant edge segments of all target category interest points are processed.
[0078] Step (403): Screen and calculate the scores of candidate segments. After determining the candidate segments, score them to identify the optimal segment. According to the score definition, the score of a candidate segment consists of two parts: the total heat value of the points of interest in the corresponding point of interest list, and the distance of the shortest route required to cover the target category of points of interest. ,in, It is necessary to include the points of interest with the highest heat value in each target category and connect these points with the shortest path. A direct way to obtain the maximum score position is to calculate the score of each candidate edge segment one by one and compare it to the current highest score For comparison. , the candidate edge segment can be directly discarded. Otherwise, use renew Although it is relatively simple to calculate the total heat value of the points of interest in the candidate edge segment, the calculation of the shortest path required to cover these points of interest is more complicated, because different access orders may generate multiple candidate routes, which will significantly increase the computational overhead. Therefore, for each candidate edge segment, we first calculate the upper bound of its score , avoiding the need to directly calculate the complex exact score. If Less than the maximum score found so far , then prune and discard the edge segment in advance, otherwise continue to calculate its exact score. Finally, calculate the exact score of the candidate edge segment based on the scoring formula , and the current maximum score In contrast, if ,renew And record the edge segment as the current optimal position; otherwise, discard the edge segment. Specifically:
[0079] Step (4031): Calculate the upper bound of the score of the candidate edge segment Without calculating the exact shortest path, the upper bound of the score is calculated by estimating the lower bound of the shortest path of the candidate edge segment. .like , then the candidate edge segment can be discarded in advance to avoid further calculation. The specific process is:
[0080] Step (40311): Generate all possible permutations of the target category interest point visit sequence covered by the edge segment;
[0081] Step (40312): Calculate the “Euclidean distance” of each access sequence, that is, the sum of the straight-line distances between adjacent interest points in the access sequence;
[0082] Step (40313): Take the minimum Euclidean distance of all access sequences as the lower bound of the shortest path distance to estimate .
[0083] Step (40314): If Less than the maximum score found so far , then prune and discard the edge segment in advance, otherwise continue to calculate its exact score.
[0084] Step (4032): Calculate the precise score of the candidate edge segment. Calculate the precise score of the candidate edge segment based on the scoring formula , and the current maximum score In contrast, if ,renew And record the edge segment as the current optimal position; otherwise, discard the edge segment. The specific process is:
[0085] Step (40321): sort the POI visit sequence in ascending order according to the Euclidean distance;
[0086] Step (40322): Calculate the actual shortest route distance for the access sequence in order (using the A* algorithm to calculate the shortest route between adjacent points of interest), and set the shortest route distance of the first sequence as the current minimum distance ;
[0087] Step (40323): If the Euclidean distance of the next access sequence is greater than the current minimum distance , then skip the remaining sequence; otherwise calculate the actual distance of the access sequence and continue to evaluate to update .
[0088] Step (40324): Calculate the exact score of the candidate edge segment based on the scoring formula , and the current maximum score In contrast, if ,renew And record the edge segment as the current optimal position; otherwise, discard the edge segment.
[0089] As one or more embodiments, step (5): after traversing all candidate subgraphs, the edge segment with the highest score is returned to the user as the global final query result. Specifically, after repeating step (4) to search all candidate subgraphs, the current maximum score stored is This is the global maximum score; extract the edge segment corresponding to this score as the query The final result; return the query result and the algorithm ends.
[0090] In Figure 2, assume that only the subgraph If the conditions are met, it becomes a candidate subgraph. The specific situation is as follows Figure 3 Next, for the subgraph , the algorithm generates relevant edge segments based on the coverage of the interest point. For example, the interest point The relevant edge segment of , and . By merging overlapping related edge segments, subgraphs are identified The set of candidate edge segments in . Assume that in the subgraph In , we identify two candidate edge segments: Covering points of interest , and , edge segment Covering points of interest , and .
[0091] Then, the algorithm calculates the score for the candidate edge segment. , rated ; For candidate edge segments , rated In the process of calculating the score, the upper bound of the score can be calculated by estimating the lower bound of the shortest path of the candidate edge segment. Covered points of interest collection , , For example, first, fix the first element to , and then find the remaining two elements and For the remaining two points of interest, there are the following permutations: fix the second element to , then the access sequence is: , , ; Fix the second element to , then the access sequence is: , , Repeat the above steps to obtain the remaining four interest point sequences: , , , , For each access sequence, the Euclidean distance between the interest points is calculated. For example, for the sequence , we calculate arrive The straight-line distance and arrive The straight-line distance of After calculating the Euclidean distance of each access sequence, take the smallest Euclidean distance among all sequences as the lower bound of the shortest route for estimating the edge segment. The upper bound of the score. At this time, if the upper bound score is less than the maximum score currently found , we can prune in advance and discard the edge segment. After all candidate edge segment scores in step (4) are completed, due to is the only candidate subgraph, and the candidate edge segment has the highest rating, so is selected as the global optimal edge segment. At this time, the edge segment The points of interest covered are , and , whose shortest route is , the path length is 5.8. Finally, the algorithm returns the edge segment and its related POI list and ratings as the entire graph 's query results.
[0092] The optimal location query solution for road networks that considers the popularity and arrival distance of points of interest proposed in the present invention improves query efficiency while ensuring query accuracy, has good scalability, and can adapt to the optimal location query needs of large-scale road networks. The flowchart of the query method (SRMaxRS-Search algorithm) of the present invention is as follows: Figure 6 shown.
[0093] The effectiveness of the present invention is demonstrated by experiments below.
[0094] 1. Experimental settings and datasets,
[0095] The hardware configuration of the experimental environment of the present invention is as follows: CPU: Intel(R) Core(TM) i7-8565U CPU @1.80GHz; memory: 16GB; operating system: 64-bit Windows 10. All algorithms are implemented in C++ language.
[0096] Dataset. This experiment uses two road network datasets, California (CA) and Oldenburg (OL). The CA dataset is real and contains real road network data and POI data; the POI data used in the OL road network is generated using a uniform distribution. We randomly generate 30 types of POIs and evenly distribute them on the road network with weights ranging from 0 to 10. Table 2 summarizes the details of the dataset, and Table 3 gives the parameters and value ranges used in the experiment. Unless otherwise specified, the following evaluations involving these parameters will use the default values.
[0097] Table 2 Statistics of road network dataset
[0098] Dataset Number of vertices Number of edges Average side length Number of POI types OL 6105 7035 73.18 30 CA 21048 21693 0.0162 63
[0099] Table 3 Summary of parameters used in the experiment
[0100]
[0101] Benchmark algorithms. We use SEG and TSRN as benchmark algorithms, and we briefly introduce these two benchmark algorithms below.
[0102] The basic idea of the SEG algorithm is to determine the optimal query location by calculating overlapping segments. For each point of interest on the road network, the edge or part of the edge within its coverage is defined as a segment. The overlapping intervals of different segments correspond to a position weight, which is defined as the sum of the scores of all segments covering the overlapping intervals. The result of the optimal location query in the road network is the segment with the highest position weight in the road network. The process of the SEG algorithm is as follows: it first uses a depth-first search to recursively generate all the segments within the coverage range for each point of interest on the road network, and stores these segments in the seg file. In the process, segments associated with the same point of interest are merged to reduce redundancy. Finally, the algorithm scans the seg file and applies the line scanning algorithm to calculate the optimal result for each edge.
[0103] The TSRN algorithm pre-calculates the distances between vertices and interest points on the road network and indexes the shortest distance from each interest point to the nearest point on each edge, thereby identifying candidate edges more quickly. The line scanning algorithm then calculates the optimal result on the screened candidate edges.
[0104] 1. Query performance evaluation,
[0105] In this example, we first compare with the baseline algorithm to demonstrate the superiority of our method (SRMaxRS-Search) in query performance, and then evaluate the effect of the pruning strategy used in our solution.
[0106] Comparison with Baseline Algorithms. We compare the query performance of SRMaxRS-Search with the baseline algorithms SEG and TSRN. Figure 4 The query performance of each algorithm under different parameter changes is shown. First, we evaluate the query time when the query radius changes. The results are shown in Figure 4 As shown in (a) and (b) in OL and CA road networks, the radius values of the two are adjusted in the experiment because of the different scales of edge weights. In the evaluation, we randomly generated 100 queries, each of which specified 5 types of interest points. The results show that regardless of the radius, Regardless of the value of , SRMaxRS-Search always significantly outperforms the two baseline algorithms. Moreover, although the query time of all algorithms decreases with As the number of nodes increases, the growth rate of SRMaxRS-Search is significantly lower than that of the suboptimal algorithm TSRN.
[0107] Next, we analyzed the impact of the total number of POIs in the road network on the query time. The results are as follows: Figure 4 As shown in (c) and (d) in As ,the query time of all algorithms increases, because more candidate edges and points of interest covered by candidate edges increase the query overhead.,However, since SRMaxRS-Search introduces a series of pruning strategies for subgraphs and candidate edges,,the query time is significantly reduced, making its performance still significantly better than the baseline,algorithm.
[0108] Pruning efficiency evaluation. To verify the effectiveness of the pruning strategy, we compared the proposed method SRMaxRS-Search with its version SRMaxRS-NoPrune that does not use the pruning strategy. In this evaluation, we tested the query radius in the OL and CA datasets respectively. and the number of points of interest in the road network The pruning efficiency when changing, the results are as follows Figure 5 As shown in Figure 2. We define the pruning rate as: the ratio of the number of edges that contain interest points of a specified category but do not require precise score calculation to the total number of edges that contain interest points of a specified category. The line in the figure represents the trend of the pruning rate, while the bar chart represents the number of candidate edge segments processed by SRMaxRS-Search and SRMaxRS-NoPrune. Figure 5 As can be seen from (a) and (b) in As the value increases, the pruning rate gradually increases. This is because as the query range expands, the pruning strategy can filter out more candidate results that do not meet the conditions. At the same time, as the query radius As the size of the edge segments increases, the gap in the number of candidate edge segments processed by SRMaxRS-Search and SRMaxRS-NoPrune also widens.
[0109] Finally, we evaluated the total number of points of interest in the road network In this set of experiments, if Figure 5 As shown in (c) and (d) in Figure 3, the pruning rate shows a similar trend. This is because the increase in the number of interest points distributed on the edge will lead to the emergence of more redundant candidate edge segments, which provides the SRMaxRS-Search algorithm with more space to exclude candidate edge segments that have no potential, thereby improving the pruning rate.
[0110] Example 2
[0111] This embodiment provides an optimal location query system that takes into account the popularity and reach distance of a point of interest.
[0112] A computer-readable storage medium stores a plurality of instructions, wherein the instructions are suitable for being loaded and executed by a processor of a terminal device for performing an optimal location query method that considers the popularity and arrival distance of a point of interest.
[0113] A terminal device includes a processor and a computer-readable storage medium, wherein the processor is used to implement various instructions; the computer-readable storage medium is used to store multiple instructions, wherein the instructions are suitable for being loaded by the processor and executing the optimal location query method considering the popularity and arrival distance of a point of interest.
[0114] The above are all preferred embodiments of the present invention, and are not intended to limit the protection scope of the present invention. Therefore, any equivalent changes made based on the structure, shape, and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. An optimal location query method considering the popularity and reach distance of a point of interest, characterized in that: include: Obtaining road network data and user query parameters; wherein the user query parameters include query radius and point of interest target category; The road network is modeled as a weighted undirected graph, where vertices represent intersections, edges represent road segments, edge weights represent the length of the road segments, and points of interest are mapped to vertices or edges of the graph; Divide the entire weighted undirected graph into several subgraphs, and construct a list of points of interest and a shortest distance table within each subgraph for each subgraph; In the filtering stage, the categories of POIs are counted according to the user query parameters, and the subgraphs that fail to fully cover the target category are marked as low-potential subgraphs; the search is expanded from the low-potential subgraphs to the adjacent subgraphs until the query radius is covered, and the POIs covered in the subgraphs and the expanded area are counted, and the subgraphs that still do not contain all the POIs of the target category after the expansion are removed; In the refinement phase, potential edge segments covering all target category interest points are identified from the candidate subgraphs, and the upper bound of the score is estimated first to quickly filter out the edge segments that are unlikely to exceed the current highest score. For the edge segments whose upper bound may exceed the current highest score, their precise scores are further calculated and the optimal result is updated. After traversing all candidate subgraphs, the edge segment with the highest score is returned to the user as the global final query result; The road network is modeled as a weighted undirected graph, including representing the road network model as a weighted undirected graph , where the vertex Represents an intersection in a road network. express and The weight of the road segment between wi,j∈W represents the road segment The length of the road network, the points of interest in the road network are mapped to the edges or vertices of the graph G; The entire weighted undirected graph is divided into several subgraphs, and a list of interest points and a shortest distance table within the subgraph are constructed for each subgraph, including , starting from any vertex, traverse the graph using a breadth-first strategy Generate several subgraphs, and ensure that the number of vertices in each subgraph is at most , different subgraphs do not share vertices but share edges. The set of subgraphs after partitioning is expressed as , where n is the number of subgraphs. If the subgraph A vertex in At least one adjacent vertex in the graph belongs to a different subgraph , then the vertex is a boundary vertex; The method of counting interest point categories according to user query parameters and marking sub-graphs that fail to completely cover the target category as low potential sub-graphs includes: , check the categories covered by each sub-graph's interest point list, if the sub-graph The category tags cannot be fully covered , it is marked as a low potential subgraph, where the scan subgraph List of points of interest , count the number of categories of interest points of the target category ,Compare and the size of the target category set ,like , then mark is a low potential subgraph, otherwise, it is marked as a candidate subgraph; The method of expanding the search from the low potential subgraph to the adjacent subgraph until the range of the query radius is covered includes abstracting the subgraph marked as low potential in the preliminary screening as a virtual vertex, connecting it to the boundary vertex of the adjacent subgraph through an external edge, and using the Dijkstra algorithm to expand the coverage area to the adjacent subgraph with each boundary vertex as the source point until the expansion range reaches the query radius. , further count the categories and numbers of POIs within its potential coverage. If it still cannot meet the user’s query conditions, the subgraph will be cut off; otherwise, it will be marked as a candidate subgraph for further calculation; The method of identifying potential edge segments covering all target category interest points from the candidate subgraph includes: Starting from, Dijkstra algorithm is used to calculate the radius Range query to record the coverage of the point of interest , according to the coverage, mark the relevant edges that meet the conditions as After generating the relevant edge segments of all interest points, the candidate edge segments covering all the interest points of the target category are identified through a scanning operation; The prioritization of estimating the upper bound of the score to quickly screen out the edge segments that are unlikely to surpass the current highest score includes calculating the scores of the candidate edge segments after determining them to identify the optimal edge segments, wherein for each candidate edge segment, first calculating the upper bound of its score , avoiding the need to directly calculate the complex exact score, if Less than the maximum score found so far , then prune and discard the edge segment in advance, otherwise continue to calculate its exact score. Finally, calculate the exact score of the candidate edge segment based on the scoring formula , and the current maximum score In contrast, if ,renew And record the edge segment as the current optimal position; otherwise, discard the edge segment; Generate relevant edges of interest points from each interest point in the candidate subgraph coverage Starting from, Dijkstra algorithm is used to calculate the radius Range query to record the coverage of the point of interest , according to the coverage, mark the relevant edge segments that meet the following conditions as Related edge segments: 1) Edges completely included in the coverage range; 2) Edge segments intersecting with the coverage range; For each related edge segment, record its starting point and end point, and use a triple Points of interest As the key, the list of related edge segments is the value, and an inverted index table of interest points-related edge segments is constructed; Identify candidate edges,After generating the relevant edges of all interest points, identify the candidate edges covering all interest points of the target category through scanning operation. First, randomly select a target category interest point , process each unprocessed related edge segment in turn: for each related edge segment , perform a linear scan from the starting point to the end point, query the overlapping part of the edge segment with other interest point related edge segments through the inverted index, and mark all related edge segments forming the overlapping part as processed. Then, for each edge segment formed by the overlapping part of the related edge segments, check whether the interest point category it covers meets the user query condition If the interest point list corresponding to the edge segment covers all target categories, it is marked as a candidate edge segment and retained. Otherwise, the edge segment is discarded and the current interest point is processed. Finally, the next interest point is selected from the unprocessed interest points, and the above operation is repeated until the relevant edge segments of all target category interest points are processed.
2. An optimal location query system considering the popularity and arrival distance of a point of interest, executing the optimal location query method considering the popularity and arrival distance of a point of interest as claimed in claim 1, comprising: The data acquisition module is configured to acquire road network data and user query parameters; wherein the user query parameters include query radius and point of interest target category; The modeling module is configured to model the road network as a weighted undirected graph, wherein vertices represent intersections, edges represent road segments, weights of edges represent lengths of road segments, and points of interest are mapped to positions on vertices or edges of the graph; The partitioning module is configured to partition the entire weighted undirected graph into a plurality of subgraphs, and construct a list of interest points and a shortest distance table within the subgraph for each subgraph; The filtering module is configured to, in the filtering stage, count the categories of interest points according to the user query parameters, mark the sub-graphs that fail to completely cover the target category as low-potential sub-graphs; expand the search from the low-potential sub-graphs to the adjacent sub-graphs until the range of the query radius is covered, count the interest points covered in the sub-graphs and the extended area, and remove the sub-graphs that still do not contain all the interest points of the target category after the extended range; A refinement module is configured to, in the refinement phase, identify potential edge segments covering all target category interest points from the candidate subgraph, prioritize estimating the upper bound of the score to quickly filter out the edge segments that are unlikely to exceed the current highest score, and further calculate the precise score of the edge segments whose upper bound may exceed the current highest score and update the optimal result; The feedback module is configured to, after traversing all candidate subgraphs, return the edge segment with the highest score as the global final query result to the user.
3. A computer-readable storage medium storing a plurality of instructions, characterized in that: The instructions are suitable for being loaded by a processor of a terminal device and executing the method according to claim 1 .
4. A terminal device, comprising a processor and a computer-readable storage medium, wherein the processor is used to implement each instruction; and the computer-readable storage medium is used to store multiple instructions, characterized in that: The instructions are suitable for being loaded by a processor and executing the method as claimed in claim 1 .
Citation Information
Patent Citations
Tourist route planning method and system supporting multi-keyword search
CN117951396A