Multi-keyword coverage optimal path query method based on deep reinforcement learning

Through a multi-keyword coverage path query method based on deep reinforcement learning, IG-Tree and H2H technologies are used to filter the candidate set of points of interest, and through the encoder-decoder model and reinforcement learning optimization strategy, the problem of difficult to balance path quality and computing time in the existing technology is solved, and efficient and high-quality path generation is achieved.

CN120011659APending Publication Date: 2025-05-16SHENYANG AEROSPACE UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510182050.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art is difficult to find an effective balance between path quality and calculation time under multi-keyword coverage, resulting in insufficient path generation quality or excessive calculation cost.

Method used

Using a deep reinforcement learning method, the road network is divided and candidate set screened through IG-Tree structure and H2H road network index technology, combined with the encoder-decoder deep learning model to explore the optimal combination of interest points, and optimize the model strategy through reinforcement learning to independently find the global optimal path.

Benefits of technology

Efficiently calculate the optimal path covering all keywords and the shortest path length in large-scale road networks. This is suitable for the path optimization requirements of complex road networks, significantly improving the quality of path generation and computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011659A_ABST
    Figure CN120011659A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of spatio-temporal data management, and particularly relates to an optimal path query method under multi-keyword coverage based on deep reinforcement learning, which comprises the following steps of: performing regional division on a large-scale road network by utilizing an IG-Tree structure, and generating a corresponding interest point candidate set in combination with an input keyword demand, thereby reducing a target range; further screening the candidate set by adopting an efficient road network index technology H2H; inputting the screened interest point candidate set into a deep learning model based on an encoder-decoder for exploring an optimal interest point combination covered by keywords; and finally, training and optimizing the model through reinforcement learning, so that the model can autonomously adjust a strategy in a road network environment and find a global optimal path.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of spatiotemporal data management, and specifically relates to an optimal path query method under multi-keyword coverage based on deep reinforcement learning, which is used to realize the shortest path query that meets specific keyword requirements in a real road network. Background Art

[0002] With the widespread popularity of GPS-enabled smart devices in daily life, the information about points of interest in the road network has become more and more detailed. This development has enabled navigation systems to go beyond calculating the shortest or fastest path for users and to further meet the diverse and personalized needs of users in travel, such as route recommendations based on different preferences, interests, or goals. Therefore, the emerging keyword-aware path planning has received widespread attention in recent years because of its excellent performance in finding the path with the minimum cost. This approach can not only optimize traditional path cost parameters such as distance and travel time, but also flexibly meet users' multiple keyword needs and cover specific points of interest or service requirements during travel. These keywords are usually related to specific points of interest or service requirements that users want to visit during their travel, allowing navigation systems to provide solutions that are more in line with users' actual needs. For example, users may want to visit several specified types of points of interest, such as restaurants, gas stations, or tourist attractions during their travel, and keyword-aware routing can plan a cost-optimal path while meeting these needs.

[0003] Existing solutions can be divided into two categories: the first category is based on path extension algorithms, which find solutions by gradually expanding existing paths. Its advantage is that since the paths are formed in order of increasing length, it can usually achieve good results in terms of solution quality. However, since the real road network may contain thousands of nodes and a large number of keywords, its computational cost is also significantly increased, which seriously limits its practical application. The second category is based on the candidate set of interest points. It uses different spatial indexing techniques to screen a set of candidate points of interest that may become the final solution from the point set V, and then generates a path based on these candidate points. This method significantly reduces the computational time by avoiding path extension, but path generation is essentially an NP-hard problem, so the quality of its solution is often lacking.

[0004] Therefore, finding an effective balance strategy between path quality and computation time remains a challenging task. The above problems can be solved if the candidate set of interest points is further filtered and a deep reinforcement learning model is used to explore the combinatorial space. Summary of the invention

[0005] In order to achieve a balance between path quality and computation time in the optimal path query problem with multi-keyword coverage, the present invention proposes an optimal path query method under multi-keyword coverage based on deep reinforcement learning, comprising the following steps:

[0006] Step 1: Use the IG-Tree structure to divide the large-scale road network into regions, and generate the corresponding POI candidate set based on the input keyword requirements; use the efficient road network indexing technology H2H to calculate the distance relationship between each POI and the starting and ending points and other POIs to further screen the candidate set, and obtain the POI candidate set including keyword information and mutual distance information;

[0007] Step 2: Establish a deep learning model based on the encoder-decoder architecture, input the screened POI candidate set into the model, and obtain the optimal POI combination; train and optimize the model through reinforcement learning so that it can autonomously adjust its strategy in the road network environment and find the global optimal path.

[0008] Furthermore, in the encoder, the attention mechanism is used to transmit the keyword information of each point of interest, and the multi-layer perceptron is used to process the distance information between the points of interest; in the decoder, the attention mechanism is used to obtain the points of interest that should be visited at each time step; at the same time, GPU-based parallel computing is used to simultaneously generate multiple paths to improve the quality of the results, and finally the distance mean of multiple paths is used as the baseline to participate in the policy gradient calculation of reinforcement learning to obtain the global optimal path.

[0009] Furthermore, in step 1, the filtering and screening candidate set is obtained by combining the IG-Tree structure with the H2H algorithm, and the specific steps include the following:

[0010] Step 1.1: IG-Tree organizes the partitions of the graph in a hierarchical manner to achieve efficient k-nearest neighbor search. The original graph is divided into a set of edge-disjoint subgraphs through an arbitrary edge-cutting algorithm; these partitions are merged into larger partitions in a hierarchical manner, where the original graph is the largest partition; the merging relationship generates a tree to organize these partitions, and the original graph is the root node of the tree; the leaf nodes of IG-Tree are points of interest within each partition for storing keyword information and are stored as bitmaps; for each keyword, a binary 0 or 1 is used to indicate whether the current vertex contains the keyword;

[0011] In order to obtain the candidate set of interest points, we first need to find a subgraph that meets the requirements of including the start point, the end point and all the keywords. Therefore, we first determine the nearest common ancestor node of the start point and the end point. Then we use its corresponding bitmap to test whether it meets the query keyword. If the interest points in the nearest common ancestor partition cannot cover the keyword set, we recursively search the parent node of the nearest common ancestor until we find a partition node that can fully meet the keyword requirements. Then, we add all the interest points that meet the keyword requirements in the partition to the candidate set.

[0012] Step 1.2: Use the shortest path index H2H algorithm to calculate the distance between two points; specifically, for all interest points of each keyword in the candidate set, first calculate the sum of the distances between each interest point and the starting point and the end point; then calculate the maximum lower bound of the distance to other interest points; then, sort the candidate point set under the same keyword in ascending order of the sum of distances, and select the first k points with the smallest sum of distances to obtain the candidate set of interest points after filtering the initial candidate set.

[0013] Furthermore, in step 2, the optimal POI combination method based on the encoder-decoder deep learning model is as follows:

[0014] Step 2.1: The encoder input consists of row node information R (represented by vector emb r Represented by), column node information C (represented by vector emb c The distance matrix is ​​derived from the filtered interest point candidate set, and the row node r i To column node c j The distance between them is represented by w(i, j). The row node vector is used to represent the keyword information of each point of interest. Its design is based on the binary representation of keywords, using 0 or 1 to indicate whether a point of interest contains a specific keyword, thereby giving the model the ability to capture the semantic features of the point of interest. The column node vector is initialized by hot vector encoding to identify the category information of the point of interest. The points of interest of the same category share the same hot vector encoding, which enables the model to effectively learn the classification features of the point of interest. Finally, the multi-layer attention mechanism F performs information transfer, which is expressed by the following formula

[0015]

[0016] Step 2.2: Reconstruct the attention mechanism to fuse the matrix information, highlight the most relevant information through dynamic weighting, and enhance the feature representation capability; specifically, the model uses the distance matrix as external information, the keywords and classification features of the points of interest as internal information, and uses a multi-layer perceptron to achieve score mixing of internal and external information, dynamically adjust the weights of different information sources, and finally extract the optimal feature representation;

[0017] Step 2.3: After the encoder completes the attention encoding of the interest point set, the decoder will iterate multiple times to determine the interest point to be visited at each time step through step-by-step decision making; specifically, the decoder first uses the row node vector as the query of the attention mechanism and the column node vector as the key and value to achieve dynamic modeling of the relationship between interest points;

[0018] Step 2.4: During the model training process, multiple solutions are generated by using the parallel computing of the GPU, and the feedback information of these solutions is used for parameter adjustment. Specifically, during the parameter training process, each point of interest is used as a starting point to explore the path separately, and multiple complete path solutions are generated in parallel. Subsequently, the mean is calculated based on the distance of all generated paths, and the mean is used as the benchmark value of the REINFORCE algorithm in reinforcement learning to participate in the policy gradient update.

[0019] Step 2.5: By simply calculating the distance from each interest point to the specified starting point, the most suitable k interest points are selected as starting point candidates; then, the encoding is performed n times in the encoding stage, each time using a different hot vector encoding as input to guide the model to analyze the current problem from different angles; through the above method, the model will generate n×k solutions, from which the globally optimal one is selected as the final output.

[0020] Furthermore, in step 2.3, the model concatenates the first visited POI, the last visited POI, and the row node vectors corresponding to the starting point and the end point to construct a vector representation of the current state and uses it as a query token; the query token and the vector of each POI in the column node are used to calculate the attention score, and the obtained score reflects the probability of each POI as the next visited node at the current time step;

[0021] The model further restricts the decoding process by directly setting the probabilities of the starting point, the end point, and other points of interest under the keywords of all visited points of interest to negative infinity, thereby excluding these candidate points from the probability distribution; through multiple iterations, the decoder gradually selects points of interest until a complete path starting from the starting point and eventually reaching the end point is constructed; and in the process of generating this path, the model ensures that one point of interest is selected under each keyword through constraints, meeting the keyword coverage requirements.

[0022] The beneficial effects of the present invention are as follows:

[0023] The present invention proposes an optimal path query method under multi-keyword coverage based on deep reinforcement learning, which can efficiently calculate the optimal path covering all keywords and with the shortest path length in a large-scale road network, and is suitable for the path optimization needs of complex road networks. The method includes: using the IG-Tree structure to divide the road network into regions, generating a candidate set of interest points in combination with keyword requirements to narrow the search range; further screening the candidate set through the H2H road network indexing technology; inputting the screened candidate set into a deep learning model based on an encoder-decoder to explore the optimal combination of keyword coverage; and optimizing the model strategy through reinforcement learning to autonomously find the global optimal path.

[0024] The encoder proposed in the present invention is specially designed to process the road network topological relationship expressed in the form of distance, which makes up for the deficiency of the original encoder in the fusion of key features. In the traditional encoder design, only the attribute information of the point of interest itself is paid attention to, while the distance relationship between the points of interest is ignored. However, in the road network space, the distance between the points of interest is a crucial decision-making factor. Therefore, the present invention innovatively introduces a multi-layer perceptron to deeply fuse the keyword information of the point of interest with the distance information. The weights are automatically learned and assigned through the model, so as to more comprehensively represent the structural characteristics in the road network.

[0025] In addition, in order to accelerate the training process of the model and make full use of the advantages of GPU in parallel computing, the present invention designs a multi-path generation strategy. In each training step, the model generates multiple paths at the same time, and these paths are used for gradient update to reduce the variance of gradient estimation. This parallel path generation and update mechanism can not only significantly reduce the fluctuations in the training process, but also accelerate the convergence speed of the model, thereby effectively improving the training efficiency and performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to facilitate understanding by those skilled in the art, the present invention is further described below with reference to the accompanying drawings.

[0027] Figure 1 A framework diagram of the multi-keyword coverage optimal path query method based on deep reinforcement learning provided by the present invention;

[0028] Figure 2 It is a road network partition structure diagram in the embodiment;

[0029] Figure 3 is a filtering strategy diagram of a candidate set of interest points in an embodiment;

[0030] Figure 4 It is a diagram of the encoder model in the embodiment;

[0031] Figure 5 A diagram of a module responsible for processing matrix information in an encoder in an embodiment;

[0032] Figure 6 It is a decoder diagram in the embodiment;

[0033] Figure 7 This is a graph of parallel path exploration using GPU in an embodiment; DETAILED DESCRIPTION

[0034] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0035] The present invention proposes an optimal path query method under multi-keyword coverage based on deep reinforcement learning, which can efficiently calculate the optimal path covering all keywords and with the shortest path length in a large-scale road network, and is suitable for the path optimization needs of complex road networks. The method includes: using the IG-Tree structure to divide the road network into regions, and generating a candidate set of interest points in combination with keyword requirements to narrow the search scope; further screening the candidate set through H2H road network indexing technology; inputting the screened candidate set into a deep learning model based on an encoder-decoder to explore the optimal combination of keyword coverage; and optimizing the model strategy through reinforcement learning to autonomously find the global optimal path. The method includes the following steps:

[0036] Step 1: Use the IG-Tree structure to divide the large-scale road network into regions to reduce the search space, and generate the corresponding POI candidate set based on the input keyword requirements and the starting and ending points, thereby narrowing the target range; secondly, use the efficient road network indexing technology H2H to calculate the distance relationship between each POI and the starting and ending points and other POIs to further screen the candidate set, and then input the keyword information and mutual distance information of the screened POI candidate set into the deep learning model.

[0037] In step 1, the method of combining the IG-Tree structure and the H2H algorithm to obtain the filtering candidate set is as follows:

[0038] Step 1.1: Use the IG-Tree algorithm to divide the road network to obtain a candidate set of interest points;

[0039] IG-Tree organizes the partitions of the graph in a hierarchical manner to achieve efficient k-nearest neighbor search. The original graph is divided into a set of edge-disjoint subgraphs through an arbitrary edge-cutting algorithm. These partitions are then merged into larger partitions in a hierarchical manner, where the original graph is the largest partition. The merge relationship generates a tree to organize these partitions, with the original graph as the root node of the tree. In order to store keyword information, the leaf nodes of IG-Tree are points of interest within each partition, and the keyword information is stored as a bitmap, where for each keyword, a binary 0 or 1 is used to indicate whether the current vertex contains the keyword.

[0040] In order to obtain the candidate set of interest points, we first need to find a subgraph that meets the requirements of including the start point, the end point and all the keywords. Therefore, we first determine the nearest common ancestor node of the start point and the end point. Then use its corresponding bitmap to test whether it meets the query keyword. If the interest points in the nearest common ancestor partition cannot cover the keyword set, recursively search the parent node of the nearest common ancestor until a partition node that can fully meet the keyword requirements is found. Subsequently, all interest points that meet the keyword requirements in the partition are added to the candidate set.

[0041] Step 1.2: The initial candidate set is too large and needs to be filtered. The shortest path index H2H algorithm is used to calculate the distance between two points. Specifically, for all points of interest for each keyword in the candidate set, the sum of the distances between each point of interest and the starting point and the end point is first calculated. Then the maximum lower bound of the distance to other points of interest is calculated. Subsequently, the set of candidate points under the same keyword is sorted in ascending order of the sum of distances, and the first k points with the smallest sum of distances are selected to form the filtered candidate set of points of interest.

[0042] Step 2: For the input provided in step 1, design a deep learning model based on the encoder-decoder architecture to explore the optimal combination of interest points. In the encoder, the attention mechanism is used to transmit the keyword information of each interest point, and the multi-layer perceptron is used to process the distance information between interest points. In the decoder, the attention mechanism is used to obtain the interest points that should be visited at each time step. At the same time, multiple paths are generated simultaneously based on the parallel computing of the GPU to improve the quality of the results. Finally, the distance mean of multiple paths is used as the baseline to participate in the policy gradient calculation of reinforcement learning.

[0043] Step 2: Based on the encoder-decoder deep learning model, the optimal POI combination method is explored as follows:

[0044] Step 2.1: The encoder input consists of row node information R (represented by vector emb r Represented by), column node information C (represented by vector emb cThe distance matrix is ​​derived from the filtered interest point candidate set, and the row node r i To column node c j The distance between them is represented by w(i, j). The row node vector is used to represent the keyword information of each point of interest. Its design is based on the binary representation of keywords, using 0 or 1 to indicate whether a point of interest contains a specific keyword, thereby giving the model the ability to capture the semantic features of the point of interest. The column node vector is initialized by hot vector encoding to identify the category information of the point of interest. The points of interest of the same category share the same hot vector encoding, which enables the model to effectively learn the classification features of the point of interest. Finally, the multi-layer attention mechanism F performs information transfer, which is expressed by the following formula

[0045]

[0046] Step 2.2: In order to fully integrate the matrix information, the attention mechanism is reconstructed to highlight the most relevant information through dynamic weighting, thereby enhancing the feature representation capability. Specifically, the model uses the distance matrix as external information, the keywords and classification features of the points of interest as internal information, and uses a multi-layer perceptron to achieve score mixing of internal and external information, dynamically adjust the weights of different information sources, and finally extract the optimal feature representation.

[0047] Step 2.3: After the encoder completes the attention encoding of the interest point set, the decoder will iterate multiple times to determine the interest point to be visited at each time step through step-by-step decision making. Specifically, the decoder first uses the row node vector as the query of the attention mechanism and the column node vector as the key and value to achieve dynamic modeling of the relationship between interest points.

[0048] Step 2.3.1: In order to track and express the current state, the model concatenates the first visited point of interest, the last visited point of interest, and the row node vectors corresponding to the starting point and the end point, constructs a vector representation of the current state, and uses it as a query token. The query token is used to calculate the attention score with the vector of each point of interest in the column node, and the obtained score reflects the probability of each point of interest being the next visited node at the current time step.

[0049] Step 2.3.2: To avoid revisiting selected POIs or violating constraints, the model further restricts the decoding process. Specifically, the probabilities of the starting point, the end point, and other POIs under the keywords of all visited POIs are directly set to negative infinity, thereby excluding these candidate points from the probability distribution. Through multiple iterations, the decoder gradually selects POIs until a complete path starting from the starting point and finally reaching the end point is constructed. More importantly, during the path generation process, the model ensures that a POI is selected under each keyword through constraints, meeting the keyword coverage requirements. The final output path not only maximizes the use of semantic and distance information between POIs, but also generates an optimized combination scheme in the semantic level and spatial dimension.

[0050] Step 2.4: GPUs have demonstrated excellent performance in the field of parallel computing. In the training process of deep learning models, compared with traditional methods, only outputting a single solution and using it to adjust parameters can easily lead to large variance in policy gradients, which in turn affects the stability and convergence speed of the model. By making full use of the parallel computing capabilities of the GPU, multiple solutions can be generated at the same time, and the feedback information of these solutions can be used for parameter adjustment, thereby improving the training efficiency and result quality of the model. Specifically, during the parameter training process, each point of interest is used as a starting point to explore the path separately, and multiple complete path solutions are generated in parallel. Subsequently, the mean is calculated based on the distance of all generated paths, and the mean is used as the benchmark value of the REINFORCE algorithm in reinforcement learning to participate in the policy gradient update.

[0051] Step 2.5: In the test phase, first, the distance from each point of interest to the specified starting point is simply calculated to select the most suitable k points of interest as starting point candidates. Subsequently, n encodings are performed in the encoding phase, each time using a different hot vector encoding as input to guide the model to analyze the current problem from different perspectives. This multiple encoding strategy helps the model to explore potential high-quality solutions in different feature spaces, thereby expanding the exploration scope of the combination space. Through the above method, the model will generate n×k solutions, from which the globally optimal one is selected as the final output.

[0052] Example:

[0053] like Figure 1 As shown in FIG. 1 , the multi-keyword coverage optimal path query method based on deep reinforcement learning includes the following steps:

[0054] First, the road network is partitioned based on the IG-Tree structure, and the results of the road network partitioning are as follows: Figure 2(a) shows that each partition has several road network nodes, and takes the query Q(v5, v9, {k1, k2, k3}) as an example, where v5 and v9 are the starting and ending points respectively, and k1-k3 are the keyword sets that need to be satisfied. Figure 2 (b) shows how IG-Tree stores keyword information, where the leaf nodes are the points of interest in each partition, and the keyword information is stored as a bitmap. Figure 2 In (b), there are three keywords {k1, k2, k3}, so each vertex in the leaf layer has a bitmap consisting of three bits. Since the keyword set of vertex v4 is {k1}, the binary linked list representation of the leaf node is 001. For the example query Q(v5, v9, {k1, k2, k3}), the nearest common ancestor (LCA) of the starting point v5 and the end point v9 is G1. By checking the binary keyword information of partition G1, it can be found that it is represented as 111, which indicates that the current partition can satisfy all keywords in the query keyword set. Therefore, all points of interest in partition G1 that satisfy the query keywords constitute the candidate set.

[0055] Next, we need to filter the points of interest under each keyword in the candidate set to a smaller scale, leaving a high-quality set of points of interest. Figure 3 The example of filtering the set of interest points under keyword k1 to 3 is shown. The distances from each interest point to the starting and ending points and the maximum lower bounds of other interest points are counted, and then sorted from small to large. Finally, the first three interest points with the shortest sum of distances are selected.

[0056] The keywords and mutual distance information of the filtered candidate set are used as the input of the model. Here, an encoder-decoder model based on the attention mechanism is used. Figure 4 As shown in Figure 1, the input of the encoder consists of a row node vector, a column node vector, and a corresponding distance matrix. The row node vector is used to represent the keyword information of each interest point, while the column node vector is initialized by hot vector encoding to identify the category information of the interest point. Interest points of the same category share the same hot vector encoding.

[0057] In order to fully integrate these three types of input information, such as Figure 5 As shown in the figure, the model uses the distance matrix as external information, the keywords and classification features of the points of interest as internal information, and uses a multi-layer perceptron to achieve score mixing of internal and external information, dynamically adjust the weights of different information sources, and finally extract the optimal feature representation.

[0058] During the decoding process, Figure 6As shown in the figure, in order to express the current state, the decoder concatenates the first visited point of interest, the last visited point of interest, and the row node vectors corresponding to the starting point and the end point to construct a vector representation of the current state and use it as a query token. The query token and the vector of each point of interest in the column node are used to calculate the attention score, and the obtained score reflects the probability of each point of interest being the next visited node at the current time step. To avoid repeated visits to selected points of interest, the probabilities of the starting point, the end point, and other points of interest under the keywords to which all visited points of interest belong are directly set to negative infinity, thereby excluding these candidate points from the probability distribution.

[0059] During the testing phase, Figure 7 As shown in the figure, first, the distance from each interest point to the specified starting point is simply calculated, and the most suitable k interest points are selected as starting point candidates. Then, in the encoding stage, n encodings are performed, each time using a different hot vector encoding as input. The model will generate n×k solutions, from which the global optimal one is selected as the final output.

[0060] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. The optimal path query method under multi-keyword coverage based on deep reinforcement learning is characterized by: The following steps are included: Step 1: Use the IG-Tree structure to divide the large-scale road network into regions, and generate the corresponding POI candidate set based on the input keyword requirements; use the efficient road network indexing technology H2H to calculate the distance relationship between each POI and the starting and ending points and other POIs to further screen the candidate set, and obtain the POI candidate set including keyword information and mutual distance information; Step 2: Establish a deep learning model based on the encoder-decoder architecture, input the screened POI candidate set into the model, and obtain the optimal POI combination; train and optimize the model through reinforcement learning so that it can autonomously adjust its strategy in the road network environment and find the global optimal path.

2. The optimal path query method under multi-keyword coverage based on deep reinforcement learning as claimed in claim 1, characterized in that: In the encoder, the attention mechanism is used to transmit the keyword information of each point of interest, and the multi-layer perceptron is used to process the distance information between the points of interest; in the decoder, the attention mechanism is used to obtain the points of interest that should be visited at each time step; at the same time, GPU-based parallel computing is used to generate multiple paths at the same time to improve the quality of the results. Finally, the distance mean of multiple paths is used as the baseline to participate in the policy gradient calculation of reinforcement learning to obtain the global optimal path.

3. The optimal path query method under multi-keyword coverage based on deep reinforcement learning according to claim 1 is characterized in that: In step 1, the filtering candidate set is obtained by combining the IG-Tree structure with the H2H algorithm, and the specific steps include the following: Step 1.1: IG-Tree organizes the partitions of the graph in a hierarchical manner to achieve efficient k-nearest neighbor search. The original graph is divided into a set of edge-disjoint subgraphs through an arbitrary edge-cutting algorithm; These partitions are merged into larger partitions in a hierarchical manner, where the original graph is the largest partition; the merged relationship generates a tree to organize these partitions, and the original graph is the root node of the tree; the leaf nodes of the IG-Tree are the points of interest within each partition for storing keyword information and are stored as bitmaps; For each keyword, use a binary 0 or 1 to indicate whether the current vertex contains the keyword; In order to obtain the candidate set of interest points, we first need to find a subgraph that meets the requirements of including the starting point, the end point and all the keywords. Therefore, we first determine the nearest common ancestor node of the starting point and the end point. Then, we use its corresponding bitmap to test whether it meets the query keywords. If the interest points in the LCA partition cannot cover the keyword set, the parent node of the LCA is recursively searched until a partition node that can fully meet the keyword requirements is found; then, all interest points that meet the keyword requirements in the partition are added to the candidate set; Step 1.2: Calculate the distance between two points using the shortest path index H2H algorithm; specifically, for all points of interest for each keyword in the candidate set, first calculate the sum of the distances between each point of interest and the starting point and the end point; then calculate the maximum lower bound of the distances to other points of interest; Subsequently, the candidate point set under the same keyword is sorted in ascending order of the sum of distances, and the first k points with the smallest sum of distances are selected to obtain a candidate set of interest points after filtering the initial candidate set.

4. The optimal path query method under multi-keyword coverage based on deep reinforcement learning according to claim 1 is characterized in that: In step 2, the optimal POI combination method based on the encoder-decoder deep learning model is as follows: Step 2.1: The encoder input consists of row node information R (represented by vector emb r Represented by), column node information C (represented by vector emb c The distance matrix is ​​derived from the filtered interest point candidate set, and the row node r i To column node c j The distance between them is represented by w(i, j). The row node vector is used to represent the keyword information of each point of interest. Its design is based on the binary representation of keywords, using 0 or 1 to indicate whether a point of interest contains a specific keyword, thereby giving the model the ability to capture the semantic features of the point of interest. The column node vector is initialized by hot vector encoding to identify the category information of the point of interest. The points of interest of the same category share the same hot vector encoding, which enables the model to effectively learn the classification features of the point of interest. Finally, the multi-layer attention mechanism F performs information transfer, which is expressed by the following formula Step 2.2: Reconstruct the attention mechanism to fuse the matrix information, highlight the most relevant information through dynamic weighting, and enhance the feature representation capability; specifically, the model uses the distance matrix as external information, the keywords and classification features of the points of interest as internal information, and uses a multi-layer perceptron to achieve score mixing of internal and external information, dynamically adjust the weights of different information sources, and finally extract the optimal feature representation; Step 2.3: After the encoder completes the attention encoding of the interest point set, the decoder will iterate multiple times to determine the interest point to be visited at each time step through step-by-step decision making; specifically, the decoder first uses the row node vector as the query of the attention mechanism and the column node vector as the key and value to achieve dynamic modeling of the relationship between interest points; Step 2.4: During the model training process, multiple solutions are generated by using the parallel computing of the GPU, and the feedback information of these solutions is used for parameter adjustment. Specifically, during the parameter training process, each point of interest is used as a starting point to explore the path separately, and multiple complete path solutions are generated in parallel. Subsequently, the mean is calculated based on the distance of all generated paths, and the mean is used as the benchmark value of the REINFORCE algorithm in reinforcement learning to participate in the policy gradient update. Step 2.5: By simply calculating the distance from each interest point to the specified starting point, the most suitable k interest points are selected as starting point candidates; then, the encoding is performed n times in the encoding stage, each time using a different hot vector encoding as input to guide the model to analyze the current problem from different angles; through the above method, the model will generate n×k solutions, from which the globally optimal one is selected as the final output.

5. The optimal path query method under multi-keyword coverage based on deep reinforcement learning as claimed in claim 4, characterized in that: In step 2.3, the model concatenates the first visited POI, the last visited POI, and the row node vectors corresponding to the starting point and the end point to construct a vector representation of the current state and uses it as a query token. The query token and the vector of each POI in the column node are used to calculate the attention score. The obtained score reflects the probability of each POI as the next visited node at the current time step. The model further restricts the decoding process by setting the probabilities of the starting point, the end point, and other points of interest under the keywords of all visited points of interest to negative infinity, thereby excluding these candidate points from the probability distribution; through multiple iterations, the decoder gradually selects points of interest until a complete path from the starting point to the end point is constructed; In the process of path generation, the model ensures that one point of interest is selected under each keyword through constraints, meeting the keyword coverage requirements.