Urban vehicle personalized path planning method and system based on adversarial inverse reinforcement learning
Through the path planning method of adversarial inverse reinforcement learning, combined with generator and discriminator training, the problems of poor adaptability and weak personalization ability caused by sparse trajectory data in large-scale urban road networks are solved, and more accurate path recommendations are achieved.
Patent Information
- Application Number
- CN202510421906.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-04-07
AI Technical Summary
The existing path planning methods are in large-scale and complex urban road networks, which have poor adaptability and weak personalization capabilities due to sparse trajectory data.
Adversarial inverse reinforcement learning is adopted, and the adversarial training of the generator and discriminator is combined with the attention mechanism to learn the urban road network environment characteristics and path preferences to generate a path that meets the user's personalized needs.
It improves the adaptability and generalization ability of path planning, can accurately capture users' personalized travel preferences, reduce path length, and is suitable for complex small-scale and large-scale urban road networks.
Smart Images

Figure CN120373586A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of path planning, and in particular, to a personalized path planning method and system for urban vehicles based on adversarial inverse reinforcement learning. Background Art
[0002] With the acceleration of urbanization and the continuous growth of the motor vehicle ownership, problems such as urban traffic congestion and low travel efficiency have become increasingly prominent. Traditional vehicle path planning methods mainly rely on static road network data and real-time traffic information (such as congestion index), and use classical algorithms to calculate the shortest or fastest path. However, these methods have the following limitations:
[0003] Chinese invention patent "A personalized path planning method and system based on federated learning and edge computing" (Patent No.: CN119417001A) records quantifying the travel experience of users in the visited area into the trust score of the path (path length, travel time, path safety, and path congestion), and optimizing the path trust score according to the path selection frequency in the user's social network to provide personalized path planning for users. By adding edge computing technology to improve the efficiency of the system in transmitting and processing data. Chinese invention patent "A digital tourist travel planning system" (Patent No.: CN119692584A) records making path recommendations based on the user's basic information (personal information, travel preferences, and historical itinerary) and the relevant information of the target scenic spot (scenic spot location, opening hours, ticket price, and transportation mode). The path planning algorithm is constructed with the set path length as the objective function and the commuting time as the constraint condition, and the planned routes are sorted. However, both of the above methods are based on external specific rules, that is, setting a unified weight formula, and do not fully explore the potential path selection preferences of users, resulting in low personalized selection of users.
[0004] Chinese invention patent "A personalized path recommendation algorithm based on a deep learning-guided A* algorithm" (Patent No.: CN115577175A) records preprocessing the data and extracting available data from the user's historical travel information; the CNN+LSTM neural network learns the implicit personalized preference features of the user from the extracted data, uses the CNN+LSTM neural network for path planning to obtain alternative paths; the evaluation criteria of the alternative paths are used as the estimation function of the A* search algorithm, which is used as the basis for selecting the next moment node in the iterative process of the A* search algorithm; the path planned by the A* search algorithm is the final solution result, which is the personalized path recommended for the user. However, the above method only uses simulation data and only conducts simple obstacle avoidance experiments on the grid, without using real road data, resulting in poor adaptability in large-scale complex urban road networks and inaccurate path planning. Summary of the Invention
[0005] The technical problem to be solved by the present invention is:
[0006] To solve the problems of poor adaptability and weak personalization ability caused by sparse trajectory data in large-scale complex road networks in cities in existing path planning methods.
[0007] The technical solution adopted by the present invention to solve the above technical problems:
[0008] The present invention provides a personalized path planning method for urban vehicles based on adversarial inverse reinforcement learning, including the following steps:
[0009] S100. Data processing, matching the processed historical trajectory data to the urban road network map; including optimizing the road network structure of the urban road network data, processing the historical trajectory data, and then mapping the processed trajectory data to the road network using a map matching algorithm;
[0010] S200. Model training, including based on adversarial inverse reinforcement learning, introducing an attention mechanism, obtaining a reward function through adversarial training of a generator and a discriminator with the goal of maximizing trajectory similarity and minimizing path cost, learning the potential path preferences, and capturing the multi-objective decision-making strategy of experts through adversarial training and inverse reasoning to generate paths;
[0011] S300. Path recommendation, making corresponding path recommendations for large-scale road networks and small-scale road networks.
[0012] Further, in step S100, the method for processing GPS data includes coordinate system conversion, data screening, data cleaning, and trajectory segmentation.
[0013] Further, in step S200, it includes
[0014] S210. Environment modeling, modeling the road network as a directed graph, represented by G=(V, E), where V is the set of nodes, representing key positions, and E is the set of edges, representing directed road segments, that is, road segments along which vehicles can travel in a specific direction;
[0015] S220. Generator design, designing a generator based on a convolutional neural network, and generating feasible actions by processing input states, destinations, path features, and environmental features;
[0016] S230. Discriminator design, designing a discriminator based on a convolutional neural network, and the discriminator extracts neighborhood information in the input features through the convolutional neural network to distinguish the trajectories generated by the generator from the real expert trajectories;
[0017] S240. Learning objective, a reward function that comprehensively considers the latent path preferences and trip lengths in historical trajectory data, including environmental feature extraction, path cost calculation, and design optimization objectives;
[0018] S250. Adversarial training, by alternately optimizing the loss functions of the generator and discriminator, gradually improving the quality of the generated data and the discrimination ability of the discriminator, making the distribution of the generated data as consistent as possible with the real data.
[0019] Further, in step S220, the convolutional layer processes the input path features. The first convolutional layer uses 20 3*3 convolutional kernels, and the second convolutional layer uses 30 2*2 convolutional kernels to extract local features; the attention mechanism further processes the features output by the convolutional layer to capture the global dependencies between path features; three fully connected layers are used to gradually extract features to generate the action probability distribution; the convolutional layer uses the LeakyReLU activation function, and the pooling layer uses the max pooling operation; the output of the generator is the probability distribution of each action, obtained through the Softmax function, and is used to select actions or calculate the logarithmic probability of actions.
[0020] Further, in step S230, first, a convolutional layer is used to extract features, and then the max pooling layer is used to reduce the size of the feature map and increase the degree of feature abstraction; the second convolutional layer further extracts high-level features. After adding a multi-head self-attention (MHSA) module after the convolutional layer of the discriminator, the feature vector obtained after the features output by the convolutional layer are processed by the MHSA module is mapped through a fully connected layer, and the Sigmoid activation function is used to calculate the score of the trajectory to distinguish the difference between the recommended path and the real path and calculate the path similarity.
[0021] Further, in step S240, specifically,
[0022] S241. Environmental feature extraction, extracting environmental features related to the trip from historical trajectory data, including road grade, number of turns, and distance to the destination;
[0023] S242. Path cost calculation, defining the path cost as the trip distance, extracting road length information from the road network, and calculating the cost of each path;
[0024] S243. Optimization objective, fusing the environmental features and the path cost, and the form of the optimization objective is as follows:
[0025] R(s,a) = R expert (s,a) + αR cost (s,a)
[0026] where, R expert (s,a) represents the reward of learning based on environmental features, Rcost (s, a) represents the path cost reward, and α represents a hyperparameter that adjusts the weight between the two rewards.
[0027] Furthermore, in step S250, an alternating training method is adopted. First, the discriminator D is fixed, and the loss function of the generator G is optimized. The optimization objective of the generator is to maximize the negative logarithm of the probability that the generated data is discriminated as real data in the discriminator. Then, the generator G is fixed, and the loss function of the discriminator D is optimized. The optimization objective of the discriminator is to maximize the ability to discriminate real data as real data. The optimization objectives of the generator and the discriminator are as follows:
[0028]
[0029] Among them, z represents the input of the generator, G(z) represents the data generated by the generator, D(·) represents the discriminator function, which can represent the probability that the input data is real data; x represents the real-world data collected, p data (x) represents the distribution of real data, and p(z) represents the distribution of data generated by the generator.
[0030] Furthermore, in step S300, it specifically includes
[0031] S310. For a small-scale road network, since the trajectory distribution is relatively dense, directly use the action selection probability output by the model for path finding;
[0032] S320. For a large-scale road network, since the trajectory distribution is relatively sparse, first use the model to obtain the selection probability of each action in the current state, and use the reciprocal of the action selection probability as the cost of path finding by the Dijkstra algorithm. The optimized model can select the overall optimal path based on the global situation.
[0033] A personalized path planning system for urban vehicles based on adversarial inverse reinforcement learning. The system has program modules corresponding to the above steps and executes the steps in the above-mentioned personalized path planning method for urban vehicles based on adversarial inverse reinforcement learning when running.
[0034] A computer-readable storage medium stores a computer program, and the computer program is configured to implement the steps of a personalized path planning method for urban vehicles based on adversarial inverse reinforcement learning when called by a processor.
[0035] Compared with the prior art, the beneficial effects of the present invention are:
[0036] The present invention relates to a method and system for personalized path planning of urban vehicles based on adversarial inverse reinforcement learning. Compared with traditional path recommendation algorithms, such as discrete choice models or experience-based shortest path algorithms, it exhibits stronger adaptability and generalization ability. Through adversarial training, the present invention can not only learn the optimal path selection in expert examples (historical trajectories), but also accurately capture the personalized travel preferences of users, overcoming the deficiency of traditional algorithms that cannot flexibly adapt to individual needs.
[0037] In traditional path planning algorithms, shortest path algorithms (such as A* algorithm, Dijkstra algorithm) focus on the optimization of path length and ignore the personalized needs of users, such as the driving habits of drivers and the types of paths preferred for travel (such as avoiding congested areas or choosing expressways); path recommendation methods based on machine learning are more widely used in multi-objective path optimization problems, but the goals of path planning are also artificially designed. However, different travelers have different cognitions of the road network and there are cognitive biases in understanding the complex factors affecting path selection. For the same starting point and destination, there are also individual preferences for travel path selection. The recommendation results of traditional path selection models are difficult to fully meet the actual needs of users.
[0038] Existing path planning models based on adversarial inverse reinforcement learning can learn potential path selection preferences from expert examples (historical trajectories), but do not consider the impact of the actual travel environment, have a strong dependence on data, and are prone to falling into the situation of local optimization during the pathfinding process, resulting in poor path recommendation results in large-scale or sparse-trajectory road networks.
[0039] The present invention improves the generator and discriminator of the generative adversarial network. When the generator learns historical trajectories, it considers the impact of the environment on path selection. The discriminator considers the impact of both path preference and path length on path selection, that is, when imitating the strategy of experts, an additional path length reward is introduced. By introducing the path length reward, the discriminator can guide the trajectory to approach the target location, rather than simply imitating expert data, reducing the situation of local optimization and improving the overall path quality. Secondly, an attention layer is added after the convolutional neural network of the generator and discriminator, strengthening the information flow in the path selection process and being able to better focus on the dynamically changing traffic environment. In addition, the present invention optimizes the path generation strategy, using the output of the path planning model based on the adversarial inverse reinforcement learning algorithm as the weight for pathfinding by the Dijkstra algorithm, reducing the dependence of the model on trajectory data, being able to retain the trajectory selection preferences of the model while effectively controlling the path length, making the trajectory structure closer to real data, and realizing the integration of model preferences and global optimization.
[0040] The experimental results show that the present invention can comprehensively consider the performance of the model in reducing the path length and approaching the user's travel preferences. With fewer training times, it can achieve better path recommendation effects, that is, it has achieved good effects in fitting the user's travel preferences and reducing the path length, improving the personalized path recommendation ability and the adaptability of the path recommendation system, making the recommended path more in line with the actual needs of drivers; and the present invention is applicable to complex small-scale cities and large-scale cities. Description of the Drawings
[0041] Figure 1 It is a flowchart of a method for personalized path planning of urban vehicles based on adversarial inverse reinforcement learning in an embodiment of the present invention;
[0042] Figure 2 It is a flowchart of path recommendation of a method for personalized path planning of urban vehicles based on adversarial inverse reinforcement learning in an embodiment of the present invention;
[0043] Figure 3 It is a structural diagram of a generator in an embodiment of the present invention;
[0044] Figure 4 It is a structural diagram of a discriminator in an embodiment of the present invention;
[0045] Figure 5 It is a road network within the scope of the Shanghai dataset in an embodiment of the present invention;
[0046] Figure 6 It is a road network within the scope of the Wuhan dataset in an embodiment of the present invention;
[0047] Figure 7 It is a matching diagram of real historical trajectory GPS on the road network in Wuhan in an embodiment of the present invention;
[0048] Figure 8 It is a result diagram of path recommendation in Wuhan in an embodiment of the present invention. Detailed Embodiments
[0049] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following detailed description of the specific embodiments of the present invention is provided in conjunction with the accompanying drawings.
[0050] Specific Embodiment 1: In combination with Figures 1 to 4 As shown, the present invention provides a method for personalized path planning of urban vehicles based on adversarial inverse reinforcement learning, including the following steps:
[0051] S100. Data processing: Optimize the road network structure of urban road network data. For example, perform coordinate system conversion, data screening, data cleaning, and trajectory segmentation on taxi GPS data, and then use the map matching algorithm to accurately map the processed trajectory data to the road network using the fastmap matching (fmm) algorithm. As Figure 5 shown in the matching result;
[0052] S200. Model training: Include research on path selection methods based on adversarial inverse reinforcement learning and introduce an attention mechanism. Through adversarial training of the generator and discriminator, obtain a robust reward function with the goal of maximizing trajectory similarity and minimizing path cost, learn the potential path preferences, and capture the multi-objective decision-making strategies of experts through adversarial training and inverse reasoning to generate paths that meet various requirements. Specifically,
[0053] S210. Environment modeling: Model the road network as a directed graph, denoted as G=(V, E), where V is the set of nodes representing key locations (such as intersections or road start and end points), and E is the set of edges representing directed road segments, that is, road segments along which vehicles can travel in a specific direction. The connectivity of road segments is determined by their spatial positions, that is, one road segment can be connected to another to form a network. Each road has a feature vector, mainly considering road grade, number of turns, and distance to the destination. These features can be mined from historical GPS trajectory data to provide references or recommendations for similar travel demands;
[0054] S220. Generator design: The present invention designs a generator based on a convolutional neural network (CNN). By processing the input state, destination, path features, and environmental features, generate feasible actions. The convolutional layers (conv1, pool, conv2) first process the input path features. The first convolutional layer uses 20 3*3 convolutional kernels, and the second convolutional layer uses 30 2*2 convolutional kernels to extract local features. The attention mechanism further processes the features output by the convolutional layer to capture the global dependencies between path features. Use three fully connected layers to gradually extract features and generate the action probability distribution. The convolutional layer uses the LeakyReLU activation function, and the pooling layer uses the max pooling operation. The output of the generator is the probability distribution of each action, obtained through the Softmax function. These probability distributions are used to select actions or calculate the logarithmic probability of actions. The generator structure is as Figure 3 shown;
[0055] S230. Discriminator design: The present invention designs a discriminator based on a convolutional neural network (CNN) to distinguish the trajectories generated by the generator from the real expert trajectories. The discriminator extracts the neighborhood information in the input features through the convolutional neural network. First, a convolutional layer is used to extract features, and then a max-pooling layer is used to reduce the size of the feature map and increase the degree of feature abstraction. The second convolutional layer further extracts high-level features. After the convolutional layer of the discriminator, a multi-head self-attention (MHSA) mechanism module is added. After the features output by the convolutional layer are processed by the MHSA module, the obtained feature vectors are mapped through a fully connected layer, and the Sigmoid activation function is used to calculate the score of the trajectory to distinguish the difference between the recommended path and the real path and calculate the path similarity. The structure of the discriminator is as Figure 4 shown;
[0056] S240. Learning objective: The present invention can learn a reward function that comprehensively considers the latent path preferences and trip lengths in historical trajectory data, so that the generated path not only conforms to the driver's path preferences but also can find the best balance among multiple objectives. Specifically,
[0057] S241. Extraction of environmental features: Extract environmental features related to the trip from historical trajectory data, including road grade, number of turns, and distance to the destination. The environmental feature matrix is a three-dimensional matrix that comprehensively describes the complexity and cost of the path and is crucial for path selection.
[0058] S242. Calculation of path cost: Define the path cost as the trip distance, extract the road length information from the road network, and calculate the cost of each path.
[0059] S243. Optimization objective: Integrate the environmental features and the path cost, and the form of the optimization objective is as follows:
[0060] R(s,a) = R expert (s,a) + αR cost (s,a)
[0061] where, R expert (s,a) represents the reward learned based on environmental features, reflecting the effectiveness of the model in fitting path preferences; R cost (s,a) represents the path cost reward, and α represents a hyperparameter that adjusts the weight between the two rewards. In the present invention, considering both the path cost and the path preferences, while not affecting the ability of the model to learn from environmental features to fit path preferences, the path cost is minimized as much as possible. Therefore, the value of α is set to 1.5.
[0062] S250. Adversarial training. The model training adopts an alternating training method. First, fix the discriminator D and optimize the loss function of the generator G. The optimization objective of the generator is to maximize the negative logarithm of the probability that the generated data is judged as real data by the discriminator. Then, fix the generator G and optimize the loss function of the discriminator D. The optimization objective of the discriminator is to maximize the ability to judge real data as real data. The optimization objectives of the generator and the discriminator are as follows:
[0063]
[0064] Among them, z represents the input of the generator, G(z) represents the data generated by the generator, D(·) represents the discriminator function, which can represent the probability that the input data is real data; x represents the data collected from the real world, p data (x) represents the distribution of real data, and p(z) represents the distribution of data generated by the generator;
[0065] By alternately optimizing the loss functions of the generator and the discriminator, gradually improve the quality of the generated data and the discrimination ability of the discriminator, so that the generated data is as consistent as possible with the distribution of real data;
[0066] S300. Route recommendation, including route recommendation algorithms for large-scale road networks and small-scale road networks;
[0067] S310. For small-scale road networks, since the trajectory distribution is relatively dense, directly use the action selection probability output by the model to find the route;
[0068] S320. For large-scale road networks, since the trajectory distribution is relatively sparse, define the "reciprocal of the action probability" output by the model as the path cost of the Dijkstra algorithm. A higher probability action means a lower cost path, so as to reduce the dependence on the data coverage of the model. First, use the model to obtain the selection probability of each action in the current state, and use the reciprocal of the action selection probability as the cost of finding the route by the Dijkstra algorithm. The optimized model can select the overall optimal path based on the global situation, retain the trajectory selection preference of the model, and effectively control the path length, making the trajectory structure closer to the real data and realizing the integration of model preference and global optimization.
[0069] Specific implementation plan two: The present invention provides an urban vehicle personalized path planning system based on adversarial inverse reinforcement learning. This system has program modules corresponding to the above steps and executes the steps in the above-mentioned method for urban vehicle personalized path planning based on adversarial inverse reinforcement learning when running.
[0070] The other combinations and connection relationships in this implementation plan are the same as those in the first specific implementation plan.
[0071] Specific Embodiment 3: The present invention provides a computer-readable storage medium storing a computer program configured to implement the steps of a personalized path planning method for urban vehicles based on adversarial inverse reinforcement learning when called by a processor.
[0072] Other combinations and connection relationships in this embodiment are the same as those in Specific Embodiment 1.
[0073] Experiment
[0074] Use small-scale Shanghai datasets and large-scale Wuhan datasets to verify the effectiveness of the model of the present invention. The road network within the scope of the Shanghai dataset is as Figure 5 shown, and the road network within the scope of the Wuhan dataset is as Figure 6 shown. The experimental datasets are shown in Table 1.
[0075] Table 1 Dataset Introduction
[0076]
[0077]
[0078] Evaluation Metrics: 1. Fitted Path Preference: Trajectory Structure Similarity ED (the smaller the better), Node Matching Degree BLUE (the larger the better), Distribution Fitting Degree JSD (the smaller the better); 2. Reduced Path Length: Average Path Length (the smaller the better).
[0079] The performance comparisons of MR-AIRL of the present invention and other path recommendation algorithms in terms of ED, BLUE, JSD, and average path length are shown in Tables 2 to 5 below. Since the average values of the metrics in the Shanghai and Wuhan experiments can more clearly reflect the experimental results of the model, the specific values of the eight groups of experiments for each metric are not listed. Calculate the average values of each metric in the Shanghai and Wuhan experiments respectively, and analyze the comprehensive performance of each path recommendation method. The comprehensive performance of the 4 performance metrics of the test data of MR-AIRL of the present invention and other path recommendation algorithms on the eight datasets in the two cities is shown in Tables 6 and 7 below.
[0080] Table 2 Result Comparison of ED for Different Path Recommendation Methods
[0081]
[0082] Table 3 Result Comparison of BLEU for Different Path Recommendation Methods
[0083]
[0084]
[0085] Table 4 Result Comparison of JSD for Different Path Recommendation Methods
[0086]
[0087] Table 5 Comparison of the average path lengths of different path recommendation methods
[0088]
[0089] Table 6 Comprehensive performance in Shanghai
[0090]
[0091] Table 7 Comprehensive performance in Wuhan
[0092]
[0093] Experimental results: The model designed by the present invention can comprehensively consider the performance of the model in terms of reducing the path length and approaching the user's travel preferences. The experimental results on the Shanghai dataset show that the MR-AIRL model performs optimally in all four indicators. In particular, there are significant improvements in path structure consistency (ED) and node matching (BLEU), and the average path length is also basically close to that of A*, with the best comprehensive performance. The experimental results on the Shanghai dataset show that the model has improved by 82.2%, 341%, and 72% respectively in trajectory structure, node matching, and distribution fitting compared with the Dijkstra algorithm, and the average path length of the recommendation has decreased by 17.8%. In trajectory structure, node matching, and distribution fitting, it has improved by 71.7%, 160%, and 54.3% respectively compared with the A* algorithm. With fewer training times, a better path recommendation effect can be achieved. Compared with the AIRL model that only considers path preferences, the model designed by the present invention has improved by 48.4%, 26.5%, and 20.1% respectively in trajectory structure, node matching, and distribution fitting, and the average path length of the recommendation has decreased by 7.8%.
[0094] The matching graph of real historical trajectories on the road network in Wuhan is as Figure 7 shown, and the result of path recommendation using the model of the present invention is as Figure 8As shown. The experimental results on the Wuhan dataset show that the model has improved by 83.74%, 184.04%, and 62.02% respectively compared with the Dijkstra algorithm in terms of trajectory structure, node matching, and distribution fitting. The average path length recommended has decreased by 13.97%. Compared with the A* algorithm, it has improved by 73.05%, 77.75%, and 55.03% respectively in terms of trajectory structure, node matching, and distribution fitting. With fewer training times, it has also achieved better path recommendation effects. Compared with the AIRL model that only considers path preferences, the model designed in the present invention has improved by 78.6%, 102.5%, and 46.9% respectively compared with the Dijkstra algorithm in terms of trajectory structure, node matching, and distribution fitting. The average path length recommended has decreased by 62.6%
[0095] Based on the above two experiments, the model of the present invention is superior to other models on both the small-scale road network in Shanghai and the large-scale road network in Wuhan. Especially when facing the path recommendation problem of large-scale road networks and data sparsity, it shows stronger adaptability.
[0096] Although the present invention is disclosed as above, the protection scope of the present invention is not limited thereto. Those skilled in the art of the present invention can make various changes and modifications without departing from the spirit and scope of the present invention, and these changes and modifications will all fall within the protection scope of the present invention.
Claims
1. A personalized path planning method for urban vehicles based on adversarial inverse reinforcement learning, characterized in that It includes the following steps: S100. Data processing, matching the processed historical trajectory data to the urban road network map; including optimizing the road network structure of the urban road network data, processing the historical trajectory data, and then mapping the processed trajectory data to the road network using a map matching algorithm; S200. Model training, including based on adversarial inverse reinforcement learning, introducing an attention mechanism, through adversarial training of a generator and a discriminator, obtaining a reward function with the goal of maximizing trajectory similarity and minimizing path cost, learning the potential path preference, and capturing the multi-objective decision-making strategy of experts through adversarial training and inverse reasoning to generate paths; S300. Path recommendation, making corresponding path recommendations for large-scale road networks and small-scale road networks.
2. The personalized path planning method for urban vehicles based on adversarial inverse reinforcement learning according to claim 1, characterized in that: In step S100, the methods for processing GPS data include coordinate system conversion, data screening, data cleaning, and trajectory segmentation.
3. The personalized path planning method for urban vehicles based on adversarial inverse reinforcement learning according to claim 1, characterized in that: In step S200, it includes S210. Environment modeling, modeling the road network as a directed graph, represented by G=(V, E), where V is the set of nodes, representing key positions, and E is the set of edges, representing directed road segments, that is, the segments along which vehicles can travel in a specific direction; S220. Generator design, designing a generator based on a convolutional neural network, generating feasible actions by processing input states, destinations, path features, and environmental features; S230. Discriminator design, designing a discriminator based on a convolutional neural network, and the discriminator extracts neighborhood information in the input features through the convolutional neural network to distinguish the trajectories generated by the generator from the real expert trajectories; S240. Learning objective, comprehensively considering the path preference latent in the historical trajectory data and the reward function of the travel length, including environmental feature extraction, path cost calculation, and design optimization objectives; S250. Adversarial training, gradually improving the quality of the generated data and the discrimination ability of the discriminator by alternately optimizing the loss functions of the generator and the discriminator, making the distribution of the generated data as consistent as possible with the real data.
4. The personalized path planning method for urban vehicles based on adversarial inverse reinforcement learning according to claim 3, wherein: In step S220, the convolutional layer processes the input path features. The first layer of convolution uses 20 3*3 convolutional kernels, and the second layer of convolution uses 30 2*2 convolutional kernels to extract local features; The attention mechanism further processes the features output by the convolutional layer to capture the global dependencies between path features; uses three fully connected layers to gradually extract features to generate an action probability distribution; the convolutional layer uses the LeakyReLU activation function, and the pooling layer uses the max pooling operation; the output of the generator is the probability distribution of each action, obtained through the Softmax function, used to select actions or calculate the logarithmic probability of actions.
5. The personalized path planning method for urban vehicles based on adversarial inverse reinforcement learning according to claim 3, characterized in that: In step S230, first, features are extracted through a convolutional layer, and then the size of the feature map is reduced and the degree of feature abstraction is increased through a max-pooling layer; the second convolutional layer further extracts high-level features. After adding a multi-head self-attention (MHSA) module after the convolutional layer of the discriminator, the features output by the convolutional layer are processed by the MHSA module, and the obtained feature vectors are mapped through a fully connected layer. The Sigmoid activation function is used to calculate the score of the trajectory, to distinguish the difference between the recommended path and the real path, and to calculate the path similarity.
6. The personalized path planning method for urban vehicles based on adversarial inverse reinforcement learning according to claim 3, characterized in that: In step S240, it specifically includes: S241. Extracting environmental features, extracting environmental features related to the itinerary from historical trajectory data, including road grade, number of turns, and distance to the destination; S242. Calculating path costs, defining the path cost as the travel distance, extracting road length information from the road network, and calculating the cost of each path; S243. Optimization objective, fusing the environmental features and the path costs, and the form of the optimization objective is as follows: R(s,a) = R expert (s,a) + αR cost (s,a) Among them, R expert (s,a) represents the reward for learning based on environmental characteristics, R cost (s,a) represents the path cost reward, and α represents a hyperparameter that adjusts the weight between the two rewards.
7. The personalized path planning method for urban vehicles based on adversarial inverse reinforcement learning according to claim 3, wherein: In step S250, an alternating training method is adopted. First, the discriminator D is fixed, and the loss function of the generator G is optimized. The optimization objective of the generator is to maximize the negative logarithm of the probability that the generated data is judged as real data in the discriminator; then the generator G is fixed, and the loss function of the discriminator D is optimized. The optimization objective of the discriminator is to maximize the ability to judge real data as real data; the optimization objectives of the generator and the discriminator are as follows respectively: minL G =-E z~p(z) [logD(G(z))] Among them, z represents the input of the generator, G(z) represents the data generated by the generator, D(·) represents the discriminator function, which can represent the probability that the input data is real data; x represents the real-world data collected, and p data (x) represents the distribution of real data, and p(z) represents the distribution of data generated by the generator.
8. A personalized path planning method for urban vehicles based on adversarial inverse reinforcement learning according to claim 1, characterized in that: In step S300, it specifically includes: S310. For a small-scale road network, since the trajectory distribution is relatively dense, the action selection probability output by the model is directly used for path finding; S320. For a large-scale road network, since the trajectory distribution is relatively sparse, first, the model is used to obtain the selection probability of each action in the current state, and the reciprocal of the action selection probability is used as the cost of the Dijkstra algorithm for path finding. The optimized model can select the overall optimal path based on the global situation.
9. A personalized path planning system for urban vehicles based on adversarial inverse reinforcement learning, characterized in that: The system has program modules corresponding to the steps of any one of the above claims 1-8, and when running, it executes the steps in the above-mentioned method for personalized path planning of urban vehicles based on adversarial inverse reinforcement learning.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is configured to implement the steps of the method for personalized path planning of urban vehicles based on adversarial inverse reinforcement learning according to any one of claims 1-8 when called by a processor.
Citation Information
Patent Citations
Personalized path recommendation algorithm of A* algorithm based on deep learning guidance
CN115577175A
Personalized path planning method and system based on federated learning and edge calculation
CN119417001A
Digital tourist travel planning system
CN119692584A
Inverse reinforcement learning method and system for enhancing authenticity of traffic simulator
CN113221469A
End-to-end sparse trajectory recovery method based on road network
CN119251525A