Taxi route recommendation method based on FastDTW time sequence clustering algorithm
Through the improved FastDTW time series clustering algorithm and hierarchical clustering technology, combined with the multi-objective optimization scheduling method, the problems of low data quality, low trajectory similarity calculation efficiency and unreasonable route recommendation in taxi operations are solved, efficient passenger destination prediction and route recommendation are achieved, and air driving rate and operation costs are reduced.
Patent Information
- Application Number
- CN202510366841.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-04
AI Technical Summary
There are problems in existing taxi operations such as low data quality, low trajectory similarity calculation efficiency, inaccurate destination prediction and unreasonable route recommendation, resulting in high air driving rates and low scheduling efficiency.
The improved FastDTW time series clustering algorithm and hierarchical clustering technology are adopted, combined with the multi-objective optimization scheduling method, and by cleaning and segmenting the GPS trajectory data, calculating the trajectory similarity, building a distance matrix and clustering, combining driving distance, time and cost factors, the optimal passenger search route is recommended.
Improve data quality and computing efficiency, accurately predict high-frequency passenger destinations, reduce air travel time, optimize vehicle resource allocation, and improve driver operation efficiency and passenger travel experience.
Smart Images

Figure CN120258272A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to trajectory data mining technology and machine learning technology in an intelligent transportation system (ITS), and particularly to a taxi route recommendation method based on the FastDTW time series clustering algorithm, a spatio-temporal model analysis method based on taxi GPS trajectory data, and combines the FastDTW time series clustering algorithm to achieve high-frequency passenger destination prediction and route optimization recommendation. Background Art
[0002] With the acceleration of the urbanization process and the continuous increase in traffic demand, taxis, as the main urban public transportation means, undertake a large number of travel tasks. However, there are many problems in the taxi operation process, such as inaccurate destination prediction by drivers, unreasonable route selection, too long empty driving time, etc., resulting in poor passenger travel experience and low driver operation efficiency.
[0003] In recent years, with the popularization of GPS technology and the development of big data technology, the GPS trajectory data of taxis has become an important data source for studying taxi operation behaviors. Through the analysis of these data, the excavation of taxi driving trajectories, destination prediction, and route optimization can be realized. However, when dealing with GPS trajectory data, the existing methods often have the following problems:
[0004] 1. Low data quality: GPS trajectory data often contains noise, missing values, and outliers. Directly using these data will lead to inaccurate data analysis results.
[0005] 2. Low efficiency in calculating trajectory similarity: Traditional trajectory similarity calculation methods (such as Euclidean distance, dynamic time warping DTW, etc.) have high computational complexity and cannot process large-scale trajectory data.
[0006] 3. Inaccurate destination prediction: Existing destination prediction methods are mostly based on simple statistical models and lack in-depth excavation of the spatio-temporal characteristics of trajectories, resulting in inaccurate prediction results.
[0007] 4. Unreasonable route recommendation: Existing route recommendation methods mostly rely on a single factor (such as the shortest path) and lack comprehensive consideration of multiple factors (such as driving time, cost, traffic conditions, etc.), resulting in unreasonable recommendation results.
[0008] Therefore, the present invention proposes an intelligent taxi route recommendation method through an improved FastDTW algorithm and hierarchical clustering technology to solve the above problems. Summary of the Invention
[0009] Objective of the Invention: Aiming at the problems of high empty driving rate and low dispatching efficiency in the existing taxi operation, the present invention provides a taxi route recommendation method based on the FastDTW time series clustering algorithm. By improving the FastDTW algorithm and hierarchical clustering technology, combined with the multi-objective optimization dispatching method, it predicts the high-frequency demand areas to guide vehicle dispatching, and at the same time recommends potential high-probability passenger-carrying destinations on the return route for drivers, reduces the empty driving time and operation cost, and improves the resource utilization efficiency of taxi companies and the operation efficiency of drivers.
[0010] Technical Solution: The present invention proposes a taxi route recommendation method based on the FastDTW time series clustering algorithm, including the following steps:
[0011] Step 1: Obtain the original taxi GPS trajectory data set from the public data platform, and perform data cleaning, data sorting, and trajectory segmentation processing based on the change of passenger-carrying status to obtain the preprocessed trajectory data set;
[0012] Step 2: Use the FastDTW algorithm to calculate the similarity between trajectories, and construct the distance matrix D between pairwise trajectories Matrix ;
[0013] Step 3: Apply the distance matrix D Matrix to the hierarchical clustering algorithm to cluster the trajectories, obtain the high-probability passenger-carrying destinations, and construct the candidate route list Route_list;
[0014] Step 4: Combine the factors of driving distance, driving time, and driving cost, call the Gaode Map API to obtain detailed route information, and recommend the optimal passenger-seeking route from the candidate route list Route_list through comprehensive scoring.
[0015] Furthermore, the specific method of the above Step 1 is as follows:
[0016] Step 1.1: Obtain the original taxi GPS trajectory data set from the public data platform, including the taxi ID number, positioning timestamp, longitude Lon, latitude Lat, and passenger-carrying status, and the passenger-carrying status includes empty car and full car;
[0017] Step 1.2: Perform data preprocessing on the data, including data cleaning, data sorting, and trajectory segmentation processing based on the change of passenger-carrying status;
[0018] Data cleaning includes removing invalid or abnormal data points, filling in missing values using interpolation methods, and removing duplicate GPS data; data arrangement includes data type conversion to convert all location timestamps into a unified time format; trajectory segmentation processing based on changes in passenger-carrying status. First, according to the passenger-carrying status, continuous trajectories are segmented into multiple independent sub-trajectories. The longitude and latitude coordinate sequences of the sub-trajectories in the 'heavy vehicle' state are extracted to represent a complete passenger-carrying trip, forming a trajectory dataset.
[0019] Further, the specific method of step 2 is as follows:
[0020] Step 2.1: Screen the trajectories near the starting point. According to historical data analysis, determine the high-frequency departure location P of taxis, set the starting point as P, and screen all the trajectory data near the starting point P from the preprocessed trajectory dataset to obtain the final trajectory dataset Data;
[0021] Step 2.2: Each trajectory in Data is represented as T n ={(lon1,lat1),...(lon i ,lat i ),...(lon n ,lat n )}, where (lon i ,lat i ) represents the longitude and latitude coordinates of the i-th trajectory point, n represents the total number of trajectory points, and T n represents a set containing n position points from (lon1,lat1) to (lon n ,lat n ), and i = 1, 2,..., n;
[0022] Step 2.3: Divide the trajectory dataset Data into training trajectories and test trajectories at a ratio of 80% training trajectories and 20% test trajectories. Let the total number of training trajectories be M and the total number of test trajectories be N;
[0023] Step 2.4: Create an empty distance matrix D Matrix , which is used to store the similarity between all trajectory pairs, with the initial value set to infinity, and the matrix size is M*M, where M is the total number of training trajectories;
[0024] Step 2.5: Introduce a banded search window, define the radius of the banded window as r, and only calculate the point pairs within the diagonal ±r range. The time complexity is reduced from O(n 2 ) to O(rn). The calculation formula is as follows: |i - j| ≤ r, where i and j are the trajectory sequence indices;
[0025] Step 2.6: The original trajectory sequence T nDownsample to T by average pooling method n ', calculate the DTW path at low resolution, constrain the high-level search range, and gradually refine to the original resolution. The calculation formula is as follows;
[0026]
[0027] Among them, lon i ' represents the average value of the i-th longitude, lat i ' represents the average value of the i-th latitude, lon 2i-1 and lon 2i respectively represent the longitudes of the (2i - 1)-th and 2i-th trajectory points, lat 2i-1 and lat 2i respectively represent the latitudes of the (2i - 1)-th and 2i-th trajectory points, and i represents the index in the sequence, ranging from 1 to n represents the total number of trajectory points;
[0028] Step 2.7: For each pair of training trajectory time series T i and T j , use the FastDTW algorithm to calculate D TW (i, j), and fill it into the corresponding initialized distance matrix D Matrix at the same time. The calculation formula is:
[0029]
[0030] Among them, D TW (i, j) represents the minimum cumulative alignment distance between the two time series T i and T j at positions i and j, Dist(i, j) represents the Euclidean distance between trajectories i and j, and D TW (i - 1, j - 1) represents the cumulative distance of aligning diagonally from the (i - 1)-th point of T i and the (j - 1)-th point of T j , D TW (i, j - 1) represents the cumulative distance of aligning horizontally from the i-th point of T i and the (j - 1)-th point of T j , and D TW (i - 1, j) represents the cumulative distance of aligning vertically from the (i - 1)-th point of T i and the j-th point of T j .
[0031] Furthermore, the specific method of step 3 is:
[0032] Step 3.1: The distance matrix D MatrixApplied to the hierarchical clustering algorithm, the average linkage method is used to construct a hierarchical clustering tree, and the average distance of all trajectory pairs between two clusters is calculated as the inter-cluster distance. For two clusters C1 and C2, the formula for calculating the inter-cluster distance is as follows;
[0033]
[0034] Among them, d avg (C1, C2) represents the average distance between clusters C1 and C2. This value is used to measure the similarity between two clusters. |C1| and |C2| represent the number of trajectories in clusters C1 and C2 respectively. T i and T j represent two training trajectories. D Matrix (i)(j) represents the element value in the i-th row and j-th column of the distance matrix D Matrix . This matrix stores the distances between all element pairs;
[0035] Step 3.2: When the average distance d avg (C1, C2) ≤ t, merge the two clusters, where t is the set distance threshold;
[0036] Step 3.3: Obtain the cluster labels of M training trajectories from the results of hierarchical clustering. Each trajectory is assigned to a specific cluster C to form the classification result of M training trajectories;
[0037] Step 3.4: Define a counting matrix B Matrix to store the number of occurrences of each clustering, and define an accumulated distance matrix R Matrix to accumulate the total distance of each clustering;
[0038] Step 3.5: For each test trajectory T test , calculate the FastDTW distance D test,i between N test trajectories and M training trajectories. Accumulate the FastDTW distance D test,i of each trajectory to the total distance R Matrix of the corresponding clustering, and update the number of the corresponding clustering in the counting matrix B Matrix :
[0039] Step 3.6: For each cluster C i , calculate its average distance Avg_dis i , and the calculation formula is as follows:
[0040]
[0041] Among them, R i is the total distance accumulated by cluster C i , and B Matrix [i] is for cluster Ci The total number of trajectories.
[0042] Step 3.7: Select the cluster C with the minimum average distance from all clusters min , and extract from it all the predicted end destination coordinate sets E = {e1, e2,... e L}, where e1, e2,... e L represent the addresses of multiple predicted destinations in the set E. Combine the end point E and the start point P to form a candidate route list Route_list = {(P→e1), (P→e2),..., (P→e L )}, and L is the total number of candidate routes.
[0043] Furthermore, the specific method of step 4 is as follows:
[0044] Step 4.1: Call the Gaode Map API to obtain the key information of the L candidate routes, including the driving cost C = {c1, c2,..., c L}, the driving time T = {t1, t2,..., t L}, the driving distance D = {d1, d2,..., d L}, and the traffic status of each section, including smooth, slow, congested, severely congested, unknown;
[0045] Step 4.2: According to the traffic status of each section, use the weighted average method to calculate the average smoothness S = [S1, S2,..., S i ..., S L of each candidate route. The calculation formula is as follows:
[0046]
[0047] where A j is the weight set for the j-th sub-section according to its traffic status. Set the smooth weight as a1, the slow weight as a2, the congested weight as a3, the severely congested weight as a4, and the unknown weight as a5. O i is the total number of sub-sections of the i-th candidate route;
[0048] Step 4.3: Perform min-max normalization on the driving cost C, driving time T, and driving distance D to obtain the normalized driving cost C norm , driving time T norm , and driving distance D norm ;
[0049] Step 4.4: Adjust the weights of the driving cost, driving time, driving distance, and smoothness to reflect the importance for the driver to select a route;
[0050] Step 4.5: Use the weighted summation method to calculate the comprehensive scores Scores of L candidate routes. The calculation formula is as follows:
[0051] Scores = w c *C norm + w t *T norm + w d *D norm + w s *S
[0052] w c + w t + w d + w s = 1
[0053] where w c , w t , w d , w s are the weights of driving cost, driving time, driving distance, and traffic smoothness respectively;
[0054] Step 4.6: Sort all candidate routes according to the comprehensive scores Scores, find the route with the highest score, and display it visually on the map for drivers' reference.
[0055] The present invention adopts the above technical solutions and has the following beneficial effects:
[0056] 1. By performing preprocessing operations such as cleaning GPS trajectory data, interpolating to fill in missing values, and segmenting trajectories, the present invention significantly improves the quality of data, providing a reliable data basis for subsequent analysis.
[0057] 2. By introducing the FastDTW algorithm, restricting the search range and multi-resolution data abstraction, the present invention significantly reduces the time complexity of trajectory similarity calculation, improves the calculation efficiency, and can process large-scale trajectory data.
[0058] 3. By mining historical trajectories, obtaining high-frequency destinations, guiding taxi companies to dynamically dispatch vehicles to areas with concentrated demand, reducing empty driving time, it can help drivers quickly match return orders, reducing empty driving mileage and fuel consumption.
[0059] 4. Based on the improved FastDTW algorithm and hierarchical clustering technology, the present invention quickly identifies the spatio-temporal distribution of high-frequency demand in the city, provides real-time dispatching decision support for taxi companies, and optimizes vehicle resource allocation.
[0060] 5. The present invention comprehensively considers driving costs, time, distance, and real-time traffic conditions, and recommends the most economically efficient dispatching route and return passenger-carrying path through a weighted scoring model, improving the passenger travel experience and the driver operation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 is the overall method flow chart of the present invention;
[0062] Figure 2 is the detailed flow block diagram of the taxi route recommendation method of the present invention;
[0063] Figure 3 is the schematic diagram of the recommended result of the optimal passenger-carrying route of the present invention;
[0064] Figure 4 is the visualization display diagram of the optimal passenger-carrying route of the present invention on the map. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0065] The following further clarifies the present invention in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent forms of modification of the present invention by those skilled in the art all fall within the scope defined by the appended claims of this application.
[0066] The present invention discloses a taxi route recommendation method based on the FastDTW time series clustering algorithm, which specifically includes the following steps:
[0067] Step 1: Obtain the original taxi GPS trajectory data set from the public data platform, and perform preprocessing operations such as data cleaning, data sorting, and trajectory segmentation processing based on the passenger-carrying status:
[0068] Step 1.1: Obtain the original taxi GPS trajectory data set of Wuhan from the public data platform, ensuring that the data includes key fields such as taxi ID number, positioning timestamp, longitude Lon, latitude Lat, and passenger-carrying status (empty car / occupied car).
[0069] Step 1.2: Perform preprocessing on the data, including data cleaning, data sorting, and trajectory segmentation processing based on changes in the passenger-carrying status.
[0070] Data cleaning includes removing invalid or out-of-range data points of longitude and latitude, filling missing values using interpolation method, and removing duplicate GPS data; data sorting includes data type conversion, converting all positioning timestamps to a unified time format; trajectory segmentation processing based on the passenger-carrying status, first dividing the continuous trajectory into multiple independent sub-trajectories according to the passenger-carrying status (empty car / occupied car), and only retaining the sub-trajectories in the 'occupied car' status to form a trajectory data set.
[0071] Step 2: Calculate the similarity between trajectories using the FastDTW algorithm and construct a distance matrix D between pairwise trajectories Matrix :
[0072] Step 2.1: Based on historical data analysis, determine the high-frequency departure location P of taxis, set the starting point as P, and from the data analysis of this dataset, P = {'No. 256, Yanhe Avenue, Hanyang District, Wuhan City'}. Centered on P, filter out all trajectory data within 300 meters nearby to obtain the final trajectory dataset Data.
[0073] Each trajectory in Data is represented as T n ={(lon1,lat1),...(lon i ,lat i ),...(lon n ,lat n )}, where (lon i ,lat i ) represents the longitude and latitude coordinates of the i-th trajectory point, n represents the total number of trajectory points, and T n represents a set containing n position points from (lon1,lat1) to (lon n ,lat n ).
[0074] Step 2.3: Divide the trajectory dataset Data into training trajectories and test trajectories, with a ratio of 80% training trajectories and 20% test trajectories. Let the total number of training trajectories be M and the total number of test trajectories be N.
[0075] Step 2.4: Create an empty distance matrix D Matrix , which is used to store the similarity between all trajectory pairs. The initial value is set to infinity, and the matrix size is M*M, where M is the total number of training trajectories.
[0076] The DTW algorithm has low running efficiency and high computational complexity. The FastDTW algorithm accelerates the running efficiency of DTW by restricting the search range and multi-resolution data abstraction.
[0077] Step 2.6: Introduce a banded search window, define the radius of the banded window as r, and only calculate the pairs of points within the diagonal range of ±r. The time complexity is reduced from O(n 2 ) to O(rn). The calculation formula is as follows:
[0078] |i - j| ≤ r
[0079] where i and j are the indices of the trajectory time series.
[0080] Step 2.7: The original trajectory sequence T n={(lon1, lat1), (lon2, lat2),...(lon n , lat n )} is downsampled to T n ' = {(lon1', lat1'), (lon2', lat2'),...(lon k ', lat k ')}, Reduce the resolution, calculate the DTW path at the low resolution, constrain the high-level search range, and gradually refine to the original resolution. The calculation formula is as follows:
[0081]
[0082] Among them, lon i ' represents the average value of the i-th longitude, lat i ' represents the average value of the i-th latitude, lon 2i-1 and lon 2i respectively represent the longitudes of the (2i - 1)-th and 2i-th trajectory points, lat 2i-1 and lat 2i respectively represent the latitudes of the (2i - 1)-th and 2i-th trajectory points, i represents the index in the sequence, from 1 to n represents the total number of trajectory points.
[0083] Step 2.8: For each pair of training trajectories T i and T j , use the FastDTW algorithm to calculate D TW (i, j), and fill it into the corresponding initialized distance matrix D Matrix at the same time. The calculation formula is:
[0084]
[0085] Among them, D TW (i, j) represents the minimum cumulative alignment distance between the two time series T i and T j at positions i and j, Dist(i, j) represents the Euclidean distance between trajectories i and j, D TW (i - 1, j - 1) represents the cumulative distance of aligning diagonally from the (i - 1)-th point of T i and the (j - 1)-th point of T j , D TW (i, j - 1) represents the cumulative distance of aligning horizontally from the i-th point of T i and the (j - 1)-th point of T j , D TW (i - 1, j) represents the cumulative distance of aligning vertically from the (i - 1)-th point of T iThe cumulative distance where the (i - 1)-th point of and T j are aligned along the column direction with the j-th point of
[0086] Step 3: Apply the distance matrix D Matrix to the hierarchical clustering algorithm to cluster the trajectories, obtain the high-probability passenger-carrying destinations, and construct a candidate route list Route_list:
[0087] Step 3.1: Apply the distance matrix D obtained in Step 2.8 Matrix to the hierarchical clustering algorithm, use the average linkage method to construct a hierarchical clustering tree, set the parameter method to average, metric to precomputed, and the distance threshold t = 1. After completing the parameter settings, perform the clustering.
[0088] Step 3.2: Obtain the cluster labels of the M training trajectories from the results of the hierarchical clustering. Each trajectory is assigned to a specific cluster C, forming the classification results of the M training trajectories.
[0089] Step 3.3: Define a counting matrix B Matrix , which is used to store the number of occurrences of each cluster, and define a cumulative distance matrix R Matrix , which is used to accumulate the total distances of each cluster.
[0090] Step 3.4: For each test trajectory T test , use Step 2.8 to calculate the FastDTW distances D between the N test trajectories and the M training trajectories test,i , and accumulate the FastDTW distances D of each trajectory test,i to the total distance R of the corresponding cluster Matrix , and update the number of the corresponding cluster in the counting matrix B Matrix .
[0091] Step 3.5: For each cluster C i , calculate its average distance Avg_dis i , and the calculation formula is as follows:
[0092]
[0093] where R i is the total distance accumulated by cluster C i , and B Matrix [i] is the total number of trajectories in cluster C i .
[0094] Step 3.6: Select the cluster C with the minimum average distance from all clusters min , and extract all the predicted end destination coordinate sets E = {e1, e2,... e L}, combine the end point E and the starting point P to form a list of candidate routes Route_list = {(P→e1), (P→e2),...,(P→e L )}, where L is the total number of candidate routes.
[0095] Step 4: Combine factors such as driving distance, driving time, and driving cost, call the Gaode Map API to obtain detailed route information, and recommend the optimal passenger-carrying route to the driver through a comprehensive scoring model:
[0096] Step 4.1: Call the Gaode Map API to obtain the key information of the L candidate routes, including the driving cost C = {c1, c2,..., c L}, driving time T = {t1, t2,..., t L}, driving distance D = {d1, d2,..., d L}, and the traffic status (smooth / slow / congested / severely congested / unknown) of each section.
[0097] Step 4.2: According to the traffic status of each section obtained in Step 4.1, use the weighted average method to calculate the average smoothness S = [S1, S2,..., S i ..., S L , and the calculation formula is as follows:
[0098]
[0099] Among them, A j is the weight set for the jth sub-section according to its traffic status (for example, set the smooth weight a1 to 1.0, slow a2 to 0.5, congested a3 to 0.2, severely congested a4 to 0.0, unknown a5 to 0.3), and O i is the total number of sub-sections of the ith candidate route.
[0100] Step 4.3: Perform min-max normalization on the driving cost C, driving time T, and driving distance D, and the calculation formula is:
[0101]
[0102] Step 4.4: Use the weighted summation method to calculate the comprehensive scores Scores of the L candidate routes, and the calculation formula is as follows:
[0103] Scores = w c *C norm + w t *T norm + w d *D norm + w s *S
[0104] w c + w t + w d + w s = 1
[0105] where w c , w t , w d , w s are the weights of driving cost, driving time, driving distance, and smoothness respectively. (For example, if the driving cost is relatively important, set w c to 0.4, w t , w d , w s to 0.2).
[0106] Step 4.5: According to the comprehensive scores of the L candidate routes calculated above, sort all the candidate routes and find the route with the highest score. In this experiment, the first candidate route has the highest score, as specifically Figure 3 shown, visualize it on the map, specifically as Figure 4 shown, and the result is for the driver's reference.
[0107] In summary, based on the taxi GPS trajectory dataset in Wuhan, the present invention has successfully realized an intelligent solution for taxi passenger-seeking routes. By introducing the FastDTW algorithm and the hierarchical clustering algorithm, combined with the Amap API, the data processing efficiency and the accuracy of passenger destination prediction have been significantly improved, and the comprehensive performance of route recommendation has been optimized. The present invention provides data-driven technical support for the optimization of urban traffic resources, improves the operation efficiency of drivers, and has significant economic and social benefits.
[0108] The above embodiments are only used to illustrate the technical concept and features of the present invention, and their purpose is to enable those who are familiar with this technology to understand the content of the present invention and implement it accordingly, and it should not be used to limit the protection scope of the present invention. Any equivalent transformation or modification made according to the spirit and essence of the present invention should be covered within the protection scope of the present invention.
Claims
1. A taxi route recommendation method based on the FastDTW time series clustering algorithm, characterized in that It includes the following steps: Step 1: Obtain the original taxi GPS trajectory dataset from the public data platform, and perform data cleaning, data sorting, and trajectory segmentation processing based on the change in the passenger-carrying status to obtain the preprocessed trajectory dataset; Step 2: Calculate the similarity between trajectories using the FastDTW algorithm and construct the distance matrix D between pairwise trajectories Matrix ; Step 3: Apply the distance matrix D Matrix to the hierarchical clustering algorithm to cluster the trajectories, obtain the high-probability passenger-carrying destinations, and construct a candidate route list Route_list; Step 4: Combine factors such as driving distance, driving time, and driving cost, call the Gaode Map API to obtain detailed route information, and recommend the optimal passenger-seeking route from the candidate route list Route_list through comprehensive scoring calculation.
2. The taxi route recommendation method based on the FastDTW time series clustering algorithm according to claim 1, characterized in that The specific method of the said Step 1 is: Step 1.1: Obtain the original taxi GPS trajectory dataset from the public data platform, including taxi ID number, positioning timestamp, longitude Lon, latitude Lat, and passenger-carrying status. The passenger-carrying status includes empty car and full car; Step 1.2: Perform preprocessing on the data, including data cleaning, data sorting, and trajectory segmentation processing based on the change in the passenger-carrying status; Data cleaning includes removing invalid or abnormal data points, filling missing values using interpolation method, and removing duplicate GPS data; data sorting includes data type conversion, converting all positioning timestamps to a unified time format; for the trajectory segmentation processing based on the change in the passenger-carrying status, first, according to the passenger-carrying status, divide the continuous trajectory into multiple independent sub-trajectories, and extract the longitude and latitude coordinate sequences of the sub-trajectories in the 'full car' status to represent a complete passenger-carrying trip, forming a trajectory dataset.
3. A taxi route recommendation method based on the FastDTW time series clustering algorithm according to claim 1, characterized in that, The specific method of the said Step 2 is: Step 2.1: Screen the trajectories near the starting point. According to historical data analysis, determine the high-frequency departure location P of taxis, set the starting point as P, and screen all the trajectory data near the starting point P from the preprocessed trajectory dataset to obtain the final trajectory dataset Data; Step 2.2: Each track of Data represents T n ={(lon1,lat1),...(lon i ,lat i ),...(lon n ,lat n )}, where (lon i ,lat i ) represents the longitude and latitude coordinates of the i-th track point, n represents the total number of track points, and T n represents a set containing n position points from (lon1,lat1) to (lon n ,lat n ), and i = 1, 2,..., n; Step 2.3: Divide the trajectory dataset Data into training trajectories and test trajectories, with a ratio of 80% training trajectories and 20% test trajectories. Let the total number of training trajectories be M and the total number of test trajectories be N; Step 2.4: Create an empty distance matrix D Matrix , which is used to store the similarity between all trajectory pairs, with the initial value set to infinity. The size of the matrix is M*M, where M is the total number of training trajectories; Step 2.5: Introduce a strip search window, define the radius of the strip window as r, and only calculate the point pairs within the diagonal range of ±r. The time complexity is reduced from O(n 2 ) to O(rn), and the calculation formula is as follows: |i - j| ≤ r, where i and j are the indexes of the trajectory sequence; Step 2.6: Downsample the original trajectory sequence T n to T n ' by average pooling method, calculate the DTW path at low resolution, constrain the high-level search range, and gradually refine to the original resolution. The calculation formula is as follows; Among them, lon i ' represents the average value of the i-th longitude, lat i ' represents the average value of the i-th latitude, lon 2i-1 and lon 2i respectively represent the longitudes of the (2i - 1)-th and 2i-th trajectory points, lat 2i-1 and lat 2i respectively represent the latitudes of the (2i - 1)-th and 2i-th trajectory points, i represents the index in the sequence, from 1 to n represents the total number of trajectory points; Step 2.7: For each pair of training trajectory time series T i and T j , use the FastDTW algorithm to calculate D TW (i, j), and fill it into the corresponding initialized distance matrix D Matrix at the same time. The calculation formula is: Among them, D TW (i,j) represents the minimum cumulative alignment distance between two time series T i and T j at positions i and j, Dist(i,j) represents the Euclidean distance between trajectories i and j, D TW (i - 1,j - 1) represents the cumulative distance of aligning diagonally from the (i - 1)-th point of T i and the (j - 1)-th point of T j , D TW (i,j - 1) represents the cumulative distance of aligning horizontally from the i-th point of T i and the (j - 1)-th point of T j , D TW (i - 1,j) represents the cumulative distance of aligning vertically from the (i - 1)-th point of T i and the j-th point of T j .
4. A taxi route recommendation method based on the FastDTW time series clustering algorithm according to claim 1, characterized in that The specific method of the said Step 3 is: Step 3.1: Apply the distance matrix D Matrix to the hierarchical clustering algorithm, construct a hierarchical clustering tree using the average linkage method, calculate the average distance of all trajectory pairs between two clusters as the inter-cluster distance. For two clusters C1 and C2, the formula for calculating the inter-cluster distance is as follows; where d avg (C1, C2) represents the average distance between clusters C1 and C2, and this value is used to measure the similarity between two clusters. |C1| and |C2| represent the number of trajectories in clusters C1 and C2 respectively, T i and T j represent two training trajectories, D Matrix (i)(j) represents the element value in the i-th row and j-th column of the distance matrix D Matrix This matrix stores the distance between all element pairs; Step 3.2: When the average distance d avg (C1, C2) between two clusters is less than or equal to t, merge the two clusters, where t is a set distance threshold; Step 3.3: Obtain the cluster labels of the M training trajectories from the results of hierarchical clustering. Each trajectory is assigned to a specific cluster C to form the classification results of the M training trajectories; Step 3.4: Define a counting matrix B Matrix to store the number of occurrences of each cluster, and define an accumulated distance matrix R Matrix to accumulate the total distance of each cluster; Step 3.5: For each test trajectory T test , calculate the FastDTW distance D between the N test trajectories and the M training trajectories test,i , and accumulate the FastDTW distance D of each trajectory test,i to the total distance R of the corresponding cluster Matrix , and update the count matrix B Matrix with the number of the corresponding cluster in it: Step 3.6: For each cluster C i , calculate its average distance Avg_dis i , and the calculation formula is as follows: Among them, R i is the total accumulated distance of cluster C i , and B Matrix [i] is the total number of trajectories of cluster C i . Step 3.7: Select the cluster C with the smallest average distance from all clusters min , and extract from it all the predicted end destination coordinate sets E = {e1, e2,... e L}, where e1, e2,... e L represent the addresses of multiple predicted destinations in the set E. Combine the end point E and the start point P to form a candidate route list Route_list = {(P → e1), (P → e2),..., (P → e L )}, where L is the total number of candidate routes.
5. A taxi route recommendation method based on the FastDTW time series clustering algorithm according to claim 1, characterized in that, The specific method of the said Step 4 is: Step 4.1: Call the Gaode Map API to obtain the key information of L candidate routes, including the driving cost C = {c1, c2,..., c L}, the driving time T = {t1, t2,..., t L}, the driving distance D = {d1, d2,..., d L}, and the traffic status of each section, including smooth, slow, congested, severely congested, unknown; Step 4.2: According to the traffic conditions of each road section, use the weighted average method to calculate the average smoothness degree S = [S1, S2,..., S i ..., S L , and the calculation formula is as follows: Among them, A j is the weight set for the j-th sub-section according to its traffic condition. The weight for smooth traffic is set as a1, the weight for slow traffic is a2, the weight for congestion is a3, the weight for severe congestion is a4, and the weight for unknown is a5. O i is the total number of sub-sections of the i-th candidate route; Step 4.3: Perform min-max normalization on the driving cost C, driving time T, and driving distance D to obtain the normalized driving cost C norm , driving time T norm , driving distance D norm ; Step 4.4: Adjust the weights of driving cost, driving time, driving distance, and traffic smoothness to reflect their importance for the driver to select a route; Step 4.5: Use the weighted summation method to calculate the comprehensive scores Scores of the L candidate routes. The calculation formula is as follows: Scores = w c *C norm +w t *T norm +w d *D norm +w s *S w c +w t +w d +w s = 1 Among them, w c , w t , w d , w s are the weights of driving costs, driving time, driving distance, and smoothness respectively; Step 4.6: Sort all the candidate routes according to the comprehensive scores Scores, find the route with the highest score, and at the same time perform visual display on the map, and the results are for the driver's reference.
Citation Information
Cited By
Multi-target trajectory prediction method and system for security unmanned aerial vehicle
CN121030246A