Vehicle Travel Time Prediction Method for Trip Link Extraction
By using the Northern Goshawk optimization algorithm and Attention-ConvLSTM model in vehicle travel time prediction, combining the convolutional long and short-term memory network and attention mechanism, the traffic theme link is extracted and travel time prediction is carried out, the problem of insufficient calculation time and prediction accuracy in the existing technology is solved, and more efficient and accurate prediction effects are achieved.
Patent Information
- Application Number
- CN202410289145.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-14
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-03-14
AI Technical Summary
The existing vehicle travel time prediction methods still have room for improvement in calculation time and prediction accuracy, and it is difficult to meet the practical application needs.
The Attention-ConvLSTM model based on the Northern Goshawk optimization algorithm is adopted, combining the convolutional long short-term memory network and attention mechanism to extract traffic theme links and make it time prediction.
The prediction accuracy of the vehicle travel link travel time is improved, the calculation time is shortened, the generalization ability and prediction performance of the model are significantly improved, and the prediction accuracy can reach more than 90%.
Smart Images

Figure CN119204277B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of vehicle travel time prediction, and particularly to a vehicle travel time prediction method for travel link extraction. Background Art
[0002] With the rapid development of China's comprehensive national strength and national economy, the living consumption level and ability of residents have been continuously improved, and the number of automobiles has also increased sharply. Vehicle travel link travel time prediction, as an important part of the intelligent transportation system, is of great significance for alleviating traffic congestion, improving road utilization rate, and optimizing urban traffic planning. In recent years, a large number of studies have been dedicated to exploring more accurate and efficient vehicle travel link travel time prediction methods in order to provide strong support for real traffic problems. The existing travel time prediction research methods mainly include parametric models and machine learning models.
[0003] In recent years, many scholars have used different methods to prove the applicability and prediction accuracy of various models in travel time prediction. At the present stage, the machine learning method is mainly used for the vehicle travel time prediction model, mainly neural network, Kalman filter, K-nearest neighbor, etc. Among them, the inheritance algorithm models mainly include random forest, gradient boosting regression tree, gradient boosting decision tree, etc. The tree ensemble algorithm has the advantages of fast speed, high prediction accuracy, and low requirements for data quality compared with traditional machine learning. However, the calculation time of these methods can still be improved, and the accuracy of predicting travel time can still be further improved. Summary of the Invention
[0004] The present invention provides a vehicle travel time prediction method for travel link extraction, which can overcome certain or some defects of the prior art.
[0005] According to a vehicle travel time prediction method for travel link extraction of the present invention, it includes the following steps:
[0006] Step S1: Obtain the trajectory data of the taxi, and then extract the traffic theme link and its classification based on similarity;
[0007] Step S2: Adopt a vehicle travel link travel time prediction model of Attention-ConvLSTM based on the Northern Goshawk optimization algorithm;
[0008] Step S3: Extract the results of the traffic theme link and obtain the travel time of the predicted vehicle travel link.
[0009] Preferably, the taxi trajectory (Traj for short) data in step S1 is the data recording the position information of the taxi within a certain period of time, which includes the empty trajectory points and the passenger-carrying trajectory points of the taxi;
[0010] There are several trajectories in the taxi trajectory dataset, that is
[0011] Traj = {Traj 1 , Traj 2 ,..., Traj n} (1)
[0012] In formula (1), Traj i represents the i-th trajectory.
[0013] A trajectory Traj consists of multiple trajectory points, x 1 , x 2 ,..., x n . Among them, a trajectory point x i is the position data point collected by the GPS device during the taxi's driving. These data points include vehicle identification (id), longitude (lng), latitude (lat), time (time), passenger-carrying status (flag), etc., and are expressed as: x i = (id i , lng i , lat i , time i , flag i ).
[0014] According to the driving trajectory of the taxi and the passenger-carrying status of the taxi, judge and extract the pick-up and drop-off points of the taxi, and extract the link information between the pick-up and drop-off points. It is necessary to construct the driving trajectory of the taxi based on the timestamp and position information in the GPS data, and extract the links between each trajectory point. The specific steps are as follows:
[0015] When the passenger-carrying status flag of the taxi changes from 0 to 1, it means that the taxi will start carrying passengers, and this piece of information is used as the pick-up point of the taxi.
[0016] When the passenger-carrying status flag of the taxi changes from 1 to 0, it means that the taxi has finished carrying passengers, and this piece of information is used as the drop-off point of the taxi.
[0017] According to the preprocessed GPS data, use formula (2) to analyze the spatial relationship between sampling points to estimate the position of the unknown point Yu, so as to construct the driving trajectory Traj of the taxi.
[0018] Extract the links between adjacent trajectory points on the driving trajectory Traj. Set a time threshold t. If the time interval between two trajectory points is less than the threshold t, it is considered that there is a link between these two trajectory points. The algorithm flow chart is as Figure 1 shown.
[0019]
[0020] In the formula, Y u is the value of the unknown point, and Y i is the value of the known point, X i is the weight for the known point, and r i is the distance between the known point and the unknown point.
[0021] Preferably, the method for extracting traffic theme links based on similarity in step S1 includes:
[0022] The trajectory similarity is an index used to measure the similarity degree between two trajectories. An optimal trajectory is found between the starting and ending points according to the similarity and used as the traffic theme link. When a taxi selects to drive along the traffic theme link, it can reach the destination as soon as possible.
[0023] First, preprocess the input trajectory coordinates. Convert the S longitude and latitude coordinates to tangent plane coordinates (x, y, z), and then convert the tangent plane coordinates (x, y, z) to coordinate data (X, Y, Z) in the Cartesian coordinate system. Convert the longitude to radian system and divide it by the radius of the earth to obtain a value less than 1 with the unit of radian, and directly convert the latitude to radian system;
[0024] Next, calculate the longitude and latitude distance between each point on the two trajectories. For points i and j on each trajectory, the Euclidean distance formula can be used, that is
[0025]
[0026] In the formula, x i and y i are the Cartesian coordinates of point i, and x j and y j are the Cartesian coordinates of point j;
[0027] Then calculate the contribution degree of each point to the entire trajectory. Calculate the similarity between the two trajectories according to the contribution degree, perform weighted averaging on these distances, and then calculate the similarity between the two trajectories.
[0028] Finally, output the similarity between the two trajectories. The closer the similarity is to 1, the more similar the two trajectories are. Consider the trajectories with similar similarities as one trajectory. Traverse all the links, count the number of times each travel link between the pick-up and drop-off points is traveled, and extract the travel links with the number of trips greater than the threshold;
[0029] The formula for calculating the contribution degree of a trajectory point to the entire trajectory:
[0030]
[0031] The formula for calculating the similarity between two trajectories:
[0032]
[0033] Wherein, n is the number of points on the trajectory, and d(i, j) is the distance between trajectory points i and j;
[0034] First, determine the partition of OD points, extract the OD points with high travel demand, extract the paths between OD partitions according to the similarity of trajectories, and extract the traffic theme links with the most trips between each OD partition;
[0035] Step1: Calculate the distances from points O and D to the center points of each cluster respectively, and match the numbers s_number and e_number of the cluster centers closest to points O and D to confirm the partition of OD points.
[0036] Step2: Group according to s_number and e_number of OD points, count the number count of OD point partitions, select the OD partitions with the number count greater than 200, and count that there are hang rows in the partitions.
[0037] Step3: Extract all OD travel links in each partition, and create tabular data guiji_ODzhuti for storing the theme links of all partitions.
[0038] Step3.1: h = 0. When h < hang, take the value of the first column in the h-th partition as a, and the value of the second column as b, make s_number = a and e_number = b in the OD partition, extract the boarding time stime, alighting time etime and license plate ID of each OD point in this partition, and count that there are hang_fenqu rows in this partition.
[0039] Step3.2: Loop through each row i of this partition, make open_time equal to the stime of the i-th row, make close_time equal to the etime of the i-th row, and C equal to the license plate ID of the i-th row. Create a list list, convert the trajectory data of the pick-up and drop-off points into an array, traverse and loop through each trajectory data item, take out the license plate number sh in the array. If C = sh, add item to list and convert list into tabular data.
[0040] Step3.3: Extract the data in the partition whose trajectory data time is greater than open_time and less than close_time, add each trajectory data in the partition to guiji_fenqu, initially i = 0, and create a list Slist for counting the number of trips of the current path.
[0041] Step 4: In the OD trajectories of the current partition, the similarity between two trajectories is calculated, and the trajectories with close similarity are regarded as one trajectory.
[0042] Step 4.1: Repeat steps 3.2-3.3, take out the trajectory data guiji_od1 of the i-th OD point, make r=i+1, count the number of times of the current route s=1, when r is less than hang_fenqu, repeat steps 3.2-3.3, take out the trajectory data guiji_od2 of the r-th OD point.
[0043] Step 4.2: Find each longitude and latitude (X j ,Y j )The latitude and longitude point closest to the latitude and longitude in guiji_od1 (X i ,X i ), the distance between the two coordinate points is calculated as dist according to the distance formula (3), and the sum of squares of each distance dist is calculated according to the formula (2), and then the square root is taken to obtain sqrt. If sqrt is less than the threshold, s=s+1.
[0044] Step 4.3: r = r + 1, repeat steps 4.1-4.2, and calculate the similarity between the i-th OD trajectory data and the r-th OD trajectory data. When r is greater than hang_fenqu, this loop ends and the value of s is added to Slist.
[0045] Step 5: i=i+1, repeat step 4 until i is greater than hang_fenqu, the loop ends.
[0046] Step 6: Extract the OD path with the most travel demand in the current OD partition, list the largest values in the list Slist, and store them in the list res, and create the table data guiji_ODzhuti1 that stores the current partition topic link.
[0047] Step 6.1: Take out the value b in res one by one, take out the value of stime in row b in the partition data, which is Otime, the value of etime in row b is Ctime, and the license plate ID in row b is C. Create a list blist so that the license plate number in the trajectory data is equal to C, and the trajectory data with time between Otime and Ctime is guiji0.
[0048] Step 6.2: Let a=b+1, repeat step 6.1, and get the trajectory data of the ath OD point guiji1. Take the square root of the sum of the longitude and latitude points of guiji0 and guiji1 and get sqrt. If sqrt is less than the threshold, add guiji0 and guiji1 to the table data guiji_ODzhuti1.
[0049] Step 6.3: a = a + 1, loop step 6.2 until a is greater than hang_fenqu, then the loop ends. Add the theme link trajectory guiji_ODzhuti1 of the current partition to guiji_Odzhuti.
[0050] Step 7: Extract traffic theme links for the next partition, set h = h + 1, and repeat steps 3.1 - 6.3 until h is greater than hang.
[0051] Preferably, in step S1, for the travel link classification based on travel purposes, the classification method is as follows:
[0052] The purpose of travel link classification is to classify links according to different attribute characteristics or uses, and distinguish travel links through the attribute characteristics or uses of their categories.
[0053] The present invention classifies traffic travel links according to the points of interest at the pick-up and drop-off points and the POI types.
[0054] First, extract the pick-up and drop-off points from the taxi GPS data, and perform clustering on the pick-up and drop-off points to obtain the pick-up hotspots.
[0055] Next, crawl the POI data within a fixed radius centered on the pick-up hotspots, select the POI type with the largest proportion as the type of the pick-up hotspots, and use the pick-up hotspot area as the starting and ending points of the travel link.
[0056] Divide the types of the starting and ending points into four types: residential, office, life service, and school.
[0057] From the actual travel situation of travelers, it can be seen that the number of passengers is the largest during the morning and evening rush hours. Because during the morning rush hour, people have a large travel demand for going to work, school, etc., and during the evening rush hour, it is a time when the travel demand for getting off work, shopping, etc. is relatively high.
[0058] Therefore, classify the travel links according to the attribute characteristics of the starting and ending points into links from residential to company, residential to school, residential to life service, from school to company, school to life service, and their round-trip links.
[0059] Preferably, the model in step S2 includes:
[0060] An input layer, which is responsible for converting the received original sequence data into a form readable by the model;
[0061] A convolutional layer, which is used to extract local features from the input data. The convolutional layer can contain multiple convolutional kernels, and each convolutional kernel corresponds to a different feature;
[0062] The LSTM layer, which has multiple LSTM cells, and each LSTM cell consists of an input gate, a forget gate, and an output gate, is used to capture long-term dependencies in the input data;
[0063] The Northern Goshawk Optimization Algorithm, which is used to optimize the parameters of the ConvLSTM layer to improve the prediction performance of the model;
[0064] The Attention mechanism, which is used to weight information at different time steps to help the model focus on more important information during prediction and improve the prediction accuracy;
[0065] The fully connected layer, which is used to perform a fully connected process on the output of the Attention-ConvLSTM to obtain the final prediction result;
[0066] The output layer, which is used to transform the prediction result of the fully connected layer into a certain value in the time series data.
[0067] Preferably, the Northern Goshawk Optimization Algorithm is mainly divided into two stages: the global search stage and the local search stage;
[0068] Global search stage: In this stage, the Northern Goshawk is used to randomly select a prey in the search space and then quickly attack it;
[0069] This process can be described by a mathematical model. For example, Equation (6) is for position update and Equation (7) is for velocity update, where the prey represents the objective function in the optimization problem, and the Northern Goshawk needs to find the solution that minimizes the objective function in the search space.
[0070] x t+1 = x t - a * (b - x t ) (6)
[0071] v t+1 = v t - c * (w - v t ) (7)
[0072] In the equations, a and c are acceleration coefficients, b is the prey position, w is the goshawk position, xt is the current goshawk position, and vt is the current goshawk velocity.
[0073] 1) Local search phase: After successfully attacking the prey, the northern goshawk will enter the local search phase. This phase can be regarded as a process in which the northern goshawk tracks and approaches the prey when the position of the prey is known. During this process, the northern goshawk will continuously update its search strategy to better approach the prey. If the simulated annealing method is adopted, a temperature parameter T can be set. In each iteration, the goshawk will decide whether to accept the new position according to the Metropolis criterion. The formula can be described as:
[0074]
[0075] In the formula, n is the number of samples in the validation set, y _i is the actual travel time, y _pred_i is the travel time predicted by the model, MSE is the mean square error between the predicted travel time and the true value of the Attention-ConvLSTM model, x new and x old are the new and old positions, and T is the temperature.
[0076] Updating the solution: After completing the local search, the northern goshawk will update its solution. This process involves attacking and pursuing the prey, and adjusting the search strategy according to the movement trajectory of the prey. Through this process, the northern goshawk can find a better solution in the search space.
[0077] Repeating the above process: The northern goshawk will continuously repeat the processes of global search, local search, and updating the solution until an optimal solution that meets the conditions is found.
[0078] Preferably, the convolutional long short-term memory network (ConvLSTM) is a deep learning model that combines the convolutional neural network (CNN) and the long short-term memory network (LSTM) and is used to process time series data. The traditional LSTM does not have a convolutional layer. The ConvLSTM adds a convolutional layer, which can effectively extract the spatial features in the input data and has strong capabilities in processing time series data and spatial data. The LSTM can only process sequences of fixed length, while the convolutional layer in the ConvLSTM can process sequence data of variable length. The structural diagrams of the convolutional layer and the LSTM layer are as Figure 3 shown, and the formula is:
[0079] i t = σ(W xi * h t-1 + b i ) (10)
[0080] f t = σ(W fh * h t-1 + b f ) (11)
[0081] o t = σ(W ho *h t + b o )(12)
[0082] h t = σ(W ch *x t + b c )(13)
[0083] Where: it is the output of the input gate, ft is the output of the forget gate, ot is the output of the output gate, ht is the output of the convolutional layer, Wxi, Wfh, Who, Wch are the weight matrices of the input gate, forget gate, output gate, and convolutional layer respectively, ht is the hidden state at the current time, ht-1 is the hidden state at the previous time, bi, bf, bo, bc are the bias vectors of the input gate, forget gate, output gate, and convolutional layer respectively, σ is the sigmoid activation function, and * is the convolution operation.
[0084] Preferably, the attention mechanism (Attention) is a strategy widely used in deep learning models to dynamically adjust the information weights at different positions when processing the input sequence. The core idea of the attention mechanism is to endow the model with the ability to automatically learn and focus on the importance of different positions in the input sequence, so as to better capture the key information related to the task. Its structure is as Figure 4 shown.
[0085] First, use a convolutional neural network (CNN) to extract features from the travel link to obtain convolutional features. Then, input the convolutional features into a long short-term memory network (LSTM) for time series analysis. Introduce the attention mechanism into the ConvLSTM model to enable the model to focus on the features of the input data that are useful for prediction and weight different spatial positions and feature channels;
[0086] In the attention mechanism, take the output sequence [h t-s+1 ,..., h t-1 , h t of each ConvLSTM as the input sequence of the Attention mechanism. An attention weight [α t-s+1 ,..., α t-1 , α t will be calculated for each position of the input sequence. Finally, multiply these weights by the corresponding input values and then sum the product results to obtain a vector [H t-s+1 ,..., H t-1 , H t representing the entire input sequence. The formula is as (14)-(16).
[0087] During this process, the model can automatically learn to ignore unimportant or redundant information based on attention weights, thereby improving the performance and efficiency of the model. Combining the Attention mechanism of ConvLSTM to predict travel time for travel links can effectively capture local information and useful features in the prediction task, and weight the information at different time steps, thus improving the accuracy and reliability of the prediction.
[0088] c t = W s *tanh(W xa *X t + W ha *h t + b a ) (14)
[0089]
[0090] Where: W is the convolution kernel weight, X is the input feature, h is the input sequence of ConvLSTM, b a is the bias term, c t is the intermediate value for calculating the attention coefficient of each feature region, α t is the attention weight, and H is the output sequence of weighted summation.
[0091] The present invention also provides a vehicle travel time prediction device for travel link extraction, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the steps of the above method are implemented.
[0092] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed, the steps of the above method are implemented.
[0093] Beneficial effects:
[0094] Traditional vehicle travel time prediction methods often have high errors and are difficult to meet the actual application requirements. Using the taxi GPS data and the crawled POI data in Chengdu, the types of pick-up and drop-off point areas of travel links are determined according to the types of POIs, and the traffic theme links with more travel times are extracted according to the similarity. Aiming at the problem of link travel time prediction, a vehicle travel link travel time prediction method based on NGO-Attention-ConvLSTM is proposed. The attention mechanism is used to capture the traffic conditions at different time scales; then, the convolutional long short-term memory network is used to extract features and model the spatio-temporal relationship of the traffic state; in order to improve the generalization ability of the model, the northern goshawk optimization algorithm is introduced to optimize the parameters of the model. The research results show that: extracting traffic theme links according to the similarity algorithm has more advantages than traditional algorithms. Its calculation time is 86.345 s, which is about 190 s and 345 s less than other algorithms respectively; compared with the LSTM model, the proposed method improves the prediction accuracy of the travel time of travel links by 30.9% and 28.4% respectively in terms of the mean absolute percentage error and R 2 index, and the prediction accuracy can reach more than 90%. It has higher accuracy and stability in the task of predicting the travel time of vehicle travel links, provides a useful reference for the actual traffic management system, and helps to further improve the traffic operation efficiency. Description of the Drawings
[0095] Figure 1 is the flowchart of the travel link extraction algorithm in the present invention;
[0096] Figure 2 is the overall structure diagram of the NGO-Attention-ConvLSTM prediction model in the present invention;
[0097] Figure 3 is the structure diagram of the ConvLSTM model in the present invention;
[0098] Figure 4 is the structure diagram of the attention mechanism in the present invention;
[0099] Figure 5 is the travel link between an OD and the extracted traffic theme link diagram in the experimental verification step of Embodiment 1;
[0100] Figure 6 is the traffic theme link classification diagram with different travel demands in the experimental verification step of Embodiment 1;
[0101] Figure 7 is the convergence curve of the NGO-Attention-ConvLSTM model in the experimental verification step of Embodiment 1;
[0102] Figure 8It is a bar chart of attention coefficients at different time steps in the experimental verification step of Embodiment 1;
[0103] Figure 9 It is a comparison chart of travel link travel time prediction results in the experimental verification step of Embodiment 1;
[0104] Figure 10 It is a picture of the comparison table of test set accuracy in the experimental verification step of Embodiment 1; Detailed implementation manner
[0105] Embodiment 1
[0106] This embodiment provides a vehicle travel time prediction method for travel link extraction; the vehicle trajectory selected in this method is a taxi trajectory; trajectory data:
[0107] Taxi trajectory (trajectory, abbreviated as Traj) data is data that records the position information of a taxi within a certain period of time, and it includes the empty - load trajectory points and passenger - load trajectory points of the taxi. There are several trajectories in the taxi trajectory dataset, that is
[0108] Traj={Traj 1 ,Traj 2 ,...,Traj n} (1)
[0109] In the formula, Traj i represents the i - th trajectory.
[0110] A trajectory Traj consists of multiple trajectory points (x 1 ,x 2 ,...,x n ). Among them, a trajectory point x i is a position data point collected by the GPS device during the taxi's driving process. These data points include vehicle identification (id), longitude (lng), latitude (lat), time (time), passenger - load status (flag), etc., and are expressed as: x i =(id i ,lng i ,lat i ,time i ,flag i ).
[0111] According to the driving trajectory of the taxi and the passenger - load status of the taxi, judge and extract the pick - up and drop - off points of the taxi, and extract the link information between the pick - up and drop - off points. It is necessary to construct the driving trajectory of the taxi based on the timestamp and position information in the GPS data, and extract the links between each trajectory point. The specific steps are as follows:
[0112] 1) When the occupancy status flag of the taxi changes from 0 to 1, it indicates that the taxi is about to pick up passengers, and this piece of information is used as the pick-up point of the taxi.
[0113] 2) When the occupancy status flag of the taxi changes from 1 to 0, it indicates that the taxi has finished picking up passengers, and this piece of information is used as the drop-off point of the taxi.
[0114] 3) According to the preprocessed GPS data, use formula (2) to analyze the spatial relationship between sampling points to estimate the position of the unknown point Yu, thereby constructing the driving trajectory Traj of the taxi.
[0115] 4) Extract the links between adjacent trajectory points on the driving trajectory Traj. Set a time threshold t. If the time interval between two trajectory points is less than the threshold t, it is considered that there is a link between these two trajectory points. The algorithm flow chart is as Figure 1 shown.
[0116]
[0117] where Y u is the value of the unknown point, Y i is the value of the known point, X i is the weight of the known point, and r i is the distance between the known point and the unknown point.
[0118] Extract traffic theme links and their classifications:
[0119] The specific process of classifying travel links based on travel purposes is as follows:
[0120] The purpose of classifying travel links is to classify the links according to different attribute characteristics or uses, and distinguish the travel links through the attribute characteristics or uses of their categories. In this embodiment, the traffic travel links will be classified according to the type of point of interest (POI) of the pick-up and drop-off points.
[0121] First, extract the pick-up and drop-off points from the taxi GPS data, cluster the pick-up and drop-off points to obtain passenger-carrying hotspots. Then, crawl the POI data within a fixed radius centered on the passenger-carrying hotspots, select the POI type with the largest proportion as the type of the passenger-carrying hotspots, use the passenger-carrying hotspot area as the starting and ending points of the travel link, and divide the types of the starting and ending points into four types: residential, office, life service, and school.
[0122] From the actual travel situations of travelers, it can be seen that the number of passengers is the largest during the morning and evening rush hours. Because during the morning rush hour, people have high travel demands such as going to work and school, and during the evening rush hour, the travel demands for getting off work and shopping are relatively high. Therefore, the travel links are divided into those from residence to company, residence to school, residence to life services, from school to company, school to life services, and their round-trip links according to the attribute characteristics of the origin and destination.
[0123] The specific process of extracting traffic theme links is as follows:
[0124] Trajectory similarity is an index used to measure the similarity degree between two trajectories. An optimal trajectory is found between the origin and destination according to the similarity, which is used as the traffic theme link. Taxis choosing to travel on the traffic theme link can enable them to reach the destination as soon as possible.
[0125] First, preprocess the input trajectory coordinates. Convert the S longitude and latitude coordinates to tangent plane coordinates (x, y, z), and then convert the tangent plane coordinates (x, y, z) to coordinate data (X, Y, Z) in the Cartesian coordinate system. Convert the longitude to radian system and divide it by the radius of the earth to obtain a value less than 1 with the unit of radian, and directly convert the latitude to radian system. Next, calculate the longitude and latitude distances between each point on the two trajectories. For points i and j on each trajectory, the Euclidean distance formula can be used, that is
[0126]
[0127] where x i and y i are the Cartesian coordinates of point i, and x j and y j are the Cartesian coordinates of point j.
[0128] Then calculate the contribution degree of each point to the entire trajectory. Calculate the similarity between the two trajectories according to the contribution degree, perform weighted averaging on these distances, and then calculate the similarity between the two trajectories. Finally, output the similarity between the two trajectories. The closer the similarity is to 1, the more similar the two trajectories are. Trajectories with similar similarities are regarded as one trajectory. Traverse all links, count the number of times each travel link is traveled between the pick-up and drop-off points, and extract the travel links with the number of trips greater than the threshold.
[0129] The formula for calculating the contribution degree of a trajectory point to the entire trajectory:
[0130]
[0131] The formula for calculating the similarity between two trajectories:
[0132]
[0133] Where n is the number of points on the trajectory, and d(i, j) is the distance between trajectory points i and j.
[0134] First, determine the partition of OD points, extract the OD points with high travel demand, extract the paths between OD partitions according to the similarity of trajectories, and extract the traffic theme links with the most trips between each pair of OD partitions.
[0135] The specific steps are as follows:
[0136] Step1: Calculate the distances from points O and D to the center points of each cluster respectively, and match the numbers s_number and e_number of the cluster center points closest to points O and D to confirm the partition of OD points.
[0137] Step2: Group according to s_number and e_number of OD points, count the number of OD point partitions count, select the OD partitions with count greater than 200, and count that there are hang rows in the partition.
[0138] Step3: Extract all OD travel links in each partition and create tabular data guiji_ODzhuti for storing the theme links of all partitions.
[0139] Step3.1: Let h = 0. When h < hang, take the value of the first column in the h-th partition as a and the value of the second column as b, make s_number = a and e_number = b in the OD partition, extract the boarding time stime, alighting time etime and license plate ID of each OD point in this partition, and count that there are hang_fenqu rows in this partition.
[0140] Step3.2: Loop through each row i in this partition, make open_time equal to the stime of the i-th row, make close_time equal to the etime of the i-th row, and C equal to the license plate ID of the i-th row. Create a list list, convert the trajectory data of the pick-up and drop-off points into an array, traverse each trajectory data item in the loop, take out the license plate number sh in the array, if C = sh, add item to list, and convert list into tabular data.
[0141] Step3.3: Extract the data in the partition whose trajectory data time is greater than open_time and less than close_time, add each trajectory data in the partition to guiji_fenqu, initially i = 0, and create a list Slist for counting the number of trips of the current path.
[0142] Step4: Calculate the similarity between every two trajectories in the OD trajectories of the current partition, and consider the trajectories with close similarity as one trajectory.
[0143] Step 4.1: Repeat steps 3.2-3.3, take out the trajectory data guiji_od1 of the i-th OD point, make r=i+1, count the number of times of the current route s=1, when r is less than hang_fenqu, repeat steps 3.2-3.3, take out the trajectory data guiji_od2 of the r-th OD point.
[0144] Step 4.2: Find each longitude and latitude (X j ,Y j )The latitude and longitude point closest to the latitude and longitude in guiji_od1 (X i ,X i ), the distance between the two coordinate points is calculated as dist according to the distance formula (3), and the sum of squares of each distance dist is calculated according to the formula (2), and then the square root is taken to obtain sqrt. If sqrt is less than the threshold, s=s+1.
[0145] Step 4.3: r = r + 1, repeat steps 4.1-4.2, and calculate the similarity between the i-th OD trajectory data and the r-th OD trajectory data. When r is greater than hang_fenqu, this loop ends and the value of s is added to Slist.
[0146] Step 5: i=i+1, repeat step 4 until i is greater than hang_fenqu, the loop ends.
[0147] Step 6: Extract the OD path with the most travel demand in the current OD partition, list the largest values in the list Slist, and store them in the list res, and create the table data guiji_ODzhuti1 that stores the current partition topic link.
[0148] Step 6.1: Take out the value b in res one by one, take out the value of stime in row b in the partition data, which is Otime, the value of etime in row b is Ctime, and the license plate ID in row b is C. Create a list blist so that the license plate number in the trajectory data is equal to C, and the trajectory data with time between Otime and Ctime is guiji0.
[0149] Step 6.2: Let a=b+1, repeat step 6.1, and get the trajectory data of the ath OD point guiji1. Take the square root of the sum of the longitude and latitude points of guiji0 and guiji1 and get sqrt. If sqrt is less than the threshold, add guiji0 and guiji1 to the table data guiji_ODzhuti1.
[0150] Step 6.3: a = a + 1, loop through Step 6.2 until a is greater than hang_fenqu, then end the loop. Add the theme link trajectory guiji_ODzhuti1 of the current partition to guiji_Odzhuti.
[0151] Step 7: Extract traffic theme links for the next partition, set h = h + 1, and repeat Steps 3.1 - 6.3 until h is greater than hang, then end.
[0152] Furthermore, the prediction process of the travel link travel time based on NGO - Attention - ConvLSTM is as follows:
[0153] The travel time series has characteristics such as complexity, volatility, and non - linearity. Using only a single model to predict the travel time of travel links will affect the accuracy of the prediction results. Considering that NGO has strong global search ability and local search ability, the Attention mechanism can automatically learn the features of input data and perform corresponding weight allocation, and ConvLSTM can effectively capture dependencies in space and time. In this embodiment, the NGO - Attention - ConvLSTM combined model is used to predict the travel time of travel links, which can more accurately predict the travel time of travel links and provide more effective travel suggestions for users.
[0154] The overall structure diagram of the NGO - Attention - ConvLSTM prediction model proposed in this embodiment is as shown in the accompanying drawings of the specification Figure 2 as shown.
[0155] The input layer of the vehicle travel link travel time prediction model based on NGO - Attention - ConvLSTM is the factor characteristic indicators that affect the travel time of urban taxis; the output layer is the predicted travel time of taxis. As Figure 2 can be seen, the model consists of 7 parts: The input layer is responsible for converting the received original sequence data into a form readable by the model;
[0156] The convolutional layer is used to extract local features from the input data. The convolutional layer can contain multiple convolutional kernels, and each convolutional kernel corresponds to a different feature; The LSTM layer has multiple LSTM units, and each LSTM unit consists of an input gate, a forget gate, and an output gate, which are used to capture long - term dependencies in the input data;
[0157] NGO is used to optimize the parameters of the ConvLSTM layer to improve the prediction performance of the model;
[0158] The Attention mechanism weights information at different time steps to help the model focus on more important information during prediction, improving the accuracy of prediction; the fully connected layer performs a fully connected process on the output of Attention-ConvLSTM to obtain the final prediction result; the output layer transforms the prediction result of the fully connected layer into a certain value in the time series data.
[0159] The Northern Goshawk Optimization Algorithm is specifically described as follows:
[0160] The Northern Goshawk Optimization Algorithm (NGO) is an optimization algorithm that simulates the hunting behavior of northern goshawks in nature and was proposed by Mohammad Dehghani et al. in 2022. The basic idea of this algorithm is to imitate the behavior of goshawks during the hunting process, including prey recognition and attack, pursuit, and prey escape and re-pursuit. The main purpose of the algorithm is to find the optimal solution in the search space to solve the optimization problem, which is mainly divided into two stages: the global search stage and the local search stage.
[0161] 2) Global search stage: In this stage, the northern goshawk randomly selects a prey in the search space and then quickly attacks it. This process can be described by a mathematical model. For example, Equation (6) is for position update and Equation (7) is for velocity update, where the prey represents the objective function in the optimization problem, and the northern goshawk needs to find the solution that minimizes the objective function in the search space.
[0162] x t+1 = x t - a * (b - x t ) (6)
[0163] v t+1 = v t - c * (w - v t ) (7)
[0164] In the formula, a and c are acceleration coefficients, b is the prey position, w is the goshawk position, xt is the current goshawk position, and vt is the current goshawk velocity.
[0165] 3) Local search stage: After successfully attacking the prey, the northern goshawk will enter the local search stage. This stage can be regarded as a process in which the northern goshawk tracks and approaches the prey when the prey position is known.
[0166] 4) In this process, the northern goshawk will continuously update its search strategy to better approach the prey. If the simulated annealing method is adopted, a temperature parameter T can be set. In each iteration, the goshawk will decide whether to accept the new position according to the Metropolis criterion. The formula can be described as:
[0167]
[0168] In the formula, n is the number of samples in the validation set, y _i is the actual travel time, y _pred_i is the travel time predicted by the model, and MSE is the mean square error between the predicted travel time and the true value by the Attention-ConvLSTM model, x new and x old are the new and old positions, and T is the temperature.
[0169] 5) Update the solution: After completing the local search, the northern goshawk will update its own solution. This process involves attacking and pursuing the prey, as well as adjusting the search strategy according to the prey's movement trajectory. Through this process, the northern goshawk can find a better solution in the search space.
[0170] 6) Repeat the above process: The northern goshawk will continuously repeat the processes of global search, local search, and updating the solution until an optimal solution that meets the conditions is found.
[0171] This embodiment also combines a convolutional long short-term memory network model. Specifically, the convolutional long short-term memory (ConvLSTM) is a deep learning model that combines a convolutional neural network (CNN) and a long short-term memory network (LSTM) for processing time series data. The traditional LSTM does not have a convolutional layer. A convolutional layer is added in ConvLSTM, which can effectively extract the spatial features in the input data and has strong capabilities in processing time series data and spatial data. LSTM can only process sequences of fixed length, while the convolutional layer in ConvLSTM can process sequence data of variable length. Its structure diagram is as Figure 3 shown, and the formula is:
[0172] i t = σ(W xi *h t-1 + b i ) (10)
[0173] f t = σ(W fh *h t-1 + b f ) (11)
[0174] o t = σ(W ho *h t + b o ) (12)
[0175] h t = σ(W ch *x t+b c ) (13)
[0176] where: i t is the output of the input gate, f t is the output of the forget gate, o t is the output of the output gate, h t is the output of the convolutional layer, W xi 、W fh 、W ho 、W ch are the weight matrices of the input gate, forget gate, output gate, and convolutional layer respectively, h t is the hidden state at the current time step, h t-1 is the hidden state at the previous time step, b i 、b f 、b o 、b c are the bias vectors of the input gate, forget gate, output gate, and convolutional layer respectively, σ is the sigmoid activation function, and * is the convolution operation.
[0177] Furthermore, this embodiment combines the attention mechanism of convolutional long short-term memory;
[0178] Specifically, the attention mechanism (Attention) is a strategy widely used in deep learning models to dynamically adjust the information weights at different positions when processing input sequences. The core idea of the attention mechanism is to endow the model with the ability to automatically learn and focus on the importance of different positions in the input sequence, so as to better capture the key information related to the task. Its structure is as Figure 4 shown.
[0179] First, use a convolutional neural network (CNN) to extract features from the travel link to obtain convolutional features. Then, input the convolutional features into a long short-term memory network (LSTM) for time series analysis. An attention mechanism is introduced into the ConvLSTM model to make the model focus on the features of the input data that are useful for prediction and weight different spatial positions and feature channels.
[0180] In the attention mechanism, the output sequence [h t-s+1 ,..., h t-1 , h t of each ConvLSTM is used as the input sequence of the Attention mechanism. An attention weight [α t-s+1 ,..., α t-1 , α t is calculated for each position of the input sequence. Finally, these weights are multiplied by the corresponding input values, and the product results are summed to obtain a vector [H t-s+1,..., H t-1 , H t , and the formulas are as shown in (13)-(15). In this process, the model can automatically learn to ignore unimportant or redundant information according to the attention weights, thereby improving the performance and efficiency of the model. Combining the Attention mechanism of ConvLSTM to predict the travel time of the travel link can effectively capture local information and useful features in the prediction task, and weight the information at different time steps, thereby improving the accuracy and reliability of the prediction.
[0181] c t = W s *tanh(W xa *X t + W ha *h t + b a ) (14)
[0182]
[0183] Where: W is the convolution kernel weight, X is the input feature, h is the input sequence of ConvLSTM, b a is the bias term, c t is the intermediate value for calculating the attention coefficient of each feature region, α t is the attention weight, and H is the output sequence of weighted summation.
[0184] Based on the foregoing, the example verification process in this embodiment is as follows:
[0185] Result analysis of extracting traffic theme links:
[0186] First, extract traffic theme links. The data selected in this embodiment is the GPS trajectory data generated by more than 14,000 taxis in Chengdu from August 3 to August 30, 2014. The trajectory data is stored in CSV format and contains information such as vehicle ID, longitude, latitude, passenger-carrying status, and time. Cluster the pick-up and drop-off point data to obtain passenger-carrying hot spots, divide Chengdu into 41 regions, and determine the types of the 41 regions as one of the four types: residential, office, life service, and school through the type of POI.
[0187] On Monday, August 4th, approximately 89,700 vehicle trajectory data and about 7.742 million trajectory points were selected. Travel links with more than 500 origin-destination partitions in passenger-carrying hotspots were chosen, and the longitude and latitude data were converted into Euclidean distances. To determine an appropriate similarity threshold, two trajectories with the same origin and destination and the same taxi travel link were selected. The similarity thresholds were set at 0.5, 0.6, and 0.7 respectively. The similarity of the two route trajectories was tested to be 0.82, 0.73, and 0.65. By comparing the similarity results under different similarity thresholds, it was found that when the similarity threshold was 0.5, the similarity of the two route trajectories was 0.82, which was higher than the similarities under the other two thresholds.
[0188] Therefore, the similarity threshold was set at 0.5 as the basis for determining whether two trajectories were similar. Taking an OD pair with the highest travel demand on Monday as an example, there were multiple paths from the origin to the destination. One path with the most taxi trips was selected as the traffic theme link, as Figure 5 shown. In Figure (a), all taxi travel links on each path between the OD pair are shown. In Figure (b), the path with the most taxi trips in one path of this OD pair is 17, and this path is used as the traffic theme link between the origin and destination.
[0189] According to the above method, 808 traffic theme links were extracted. Among them, as Figure 6 shown, there were 103 red traffic theme links with residential areas and companies as the origin and destination, 262 green traffic theme links with residential areas and schools as the origin and destination, 99 blue traffic theme links with residences and living service areas as the origin and destination, 133 purple traffic theme links with residences and residences as the origin and destination, 92 black traffic theme links with schools and companies as the origin and destination, and 119 brown traffic theme links with schools and life as the origin and destination.
[0190] Subsequently, the calculation efficiency analysis of traffic theme links was extracted;
[0191] In this embodiment, the calculation efficiency comparison analysis will be carried out according to the similarity algorithm, the graph density peak clustering algorithm of Zhou Yang, and Wang Shaofan's method of inferring travel links based on the distinction of travel endpoints using mobile phone GPS data. The results are shown in Table 1. The similarity algorithm proposed in this embodiment can effectively reduce the calculation complexity when dealing with high-dimensional data and large-scale data sets. The similarity algorithm shows higher calculation efficiency and has an advantage in the calculation efficiency of extracting traffic theme links. This is because the similarity algorithm can find the correlation between data points faster when dealing with complex data, and compared with other algorithms, the calculation complexity of the similarity algorithm is more controllable.
[0192] Table 1 Comparison of calculation efficiency between similarity algorithm and traditional algorithms
[0193] Algorithm Name Computation Time / s Graph Density Peak Clustering Algorithm 376.1532 Infer Travel Links Based on Stop Points to Distinguish Travel Endpoints 431.2317 Similarity Algorithm 86.345
[0194] Analysis of Travel Time Prediction Results
[0195] Experimental Data and Parameter Settings
[0196] The morning, noon, and evening peak hours are periods with relatively high travel demands of urban residents. The travel purposes and routes of travelers are relatively fixed. The taxi GPS data during peak hours is more representative than that during other periods. Therefore, only the taxi trajectory data for three peak hours in Chengdu are selected: 7:00 - 9:00 in the morning peak, 12:00 - 14:00 at noon, and 19:00 - 21:00 in the evening peak. The time interval is selected as 15 minutes, that is, each daily dataset includes 24 time intervals. For the current time, the average travel times for the previous 15 minutes, 30 minutes, 40 minutes, and 60 minutes of each road section are calculated. The types of destinations are four categories: residential, office, life service, and school. 70% of the data in the dataset is selected for training the model, and 30% of the data is used for testing.
[0197] Convert the input features into numerical data and perform normalization processing; use a convolutional neural network to extract features, set the convolutional kernel size to 3*3, and the number of output channels to 16; the pooling kernel size is 2*2; the hidden layer dimension of the LSTM is 64, and the number of iterations is 50; the attention weight dimension is 64, and the number of neurons in the attention layer is 16; the number of output nodes in the fully connected layer is 1, and the Sigmoid activation function is used.
[0198] The learning factor of the Northern Goshawk is set to 0.7, the inertia weight value is 0.9, the range of the speed is between 0.01 and 0.1, the population size is taken as 40, and the maximum number of iterations is 50 times. After being optimized by the NGO algorithm, the batch size, number of iterations, filter size, and number of convolutional kernels in the Attention-ConvLSTM model are 64, 50, 3*3, and 32 respectively. The Mean Squared Error (MSE) is selected as the optimization objective function of the NGO for the Attention-ConvLSTM. MSE reflects the average squared difference between the travel time prediction results of the model and the actual values. The convergence curve of the model is as Figure 7 shown. When the number of iterations reaches 17 times, the MSE approaches 3.28%.
[0199] To enable the model to better capture information related to the prediction target, the attention mechanism is introduced. It can adjust the corresponding weights of the features, which helps to further optimize and adjust the model, thereby improving the real-time performance and accuracy of the travel link travel time prediction. The attention mechanism can automatically learn the correlation between different time steps in the input sequence and assign higher weights to important time steps, such asFigure 8 As shown, the attention coefficients at the 3rd, 5th, 14th, and 15th time steps are relatively large, indicating a higher level of attention from the model, while the attention to the 7th and 9th is relatively low.
[0200] 3.2.2 Result Analysis
[0201] To verify the accuracy of the NGO-Attention-ConvLSTM model in predicting the travel link travel time of vehicles, an unoptimized Attention-ConvLSTM prediction model, a ConvLSTM model without an attention mechanism, and a commonly used LSTM model were selected for comparison. The models were used to predict the travel time of 70 randomly selected paths in the test set data for the next 15 minutes of vehicle travel links. The comparison results between the true values and predicted values of various models are as Figure 9 shown, and from Figure 9 (a), it can be seen that the abscissa represents each path, the ordinate represents the travel time, the red dashed line is the true value of the path travel time, and the blue dashed line is the predicted value of the path travel time. It can be clearly shown that the model used in this embodiment can achieve a prediction accuracy of over 90% for the travel time of the path. From Figure 9 (b), it can be seen that the prediction result of the LSTM model for the travel time of the vehicle travel link is relatively poor. There is a certain gap between the predicted value and the true value of the travel link travel time for the Attention-ConvLSTM and ConvLSTM models. The travel time prediction result of the NGO-Attention-ConvLSTM model is closest to the true value, with the best prediction effect, and this model has high value in practical applications.
[0202] To quantify the prediction ability of each model for the travel link travel time and determine the effect of the model in practical applications, the Mean Absolute Error (MAE), Mean Relative Error (MRE), Mean Absolute Percentage Error (MAPE), Root Mean Squared Error (RMSE), and Coefficient of Determination (R 2 ) were used. Among them, MAE and MAPE focus on absolute errors, while MRE and RMSE focus on relative errors. They are all used to measure the difference between the predicted value and the actual value, and R 2 is used to evaluate the closeness of the predicted data to the true data. The five evaluation indicators can be expressed as:
[0203]
[0204] In the formula, Y i represents the true value, represents the predicted value, represents the average value, and n represents the number of samples.
[0205] Combined with Figure 10 , the table shows the average accuracy of each model in 5 experiments. The prediction accuracy of the NGO-Attention-ConvLSTM model is the highest. In terms of RMSE, this model reduces by 187.346s, 127.238s, and 109.201s respectively. In terms of R 2 , the coefficient of determination of this model is closest to 1, being 28.4%, 19.9%, and 7.2% higher than other models respectively. The unoptimized Attention-ConvLSTM prediction model has certain advantages over the ConvLSTM model without the attention mechanism and the commonly used LSTM model in various indicators, indicating that the attention mechanism helps to improve the prediction accuracy. It can be seen that the model of this embodiment can be used as an optimal model in practical applications.
[0206] Therefore, for the problem of predicting the travel time of vehicle travel links, through the analysis of traffic data, traffic theme links are extracted based on similarity, providing strong support for subsequent travel time prediction. The vehicle travel link travel time prediction model based on the Northern Goshawk optimization algorithm of Attention-ConvLSTM is adopted. The convolutional long short-term memory network helps to extract the spatio-temporal characteristics of the link and enhance the expression ability of the model; the introduction of the attention mechanism enables the model to better focus on the key information related to the predicted target link; the Northern Goshawk optimization algorithm has good global search ability. Optimizing the model can find the optimal weights of the model faster, improve the convergence speed and prediction performance of the model, and thus improve the prediction accuracy. By comparing the error between the prediction result and the actual value, it is found that the evaluation indicators of this model are better than other methods, and it can accurately predict the travel time of vehicle travel links.
[0207] It is easy to understand that those skilled in the art can combine, split, recombine, etc. the embodiments of the present application based on one or several embodiments provided by the present application to obtain other embodiments, and these embodiments do not exceed the protection scope of the present application.
[0208] The above schematically describes the present invention and its implementation manners. This description is not restrictive. What is shown in the embodiments is only part of the implementation manners of the present invention, and the actual structure is not limited thereto. Therefore, if those of ordinary skill in the art are inspired by it and design similar structural manners and embodiments without creative efforts without departing from the purpose of the present invention's creation, they should all fall within the protection scope of the present invention.
Claims
1. A vehicle travel time prediction method for travel link extraction, characterized in that: The following steps are involved: Step S1, obtaining the vehicle trajectory data of the taxi, and then extracting the traffic theme links and their classification; Step S2: constructing a vehicle travel link travel time prediction model based on NGO-Attention-ConvLSTM; Step S3, extracting the results of the traffic theme link and obtaining the predicted travel time of the vehicle travel link; The vehicle trajectory of the taxi in step S1, Traj data, is data recording the location information of the taxi in a certain period of time, which includes the unloaded trajectory points and the loaded trajectory points of the taxi; There are several trajectories in the taxi trajectory dataset, namely Path={Path1,Path2,...,Path n } (1) (1) Where Traj i represents the i-th trajectory; A trajectory Traj consists of multiple trajectory points, x1, x2, ..., x n , composed of, where a trajectory point x i It is the location data points collected by the GPS device during the taxi's driving process. These data points include vehicle identification ID, longitude lng, latitude lat, time time, and passenger status flag, expressed as: x i =(id i , lng i ,lat i , time i , flag i ); According to the driving trajectory of the taxi and the passenger status of the taxi, the pick-up and drop-off points of the taxi are judged and extracted, and the link information between the pick-up and drop-off points is extracted; First, based on the timestamp and location information in the GPS data, the taxi's driving trajectory is constructed, and the links between the various trajectory points are extracted from it. The specific steps are as follows: When the taxi's passenger-carrying status flag changes from 0 to 1, it means that the taxi will start to pick up passengers, and this information will be used as the taxi's pick-up point; When the taxi's passenger-carrying status flag changes from 1 to 0, it means the taxi has finished carrying passengers, and this information is used as the taxi's drop-off point; According to the preprocessed GPS data, the spatial relationship between the sampling points is analyzed using formula (2) to estimate the position of the unknown point Yu and construct the taxi's driving trajectory Traj; Extract the links between adjacent trajectory points on the driving trajectory Traj, set a time threshold t, and if the time interval between two trajectory points is less than the threshold t, it is considered that there is a link between the two trajectory points; Where Y u is the value of the unknown point, Y i is the value of the known point, X i is the weight of the known point, r i is the distance between the known point and the unknown point; The method for extracting traffic theme links in step S1 is based on similarity, and specifically includes: First, the input trajectory coordinates are preprocessed, the S longitude and latitude coordinates are converted to the tangent plane coordinates (x, y, z), and then the tangent plane coordinates (x, y, z) are converted to the coordinate data (X, Y, Z) in the Cartesian coordinate system, the longitude is converted to radians and divided by the radius of the earth to obtain a value less than 1 in radians, and the latitude is directly converted to radians; Next, calculate the longitude and latitude distance between each point on the two trajectories. For each point i and j on each trajectory, use the Euclidean distance formula, that is, In the formula, x i and i is the Cartesian coordinate of point i, x j and j are the Cartesian coordinates of point j; Then calculate the contribution of each point to the entire trajectory, calculate the similarity between the two trajectories based on the contribution, take the weighted average of these distances, and then calculate the similarity between the two trajectories. Finally, the similarity between the two trajectories is output, all links are traversed, the number of times each travel link is walked between the boarding and alighting points is counted, and the travel links with a travel number greater than the threshold are extracted; The formula for calculating the contribution of a trajectory point to the entire trajectory is: The formula for calculating the similarity between two trajectories is: Where n is the number of points on the trajectory, d(i, j) is the distance between trajectory points i and j; First, determine the partitions of OD points, extract OD points with more travel demand, extract the paths between OD partitions according to the similarity of trajectories, and extract the traffic theme links with the most travel between each OD partition; The classification method of the traffic theme link in step S1 is performed based on the travel purpose, and the classification method is as follows: First, we extract the pick-up and drop-off points from the taxi GPS data and cluster the pick-up and drop-off points to get the passenger hotspots. Next, crawl the POI data with a fixed radius around the passenger hotspot as the center, select the POI type with the largest proportion as the type of passenger hotspot, and use the passenger hotspot area as the starting and ending point of the travel link. The types of starting and ending points are divided into four types: residential, office, life service and school; Travel links are divided into links from residence to company, residence to school, residence to life services, from school to company, school to life services, and round-trip links according to the attribute characteristics of the starting and ending points.
2. The vehicle travel time prediction method for travel link extraction according to claim 1 is characterized in that: The model in step S2 includes: The input layer is used to convert the received raw sequence data into a form readable by the model; Convolutional layer, used to extract local features from input data. A convolutional layer can contain multiple convolution kernels, each of which corresponds to a different feature. LSTM layer, which has multiple LSTM units. Each LSTM unit consists of an input gate, a forget gate, and an output gate to capture long-term dependencies in the input data. Northern Goshawk optimization algorithm, used to optimize the parameters of the ConvLSTM layer; Attention mechanism, which is used to weight information at different time steps to help the model focus on more important information when making predictions; The fully connected layer is used to fully connect the output of Attention-ConvLSTM to obtain the final prediction result; The output layer is used to transform the prediction results of the fully connected layer into a numerical value in the time series data.
3. The vehicle travel time prediction method for travel link extraction according to claim 2 is characterized in that: The northern goshawk optimization algorithm is mainly divided into two stages: global search stage and local search stage; Global search phase: In this phase, the northern goshawk is used to randomly select a prey in the search space and then quickly attack it; This process is described by a mathematical model, such as equation (6) for position update and equation (7) for velocity update, where the prey represents the objective function in the optimization problem, and the northern goshawk needs to find a solution that minimizes the objective function in the search space; x t+1 =x t -a*(b-x t ) (6) v t+1 =v t -c*(w-v t ) (7) Where a and c are acceleration coefficients, b is the position of the prey, w is the position of the goshawk, and x is t is the current position of the goshawk, v t is the current speed of the goshawk; Local search phase: After successfully attacking the prey, the northern goshawk will enter the local search phase. Using the simulated annealing method, a temperature parameter T is set. In each iteration, the goshawk will decide whether to accept the new position according to the Metropolis criterion. The formula can be described as: In the formula, n is the number of samples in the validation set, y _i is the actual travel time, y _pred_i is the travel time predicted by the model, MSE is the mean square error between the predicted travel time and the true value of the Attention-ConvLSTM model, x new and x old are the new and old positions and T is the temperature.
4. The vehicle travel time prediction method for travel link extraction according to claim 3 is characterized in that: The formulas for the convolutional layer and LSTM layer are: I t =σ(W xi *h t-1 +b i ) (10) f t =σ(W fh *h t-1 +b f ) (11) the t =σ(W ho *h t +b o ) (12) h t =σ(W ch *x t +b c ) (13) Where: i t is the output of the input gate, f t is the output of the forget gate, o t is the output of the output gate, h t is the output of the convolutional layer, W xi , W fh , W ho , W ch They are the input gate, forget gate, output gate, and weight matrix of the convolutional layer, h t is the hidden state at the current moment, h t-1 is the hidden state at the previous moment, b i 、b f 、b o 、b c They are the input gate, forget gate, output gate, and bias vector of the convolution layer. σ is the sigmoid activation function, and * is the convolution operation.
5. The vehicle travel time prediction method for travel link extraction according to claim 4 is characterized in that: Introducing the attention mechanism into the ConvLSTM model so that the model focuses on the features of the input data that are useful for prediction and weights different spatial positions and feature channels; In the attention mechanism, the output sequence [h t-s+1 , ..., h t-1 ,h t ] as the input sequence of the Attention mechanism, For each position in the input sequence, an attention weight [α t-s+1 , ..., α t-1 , α t ], Finally, these weights are multiplied by the corresponding input values, and the products are summed to obtain a vector [H t-s+1 , ..., H t-1 , H t ], formulas as (14)-(16); c t =W s *tanh(W xa *X t +W ha *h t +b a ) (14) Where: W is the convolution kernel weight, X is the input feature, h is the input sequence of ConvLSTM, b a is the bias term, c t To calculate the median value of the attention coefficient for each feature region, α t is the attention weight, and H is the output sequence of the weighted sum.
6. A vehicle travel time prediction device for travel link extraction, comprising a memory and a processor, wherein a computer program is stored in the memory, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Modeling method for optimizing thermal error of electric spindle of KELM neural network by northern eagle algorithm
CN117077509A
Method for bus arrival time prediction when lacking data
WO2023029234A1