Private car commuting travel destination prediction method based on checkpoint data

By adopting dynamic weight optimization clustering method and XGBoost prediction model in the prediction of private car commuting destinations, the problem of limited prediction accuracy in the prior art is solved, and more efficient commuting individual identification and travel destination prediction are achieved.

CN120146267APending Publication Date: 2025-06-13SOUTHEAST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510201171.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

When the prior art uses traffic monitoring data to predict the destination of private car commuting travel, there are limitations of static feature extraction and the problem that the prediction model and feature extraction process are independent of each other, resulting in limited prediction accuracy.

Method used

The unsupervised clustering method with dynamic weight optimization combined with the commuting destination prediction model based on XGBoost, realizes data-driven adaptive adjustment of clustering indicator weights and two-way optimization of clustering results and prediction models.

Benefits of technology

It reduces the subjectivity of manual feature selection, improves prediction accuracy, and can more accurately identify commuting individuals and predict their travel destinations. It is of great significance to traffic flow management and personalized travel services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146267A_ABST
    Figure CN120146267A_ABST
Patent Text Reader

Abstract

The invention discloses a private car commuting travel destination prediction method based on checkpoint data. The method comprises the following steps: step 1, acquiring longitude and latitude of a checkpoint point location; step 2, collecting and preprocessing data of vehicles passing through a bayonet; step 3, dividing a single travel; step 4, clustering is carried out; step 5, constructing and training a commuting travel destination prediction model: step 5-1, constructing the commuting travel destination prediction model; step 5-2, training a commuting travel destination prediction model; step 6, updating the dynamic weight; and step 7, predicting. According to the invention, a dynamic weight optimization unsupervised clustering method is combined with an XGBoost-based commuting travel destination prediction model, so that data-driven clustering index weight adaptive adjustment and bidirectional optimization of a clustering result and the prediction model are realized, space-time travel modes of individual vehicles can be accurately separated from checkpoint data; and the clustering structure can better fit the prediction task through the feedback of the prediction result, so that the prediction accuracy of the commuting category is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of traffic data analysis and traffic prediction, and particularly to a method for predicting the commuting destinations of private cars based on checkpoint data. Background Art

[0002] With the rapid development of urban transportation, the monitoring devices on the roads, such as cameras and sensors, are increasing day by day, and these devices continuously collect a large amount of vehicle data. However, most of the existing studies only focus on using traffic monitoring data for local applications such as traffic flow estimation, vehicle arrival pattern inference, and intersection queue length evaluation, and there is relatively insufficient research on how to deeply mine individual travel destination information from these data.

[0003] Previous travel behavior prediction models mostly rely on vehicle GPS trajectories, land use and POI data, or taxi trip record data to establish an aggregate model for the travel behavior of travel groups within a certain range or for certain given OD, and have not fully utilized the potential of checkpoint data in focusing on a single individual to mine its travel patterns.

[0004] The existing methods generally face the following problems: (1) Limitations in static feature extraction: Most methods use fixed rules to extract travel pattern features, making it difficult to adapt to changes in travel behavior. (2) Independence between the prediction model and the feature extraction process: The existing methods usually perform feature extraction and clustering first, and then make predictions, lacking a closed-loop mechanism for model training and feature optimization, resulting in limited prediction accuracy.

[0005] In urban traffic management, accurately predicting the commuting destinations of private cars is of great significance for optimizing traffic flow distribution, alleviating traffic congestion, pre-providing travel road information, and providing personalized travel services for different individuals. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to provide a method for predicting the commuting destinations of private cars based on checkpoint data in view of the above-mentioned deficiencies of the prior art. The method for predicting the commuting destinations of private cars based on checkpoint data adopts an unsupervised clustering method with dynamic weight optimization combined with a commuting destination prediction model based on XGBoost to achieve data-driven adaptive adjustment of the weights of clustering indicators and two-way optimization of the clustering results and the prediction model, reduce the subjectivity of manual feature selection, and improve the prediction accuracy.

[0007] To solve the above technical problems, the technical solution adopted by the present invention is:

[0008] A method for predicting the commuting destinations of private cars based on checkpoint data, comprising the following steps.

[0009] Step 1. Obtain the longitude and latitude of the checkpoint locations: First, determine the area to be studied, and then, based on an external map platform, obtain the number M of road checkpoints in the area to be studied and the longitude and latitude of the location of each checkpoint.

[0010] Step 2. Collect and preprocess the vehicle passing data at the checkpoints: Obtain the vehicle passing data of the M checkpoints in the area to be studied within the set time period T d days, randomly select Q private cars as the research objects from them, and preprocess the vehicle passing data of the research objects; where Q ≥ 1000.

[0011] Step 3. Divide single trips: By comparing the time interval between two adjacent pieces of vehicle passing data at the checkpoints with the set travel time and driving speed threshold between the checkpoint locations, single trips are identified, and then the spatio-temporal sequence set TS i of a single vehicle's single trip is extracted; where i is the i-th private car, and 1 ≤ i ≤ Q.

[0012] Step 4. Clustering: Use the Monte Carlo K-Means++ algorithm with an embedded dynamic weight allocation mechanism to cluster the spatio-temporal sequence set TS i of each private car extracted in Step 3, and screen out P private cars that conform to the characteristics of commuting trips.

[0013] Step 5. Build and train a commuting trip destination prediction model, including the following steps.

[0014] Step 5-1. Build a commuting trip destination prediction model: Build a commuting trip destination prediction model based on XGBoost machine learning; the input of the commuting trip destination prediction model is the spatio-temporal sequence of the first trip of a private car in the T d -1 days before the set time period, and the output of the commuting trip destination prediction model is the trip destination of the private car on the T d th day of the set time period.

[0015] Step 5-2. Train the commuting trip destination prediction model: Use the spatio-temporal sequences of the first trips of the P private cars in the T d -1 days before the set time period as the input of the commuting trip destination prediction model, and the trip destinations of the P private cars on the T d th day of the set time period as the output of the commuting trip destination prediction model, train the commuting trip destination prediction model built in Step 5-1, and calculate the prediction accuracy rate.

[0016] Step 6, Update the dynamic weight: Compare the prediction accuracy rate calculated in Step 5-2 with the set accuracy threshold; when the prediction accuracy rate calculated in Step 5-2 is less than the set accuracy threshold, use the prediction result of Step 5-2 to correct and update the dynamic weight in Step 4, and repeat Steps 4 to 6 until the prediction accuracy rate calculated in Step 5-2 reaches the set accuracy threshold, and use the trained commuting travel destination prediction model at this time as the optimal commuting travel destination prediction model.

[0017] Step 7, Prediction: Collect the spatio-temporal sequence of the first trip of the private cars to be predicted in the area to be studied within the first T - 1 days of the set time period, and input the collected spatio-temporal sequence of the first trip into the optimal commuting travel destination prediction model, so as to predict the travel destination of the private cars to be predicted on the Tth day of the set time period. d - 1 day, and input the collected spatio-temporal sequence of the first trip into the optimal commuting travel destination prediction model, so as to predict the travel destination of the private cars to be predicted on the Tth day of the set time period. d day.

[0018] In Step 4, the method of using the Monte Carlo K-Means++ algorithm for clustering includes the following steps.

[0019] Step 4-1, Determine and calculate the clustering indicators: Determine a clustering indicators for each private car, and calculate the a clustering indicator values k, k, ……, k, ……, k of each private car according to the single-trip spatio-temporal sequence set TS extracted in Step 3. i , and calculate the a clustering indicator values k of each private car 1 , k 2 , ……, k m , ……, k a .

[0020] Step 4-2, Standardization: Normalize the a clustering indicator values of each private car obtained in Step 4-1 into dimensionless data, so as to form k', k', ……, k', ……, k'. 1 , k' 2 , ……, k' m , ……, k' a .

[0021] Step 4-3, Assign initial weights: Assign an initial weight to each clustering indicator.

[0022] Step 4-4, Determine the optimal number of clusters K, which specifically includes the following steps.

[0023] Step 4-4A, One-round clustering: Use the a clustering indicators, the initial weights of each clustering indicator assigned in Step 4-3, and the k', k', ……, k', ……, k' corresponding to Q private cars 1 , k' 2 , ……, k' m , ……, k' a, all are input into the Monte Carlo K-Means++ algorithm, perform B clustering operations, and obtain B clustering numbers C, B clustering dispersion degrees SSE, and B inter-cluster separabilities DB.

[0024] Step 4-4B. Determine the optimal clustering number K: Compare the B clustering dispersion degrees SSE and the inter-cluster separability DB. When the comprehensive performance of both the clustering dispersion degree SSE and the inter-cluster separability DB is optimal, the corresponding clustering number C is used as the optimal clustering number K.

[0025] Step 4-5. Assign dynamic weights: Keep the optimal clustering number K obtained in Step 4-4 as the output of the Monte Carlo K-Means++ algorithm unchanged. Take a clustering metrics, the dynamic weights of each clustering metric, and k' corresponding to Q private cars 1 、k' 2 、……、k' m 、……、k' a , as the output of the Monte Carlo K-Means++ algorithm, perform several clustering operations, so as to optimize and adjust the dynamic weights of each clustering metric input in the Monte Carlo K-Means++ algorithm until convergence.

[0026] Step 4-6. Screening: Use the Monte Carlo K-Means++ algorithm containing the converged dynamic weights of each clustering metric for clustering, and obtain the clustering results of K cluster centers; then, screen out a clustering result of a cluster center that conforms to the commuting travel characteristics from the K clustering results of cluster centers, which contains P private cars.

[0027] In Step 4-5, the method for assigning dynamic weights includes the following steps.

[0028] Step 4-5A. Calculate the mutual information MI of clustering metrics m : Take the normalized a clustering metric values of each private car as row data, then the normalized a clustering metric values of Q private cars form a clustering metric value array of Q rows and a columns.

[0029] According to the clustering metric value array, calculate the mutual information MI between the m-th clustering metric and the other m - 1 clustering metrics m .

[0030] Step 4-5B. Calculate the clustering stability score S m : Randomly extract subsets from the clustering metric value array, and use the clustering number C obtained in Step 4-4 to perform several clustering operations, and then obtain the clustering stability score S of the m-th clustering metric m .

[0031] Step 4-5C. Assign dynamic weights: According to the assigned weights in the (t - 1)-th clustering The mutual information MI of clustering metrics m and the clustering stability score S m , dynamically adjust the allocation weight of the m-th clustering metric at the t-th clustering

[0032] Step 4-5D, Weight Convergence: When the weight changes of all clustering metrics in two adjacent iterations are less than the weight threshold δ, stop the iteration and obtain the dynamic weights of each clustering metric after convergence.

[0033] In Step 4-5C The calculation formula of is:

[0034]

[0035] In the formula, α is a smoothing coefficient that controls the influence degree of historical weights on the current weight, and its value range is (0, 1).

[0036] β is an influence control coefficient of mutual information on weights, and its value range is (0, 1).

[0037] 5. The method for predicting the commuting travel destination of private cars based on bayonet data according to claim 4, characterized in that: in Step 5-2, during the process of training the commuting travel destination prediction model, assume that among the samples with correct predictions, the clustering metric k m The mean value is Among the samples with incorrect predictions, the clustering metric k m The mean value is Then in Step 6, assume that the allocation weight of the m-th clustering metric at the t-th clustering The value after dynamic update is Then The calculation formula of is:

[0038]

[0039] Where:

[0040]

[0041] In the formula, C is a weight correction constant, which is an empirical constant.

[0042] γ is a feedback coefficient, and its value range is (0, 1).

[0043] Is The weight value after dynamic update.

[0044] In Step 4-1, a = 6, and the 6 clustering metric values are: the average first passing bayonet time k 1 、the daily average travel times k 2, total number of travel days \(k\) 3 , maximum number of same days for the first travel destination \(k\) 4 , travel frequency during peak hours \(k\) 5 and degree of departure time fluctuation \(k\) 6 ; where \(k\) 1 \(\in [0, 24]\), \(k\) 2 \(\in N\) and \(k\) 2 \(> 0\), \(k\) 3 \(\in [0, T d \) and \(k\) 3 \(\in Z\), \(k\) 4 \(\in [0, T d \) and \(k\) 4 \(\in Z\), \(k\) 5 \(\in [0, 1]\), \(k\) 6 \(> 0\). In step 3, the method for dividing a single trip includes the following steps:

[0045] Step 3-1, set the trip time threshold: Set the trip time threshold between adjacent two checkpoint positions. The trip time threshold includes the minimum trip time threshold \(T 1 and the maximum trip time threshold \(T 2 ; where \(T 1 \lt T 2 .

[0046] Step 3-2, set the trip speed threshold: Set the trip speed threshold between adjacent two checkpoint positions. The trip speed threshold includes the minimum trip speed threshold \(v 1 and the maximum trip speed threshold \(v 2 ; where \(v 1 \lt v 2 .

[0047] Step 3-3, mark the first trajectory point: The position of each private car at each checkpoint is called a trajectory point. Then each vehicle has a sequence of trajectory points; for the sequence of trajectory points of the same vehicle, mark the travel status of the first trajectory point as "1".

[0048] Step 3-4, traverse all the remaining trajectory points in the sequence of trajectory points of the same vehicle and mark them. Let the time interval between the current trajectory point and the previous trajectory point be \(T\), then:

[0049] A. When \(T \leq T 1 \), then mark the travel status of the current trajectory point as "0".

[0050] B. When \(T 1 \lt T \lt T 2 , and \(v 1 \leq v \leq v 2 \), then mark the travel status of the current trajectory point as "0"; where \(v\) is the current speed of the current vehicle.

[0051] C. When T 1 <T < T 2 and v < v 1 or v > v 2 then the current trajectory point is considered to belong to a new trip; at this time, the marking method of the travel status is as follows: check whether there is a process point with the travel status marked as "0" in the previous trip record. If there is, mark the travel status of the previous trajectory point as "1" and the travel status of the current trajectory point as "2", serving as the starting point of the second trip; otherwise, delete the previous trajectory point with the travel status marked as "1" and re-number the current trajectory point as "1".

[0052] D. When T > T 2 then mark the travel status of the current trajectory point as "2".

[0053] After all the trajectory points of the same vehicle are traversed and marked, a trajectory point travel status sequence in the form of 1, 0, 0,..., 1, 2, 0, 0,..., 2,... is formed; where, "1, 0, 0,..., 1" represents the first trip; "2, 0, 0,..., 2" represents the second trip, and so on.

[0054] Step 3 - 5. Repeat Step 3 - 4 to Step 3 - 5 to complete the trajectory point travel status sequences of all vehicles.

[0055] In Step 5 - 1, the model optimization parameters in the commuting travel destination prediction model are determined using the Bayesian optimization method.

[0056] In Step 5 - 1, the Bayesian optimization method includes the following steps:

[0057] Step 5 - 1A. Construct the objective function f(θ) of the model parameter θ, and the specific expression is:

[0058] f(θ) = Acc - λ·depth

[0059] In the formula, Acc is the cross - validation accuracy rate, depth is the depth of the tree, and λ is the model complexity.

[0060] Step 5 - 1B. Solve the model optimization parameter θ * : Search for the model optimization parameter θ based on Gaussian process regression * and the specific expression is:

[0061] θ * = argmax f(θ).

[0062] The external map platform in Step 1 is Amap.

[0063] The present invention has the following beneficial effects:

[0064] (1) The present invention constructs a complete application framework for road network checkpoint data, namely "data processing - commuting mode clustering - destination prediction and feedback optimization", to explore the application value of massive checkpoint data. The present invention uses the Monte Carlo K-Means++ method with an embedded dynamic weight allocation mechanism to achieve more accurate identification of commuting individuals, realizes the adaptive adjustment of the weights of clustering indicators driven by data and the two-way optimization of clustering results and prediction models, reduces the subjectivity of manual feature selection, improves the prediction accuracy, and is of great significance for further optimizing traffic flow distribution, alleviating traffic congestion, pre-providing travel road information, and pushing personalized travel services for different individuals.

[0065] (2) Adopt a dynamic feature weight optimization method based on mutual information (MI) and feature stability (S), so that the feature weights can be dynamically adjusted according to the data distribution, and combine the Monte Carlo loop to determine the optimal number of clusters, improving the reliability of the clustering results.

[0066] (3) Establish a closed-loop optimization mechanism for clustering and prediction, improve the prediction accuracy using the XGBoost prediction model, and use the prediction error to optimize the feature weights, so that the clustering features can better match the prediction task.

[0067] (4) Use Gaussian process regression for hyperparameter optimization, which has higher computational efficiency compared to traditional grid search. Description of the Drawings

[0068] Figure 1 Shows the flow chart of the private car commuting destination prediction method based on checkpoint data of the present invention.

[0069] Figure 2 Shows the flow chart of trip division in the present invention.

[0070] Figure 3 Shows the graph of the change of SSE and DB index of the KMeans++ clustering algorithm with the number of clusters in the present invention. Among them, (a) is the trend graph of SSE change; (b) is the trend graph of DB change.

[0071] Figure 4 Shows the graph of the change of the prediction accuracy of the XGBoost algorithm with the number of iterations in the present invention; among them, Run1 to Run6 correspond to 6 iterations respectively.

[0072] Figure 5 Shows the graph of the number of commuting vehicles and the prediction accuracy with the number of weight optimizations in the present invention.

[0073] Figure 6 Shows the graph of the change of the clustering index weights with the number of weight optimizations in the present invention. Specific Embodiment

[0074] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific preferred embodiments.

[0075] As Figure 1 shown, a method for predicting the commuting destination of private cars based on bayonet data includes the following steps.

[0076] Step 1: Obtain the longitude and latitude of bayonet points. First, determine the area to be studied, and then, based on an external map platform, obtain the number M of road network bayonets in the area to be studied and the longitude and latitude of the point of each bayonet.

[0077] In this embodiment, the area to be studied is preferably exemplified by Shinan District, Qingdao City, and the external map platform is preferably Amap, but it can also be other known map platforms such as Baidu Map.

[0078] Step 2: Collection and preprocessing of bayonet vehicle passing data. Randomly select Q private cars from the area to be studied. Then, collect the vehicle passing data of Q private cars passing through M bayonets within the set time period T d days and perform data preprocessing; where Q≥1000, preferably Q = 10000.

[0079] In this embodiment, in Shinan District, Qingdao City, randomly select Q = 10000 cars with a cycle T d = 15 days of road network bayonet records as the research object, and the data structure is shown in Table 1.

[0080] Table 1 Road network bayonet data structure in Qingdao

[0081]

[0082] In this embodiment, the preprocessing of the bayonet data collected in Shinan District, Qingdao City is preferably as follows:

[0083] A. Group the bayonet data by license plate number, sort them within the group in the order of detection time, and then import them into the computer database for storage and management; then use python to call the original data for processing, eliminate abnormal data, redundant data and repair missing data.

[0084] B. Use the Amap platform to complete the longitude and latitude data of the bayonets based on the bayonet location names, and use the DSCAN algorithm (eps = 30, min_samples = 1) to cluster the bayonets and re-number them to obtain M = 661 bayonet numbers.

[0085] Step 3: Single trip division. By comparing the time interval between two adjacent bayonet vehicle passing data with the set travel time threshold between bayonet positions, single trips are identified, and then the single-vehicle single-trip spatio-temporal sequence set TS is extracted.i ; where i is the i-th private car, and 1 ≤ i ≤ Q.

[0086] As Figure 2 shown, the method for dividing a single trip preferably includes the following steps.

[0087] Step 3-1: Set the trip time threshold: Set the trip time threshold between two adjacent checkpoint positions. The trip time threshold includes the minimum trip time threshold T 1 and the maximum trip time threshold T 2 ; where T 1 < T 2 .

[0088] The above time thresholds T 1 and T 2 are determined by observing the variation of the number of trips with the time threshold. T 1 can be set to 10 minutes, 15 minutes, 20 minutes, 25 minutes, and 30 minutes. T 2 usually directly takes a larger value. In this embodiment, preferably T 1 = 15 min, T 2 = 120 min.

[0089] Step 3-2: Set the trip speed threshold: Set the trip speed threshold between two adjacent checkpoint positions. The trip speed threshold includes the minimum trip speed threshold v 1 and the maximum trip speed threshold v 2 ; where v 1 < v 2 . In this embodiment, preferably v 1 = 20 km / h, v 2 = 80 km / h.

[0090] Step 3-3: First trajectory point marking: The position of each private car at each checkpoint is called a trajectory point. Then each vehicle has a sequence of trajectory points; for the sequence of trajectory points of the same vehicle, the travel status of the first trajectory point is marked as "1".

[0091] Step 3-4: Traverse all the remaining trajectory points in the sequence of trajectory points of the same vehicle and mark them. Let the time interval between the current trajectory point and the previous trajectory point be T. Then:

[0092] A. When T ≤ T 1 , then mark the travel status of the current trajectory point as "0".

[0093] B. When T 1 < T < T 2 , and v 1 ≤ v ≤ v 2If so, mark the travel status of the current trajectory point as "0"; where v is the current speed of the current vehicle.

[0094] C. When T 1 <T < T 2 , and v < v 1 or v > v 2 , it is considered that the current trajectory point belongs to a new trip; at this time, the marking method of the travel status is as follows: determine whether there is a process point with the travel status marked as "0" in the previous trip record. If so, mark the travel status of the previous trajectory point as "1" and the travel status of the current trajectory point as "2", as the starting point of the second trip; otherwise, delete the previous trajectory point with the travel status marked as "1" and re-number the current trajectory point as "1".

[0095] D. When T > T 2 , mark the travel status of the current trajectory point as "2".

[0096] After all the trajectory points of the same vehicle are traversed and marked, a trajectory point travel status sequence in the form of 1, 0, 0,..., 1, 2, 0, 0,..., 2,... is formed; where "1, 0, 0,..., 1" represents the first trip; "2, 0, 0,..., 2" represents the second trip, and so on.

[0097] Step 3-5: Repeat Step 3-3 to Step 3-5 to complete the travel status sequences of the trajectory points of all vehicles, that is, the single-vehicle single-trip spatio-temporal sequence set TS i .

[0098] In this embodiment, taking the checkpoint data on January 12, 2022 as an example, the processed data structure is shown in Table 2

[0099] Table 2 Processed Checkpoint Data Structure

[0100]

[0101] Step 4: Clustering: Use the Monte Carlo K-Means++ algorithm with an embedded dynamic weight allocation mechanism to cluster the single-vehicle single-trip spatio-temporal sequence set TS i extracted for each private car, and screen out P private cars that meet the characteristics of commuting trips.

[0102] The method of clustering using the Monte Carlo K-Means++ algorithm preferably includes the following steps.

[0103] Step 4-1: Determine and calculate the clustering index: Each private car determines a clustering indexes, and according to the single-vehicle single-trip spatio-temporal sequence set TS i, the a clustering index values k of each private car are calculated 1 and k 2 and ……, k m and ……, k a . In this application, preferably a = 6, and the 6 clustering indexes and their value ranges are as shown in the following table.

[0104]

[0105]

[0106] In this embodiment, taking the data of the first ten sampled cars as an example, the clustering index values are shown in Table 3.

[0107] Table 3 Example of the characteristic index set of the sampled vehicles

[0108]

[0109] Step 4-2, Standardization: The a clustering index values of each private car obtained in Step 4-1 are all normalized into dimensionless data, thus forming k' 1 and k' 2 and ……, k' m and ……, k' a .

[0110] Furthermore, the normalization formula of k' m is preferably:

[0111]

[0112] Where:

[0113]

[0114] In the formula, k m,i is the original index value of the m-th clustering index k m of the i-th private car.

[0115] is the average value of all the m-th clustering index k m original index values of Q private cars.

[0116] σ m is the standard deviation of all the m-th clustering index k m original index values of Q private cars.

[0117] In this embodiment, taking the data of the first ten sampled cars as an example, the average value and standard deviation of the 6 clustering indexes are shown in Table 4.

[0118] Table 4 Average value of the clustering indexes of the sampled vehicles

[0119]

[0120] Step 4-3, Assign initial weights: Assign an initial weight to each clustering index. The preferred calculation formula is:

[0121]

[0122] Step 4-4, Determine the optimal number of clusters K, which specifically includes the following steps.

[0123] Step 4-4A, One-time loop clustering: Input a clustering indices, the initial weights of each clustering index assigned in Step 4-3, and k' corresponding to Q private cars 1 、k' 2 、……、k' m 、……、k' a into the Monte Carlo K-Means++ algorithm for B times of clustering, and obtain B numbers of clusters C, B cluster dispersion degrees SSE, and B inter-cluster separabilities DB.

[0124] The calculation methods of the above SSE and DB are prior arts, and the preferred calculation formulas are:

[0125]

[0126] Wherein:

[0127]

[0128] In the formula, G c represents the set of data point indices of cluster C, z p represents the feature vector of data point p, represents the center point of cluster C, and r c represents the average distance from the data points inside each cluster C to the cluster center.

[0129] In this embodiment, using the Monte Carlo K-Means++ method, SSE is calculated by calling kmeans.inertia_, and DB is calculated by calling davies_bouldin_score(). Preferably, B = 50, that is, the random_state is set to 50 to ensure 50 clustering loops are run.

[0130] Step 4-4B, Determine the optimal number of clusters K: Compare the B cluster dispersion degrees SSE and the inter-cluster separability DB, and take the number of clusters C corresponding to the optimal comprehensive performance of both the cluster dispersion degree SSE and the inter-cluster separability DB as the optimal number of clusters K.

[0131] The comprehensive performance of the above SSE and the separation measure DB between clusters is the best. In this embodiment, it mainly refers to that the decline of SSE slows down and the DB index is the smallest. By comprehensively comparing the SSE and the DB index, as Figure 3 shown, the optimal number of clusters K = 3 is finally determined and clustering is performed accordingly.

[0132] In this embodiment, for 50 clustering cycles, by comprehensively comparing the SSE and the DB index, as Figure 3 shown, the optimal number of clusters K = 3 is finally determined and clustering is performed accordingly, and the clustering results (after inverse normalization) are shown in Table 5.

[0133] Table 5 Monte Carlo K-Means++ Clustering Center Results of Private Car Travel Characteristics

[0134]

[0135] Step 4-5, Assign dynamic weights: Keep the optimal number of clusters K obtained in Step 4-4 unchanged as the output of the Monte Carlo K-Means++ algorithm. Take a clustering metrics, the dynamic weights of each clustering metric, and the k' corresponding to Q private cars 1 、k' 2 、……、k' m 、……、k' a , as the output of the Monte Carlo K-Means++ algorithm, and perform clustering several times, so as to optimize and adjust the dynamic weights of each clustering metric input in the Monte Carlo K-Means++ algorithm until convergence.

[0136] The above method for assigning dynamic weights preferably includes the following steps.

[0137] Step 4-5A, Calculate the mutual information MI of the clustering metrics m : Take the a clustering metric values of each private car after normalization as row data, then the a clustering metric values of Q private cars after normalization together form a clustering metric value array with Q rows and a columns; According to the clustering metric value array, calculate the mutual information MI m between the m-th clustering metric and the other m-1 clustering metrics.

[0138] Furthermore, the calculation formula of the above mutual information MI m is preferably:

[0139]

[0140] Where:

[0141]

[0142] In the formula, k n,iis the nth clustering index k of the ith private car n The original index value of

[0143] Step 4-5B. Calculate the clustering stability score S m : Randomly select a subset (such as 80% of the data) from the clustering index value array, and perform clustering several times using the number of clusters C obtained in Step 4-4. Statistically analyze the contribution of each feature (i.e., clustering index) to the clustering result of data points during different sampling and clustering operations, and then obtain the clustering stability score S m of the mth clustering index. The calculation formula is preferably:

[0144]

[0145] Where:

[0146]

[0147] In the formula, S(k m,i ) is the clustering stability score of the mth clustering index of the ith private car.

[0148] is the indicator function, indicating whether the difference between the data point k m,i and the mean value m of the kindex of the category it belongs to during the lth round of clustering is within the range of ∈. If it is within the range, the function value is 1; otherwise, it is 0.

[0149] Step 4-5C. Assign dynamic weights: According to the assignment weights clustering index mutual information MI m and the clustering stability score S m during the (t - 1)th clustering, dynamically adjust the assignment weight of the mth clustering index during the tth clustering. The calculation formula is:

[0150]

[0151] In the formula, α is the smoothing coefficient that controls the influence degree of the historical weight on the current weight, and its value range is (0, 1).

[0152] β is the influence control coefficient of the mutual information on the weight, and its value range is (0, 1).

[0153] Step 4-5D. Weight convergence: When the weight changes of all clustering indexes in two adjacent iterations are less than the weight threshold δ, stop the iteration. If the convergence standard is not reached, perform at most L rounds of updates to obtain the dynamic weights of each clustering index after convergence.

[0154] In this embodiment, the clustering metrics MI (mutual information) and S (feature stability) are calculated. The smoothing coefficient a = 0.5, the mutual information influence factor β = 0.5, Q = 10000 is known, and the number of repeated clustering times L = 50 is set. The calculation process of the weights after normalizing each clustering metric is shown in Table 6.

[0155] Table 6 Operation Results of the Dynamic Weight Allocation Mechanism

[0156]

[0157] Step 4-6, Screening: Use the Monte Carlo K-Means++ algorithm with the dynamic weights of each clustering metric after convergence for clustering, and obtain the clustering results of K cluster centers; then, screen out a clustering result of the cluster center that conforms to the commuting travel characteristics from the clustering results of the K cluster centers, which contains P private cars.

[0158] In this application, to conform to the commuting travel characteristics, the following three requirements A-C need to be met simultaneously:

[0159] A. The time k when passing the checkpoint for the first time 1 Must be between 7:00 and 9:00 to ensure that the vehicle travel time conforms to the morning rush hour commuting mode.

[0160] B. The number of travel days k 3 Cannot be less than the total number of working days during the detection period minus 2 to ensure that the vehicle has commuting behavior on most working days.

[0161] C. The maximum number of same days for the first travel destination k 4 Must be ranked among the top two of all cluster centers to ensure that the commuting destinations of the vehicles have high stability.

[0162] If there are multiple cluster centers that meet the above conditions, then select the cluster center with the smallest fluctuation degree k 6 of the departure time to ensure the most stable commuting behavior. The final commuting category vehicles are composed of the vehicles with this as the cluster center. If there is no cluster center that meets the above conditions, the clustering algorithm is re-run.

[0163] In this embodiment, the cluster center that meets the above conditions is Cluster 1 in Table 5. At this time, P = 1685.

[0164] Step 5, Construct and train a commuting travel destination prediction model, including the following steps.

[0165] Step 5-1, Construct a commuting travel destination prediction model: Construct a commuting travel destination prediction model based on XGBoost machine learning; the input of the commuting travel destination prediction model is the first T of private cars in a set time period d- The spatio-temporal sequence of the first trip on the first day, and the output of the commuting trip destination prediction model is the trip destination of private cars on the Tth day in the set time period. d day.

[0166] In this application, the model optimization parameters in the commuting trip destination prediction model are preferably determined by using the Bayesian optimization method, and the determination method includes the following steps.

[0167] Step 5-1A: Construct the objective function f(θ) of the model parameter θ, and the specific expression is:

[0168] f(θ) = Acc - λ·depth

[0169] In the formula, Acc is the cross-validation accuracy, depth is the depth of the tree, and λ is the model complexity;

[0170] Step 5-1B: Solve the model optimization parameter θ * : Search for the model optimization parameter θ based on Gaussian process regression * , and the expression is:

[0171] θ * = arg max f(θ).

[0172] In this embodiment, the test set accuracy is set as the objective function, the search ranges of each parameter are specified as follows in the table, a study object is created to execute the optimization process and obtain the best parameters. After inputting the historical spatio-temporal sequence of the vehicle, the future trip destination checkpoint number is output based on the XGBoost prediction algorithm.

[0173]

[0174]

[0175] Step 5-2: Train the commuting trip destination prediction model: Use the spatio-temporal sequence of the first trips of P private cars in the first T - 1 days in the set time period as the input of the commuting trip destination prediction model, and use the trip destinations of P private cars on the Tth day in the set time period as the output of the commuting trip destination prediction model, train the commuting trip destination prediction model constructed in step 5-1, and count the prediction accuracy. d - 1 day as the input of the commuting trip destination prediction model, and use the trip destinations of P private cars on the Tth day in the set time period as the output of the commuting trip destination prediction model, train the commuting trip destination prediction model constructed in step 5-1, and count the prediction accuracy. d day as the output of the commuting trip destination prediction model, train the commuting trip destination prediction model constructed in step 5-1, and count the prediction accuracy.

[0176] In this application, during the process of training the commuting trip destination prediction model, it is assumed that among the samples with correct predictions, the mean value of the clustering index k m is Among the samples with incorrect predictions, the mean value of the clustering index k m is Then in step 6, let the assignment weight of the mth clustering index at the tth clustering be The value after dynamic update is Then The calculation formula of

[0177]

[0178] Where:

[0179]

[0180] In the formula, C is the weight correction constant, which is an empirical constant.

[0181] γ is the feedback coefficient, and its value range is (0, 1).

[0182] is The weight value after dynamic update.

[0183] In the present invention, is used as the importance of the clustering index in the prediction algorithm. If a certain clustering index has a great impact on the prediction task (that is, is large), then its weight will increase, making the clustering result better able to distinguish between correctly and incorrectly predicted samples. Feedback to the clustering algorithm to readjust the weights.

[0184] In this embodiment, based on the clustering result, the Bayesian optimization XGBoost prediction method is adopted for the private car category (the first category) with commuting travel characteristics. The spatio-temporal sequence data of the first trips of the first-class vehicles in the previous 14 days is used as the training set X_train. The specific input data format is shown in Table 7. The XGBClassifier training model is used to predict the corresponding checkpoint number y_train of the first commuting travel destination on the 15th day.

[0185] Table 7 Example of input to the XGBoost prediction model

[0186]

[0187]

[0188] When the prediction algorithm is executed for the first time, the following optimal parameter combination is obtained:

[0189] max_depth = 8, min_child_weight = 1, subsample = 0.7352, colsample_bytree = 0.7299, learning_rate = 0.0040, gamma = 2.0301, alpha = 0.4906.

[0190] As the number of iterations (iteration = 50) increases, the change in accuracy is as follows Figure 4 shown, the accuracy of the first prediction is 66%.

[0191] Next, calculate the mean of the clustering indicators and the difference between the indicator values of the two sets of "correctly predicted U correct " and "incorrectly predicted U wrong " After standardization, readjust the weights and perform normalization and use it as the initial weight for the next clustering. The value of γ is 0.5, and the operation results are shown in Table 8.

[0192] Table 8 Classification of prediction results and weight update

[0193]

[0194] Step 6, update the dynamic weight: Compare the prediction accuracy statistically in Step 5-2 with the set accuracy threshold; when the prediction accuracy statistically in Step 5-2 is less than the set accuracy threshold, use the prediction results in Step 5-2 to correct and update the dynamic weight in Step 4, and repeat Steps 4 to 6 until the prediction accuracy statistically in Step 5-2 reaches the set accuracy threshold, and use the commuting travel destination prediction model trained at this time as the best commuting travel destination prediction model.

[0195] Calculate the importance of the clustering indicators in the prediction algorithm and feedback it to the clustering algorithm to readjust the weights and run the clustering. As the number of runs of the prediction algorithm increases, the number of vehicles in the commuting category and the prediction accuracy of the travel destination are as follows Figure 5 shown, and the change in the weights of the clustering indicators is as follows Figure 6 shown. Taking 85% as the accuracy threshold, after 6 times of feedback to optimize the weights, the prediction accuracy of the travel destinations of the updated P = 1723 commuting vehicles reaches 86.1%.

[0196] Step 7, prediction: Collect the spatio-temporal sequence of the first trip of the private cars to be predicted in the area to be studied in the T d -1 days before the set time period, and input the collected spatio-temporal sequence of the first trip into the best commuting travel destination prediction model, so as to predict the travel destination of the private cars to be predicted on the T d th day of the set time period.

[0197] Through the processing and analysis of the vehicle passing data at the road network checkpoints, by mining the commuting spatio-temporal characteristics in the checkpoint data, and combining dynamic weight allocation and prediction models, the present invention realizes the accurate prediction of the commuting travel destinations of individual private cars, providing valuable decision-making support for traffic management and travelers.

[0198] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all fall within the protection scope of the present invention.

Claims

1. A method for predicting the destination of a private car commuting trip based on checkpoint data, characterized in that: The steps include: Step 1, obtain the longitude and latitude of the checkpoints: first determine the area to be studied, and then obtain the number M of road network checkpoints in the area to be studied and the longitude and latitude of each checkpoint based on the external map platform; Step 2: Data collection and preprocessing of vehicles passing through checkpoints: Obtain the data of M checkpoints in the area to be studied within a set time period T. d Days of vehicle passing data, Q private cars are randomly selected as research objects, and the vehicle passing data of the research objects are preprocessed; where Q ≥ 1000; Step 3: Single trip division: By comparing the time interval between two adjacent checkpoints and the set travel time threshold between the checkpoints, a single trip is identified, and then the single-vehicle single-trip spatiotemporal sequence set TS is extracted. i ; Where i is the i-th private car, and 1≤i≤Q; Step 4: Clustering: Using the Monte Carlo K-Means++ algorithm embedded with a dynamic weight allocation mechanism, the single trip spatiotemporal sequence set TS of each private car extracted in step 3 is clustered. i , clustering is performed to screen out P private cars that meet the commuting travel characteristics; Step 5: Build and train a commuting destination prediction model, including the following steps: Step 5-1: Construct a commuting destination prediction model: Construct a commuting destination prediction model based on XGBoost machine learning; the input of the commuting destination prediction model is the number of private cars in the previous T period of the set time. d The output of the commuting travel destination prediction model is the time-space sequence of the first trip of the day. d The travel destination of the day; Step 5-2: Train the commuting destination prediction model: train P private cars in the first T of the set time period d The spatiotemporal sequence of the first trip on day -1 is used as the input of the commuting travel destination prediction model, and the P private cars are d The travel destination of the day is used as the output of the commuting travel destination prediction model, the commuting travel destination prediction model constructed in step 5-1 is trained, and the prediction accuracy is calculated; Step 6, update dynamic weight: compare the prediction accuracy calculated in step 5-2 with the set accuracy threshold; when the prediction accuracy calculated in step 5-2 is less than the set accuracy threshold, use the prediction result of step 5-2 to correct and update the dynamic weight of step 4, and repeat steps 4 to 6 until the prediction accuracy calculated in step 5-2 reaches the set accuracy threshold, and use the commuting travel destination prediction model trained at this time as the best commuting travel destination prediction model; Step 7: Prediction: Collect the predicted private cars in the study area before the set time period. d -1 day first trip time-space sequence, and input the collected first trip time-space sequence into the optimal commuting travel destination prediction model, so as to predict the private car to be predicted in the set time period T d The travel destination of the day.

2. The method for predicting the destination of a private car commuting trip based on checkpoint data according to claim 1 is characterized by: In step 4, the method of clustering using the Monte Carlo K-Means++ algorithm includes the following steps: Step 4-1, determine and calculate clustering index: determine a clustering index for each private car, and use the single trip spatiotemporal sequence set TS of each private car extracted in step 3 i , calculate a clustering index value k1, k2, ..., k for each private car m ,……,k a ; Step 4-2, standardization: The a clustering index values ​​of each private car obtained in step 4-1 are normalized into dimensionless data, thus forming k'1, k'2, ..., k' m , ..., k' a ; Step 4-3, assign initial weights: assign an initial weight to each clustering indicator; Step 4-4: Determine the optimal number of clusters K, specifically including the following steps: Step 4-4A, one-cycle clustering: a clustering index, the initial weight of each clustering index assigned in step 4-3, and k'1, k'2, ..., k' corresponding to Q private cars m , ..., k' a , are input into the Monte Carlo K-Means++ algorithm, clustered B times, and the number of B clusters C, the degree of dispersion of B clusters SSE and the separation between B clusters DB are obtained; Step 4-4B, determine the optimal number of clusters K: compare the degree of dispersion SSE of B clusters and the separation between clusters DB, and take the number of clusters C corresponding to the optimal comprehensive performance of the degree of dispersion SSE of clusters and the separation between clusters DB as the optimal number of clusters K; Step 4-5, assign dynamic weights: The optimal number of clusters K obtained in step 4-4 is used as the output of the Monte Carlo K-Means++ algorithm and remains unchanged, and the a clustering indicators, the dynamic weight of each clustering indicator, and k'1, k'2, ..., k' corresponding to Q private cars are assigned. m , ..., k' a , as the output of the Monte Carlo K-Means++ algorithm, several clusterings are performed to optimize and adjust the dynamic weight of each clustering indicator input in the Monte Carlo K-Means++ algorithm until convergence; Step 4-6, screening: Use the Monte Carlo K-Means++ algorithm with dynamic weights for each clustering indicator after convergence to perform clustering and obtain K cluster center clustering results; then, screen out a cluster center clustering result that meets the commuting travel characteristics from the K cluster center clustering results, which includes P private cars.

3. The method for predicting the destination of private car commuting based on checkpoint data according to claim 2 is characterized by: In step 4-5, the method for allocating dynamic weights includes the following steps: Step 4-5A: Calculate the mutual information MI of clustering indicators m : The a clustering index values ​​after normalization and standardization of each private car are taken as row data, then the a clustering index values ​​after normalization and standardization of Q private cars are formed into a clustering index value array of Q rows and a columns; According to the clustering index value array, the mutual information MI between the mth clustering index and the other m-1 clustering indexes is calculated. m ; Step 4-5B: Calculate the cluster stability score S m : Randomly extract a subset from the clustering index value array, and use the number of clusters C obtained in step 4-4 to perform several clusterings, and then obtain the clustering stability score S of the mth clustering index m ; Step 4-5C: Assign dynamic weights: According to the weights assigned at the t-1th clustering Clustering index mutual information MI m and cluster stability score S m , dynamically adjust the allocation weight of the mth clustering index during the tth clustering Step 4-5D, weight convergence: When the weight changes of all clustering indicators in two consecutive iterations are less than the weight threshold δ, the iteration is stopped and the dynamic weight of each clustering indicator after convergence is obtained.

4. The method for predicting the destination of private car commuting based on checkpoint data according to claim 3 is characterized by: In step 4-5C, The calculation formula is: In the formula, α is the smoothing coefficient that controls the influence of historical weight on current weight, and its value range is (0,1); β is the control coefficient of the influence of mutual information on weight, and its value range is (0,1).

5. The method for predicting the destination of private car commuting based on checkpoint data according to claim 4 is characterized by: In step 5-2, in the process of training the commuting travel destination prediction model, it is assumed that in the samples with correct predictions, the clustering index k m The mean value of In the sample with incorrect prediction, the clustering index k m The mean value of Then in step 6, let the distribution weight of the mth clustering index at the tth clustering be The dynamically updated value is but The calculation formula is: in: In the formula, C is the weight correction constant, which is an empirical constant; γ is the feedback coefficient, and its value range is (0,1); for Dynamically updated weight value.

6. The method for predicting the destination of a private car commuting trip based on checkpoint data according to claim 2 is characterized by: In step 4-1, a=6, and the six clustering index values ​​are: average first checkpoint time k1, average daily trip number k2, total trip days k3, maximum number of days with the same first trip destination k4, peak travel frequency k5, and departure time fluctuation k6; where k1∈[0,24], k2∈N and k2>0, k3∈[0,T d ] and k3∈Z, k4∈[0,T d ] and k4∈Z, k5∈[0,1], k6>0.

7. The method for predicting the destination of a private car commuting trip based on checkpoint data according to claim 1 is characterized by: In step 3, the method for dividing a single trip includes the following steps: Step 3-1, setting the travel time threshold: setting the travel time threshold between two adjacent checkpoint positions, the travel time threshold includes a minimum travel time threshold T1 and a maximum travel time threshold T2; wherein T1<T2; Step 3-2, setting the travel speed threshold: setting the travel speed threshold between two adjacent bayonet positions, the travel speed threshold includes a minimum travel speed threshold v1 and a maximum travel speed threshold v2; wherein v1<v2; Step 3-3, marking the first track point: each private car at each checkpoint is called a track point, and each vehicle has a track point sequence; for the track point sequence of the same vehicle, the travel state of the first track point is marked as "1"; Step 3-4: Traverse all the remaining trajectory points in the same vehicle trajectory point sequence and mark them. Let the time interval between the current trajectory point and the previous trajectory point be T. Then: A. When T ≤ T1, mark the travel status of the current trajectory point as "0". B. When T1 < T < T2 and v1 ≤ v ≤ v2, mark the travel status of the current trajectory point as "0", where v is the current speed of the current vehicle. C. When T1 < T < T2 and v < v1 or v > v2, it is considered that the current trajectory point belongs to a new trip. At this time, the marking method of the travel status is as follows: Determine whether there is a process point with a travel status marked as "0" in the previous trip record. If so, mark the travel status of the previous trajectory point as "1" and the travel status of the current trajectory point as "2" as the starting point of the second trip; otherwise, delete the previous trajectory point with a travel status marked as "1" and re-number the current trajectory point as "1". D. When T > T2, mark the travel status of the current trajectory point as "2". After traversing and marking all the trajectory points of the same vehicle, a trajectory point travel status sequence in the form of 1, 0, 0,..., 1, 2, 0, 0,..., 2,... is formed. Among them, "1, 0, 0,..., 1" represents the first trip; "2, 0, 0,..., 2" represents the second trip, and so on. Step 3-5: Repeat Step 3-4 to Step 3-5 to complete the trajectory point travel status sequences of all vehicles.

8. The method for predicting the destination of a private car commuting trip based on checkpoint data according to claim 1 is characterized by: In Step 5-1, the model optimization parameters in the commuting travel destination prediction model are determined using the Bayesian optimization method.

9. The method for predicting the destination of a private car commuting trip based on checkpoint data according to claim 8 is characterized by: In Step 5-1, the Bayesian optimization method includes the following steps: Step 5-1A: Construct an objective function f(θ) for the model parameter θ, and the specific expression is: f(θ) = Acc - λ·depth In the formula, Acc is the cross-validation accuracy, depth is the depth of the tree, and λ is the model complexity. Step 5-1B: Solve the model optimization parameter θ * : Optimize the parameter θ based on Gaussian process regression search model * , the specific expression is: i * = arg maxf(θ).

10. The method for predicting the destination of private car commuting based on checkpoint data according to claim 1, characterized in that: The external map platform in Step 1 is Amap.