Bus arrival time robust prediction method for arrival information missing sparse GPS data
By optimizing the geographic hash grid and multi-scale spatiotemporal network diagram, combined with machine learning models, the prediction problem of bus arrival time in scenarios of missing information or missing GPS data is solved, and high-precision prediction in a low data quality environment is achieved.
Patent Information
- Application Number
- CN202510828948.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-20
AI Technical Summary
The prior art is difficult to accurately predict the arrival time of bus vehicles in scenarios such as missing information or missing GPS data, especially in suburban areas and complex environments, resulting in inconvenience in bus route management and passenger travel planning.
The geographic hash grid is optimized through the K-means clustering algorithm and the pigeon flock optimization algorithm, combined with multi-scale spatiotemporal network diagram and Xgboost model, and the existing GPS data and urban point of interest data are used to construct a robust bus arrival time prediction method.
In a low data quality environment, the accuracy and robustness of bus arrival time prediction is improved, the cost is reduced, and the applicability is strong, and it is suitable for real-time prediction of multiple bus routes.
Smart Images

Figure CN120356359A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent transportation, and more specifically, to a robust prediction method for bus arrival duration facing sparse GPS data with missing arrival information. Background Art
[0002] Although the public transportation network has covered a vast area of the city, it has not yet fully achieved seamless connection across the entire region. For areas that cannot be covered by urban rail transit, such as suburban areas and areas between rail lines in the city, public bus lines play a crucial role in ensuring the basic travel needs of the general public. Providing accurate bus arrival duration prediction services can optimize the on-time rate of public bus lines, improve the level of refined management, help passengers reasonably plan their travel time, and improve the convenience of using public transportation.
[0003] There are differences in the data reported by bus vehicles in different cities. In the bus system of Beijing, due to the need to swipe cards for getting on and off, vehicle arrival data can be obtained, but in Shanghai, only the card is swiped for getting on, resulting in missing arrival information; the GPS packet loss rate exceeds 40% in complex environments such as tunnels. Usually, the arrival duration prediction model lacks the mining of the actual complex environment and geographical space network of vehicle operation, and it is difficult to ensure the accurate prediction of the arrival duration in such low-data-quality environments.
[0004] Therefore, there is an urgent need to construct a method that can ensure the accuracy of bus arrival duration prediction in scenarios of missing information or missing GPS data, and improve the robustness of bus arrival duration prediction in low-data-quality environments. Summary of the Invention
[0005] Aiming at the deficiencies of the existing technology, the purpose of the present invention is to provide a robust prediction method for bus arrival duration facing sparse GPS data with missing arrival information.
[0006] To achieve the above purpose, the present invention provides the following technical solutions:
[0007] A robust prediction method for bus arrival duration facing sparse GPS data with missing arrival information, and the specific process is as follows:
[0008] Step 1: Cluster to obtain a common arrival and stop geographical hash network.
[0009] As Figure 2 shown, use the K-means clustering algorithm to cluster the historical GPS data of vehicles near bus stops, and identify the geographical hash grid groups where vehicles of different bus lines arrive and stop;
[0010] Step 2: Construct a geographical space bus vehicle driving network diagram.
[0011] According to the GPS data of vehicle travel, the longitude and latitude of bus stop positions, and the longitude and latitude of urban points of interest (POIs), encode them into geohash grids with different precisions, and construct a multi-scale geospatial bus vehicle travel network diagram, so that different bus lines can share the same geohash grid. Even if the vehicle GPS data reporting is sparse and missing, a series of geohash grids can be matched from the network diagram according to the bus stop direction through the current vehicle position;
[0012] Step 3: Construct the attributes of the relationship edges of the network diagram.
[0013] Construct the travel duration between bus stops of the bus vehicle through the historical travel GPS data of the vehicle at different times and spaces as the attribute of the edge of the network diagram;
[0014] Step 4: Train the model to achieve the prediction of arrival time.
[0015] Construct the arrival time data based on the historical GPS data of the vehicle, and use the nodes and edges of the constructed multi-scale spatio-temporal network to train the model to learn the historical travel patterns of the vehicle and predict the arrival time of the bus line in real time;
[0016] Furthermore, the specific process of Step 1 is as follows:
[0017] Adopt the pigeon flock optimization algorithm to optimize the initial clustering centers of K-means clustering, improve the clustering effect, and enhance the representativeness of the selected geohash grids;
[0018] Apply the statistical importance analysis method to perform variance analysis on each grid, and select the geohash grid that can represent the characteristics of the group;
[0019] Evaluate the grid representativeness through cross-validation, screen the historical vehicle GPS data within 300 meters of the bus stop as the candidate set of geohash grids for selecting the stopping sites, and use it as the judgment basis for the arrival of the bus vehicle.
[0020] Furthermore, the specific process of constructing the geospatial bus vehicle travel network diagram in Step 2 includes:
[0021] Data collection: Collect the GPS trajectory data of the bus vehicle, including timestamp, longitude and latitude, speed, and direction; at the same time, collect the longitude and latitude of the bus stop and the longitude and latitude of the POI;
[0022] Grid encoding: Use the Geohash algorithm to encode the longitude and latitude into a string with a length of n to form geohash grids of different sizes, where n is the grid precision, and the value range is ;
[0023] Multi-scale grid generation: Generate a multi-scale grid system, such as grids with 5-digit, 6-digit, and 7-digit precisions, to form a set of Geohash grids with different spatial resolutions;
[0024] Trajectory mapping: Map the GPS data reported by the vehicle driving trajectory to the Geohash grid corresponding to the precision (Geohash strings of different digits) in time series to form a sequence of trajectory points:
[0025] ;
[0026] Among them, ;
[0027] Spatio-temporal network graph construction: Construct a spatio-temporal network graph, where the nodes are all the Geohash grids that appear, including the GPS trajectory network, bus stop grids, and POI grids; if the vehicle GPS trajectory grid appears continuously in the time series more than the threshold, a spatio-temporal relationship edge is formed; if the vehicle GPS trajectory grid is the same as the bus stop or POI grid, a matching relationship edge is formed.
[0028] Furthermore, the attributes of the network graph relationship edges in step three are constructed, and the specific process includes:
[0029] Travel time calculation: According to the adjacent Geohash grid pairs in the vehicle historical GPS trajectory sequence, extract their timestamps and calculate the travel duration , and the specific calculation formula is as follows:
[0030] ;
[0031] Among them, The vehicle GPS data at time and the vehicle GPS data at time
[0032] are the two pieces of data before and after in the GPS trajectory sequence; Statistical analysis: For grids of different scales, respectively calculate the mean and standard deviation
[0033] ;
[0034] Among them, is the data length, is the travel duration;
[0035] Spatio-temporal feature acquisition: Divide by Geohash grid to obtain historical weather; divide by timestamp into morning rush hour (07:00 - 09:00) and evening rush hour (17:00 - 19:00); mark by date as working day, weekend, or legal holiday;
[0036] Attribute storage: Adopt a sliding window mechanism to statistically calculate the mean and standard deviation of influencing factors such as weather, morning and evening rush hours, and holidays, and store them as edge attributes.
[0037] Furthermore, in step four, a model is used to predict the arrival time of bus vehicles. The specific process is as Figure 1 shown:
[0038] During prediction, first, based on the real-time GPS data obtained by the vehicle and the GPS data of the target bus stop to be predicted, node matching is performed in the pre-constructed multi-scale spatio-temporal network graph. After accurately locating the corresponding nodes, the attribute information of these nodes themselves and the attribute data of the edges between the nodes are extracted as the input of the model, and finally, the predicted value of the arrival time of the target bus stop to be predicted is output, realizing the effective prediction of the real-time arrival time of buses in a low data quality environment.
[0039] Compared with the prior art, the present invention has the following beneficial effects:
[0040] 1. Data-driven innovation: It does not rely on additional hardware devices, only uses existing GPS data, bus stop locations, and POI data, reducing costs and having strong applicability;
[0041] 2. Algorithm fusion advantages: The pigeon flock optimization algorithm is combined to optimize the clustering process, improving the accuracy of arrival grid judgment; the multi-scale spatio-temporal network graph fully considers the spatial and temporal dimensions, enhancing data feature expression;
[0042] 3. Accurate prediction: Through multi-source data feature extraction and machine learning model training, the prediction accuracy of bus arrival time in a low data quality environment is effectively improved, providing reliable support for traffic scheduling and public travel. Description of the Drawings
[0043] Figure 1 It is a schematic diagram of a robust prediction method for bus arrival time facing sparse GPS data with missing arrival information;
[0044] Figure 2 It is a schematic diagram for the present invention to judge whether a vehicle arrives based on vehicle GPS data. Detailed Embodiments
[0045] Example 1, referring to Figure 1 shown, the robust prediction method for bus arrival time facing sparse GPS data with missing arrival information in this embodiment is as follows:
[0046] Step 1: Cluster to obtain common arrival and docking geographic hash grids.
[0047] When passengers get on and off the bus at the arrival station, the position of the bus cannot move. However, since there is no device to collect data when the bus door opens and closes for passengers to get on and off in Shanghai, and there is no data on passengers swiping their cards when getting off, and the bus reports GPS data sparsely and with a delay, this environment with low data quality makes it difficult to accurately judge whether the bus has arrived at the station.
[0048] As Figure 2 shown: Use the K-means clustering algorithm to cluster the historical GPS data of vehicles near the bus station to identify the geographical hash grid groups where vehicles on different bus lines arrive and stop. The specific method is as follows:
[0049] 1. Optimize the initial clustering centers of K-means using the pigeon flock optimization algorithm to improve the clustering effect and enhance the representativeness of the selected geographical hash grids. The specific process includes:
[0050] 1.1 Data preprocessing: The input data is the feature vector of the geographical hash grid , and the target variable is to optimize the initial clustering centers of K-means , so that the initial centers are distributed as much as possible in the core areas of different data clusters. Encode each initial center candidate solution as a vector with a dimension of ( is the feature dimension);
[0051] 1.2 Initialize the pigeon flock population: Randomly select samples from the dataset as the initial candidate centers, ensuring that the centers in each individual are different from each other;
[0052] 1.3 Global exploration: Use the within-cluster sum of squares of K-means as the optimization objective to measure the clustering quality. The specific calculation formula is:
[0053] ;
[0054] Among them, is the sample set assigned to the th cluster, and the goal is to minimize ; For each individual , update the position according to the current position , the global optimal position and its own historical optimal position . The specific calculation formula is as follows:
[0055] ;
[0056] Among them is a random number within, is the number of iterations;
[0057] 1.4 Local development: Conduct a fine search near the global optimal region. Each individual adjusts its position according to its own movement information. The specific calculation formula is as follows:
[0058] ;
[0059] where is the scaling factor, is the mean vector of the current position of the individual.
[0060] 1.5 Iterative optimization: Repeat steps 3 - 4 until the change rate of the within - cluster sum of squares is less than 1% for 5 consecutive generations, and output the globally optimal individual as the initial clustering center of K - means;
[0061] 2. Apply the statistical importance analysis method to perform an analysis of variance for each grid. Calculate the difference in the eigenvalues of the grid between its own group and other groups in terms of the grid unit. Select the geohash grid that can represent the characteristics of the group. The selection criteria include being close to the bus stop, having a high GPS point density, and at the same time avoiding the redundancy of the reported GPS data before and after the arrival of the bus;
[0062] 3. Evaluate the grid representativeness through cross - validation. Screen the historical vehicle GPS data within 300 meters of the bus stop as the candidate set of geohash grids for selecting the stopping sites. The screened stopping site grids that can match more than 90% of the historical vehicle arrival GPS data are used as the judgment basis for the arrival of the bus, and try to restore the real parking position of the bus and the real boarding and alighting experience of the passengers as much as possible.
[0063] Step 2: Construct a geospatial bus vehicle travel network diagram.
[0064] According to the GPS data of vehicle travel, the longitude and latitude of bus stops, and the longitude and latitude of urban points of interest, encode them into geohash grids with different precisions to construct a multi - scale geospatial bus vehicle travel network diagram, so that different bus lines can share the same geohash grid. Even if the vehicle GPS data reporting is sparse and missing, a series of geohash grids can be matched from the network diagram in the direction of the bus stop according to the current position of the vehicle. The specific method is as follows:
[0065] Data collection: Collect bus vehicle GPS trajectory data, including timestamp, longitude and latitude, speed, and direction; at the same time, collect the longitude and latitude of bus stops and the longitude and latitude of POIs;
[0066] Grid encoding: Use the Geohash algorithm to encode the longitude and latitude into a string of length n to form geohash grids of different sizes, where n is the grid precision, and the value range is ;
[0067] Multi-scale grid generation: Generate a multi-scale grid system, such as grids with 5-digit, 6-digit, and 7-digit precisions, to form a set of GeoHash grids with different spatial resolutions;
[0068] Trajectory mapping: Map the GPS data reported by the vehicle driving trajectory to the GeoHash grids with corresponding precisions (GeoHash strings of different digits) in time series to form a sequence of trajectory points:
[0069] ;
[0070] Among them, ;
[0071] Space-time network graph construction: Construct a space-time network graph, where the nodes are all the GeoHash grids that appear, including the GPS trajectory network, bus stop grids, and POI grids; if the vehicle GPS trajectory grid appears continuously in the time series more than the threshold, a space-time relationship edge is formed; if the vehicle GPS trajectory grid is the same as the bus stop or POI grid, a matching relationship edge is formed;
[0072] Step 3: Construct the attributes of the relationship edges in the network graph.
[0073] Construct the driving duration between bus stops of the bus vehicle as the attribute of the network graph edge through the historical driving GPS data of the vehicle in different space-time. Specifically as follows:
[0074] Driving time calculation: Extract the timestamps according to the adjacent GeoHash grid pairs in the vehicle historical GPS trajectory sequence and calculate the driving duration , and the specific calculation formula is as follows:
[0075] ;
[0076] Among them, The vehicle GPS data at time and the vehicle GPS data at time
[0077] are the two consecutive data in the GPS trajectory sequence; Statistical analysis: For grids of different scales, respectively calculate the mean value
[0078] and the standard deviation;
[0079] Among them, is the data length, is the driving duration;
[0080] Spatio-temporal feature acquisition: Divide by geographical hash grid to obtain historical weather; divide the morning peak (07:00 - 09:00) and evening peak (17:00 - 19:00) by timestamp; mark the date as working day, weekend or legal holiday;
[0081] Attribute storage: Adopt a sliding window mechanism to statistically calculate the mean and standard deviation of influencing factors such as weather, morning and evening peaks, and holidays, and store them as edge attributes;
[0082] Step 4: Train a model to achieve the prediction of arrival time.
[0083] Construct arrival time data based on the historical GPS data of vehicles. Utilize the nodes and edges of the constructed multi-scale spatio-temporal network to train the model to learn the historical driving patterns of vehicles and predict the arrival time of bus lines in real time. The specific steps are as follows:
[0084] Training data construction: Use the time series of historical vehicle GPS data. According to what is described in Step 1, filter the GPS data whose distance from the target bus stop is less than 300 meters and whose geographical hash grid of the stopping site is hit, calculate the time of the GPS data reaching the target bus stop, and use the remaining GPS data as the input of the training data set. The target is to predict the arrival time ;
[0085] Model construction: Select the Xgboost model. The selection of each parameter is centered around balancing the model performance, avoiding overfitting and ensuring computational efficiency. The maximum tree depth max_depth is set to 8, which can prevent overfitting caused by too deep tree depth while ensuring that the model can capture complex features of the data; the learning_rate is taken as 0.1, enabling the model to converge quickly during training and not easily skip the optimal solution; considering both prediction accuracy and computational resources, n_estimators is set to 500; select reg:absoluteerror as the loss function and mae as the evaluation function, which is more robust to outliers that may appear in the prediction of bus arrival time; subsample is set to 0.8, and each tree is trained by randomly sampling 80% of the samples, effectively reducing the model variance, simulating the sparse situation of GPS data, and improving the generalization ability of the model;
[0086] Model training: Use and the node attributes of the corresponding geographical hash grid, such as encoding the id, scale size, administrative region where it is located, etc.; the attributes of the relationship edge, such as the mean and standard deviation of the historical driving duration, etc. as the input of the model. The output of the model is , and use the Mean Absolute Error (MAE) to evaluate the performance of the model. The specific calculation method is as follows:
[0087] ;
[0088] where m is the size of the data set, is the true arrival time, is the predicted arrival time;
[0089] Real-time prediction: As Figure 1 shown, during prediction, based on the real-time GPS data of the vehicle and the GPS data of the bus stop to be predicted, the nodes in the multi-scale spatio-temporal network graph are matched, and the attribute data of the nodes and relationship edges are obtained as the input of the model, and the predicted arrival duration value of the bus stop to be predicted is obtained.
[0090] The real-time bus arrival duration prediction scheme provided by the present invention uses a clustering algorithm combined with an optimization algorithm to obtain the commonly used parking geographical hash grid of the vehicle, which is used to judge whether the vehicle arrives at the bus stop, making up for the lack of information data on vehicle arrival in the original reported data in the application scenario of the present invention; in view of the sparse or missing reported data of vehicle GPS data, the historical GPS data of the vehicle and the urban point-of-interest data are fully utilized to construct the nodes and relationship edges of the multi-scale geographical space network graph, and combined with influencing factors such as dynamic weather and holidays, the prediction accuracy of the bus arrival duration in the low data quality environment is effectively improved.
[0091] An experiment was conducted to compare the historical vehicle GPS data sets of 3 bus lines in Shanghai, using the same decision tree model constructed by the traditional method and the method of the present invention. In the case of simulating the poor quality of the reported data in the real operation scenario and performing tests on the downsampled 30% test set, the average absolute error of the latter decreased by about 22% compared with the former, indicating the feasibility of the scheme of the present invention in the actual application scenario.
[0092] The above formulas are all dimensionless and take their numerical values for calculation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0093] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains a collection of one or more available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0094] It should be understood that in various embodiments of the present application, the magnitudes of the serial numbers of the above processes do not mean the order of execution is prior or subsequent. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0095] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0096] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0097] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.
[0098] If the above functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs and other various media that can store program codes.
[0099] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A robust prediction method for bus arrival duration facing sparse GPS data with missing arrival information, characterized in that The method process is as follows: Step 1: Cluster to obtain the common arrival and stop GeoHash network. Use the K-means clustering algorithm to identify the GeoHash grid groups where vehicles of different bus lines arrive and stop. Step 2: Construct a geospatial bus vehicle travel network diagram. Collect bus vehicle operation data, encode to generate multi-scale GeoHash grids, map the GPS trajectories to the grids to form sequences, and construct a spatio-temporal network diagram. Step 3: Construct the attributes of the relationship edges in the network diagram. Through the historical driving GPS data of vehicles, construct the travel duration of bus vehicles between bus stops as the attribute of the edges in the network diagram. Step 4: Train the model to achieve arrival duration prediction. Based on the historical GPS data of vehicles, construct arrival duration data, and use the nodes and edges of the multi-scale spatio-temporal network to train the model to learn the historical driving patterns and predict the bus arrival duration in real time.
2. The robust prediction method for bus arrival time of sparse GPS data with missing arrival information according to claim 1, wherein, The specific process of Step 1 includes: Use the pigeon flock optimization algorithm to optimize the initial clustering centers of K-means clustering. Apply the statistical importance analysis method to perform variance analysis on each grid, and select the GeoHash grid that can represent the characteristics of the group. Evaluate the grid representativeness through cross-validation, screen the historical vehicle GPS data within 300 meters of the bus stop as the candidate set of GeoHash grids for the selected stop sites, and use it as the judgment basis for bus vehicle arrivals.
3. The robust prediction method for bus arrival duration facing sparse GPS data with missing arrival information according to claim 2, characterized in that Use the pigeon flock optimization algorithm to optimize the initial clustering centers of K-means clustering. The specific process includes: Data preprocessing: The input data is the feature vector of the geohash grid , and the goal is to optimize the initial clustering centers of K-means . Each initial center candidate solution is encoded as a vector with a dimension of , where is the feature dimension; Initialize the pigeon flock population: Construct the initial candidate centers. Global exploration: Use the within-cluster sum of squares of K-means as the optimization objective. For each individual, update the position according to the current position, the global optimal position, and its own historical optimal position. Local exploitation: Perform fine search near the global optimal area, and each individual adjusts its position according to its own movement information. Iterative optimization: Repeat global exploration and local exploitation until the change rate of the within-cluster sum of squares is less than 1% for 5 consecutive generations, and output the global optimal individual as the initial clustering center of K-means.
4. The robust prediction method for bus arrival duration facing sparse GPS data with missing arrival information according to claim 1, characterized in that The specific process of constructing the geospatial bus vehicle travel network diagram in Step 2 includes: Data collection: Collect bus vehicle GPS trajectory data, the longitude and latitude of bus stops, and the longitude and latitude of POIs. Grid encoding: Encode the longitude and latitude into a string of length n, where n is the grid precision. Multi-scale grid generation: Generate a multi-scale grid system to form a set of GeoHash grids with different spatial resolutions. Trajectory mapping: Map the GPS data reported by the vehicle travel trajectory to the GeoHash grids of the corresponding precision in time series to form a trajectory point sequence. Spatio-temporal network diagram construction: Construct a spatio-temporal network diagram, where the nodes are all the GeoHash grids that appear; determine the spatio-temporal relationship edges and the matching relationship edges.
5. The robust prediction method for bus arrival time based on sparse GPS data with missing arrival information according to claim 1, characterized in that The method for determining the attributes of the relationship edges in the network diagram in Step 3 is as follows: Travel time calculation: Calculate the travel duration according to the historical GPS trajectory sequence of the vehicle. Statistical analysis: For grids of different scales, respectively calculate the mean and standard deviation of the travel time. Spatio-temporal feature acquisition: Obtain the historical weather according to the GeoHash grid division. Divide the morning peak and evening peak according to the timestamp; mark the date as a weekday, weekend, or legal holiday. Attribute storage: Adopt a sliding window mechanism to calculate the mean and standard deviation of the impact factors of weather, morning and evening rush hours, and holidays, and store them as the attributes of edges.
6. The bus arrival duration robust prediction method for sparse GPS data with missing arrival information according to claim 1, wherein The specific process of using the trained model to predict the arrival time of bus vehicles in Step 4 is as follows: First, based on the real-time GPS data obtained by the vehicle and the GPS data of the target bus stop to be predicted, perform node matching in the pre-constructed multi-scale spatio-temporal network graph. After accurately locating the corresponding nodes, then extract the attribute information of these nodes themselves and the attribute data of the edges representing the relationships between the nodes as the input of the model, and finally output the predicted value of the arrival time of the target bus stop to be predicted.
Citation Information
Patent Citations
Method and device for predicting total driving time of bus from starting point to ending point
CN110570678A
Trajectory prediction method and system
CN110909106A
Designated driver scheduling method based on order prediction
CN113256015A
Method for predicting arrival time of bus with missing data
CN113470365A
Bus arrival time uncertainty visualization method and system, device and medium
WO2024125253A1