Intercity bus route recommendation method and device based on mobile signaling data
By preprocessing, aggregating and clustering mobile signaling data, we generate intercity bus recommended routes, which solves the problem that existing intercity bus routes cannot meet passengers' actual travel needs, and achieves more efficient operations and stronger attractiveness.
Patent Information
- Application Number
- CN202510281279.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-27
AI Technical Summary
The recommendation and planning of existing intercity bus routes rely on experience data and limited survey information, which cannot effectively meet passengers' actual travel needs, resulting in insufficient operational efficiency and attractiveness.
The intercity bus line recommendation method based on mobile signaling data is adopted, and the mobile signaling data is preprocessed, aggregated and clustered, and the user's travel trajectory chain is generated and the intercity bus line is recommended. The specific steps include preprocessing the mobile signaling data to obtain intercity travel data, aggregation processing generates user travel trajectory chains, and generating intercity bus recommended routes through DBSCAN, K-MEANS or BERT+K-MEANS clustering algorithms.
By using the location and time information of mobile signaling data, the actual flow of people can be more comprehensively reflected, the route planning deviations caused by fixed macro factors can be avoided, the travel habits and needs of different groups of people can be accurately captured, and the operational efficiency and attractiveness of intercity buses can be improved.
Smart Images

Figure CN120219136A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a route recommendation method in the field of transportation technology, and more particularly to an intercity bus route recommendation method based on mobile signaling data, and also to an intercity bus route recommendation device based on mobile signaling data. Background Art
[0002] With the acceleration of the urbanization process and the coordinated development of regional economies, the flow of people between cities is becoming increasingly frequent. As an important link connecting different cities, the convenience of intercity transportation plays a crucial role in regional development and people's travel experience. Traditional intercity transportation methods mainly include railways, long-distance passenger transport, etc. However, as a flexible and wide-coverage transportation form, bus routes also have non-negligible potential in intercity transportation.
[0003] Currently, the recommendation and planning of intercity bus routes mostly rely on empirical data and limited survey information. For example, the transportation department determines the route direction based on macroscopic factors such as the main population aggregation points and economic activity centers between cities, or understands the travel needs of passengers through online questionnaires. However, these methods have many limitations. On the one hand, empirical data is difficult to comprehensively reflect the actual flow of people. The population flow pattern between cities is complex and changeable, affected by various factors such as time, season, and special events. Relying solely on fixed macroscopic factors is likely to lead to deviations between route planning and actual needs. On the other hand, the sample size of small-scale questionnaires is limited, and there is a problem of sample bias, which cannot accurately capture the travel habits and needs of different groups of people. Therefore, the existing technology fails to effectively mine valuable information to support the scientific planning and accurate recommendation of intercity bus routes, resulting in intercity bus routes being unable to better meet the actual travel needs of passengers, and reducing the operation efficiency and attractiveness of intercity buses. Summary of the Invention
[0004] To solve the existing technical problems, the present invention provides an intercity bus route recommendation method and device based on mobile signaling data, which solves the problem that the existing intercity bus routes cannot meet the actual travel needs of passengers.
[0005] The present invention is implemented by the following technical solutions: An intercity bus route recommendation method based on mobile signaling data, which includes the following steps:
[0006] S1: Preprocess the mobile signaling data to obtain intercity travel data, and execute step S2;
[0007] S2: Aggregate the intercity travel data to generate user travel trajectory chains, and execute step S3;
[0008] S3: Cluster the user's travel trajectory chain to generate recommended intercity bus routes;
[0009] Among them, the aggregation processing method of the intercity travel data includes the following steps:
[0010] S2.1: Perform grid processing on the base station H3 of the intercity travel data, and execute step S2.2;
[0011] S2.2: Associate the target user signaling data with the H3 grid coding table, and execute step S2.3;
[0012] S2.3: Determine whether the user stays in the H3 grid. If so, calculate the continuous stay time of the user in the H3 grid, and execute step S2.4;
[0013] S2.4: Determine whether the continuous stay time is lower than a preset time. If so, delete the corresponding user data.
[0014] As a further improvement of the above solution, the step S1 includes:
[0015] Screen out the user data of intercity travel from the mobile signaling data;
[0016] Delete ping-pong drift data, duplicate or invalid data in the user data, and convert the timestamp and coordinate type to screen out the intercity OD data to obtain the intercity travel data.
[0017] As a further improvement of the above solution, in step S2.3, first sort the stay time of the user in a single grid to obtain a sorting value, and count the time of entering the grid earliest. Then, according to the sorting value and the time of entering the grid earliest, judge whether each subsequent moment belongs to the time of the first moment. Finally, judge whether the user stays in the H3 grid according to the time of the first moment.
[0018] Furthermore, the time T start The calculation formula of t is:
[0019] T start = T - RK * 1
[0020] Among them, T start Is the time of the first moment, T is the time of entering the grid earliest, and RK is the sorting value.
[0021] As a further improvement of the above solution, in step S3, the clustering of the user travel trajectory chain is implemented by one or more of the DBSCAN clustering algorithm, the K-MEANS clustering algorithm, and the BERT+K-MEANS clustering algorithm.
[0022] Further, the BERT+K-MEANS clustering algorithm includes the following steps:
[0023] Train the user travel trajectory chain through the BERT model;
[0024] Construct a vocabulary based on the grid codes of the H3 grid and encode it;
[0025] Convert the user travel chain data into Embedding vectors;
[0026] Verify and evaluate the Embedding vectors, and use the K-MEANS clustering algorithm for clustering based on the converted Embedding vectors.
[0027] Furthermore, the training method of the user travel trajectory chain through the BERT model includes the following steps;
[0028] Convert the H3 grid code passed by the user travel trajectory chain into a basic text processing unit, and encode the basic text processing unit;
[0029] Train the encoded data using the Bert model and convert the data into Embedding vectors;
[0030] Randomly extract user travel chain samples, and evaluate the Embedding vectors by comparing the similarity between the user travel trajectory chain before conversion and the user travel trajectory chain after being converted into Embedding vectors.
[0031] As a further improvement of the above solution, the calculation formula for the continuous stay time is:
[0032] T stay =T max -T min
[0033] Wherein, T stay is the continuous stay time, T max is the last moment when the user stays in the grid, and T min is the starting moment when the user stays in the grid.
[0034] As a further improvement of the above solution, step S2 further includes: first aggregating the user data to the H3 grid, then sorting the H3 grid according to the time series, and finally generating the user travel trajectory chain according to the sorted H3 grid.
[0035] The present invention also provides an intercity bus line recommendation device based on mobile signaling data, which includes:
[0036] A preprocessing module for preprocessing mobile signaling data to obtain intercity travel data;
[0037] An aggregation processing module for aggregating the intercity travel data to generate a user travel trajectory chain; and
[0038] A clustering module for clustering the user travel trajectory chain to generate recommended intercity bus routes;
[0039] Wherein, the method for the aggregation processing module to aggregate the intercity travel data includes the following steps:
[0040] S2.1: Perform grid processing on the base station H3 of the intercity travel data, and execute step S2.2;
[0041] S2.2: Associate the target user signaling data with the H3 grid coding table, and execute step S2.3;
[0042] S2.3: Determine whether the user stays in the H3 grid. If so, calculate the continuous stay time of the user in the H3 grid, and execute step S2.4;
[0043] S2.4: Determine whether the continuous stay time is lower than a preset time. If so, delete the corresponding user data.
[0044] The intercity bus route recommendation method and device based on mobile signaling data of the present invention use mobile signaling data to recommend and plan intercity bus routes, make full use of a large amount of location information and time information contained in the mobile signaling data, have a larger data sampling volume, a wider coverage, and higher data accuracy, can comprehensively reflect the actual personnel flow situation, avoid the deviation problem between the route planning and the actual demand caused by fixed macroscopic factors, can accurately capture the travel habits and demands of different groups of people, can better meet the actual travel needs of passengers, and improve the operation efficiency and attractiveness of intercity buses. Description of the Drawings
[0045] Figure 1 It is a flowchart of the intercity bus route recommendation method based on mobile signaling data according to Embodiment 1 of the present invention;
[0046] Figure 2 It is a distribution diagram of the user stay time in the base station in the intercity bus route recommendation method based on mobile signaling data according to Embodiment 2 of the present invention;
[0047] Figure 3 It is a schematic diagram of the user travel chain in Embodiment 2 of the present invention;
[0048] Figure 4 It is a DBSCAN clustering effect diagram in Embodiment 2 of the present invention;
[0049] Figure 5 It is the clustering effect diagram of 0 ≤ label ≤ 1000 in Embodiment 2 of the present invention;
[0050] Figure 6 It is the clustering effect diagram of 1000 < label ≤ 2000 in Embodiment 2 of the present invention;
[0051] Figure 7 It is the clustering effect diagram of label > 2000 in Embodiment 2 of the present invention;
[0052] Figure 8 It is the schematic diagram of the overlapping trajectories of User 1 and User 3 in Embodiment 2 of the present invention;
[0053] Figure 9 It is the schematic diagram of the overlapping trajectories of User 2 and User 3 in Embodiment 2 of the present invention;
[0054] Figure 10 It is the schematic diagram of the travel chains of different users in the same category in Embodiment 2 of the present invention. Detailed implementation manners
[0055] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0056] Embodiment 1
[0057] Please refer to Figure 1 , this embodiment provides an intercity bus line recommendation method based on mobile signaling data. This intercity bus line recommendation method mainly includes three steps, namely steps S1 - S3, and in some other embodiments, it may also include some other steps.
[0058] Step S1: Preprocess the mobile signaling data to obtain intercity travel data. Step S1 includes the following sub - steps: Screen out the user data of intercity travel from the mobile signaling data; Delete the ping - pong drift data, duplicate or invalid data in the user data, and convert the time stamp and coordinate type, and screen out the intercity OD data to obtain the intercity travel data. If the data does not meet the requirements after preprocessing, further processing is required.
[0059] Step S2: Aggregate the intercity travel data to generate the user travel trajectory chain. In this embodiment, first aggregate the user data to the H3 grid, then sort the H3 grid according to the time series, and finally generate the user travel trajectory chain based on the sorted H3 grid. Specifically, the method for aggregating the intercity travel data includes the following steps S2.1 - S2.4.
[0060] S2.1: Grid processing of base station H3 for intercity travel data. Here, for the associated base stations, the H3 grid system is used for grid processing. Each base station is mapped to a specific hexagonal grid according to its geographical location and is assigned a unique base station grid code. The H3 grid is a hierarchical spatial data structure used to divide the Earth's surface into hexagonal grids. This division method is similar to the structure of a honeycomb. The size of each hexagonal grid unit is basically the same, and different H3 grid precisions can be selected according to needs. There are 15 levels of H3 grid precision, which determine the size of the grid unit. The higher the precision level, the smaller the H3 grid unit, and the finer the division. Part of the H3 grid precision is shown in Table 1. Since this embodiment analyzes user trajectories and does not require particularly high positioning accuracy, while the data volume of mobile signaling data of mobile phones is too large, the base stations corresponding to the signaling data are processed by H3 grid.
[0061] Table 1 Partial H3 Grid Precision Table
[0062]
[0063] S2.2: Associate the signaling data of the target user with the H3 grid coding table. Through the user information and base station information contained in the signaling data, the signaling data of each user is associated with the H3 grid to which the corresponding base station belongs, so as to aggregate each user onto the H3 grid and complete the aggregation operation of users in the same area.
[0064] S2.3: Determine whether the user stays in the H3 grid. If so, calculate the continuous stay time of the user in the H3 grid. After the H3 grid processing of the base stations, a single H3 grid area will cover multiple base stations. When a user travels intercity, there will be a base station handover phenomenon during the movement process. This base station handover phenomenon may cause a user to enter and leave the same H3 grid multiple times, resulting in multiple stay time records in the same grid.
[0065] In this embodiment, first sort the stay time of the user in a single grid to obtain a sorting value, and count the time of the earliest entry into the grid. Then, based on the sorting value and the time of the earliest entry into the grid, determine whether each subsequent moment belongs to the time of the first moment. Finally, based on the time of the first moment, determine whether the user stays in the H3 grid. Among them, the time Tstart The calculation formula of t is:
[0066] T start = T - RK * 1 (1)
[0067] Where, T startLet \(t_0\) be the time of the first moment, \(T\) be the time when entering the grid earliest, and \(RK\) be the sorting value. Sort the staying time of the user in a certain grid, with the sorting time interval accurate to 1 minute, and use the above formula (1) to judge whether each subsequent moment belongs to the time of the first moment.
[0068] In this embodiment, the calculation formula for the continuous staying time is:
[0069] \(T\) stay \(=\) max \(T\) min \((2)\)
[0070] where \(T\) stay is the continuous staying time, \(T\) max is the last moment when the user stays in the grid, and \(T\) min is the starting moment when the user stays in the grid.
[0071] As shown in Table 2, the user enters a certain H3 grid multiple times, and the grid code is 883082401fffff. Sort the time when the user enters the grid at 1-minute intervals. As can be seen from Table 2, the user enters this grid 6 times on February 3, 2024. According to the calculations of formula (1) and formula (2), the continuous staying time \(T\) stay of this user in this grid is 4 minutes.
[0072] Table 2 Example of the discrimination logic for the continuous staying time of the grid
[0073]
[0074] In this embodiment, sort the connection times of the user in the H3 grid according to the time series. After sorting, filter out the data where \(T\) stay is greater than or equal to 5 minutes according to S2.3.1 and S2.3.2. In some other embodiments, the size of the preset time is set according to actual needs and can be adaptively changed according to different situations such as different cities and different seasons.
[0075] S2.4: Determine whether the continuous stay time is lower than a preset time. If so, delete the corresponding user data. If a user stays in a certain grid for too short a time, then such data loses its meaning for analyzing the user's trajectory. Therefore, it is very necessary to determine whether the user stays in the grid and the user's continuous stay time in the grid. After the data is further aggregated, the H3 grids of the user sorted by time series are obtained. These grids are numbered and connected in chronological order to obtain the user travel trajectory chain. Calculate the total mileage of the user travel trajectory chain, and filter out the trajectories with a total travel mileage greater than or equal to 30 km. Similarly, in other embodiments, the screening criteria for the total travel mileage are also different, which will not be elaborated here.
[0076] Step S3: Cluster the user travel trajectory chain to generate an intercity bus recommended route. In this embodiment, the clustering of the user travel trajectory chain is implemented by one or more of the DBSCAN clustering algorithm, the K-MEANS clustering algorithm, and the BERT+K-MEANS clustering algorithm. In some other embodiments, other clustering algorithms can also be used for implementation. There are mainly three common clustering algorithms, namely the DBSCAN clustering algorithm, the K-MEANS clustering algorithm, and the hierarchical clustering algorithm. For the hierarchical clustering algorithm, hierarchical clustering is divided into two ways: agglomerative clustering and divisive clustering. However, whether it is the gradual merging of agglomerative clustering or the gradual splitting of divisive clustering, both face the problem of high computational complexity. And the data involved in this embodiment - mobile signaling data - has a very large amount. Therefore, if the hierarchical clustering is used to process the mobile signaling data, the computational complexity will be extremely high and the amount of calculation will be very large. Therefore, only the DBSCAN clustering and the K-MEANS clustering algorithms are considered in this embodiment.
[0077] DBSCAN is a density-based clustering algorithm. It realizes clustering by defining parameters such as the neighborhood radius (eps) and the minimum number of points (Minpts) to determine the density of data points. If the number of data points in the neighborhood of a point is greater than or equal to the minimum number of points, it is a core point; the non-core points in the neighborhood of the core point are border points; those that are neither core nor border are noise points. Based on the core points, clusters are expanded outward, and the data points with connected densities are divided into the same cluster, thereby discovering clusters of any shape and at the same time being able to identify noise points. K-MEANS is a partitioning-based clustering algorithm. It first selects K initial center points (i.e., the centroids of the clusters), and assigns each data point to the cluster center closest to it according to the distance between each data point and the K center points. Then, the new cluster center is calculated based on the mean of all data points in the cluster. Finally, the calculation of the new cluster center is repeated until the cluster center no longer changes or changes very little, and finally K clustering results are obtained.
[0078] In addition, the BERT+K-MEANS clustering algorithm includes the following steps: training the user travel trajectory chain through the BERT model; constructing a vocabulary based on the grid codes of the H3 grid and encoding; converting the user travel chain data into Embedding vectors; validating and evaluating the Embedding vectors, and clustering using the K-MEANS clustering algorithm based on the converted Embedding vectors.
[0079] Specifically, the training method of the user travel trajectory chain by the BERT model in this embodiment includes three steps: (1) customizing the vocabulary; (2) training the Bert model; (3) evaluating the converted Embedding. Specifically, first, convert the H3 grid encoding passed by the user travel trajectory chain into a basic text processing unit, and encode the basic text processing unit; then, use the Bert model to train the encoded data and convert the data into Embedding vectors; finally, randomly extract user travel chain samples, and evaluate the Embedding vectors by comparing the similarity between the user travel trajectory chain before conversion and the user travel trajectory chain converted into Embedding vectors. Cluster the converted Embedding vectors to obtain a clustering result. After verifying the clustering result, the recommended bus route can be obtained.
[0080] The core of the Bert model in this embodiment lies in its bidirectional encoding characteristics and the Transformer architecture it uses. That is, Bert is based on the Transformer architecture and adopts a multi-layer bidirectional attention mechanism, which can simultaneously focus on all token information in the context of the text. During pre-training, through the masked language model (MLM) task, randomly mask the text tokens and let the model predict bidirectionally according to the context, comprehensively capture semantic relationships and syntactic structures, understand the text semantics more accurately and deeply, and at the same time can better capture long-distance dependencies.
[0081] The core goal of this embodiment is to process the user travel chain data. Considering that the user travel chain has obvious temporal characteristics, that is, the user is connected in chronological order at each node of the travel chain, it is very necessary to analyze the front and back locations sorted by time for each user to discover travel patterns. The Bert model has good bidirectional encoding ability and can well understand the context semantics, which is very suitable for the task of the user travel chain with temporal characteristics to be processed in this embodiment. Therefore, this embodiment considers using the Bert model to train the data and convert the user travel chain data into Embedding vectors, so as to better prepare for the next clustering.
[0082] Among them, the Embedding vector is a method of converting high-dimensional sparse data (such as words in text) into a low-dimensional, dense real vector representation. This technology effectively reduces the dimension of the data by mapping words into a continuous vector space. Specifically, the Embedding vector maps the words in the vocabulary to a space with a fixed dimension (common dimensions are 50, 100, 256, etc., and 256 dimensions are selected in this embodiment), thus significantly reducing the computational complexity and storage cost.
[0083] The intercity bus line recommendation method based on mobile signaling data in this embodiment uses mobile signaling data to recommend and plan intercity bus lines, making full use of the large amount of location information and time information contained in mobile signaling data. The sampled data volume is larger, the coverage is wider, and the data accuracy is higher. It can comprehensively reflect the actual personnel flow situation, avoid the deviation problem between line planning and actual needs caused by fixed macro factors, accurately capture the travel habits and needs of different people, better meet the actual travel needs of passengers, and improve the operation efficiency and attractiveness of intercity buses.
[0084] Embodiment 2
[0085] Please refer to Figures 2 - 10 , this embodiment provides an intercity bus line recommendation method based on mobile signaling data. This embodiment conducts an example analysis on the basis of Embodiment 1. This embodiment selects the mobile signaling data of Anhui Province and takes Hefei City as the center to mine the mobile signaling data and analyze the user's travel chain to obtain the user's recommended routes.
[0086] In step S1, data samples are collected. The daily data volume of mobile signaling data is about 8 billion per day. With the help of the Spark distributed computing framework and the distributed Hive database for processing, after removing unreasonable data such as missing data, duplicate data, and drift data in the data, the intercity travel data with the starting point in Hefei City and arriving at different cities except Hefei on the same day is screened. The data that meets the rules after screening is about 520,000. The data sample is shown in Table 3.
[0087] Table 3 Sample data of user-connected base stations
[0088]
[0089] Sample the cleaned data, and then statistically analyze the residence time of each base station where the signaling data is located. The obtained base station residence time is as Figure 2 shown. From Figure 2It can be seen that the residence time of users at the base station is concentrated within 20 minutes, and the number of people staying at a certain base station for a long time is too small. Given that the base stations to which the mobile signaling data of the vast majority of users belong should not change during the night, the residence time of the user signaling data at the base station should be very long. Therefore, this data does not conform to the actual situation. Through further analysis and research, it is found that when the user is in a stationary state, the base station connected by the user is dynamically changing. Therefore, further aggregation processing of the signaling data is required.
[0090] In step S2, the data is further aggregated. The base station H3 is gridified. The base station is gridified by H3 to obtain the corresponding grid code. Some sample data is shown in Table 4. For the H3 grid, the higher the precision, the smaller the area of the hexagon, and the finer the division. In this embodiment, considering the analysis of user trajectory data, the precision requirement is not particularly high. Therefore, the H3 grid precision considers levels 6, 7, and 8.
[0091] Table 4 Sample data of H3 grid coding
[0092]
[0093] Furthermore, the user is associated with the grid information according to the user information. The user is aggregated onto the H3 grid according to the longitude and latitude of the base station to which the existing user signaling data belongs.
[0094] Next, calculate the continuous residence time of the user within the H3 grid. First, judge whether the user stays within the grid according to formula (1) in Embodiment 1, and then calculate the continuous residence time of the user within the grid according to formula (2) in Embodiment 1. After calculation, filter out the data with T stay ≥5min, and sort the H3 grids according to the connection grid time series.
[0095] In step S3, generate the user travel trajectory. After further aggregating and processing the data in the previous step, sort and connect the H3 grids passed by the user according to the time series to generate the user travel trajectory. Filter out the travel chains with a total mileage greater than or equal to 30 km as the finally obtained user travel trajectory. The sample diagram of the user travel chain is as Figure 3 shown.
[0096] For the user travel chain, select the longitude and latitude of the grid center of the travel chain and use the DBSCAN algorithm for clustering.
[0097] Parameter setting. First, sample part of the data to explore whether the clustering effect of the DBSCAN algorithm is good. Uniformly sample 100,000 people from the existing data, cluster the user travel chains of these data, set the scanning radius eps = 0.001, and the minimum number of neighborhood samples Minpts = 10.
[0098] Clustering results. After clustering using the DBSCAN algorithm, the number of clusters is 3226 and the number of noise points is 10282. From the clustering results, there are too many cluster categories and too many noise points. From the overall clustering effect, Figure 4 As shown in the figure, the clustering effect is not ideal. The clustering results are further classified and displayed, showing the data of 0≤label≤1000, 1000<label≤2000 and label>2000 respectively, as shown in the figure. Figure 5 , Figure 6 , Figure 7 As shown in the figure, the distribution of data points is uneven. For example, in some areas, the data points are very dense, forming relatively compact clusters (such as green and cyan areas), while in other areas, the data points are sparse, and these sparse points may belong to a certain category, but cannot be divided when clustering. At the same time, it can be seen from the figure that data points of different categories have a lot of overlap and intersection in some areas. For example, the green and cyan data points are intertwined in some areas, which shows that the boundaries between clusters are not clear.
[0099] Bert+K-MEANS clustering. Considering that DBSCAN clustering effect is not good, the K-MEANS clustering algorithm is considered. The K-MEANS clustering algorithm assumes that the data is spherically distributed and the cluster center is clear. For the user travel chain obtained after processing the mobile signaling data, it does not conform to the spherical distribution, and the travel chain distribution is relatively scattered, which does not meet the basic requirements of the K-MEANS algorithm for data. If the clustering effect is used directly, it is likely to be unsatisfactory. Therefore, in this embodiment, the deep learning model Bert is considered for training, and then K-MEANS clustering is used.
[0100] Bert model training process. In this embodiment, the Bert model training process is mainly as follows: first, a custom vocabulary and encoder are constructed, that is, the H3 grids that the user's travel chain passes through are converted into a vocabulary list using a specific segmentation method to obtain the basic unit (token) of text processing, such as segmenting the grid code "8528347ffffffff" into "8528347" and "###fffffffff". These tokens are then encoded (partial encoding is shown in Table 5), and after encoding, the Bert model is used for training, and finally the user's travel chain data is converted into an Embedding vector, and then the converted Embedding vector is evaluated and verified.
[0101] Table 5. Grid coding vocabulary
[0102]
[0103] Verification and evaluation of the converted Embedding vectors. In this embodiment, a part of the trip chain data is randomly selected to verify the relationship between the user trajectory coincidence degree before the vector conversion and the user trajectory coincidence degree after being converted into Embedding vectors.
[0104] User trajectory coincidence degree in the trip chain data sample. The user trip chains of some samples are tested as shown in Table 6. Table 6 shows the trip chain data of three users. It can be seen from Table 6 that the trip trajectory grids of User 1, User 2, and User 3 are sorted according to the time series, and the orders are as follows:
[0105] (1) 85283473ffffffff → 85283477ffffffff → 8528347bffffffff → 8528347bffffffff → 8928347bffffffff
[0106] (2) 85273473ffffffff → 85283487ffffffff → 8528348bffffffff → 8524348bffffffff
[0107] (3) 85283473ffffffff → 85283477ffffffff → 8528348bffffffff
[0108] Table 6 User trip chain trajectory test samples
[0109]
[0110]
[0111] From Figure 8 and Figure 9 it can be seen that there is a trajectory coincidence between User 1 and User 3, there is a trajectory coincidence between User 2 and User 3, and the trajectory coincidence degree between User 1 and 3 is higher than that between User 2 and 3. There is no estimated coincidence between User 1 and User 2.
[0112] User trajectory coincidence degree after being converted into Embedding vectors. After converting the above user trajectories into 256-dimensional Embedding vectors through the Bert model, the cosine similarity between every two users is calculated, and the calculation results are shown in Table 7. It can be seen from Table 7 that the cosine similarity between User 1 and 3 after being converted into Embedding vectors is 0.563, and the cosine similarity between User 2 and 3 is 0.361, which is consistent with the user trajectory similarity in S5.3.1, indicating that the training effect of the Bert model is good.
[0113] Table 7 Cosine similarity between users after being converted into Embedding vectors
[0114]
[0115] K - MEANS clustering. After converting the user travel trajectory data into Embedding vectors, use the K - MEANS clustering algorithm to cluster the Embedding vectors, then extract the routes based on the clustering results, and finally obtain 90 optimal clustering categories.
[0116] Finally, verify the clustering results. Compare the travel chains of different users in a randomly selected category of the clustering results to verify the clustering results. As shown in Table 8, extract the travel chains of two randomly selected users (denoted as User 1 and User 2 respectively) in category 16 of the clustering results. Visualize the travel chain results (the display graph is a partial travel chain), as Figure 10 shown, where the red H3 grid represents the travel chain data of User 1, and the green H3 grid represents the travel chain data of User 2. It can be seen from the figure that the travel chains of User 1 and User 2 are basically on the same trajectory, and the aggregation effect is good.
[0117] Table 8 Travel chains of different users in the same category
[0118]
[0119] Example 3
[0120] This example provides an inter - city bus line recommendation device based on mobile signaling data. The device includes a pre - processing module, an aggregation processing module, and a clustering module.
[0121] The pre - processing module is used to pre - process the mobile signaling data to obtain inter - city travel data.
[0122] The aggregation processing module is used to perform aggregation processing on the inter - city travel data to generate user travel trajectory chains.
[0123] Among them, the method for the aggregation processing module to aggregate and process the inter - city travel data includes the following steps:
[0124] S2.1: Perform H3 grid processing on the base stations of the inter - city travel data; S2.2: Associate the target user signaling data with the H3 grid coding table; S2.3: Determine whether the user stays in the H3 grid. If so, calculate the continuous stay time of the user in the H3 grid; S2.4: Determine whether the continuous stay time is lower than a preset time. If so, delete the corresponding user data.
[0125] The clustering module is used to cluster the user travel trajectory chains to generate inter - city bus recommended routes.
[0126] The above three modules can be specifically designed with reference to the steps of the intercity bus line recommendation method based on mobile signaling data in Embodiment 1 or Embodiment 2, and will not be elaborated here.
[0127] Embodiment 4
[0128] This embodiment provides a computer terminal, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the intercity bus line recommendation method based on mobile signaling data in Embodiment 1 or Embodiment 2 are implemented.
[0129] When the method in Embodiment 1 or Embodiment 2 is applied, it can be applied in the form of software. For example, it can be designed as an independently running program and installed on a computer terminal. The computer terminal can be a computer, a smart phone, a control system, and other Internet of Things devices, etc. The method in Embodiment 1 or Embodiment 2 can also be designed as an embedded running program and installed on a computer terminal, such as installed on a single-chip microcomputer.
[0130] Embodiment 5
[0131] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the intercity bus line recommendation method based on mobile signaling data in Embodiment 1 or Embodiment 2 are implemented.
[0132] When the method in Embodiment 1 or Embodiment 2 is applied, it can be applied in the form of software. For example, it can be designed as an independently running program on a computer-readable storage medium. The computer-readable storage medium can be a USB flash drive, designed as a USB key, and designed as a program that starts the whole method through external triggering through the USB flash drive.
[0133] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for recommending intercity bus routes based on mobile signaling data, characterized in that: It includes the following steps: S1: pre-process the mobile signaling data to obtain inter-city travel data, and execute step S2; S2: Aggregate the inter-city travel data to generate a user travel trajectory chain, and execute step S3; S3: Clustering the user travel trajectory chain to generate intercity bus recommended routes; The inter-city travel data aggregation processing method includes the following steps: S2.1: Grid processing of the base station H3 of the inter-city travel data, and execution of step S2.2; S2.2: Associating the target user signaling data with the H3 grid code table, and executing step S2.3; S2.3: Determine whether the user stays in the H3 grid. If yes, calculate the continuous stay time of the user in the H3 grid and execute step S2.4; S2.4: Determine whether the continuous stay time is less than a preset time, and if so, delete the corresponding user data.
2. The intercity bus route recommendation method based on mobile signaling data according to claim 1, characterized in that: The step S1 comprises: Filtering inter-city travel user data from the mobile signaling data; Ping-pong drift data, duplicate or invalid data are deleted from the user data, and the timestamp and coordinate type are converted to filter out the inter-city OD data to obtain the inter-city travel data.
3. The intercity bus route recommendation method based on mobile signaling data according to claim 1, characterized in that: In step S2.3, the user's residence time in a single grid is first sorted to obtain a sorting value, and the earliest time of entering the grid is counted. Then, based on the sorting value and the earliest time of entering the grid, it is determined whether each subsequent moment belongs to the first moment time. Finally, based on the first moment time, it is determined whether the user stays in the H3 grid.
4. The intercity bus route recommendation method based on mobile signaling data as claimed in claim 3, characterized in that: The time T start The calculation formula of t is: T start =T-RK*1 Among them, T start is the first moment time, T is the earliest time to enter the grid, and RK is the ranking value.
5. The intercity bus route recommendation method based on mobile signaling data according to claim 1, characterized in that: In step S3, clustering of the user travel trajectory chain is implemented by one or more of a DBSCAN clustering algorithm, a K-MEANS clustering algorithm, and a BERT+K-MEANS clustering algorithm.
6. The intercity bus route recommendation method based on mobile signaling data according to claim 5, characterized in that: The BERT+K-MEANS clustering algorithm includes the following steps: Training the user travel trajectory chain through the BERT model; Build vocabulary and code based on the grid code of H3 grid; Convert the user travel chain data into an Embedding vector; The Embedding vector is verified and evaluated, and the K-MEANS clustering algorithm is used for clustering based on the converted Embedding vector.
7. The intercity bus route recommendation method based on mobile signaling data according to claim 6, characterized in that: The method for training the user travel trajectory chain through the BERT model includes the following steps: Convert the H3 grid codes that the user travel trajectory chain passes through into text processing basic units, and encode the text processing basic units; Use the Bert model to train the encoded data and convert the data into Embedding vectors; User travel chain samples are randomly selected, and the Embedding vector is evaluated by comparing the similarity between the user travel trajectory chain before conversion and the user travel trajectory chain after conversion to the Embedding vector.
8. The intercity bus route recommendation method based on mobile signaling data according to claim 1, characterized in that: The calculation formula of the continuous residence time is: T stay =T max -T min Among them, T stay is the continuous residence time, T max is the last time the user stayed in the grid, T min The starting time when the user stays in the grid.
9. The intercity bus route recommendation method based on mobile signaling data according to claim 1, characterized in that: The step S2 also includes: first aggregating the user data into the H3 grid, then sorting the H3 grid according to the time series, and finally generating the user travel trajectory chain according to the sorted H3 grid.
10. An intercity bus route recommendation device based on mobile signaling data, characterized in that: It includes: A preprocessing module, which is used to preprocess the mobile signaling data to obtain inter-city travel data; An aggregation processing module, which is used to aggregate the inter-city travel data to generate a user travel trajectory chain; as well as A clustering module, which is used to cluster the user travel trajectory chains to generate intercity bus recommended routes; The method for aggregating and processing the inter-city travel data by the aggregation processing module comprises the following steps: S2.1: Grid processing of the base station H3 of the inter-city travel data, and execution of step S2.2; S2.2: Associating the target user signaling data with the H3 grid code table, and executing step S2.3; S2.3: Determine whether the user stays in the H3 grid. If yes, calculate the continuous stay time of the user in the H3 grid and execute step S2.4; S2.4: Determine whether the continuous stay time is less than a preset time, and if so, delete the corresponding user data.