Analysis Method of Residents' Travel Characteristics Based on Spatiotemporal Semantic Clustering of OD Flows
Through the method based on OD flow to spatiotemporal semantic clustering, combined with taxi OD data, POI data and Weibo check-in data, the problem of existing technology being difficult to deeply explore residents' travel characteristics is solved, and spatiotemporal semantic clustering analysis is realized, which improves the accuracy and depth of the analysis.
Patent Information
- Application Number
- CN202310429880.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-21
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-04-21
AI Technical Summary
The existing technology is difficult to deeply explore residents' travel characteristics and cannot effectively combine space-time and semantic information for cluster analysis.
Using the method based on OD flow to spatiotemporal semantic clustering, the POI access probability is calculated by obtaining taxi OD data, POI data and Weibo check-in data, extracting OD flow to semantics, constructing spatiotemporal semantic similarity measurement rules, and improving the DBSCAN algorithm for clustering analysis.
It has achieved a deep exploration of residents' travel characteristics, and can effectively combine space-time and semantic information and semantic information for cluster analysis, which has improved the accuracy and depth of travel characteristic analysis.
Smart Images

Figure CN116541737B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of spatial information technology, and particularly relates to a method for analyzing residents' travel characteristics based on spatio-temporal semantic clustering of OD flows. Background Art
[0002] Analyzing residents' travel characteristics provides important scientific basis for guiding urban spatial planning, formulating urban management policies, predicting traffic conditions, preventing and monitoring the spread of the epidemic, etc., and is also an important way to solve and alleviate these increasingly prominent urban problems.
[0003] Taxis are one of the main means of transportation for urban residents to travel. Taxi data effectively records the spatio-temporal information of residents' travel and can be widely used in mining residents' travel characteristics. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a method for analyzing residents' travel characteristics based on spatio-temporal semantic clustering of OD flows, which can deeply mine residents' travel characteristics.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions:
[0006] A method for analyzing residents' travel characteristics based on spatio-temporal semantic clustering of OD flows, comprising the following steps:
[0007] Step S1: Obtain taxi OD data, POI data, and Weibo check-in data, and define the destination areas of residents' travel;
[0008] Step S2: Calculate the POI access probability within each destination area of travel, and construct a corpus based on the POI access probability;
[0009] Step S3: Based on the corpus, use the GloVe model to extract the semantic of taxi OD flows;
[0010] Step S4: Based on the semantic of OD flows, construct a spatio-temporal semantic similarity measurement rule for OD flows;
[0011] Step S5: Improve the DBSCAN algorithm based on the spatio-temporal semantic similarity measurement rule, construct a density-based spatio-temporal semantic clustering algorithm for OD flows, and obtain residents' travel characteristics according to the density-based spatio-temporal semantic clustering algorithm for OD flows.
[0012] Further, the specific content of step S1 is as follows:
[0013] Obtain taxi OD data, POI data, and Weibo check-in data, and perform unified coordinate system, remove outliers, delete duplicate values on the above data, and extract time, space, and text information from the Weibo check-in data;
[0014] Use taxi OD data and POI data to count the relationship between the POI category ratio at the alighting location and the walking distance threshold, and define the destination area of residents' trips.
[0015] Furthermore, the calculation formula for the POI access probability is as follows:
[0016]
[0017] Where P r (P i ,(x,y),t) represents the probability that a taxi passenger alights at location (x,y) at time t and visits P i ; d((x,y),P i ) represents the distance between the passenger's alighting location and the candidate POI (P i ), β is the distance decay coefficient,; P r (P i ,(x,y),t) ranges from 0 to 1, and the sum of the access probabilities of all POIs is equal to 1; P i (T s-e ) represents the total number of times that a POI of type i is visited during the time period T s-e ; P j (T s-e ) represents the total number of times that a POI of type j is visited during the time period T s-e .
[0018] Furthermore, the corpus is constructed based on the POI access probability as follows:
[0019] (a) Construct a set A of POIs within the destination area, A = {P 1 ,P 2 …,P n};
[0020] (b) If the taxi passenger's alighting time T i is not within the business hours T i - T e of P s , then construct a set B = {P i ,…P o}, the access probability of the POIs in set B is 0, let set C = A - B, and remove the POIs in set B from set A;
[0021] (c) Calculate the access probability of the POIs in set C according to formula (2), and re - sort set C in descending order of the POI access probability. Set C is the corpus based on the dynamic POI access probability.
[0022] Further, the spatio-temporal semantic similarity measurement of the OD flow direction includes that the OD flow directions are close to each other in space, the lengths and directions of the OD flow directions are similar, the time of the OD flow directions is similar, and the OD flow directions have the same semantics, which are specifically as follows:
[0023] (a) Spatial similarity measurement of OD flow direction
[0024] Therefore, the adjacent flow directions are screened out by using the midpoint of the target flow direction and the k-nearest neighbor algorithm, and then the spatial similarity relationship between the two flow directions is quantified according to the following formula
[0025] Set dislimit as a parameter that changes with the length of the flow direction, and the definition is as follows:
[0026]
[0027] Among them, min(len i ,len j ) represents the smaller flow direction length value between the flow direction f i and the flow direction f j , k is a parameter greater than or equal to 2.83; the similarity between the OD flow directions is measured by calculating the ratio R of the distance between the OD points of the two flow directions to the parameter dislimit
[0028]
[0029]
[0030]
[0031] Among them, R O and R D respectively represent the spatial dissimilarity between the O point and the D point of the flow direction f i and f j . The value range of R is 0-1, and the smaller its value, the higher the spatial similarity between f i and f j ; dis() represents the Euclidean distance between two points, and sim O , sim D represent the similarity between the O point and the D point;
[0032] (b) Temporal similarity measurement of OD flow direction
[0033] By calculating the time interval of the alighting points of two OD flow directions, it is judged whether the two OD flow directions are similar in time, and the temporal similarity between the OD flow directions is measured according to the following formula:
[0034]
[0035]
[0036] Among them, dis(Dt i , Dt j ) represents the alighting time interval between f i and f j . Timelimit is a time threshold parameter set according to actual conditions. If dis(Dt i , Dt j ) ≤ timelimit, it means that the flow directions f i and the flow direction f j are similar in time. Sim t represents the temporal similarity between the flow directions f i and f j ;
[0037] (c) Semantic similarity measurement of OD flow directions
[0038] Judge the similarity of OD flow direction semantics based on the obtained OD flow direction semantics. The formula is as follows:
[0039]
[0040] Among them, S i , S j are the semantics of the flow directions f i and f j respectively. Sim s represents the similarity between two OD flow directions. If the semantics of the two OD flow directions are the same;
[0041] Combine the above-obtained OD flow direction spatio-temporal semantic similarity measurement parameters sim O , sim D , sim t , sim s to calculate the OD flow direction spatio-temporal semantic similarity:
[0042]
[0043] Among them, sim(f i , f j ) represents the spatio-temporal semantic similarity between the flow directions f i and the flow direction f j ;
[0044] Furthermore, the density-based spatio-temporal semantic clustering algorithm is as follows:
[0045] (a) Input OD flow direction data F = {f 1 , f 2 , …, f n}, time threshold timelimit, distance parameter k, density threshold minpts, radius parameter Eps;
[0046] (b) Calculate the length of the OD flow direction, mark all OD flow directions as unvisited, calculate the distance threshold dislimit based on the length of the OD flow direction and the parameter k, and determine its Eps-neighborhood according to the distance threshold;
[0047] (c) Traverse each OD flow direction and its Eps-neighborhood to find the flow direction f whose similarity parameter meets the conditions i , and the number of f i in the Eps-neighborhood is greater than minpts, mark f i as the core flow and add it to the similarity flow direction cluster c i ; add the non-core flows in the Eps-neighborhood to the noise set n i ; after traversing all the OD flow direction data, output the OD flow direction clusters C with spatio-temporal semantic similarity and the noise set N.
[0048] The present invention has the following beneficial effects compared with the prior art:
[0049] 1. The present invention trains the GloVe model based on the dynamic POI access probability ranking corpus, improves the accuracy of flow direction semantic extraction, and effectively extracts the semantics of residents' travel
[0050] 2. The present invention improves the spatio-temporal semantic clustering method of OD flow directions, combines the time, space and semantic information of OD flow directions, calculates the spatio-temporal semantic similarity of OD flow directions, realizes the spatio-temporal semantic clustering of OD flow directions, and applies this algorithm to mining the travel characteristics of residents, and can deeply mine the travel characteristics of residents. Brief Description of the Drawings
[0051] Figure 1 is the flowchart of the method of the present invention;
[0052] Figure 2 is the relationship between the proportion of POI categories and the walking distance threshold in an embodiment of the present invention;
[0053] Figure 3 is an example diagram of similar flow directions in an embodiment of the present invention;
[0054] Figure 4 is the similarity between OD flow directions with the same value and different OD flow direction lengths in an embodiment of the present invention;
[0055] Figure 5 is the silhouette coefficient value in an embodiment of the present invention;
[0056] Figure 6It is the spatio-temporal clustering result of four types of residents' travel semantics in an embodiment of the present invention. Detailed implementation manners
[0057] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0058] Please refer to Figure 1 , the present invention provides a method for analyzing residents' travel characteristics based on spatio-temporal semantic clustering of OD flows, including the following steps:
[0059] Step S1: Obtain taxi OD data, POI data, and Weibo check-in data, and define the destination area of residents' travel;
[0060] Step S2: Calculate the POI access probability within each destination area of travel, and construct a corpus based on the POI access probability;
[0061] Step S3: Based on the corpus, use the GloVe model to extract the semantics of taxi OD flows;
[0062] Step S4: Based on the semantics of OD flows, construct a spatio-temporal semantic similarity measurement rule for OD flows;
[0063] Step S5: Improve the DBSCAN algorithm based on the spatio-temporal semantic similarity measurement rule, construct a density-based spatio-temporal semantic clustering algorithm for OD flows, and obtain residents' travel characteristics according to the density-based spatio-temporal semantic clustering algorithm for OD flows.
[0064] In this embodiment, step S1 is specifically as follows:
[0065] Obtain taxi OD data, POI data, and Weibo check-in data, and perform unified coordinate system, outlier removal, and duplicate value deletion processing on the above data, and extract time, space, and text information from the Weibo check-in data;
[0066] Use taxi OD data and POI data to statistically analyze the relationship between the POI category ratio at the alighting location and the walking distance threshold, and define the destination area of residents' travel.
[0067] Preferably, when the walking distance threshold reaches a certain value, the POI category approaches 100% and no longer changes (as shown in the appendix Figure 2 ), then define the area within 250 meters of the residents' alighting location as the destination area of residents' travel.
[0068] In this embodiment, use Weibo check-in data to calculate the probability that a POI is visited in different time periods. The formula for calculating the POI access probability is as follows:
[0069]
[0070] Among them, P r (P i , t) represents the probability that a taxi passenger visits P at time t i . P i (T s-e ) represents the total number of times that a POI of type i is visited during the time period T s-e . represents the total number of times that all types of POIs are visited during the time period T s-e . There is a distance decay effect between the passenger's drop-off location and the access probability of the POI. The closer the POI is to the passenger's drop-off location, the higher the access probability, and the drop-off time and the location of the POI are independent of each other. Therefore, the calculation formula for the POI access probability is improved as follows:
[0071]
[0072] Among them, P r (P i , (x, y), t) represents the probability that a taxi passenger gets off at the location (x, y) at time t and visits P i ; d((x, y), P i ) represents the distance between the passenger's drop-off location and the candidate POI (P i ), and β is the distance decay coefficient; the value range of P r (P i , (x, y), t) is from 0 to 1, and the sum of the access probabilities of all POIs is equal to 1; P i (T s-e ) represents the total number of times that a POI of type i is visited during the time period T s-e ; P j (T s-e ) represents the total number of times that a POI of type j is visited during the time period T s-e .
[0073] In this embodiment, a corpus is constructed based on the POI access probability, specifically as follows:
[0074] (a) Construct a set A of POIs within the destination area, A = {P 1 , P 2 …, P n};
[0075] (b) If the taxi passenger's drop-off time T i is not within the business hours T i - T s of P e , then construct a set B = {P i , … P o}, the POI access probability in set B is 0. Let set C = A - B, and remove the POIs in set B from set A;
[0076] (c) Calculate the access probability of the POIs in set C according to formula (2), and re - sort set C in descending order of the POI access probability. Set C is the corpus based on the dynamic POI access probability.
[0077] In this embodiment, each taxi OD flow can be expressed as f i = O i , Ot i , D i , Dt i , S i , where O i =(ox i , oy i )、Ot i are the longitude and latitude coordinates and time of the taxi passenger's boarding point, D i =(dx i , dy i )、Dt i are the longitude and latitude coordinates and time of the taxi passenger's alighting point, and S i is the OD flow semantics of f i .
[0078] The OD flow spatio - temporal semantic similarity measurement rules mainly measure the similarity between OD flows from the following aspects: (a) The OD flows are close to each other in space; (b) The lengths and directions of the OD flows are similar; (c) The times of the OD flows are similar; (d) The OD flows have the same semantics.
[0079] As shown in the appendix Figure 3 , different colors represent different flow semantics. f 1 and f 3 are similar in the semantics and length of the flow, but not similar in the direction of the flow. f 1 and f 4 , f 5 are similar in the length and direction of the flow, but not similar in the semantics of the flow. f 1 and f 6 are not similar in the length, direction, and semantics of the flow. Only f 2 and f 1 are similar. Specifically as follows:
[0080] (a) Measurement of the spatial similarity of OD flows
[0081] In the above rules, similar flow directions are close in space. Therefore, the midpoint of the target flow direction and the k-nearest neighbor algorithm are used to screen out adjacent flow directions, and then the spatial similarity relationship between two flow directions is quantified according to the following formula.
[0082] As attached Figure 4 As shown in the figure, even if the flow direction is obviously different, additional Figure 4 (b) and Figure 4 (c) Determine f i and f j For the case of similar flow directions, the length of the flow direction must be greater than 2dislimit / sin45° (≈2.83dislimit) to ensure that the angle between the two flow directions is less than 45 degrees. Therefore, dislimit is set as a parameter that changes with the length of the flow direction, defined as follows:
[0083]
[0084] Among them, min(len i ,len j ) indicates the flow direction f i and flow direction f j The smaller flow length value in the equation is k, which is a parameter greater than or equal to 2.83. The similarity between OD flow directions is measured by calculating the ratio R between the distance between two flow OD points and the parameter dislimit.
[0085]
[0086]
[0087]
[0088] Among them, R O and R D Respectively represent the flow direction f i and f j The spatial dissimilarity between point O and point D, R ranges from 0 to 1. The smaller the value, the greater the difference between points O and D. i and f j The higher the spatial similarity between them; dis() represents the Euclidean distance between two points, sim O ,sim D Indicates the similarity between point O and point D;
[0089] (b) Temporal similarity measure of OD flow
[0090] Determine whether two OD flows are similar in time by calculating the time interval between the alighting points of the two OD flows, and measure the temporal similarity between OD flows according to the following formula:
[0091]
[0092]
[0093] Where dis(Dt i , Dt j ) represents the alighting time interval of flows f i and f j , timelimit is a time threshold parameter set according to actual conditions. If dis(Dt i , Dt j ) ≤ timelimit, it means that flows f i and flow f j are similar in time. sim t represents the temporal similarity of flows f i and f j ;
[0094] (c) Semantic similarity measurement of OD flows
[0095] Judge the similarity of OD flow semantics based on the obtained OD flow semantics. The formula is as follows:
[0096]
[0097] Where S i , S j are the semantics of flows f i and f j respectively. sim s represents the similarity between two OD flows. If the semantics of the two OD flows are the same;
[0098] Combine the above-obtained OD flow spatio-temporal semantic similarity measurement parameters sim O , sim D , sim t , sim s to calculate the spatio-temporal semantic similarity of OD flows:
[0099]
[0100] Where sim(f i , f j ) represents the spatio-temporal semantic similarity between flow f i and flow f j .
[0101] In this embodiment, as shown in Table 1, the density-based spatio-temporal semantic clustering algorithm is specifically shown in Appendix 1
[0102]
[0103]
[0104] Example 1:
[0105] In this embodiment, data for a working day in a certain day of 2020 is selected from taxi order data in a large domestic city. There are 17,585 records for 5,060 taxis in total. The POI data includes 13 first-level categories such as dining and office, 101 second-level categories, and 393 third-level categories, with a total of 62,997 POI data records for mining and analyzing the effectiveness of residents' travel characteristics.
[0106] In this experiment, the destination area of the present invention is centered on the passenger's drop-off location and formed with a buffer threshold of 250 meters. When training the GloVe model, the dimension of the word vector is set to 128, the co-occurrence window is set to 5, and the number of iterations is set to 10. When using the K-means clustering algorithm to cluster the feature vectors of the destination area, the silhouette coefficients for K values from 2 to 13 are calculated. As shown in the appendix Figure 5 When K = 3 and K = 7, the silhouette coefficient values are the highest and the clustering effect is the best. Since 3 clusters are not sufficient to reveal the diversity of residents' travel semantics, in this experiment, K = 7 is selected as the ideal K value for further analysis and verification, and 7 types of residents' travel semantics are extracted in total. To further verify the effectiveness of the proposed OD flow spatio-temporal clustering method, in this experiment, 4 types of residents' travel semantics, namely "return home for commuting", "travel for transportation", "work commuting", and "travel for medical treatment", are selected for OD flow spatio-temporal semantic clustering analysis to further discover typical residents' travel patterns (as shown in the appendix Figure 6 To deeply understand the travel characteristics and mobility of residents.
[0107] The above are only the preferred embodiments of the present invention. All equivalent changes and modifications made according to the scope of the patent application of the present invention shall fall within the scope covered by the present invention.
Claims
1. A method for analyzing residents' travel characteristics based on spatio-temporal semantic clustering of OD flows, characterized in that, it includes the following steps: Step S1: Obtain taxi OD data, POI data, and microblog check-in data, and define the regional areas of residents' travel destinations; Step S2: Calculate the POI access probabilities within each regional area of travel destinations, and construct a corpus based on the POI access probabilities; Step S3: Based on the corpus, use the GloVe model to extract the semantic of taxi OD flows; Step S4: Based on the semantic of OD flows, design a spatio-temporal semantic similarity measurement rule for OD flows; Step S5: Construct a density-based spatio-temporal semantic clustering algorithm for OD flows, and obtain residents' travel characteristics according to the density-based spatio-temporal semantic clustering algorithm for OD flows; The spatio-temporal semantic similarity measurement of the OD flows includes that the OD flows are close to each other in space, the lengths and directions of the OD flows are similar, the time of the OD flows is similar, and the OD flows have the same semantic, specifically as follows: (a) Spatial similarity measurement of OD flows Use the midpoint of the target flow and the k-nearest neighbor algorithm to screen out adjacent flows, and then quantify the spatial similarity relationship between two flows according to the following formula Set doslimit as a parameter that changes with the length of the flow, and the definition is as follows: Among them, min(len i , len j ) represents the smaller flow length value between flow direction f i and flow direction f j . k is a parameter greater than or equal to 2.
83. The similarity between OD flow directions is measured by calculating the ratio R of the distance between two OD points of the flow direction to the parameter dislimit Among them, R O and R D respectively represent the spatial dissimilarity between the O point and the D point of the flow directions f i and f j . The value range of R is 0 - 1. The smaller the value, the higher the spatial similarity between f i and f j . dis() represents the Euclidean distance between two points, and sim O , sim D represent the similarity between the O point and the D point; (b) Time similarity measurement of OD flows Judge whether two OD flows are similar in time by calculating the time interval of the alighting points of the two OD flows, and measure the time similarity between the OD flows according to the following formula: Among them, dis(Dt i , Dt j ) represents the alighting time interval between f i and f j . timelimit is a time threshold parameter set according to actual conditions. If dis(Dt i , Dt j ) ≤ timelimit, it means that the flow direction f i and the flow direction f j are similar in time; sim t represents the time similarity between the flow direction f i and f j ; (c) Semantic similarity measurement of OD flows Judge the similarity of the semantic of OD flows according to the obtained semantic of OD flows, and the formula is as follows: Among them, S i , S j are the semantics of the flows f i and f j respectively. sim s represents the similarity between two OD flows. If the semantics of two OD flows are the same, then assign 1 to sim s ; Combine the OD flow spatio-temporal semantic similarity measurement parameters sim O 、sim D 、sim t 、sim s to calculate the OD flow spatio-temporal semantic similarity: Among them, sim(f i , f j ) represents the spatio-temporal semantic similarity between flow f i and flow f j ; The density-based spatio-temporal semantic clustering algorithm for OD flows is specifically as follows: (a) Input OD flow data F = {f 1 , f 2 , …, f n}, time threshold timelimit, distance parameter k, density threshold minpts, radius parameter Eps; (b) Calculate the length of the OD flow, mark all OD flows as unvisited, calculate the distance threshold dislimit through the length of the OD flow and the parameter k, and determine the Eps-neighborhood of the OD flow according to the distance threshold; (c) Traverse each OD flow direction and its Eps-neighborhood to find the flow direction f whose similarity parameter meets the conditions i , and the number of f i in the Eps-neighborhood should be greater than minpts, mark f i as the core flow and add it to the similar flow direction cluster c i . Add the non-core flows within the Eps-neighborhood to the noise set n i . After traversing all OD flow direction data, output the OD flow direction clusters C with spatiotemporal semantic similarity and the noise set N.
2. The method for analyzing residents' travel characteristics based on spatio-temporal semantic clustering of OD flows according to claim 1, characterized in that, the specific content of Step S1 is: Obtain taxi OD data, POI data, and microblog check-in data, and perform unified coordinate system, outlier removal, and duplicate value deletion processing on the above data, and extract time, space, and text information from the microblog check-in data; Use the taxi OD data and POI data to count the relationship between the POI category ratio of the alighting location and the walking distance threshold, and define the regional areas of residents' travel destinations.
3. The method for analyzing residents' travel characteristics based on spatio-temporal semantic clustering of OD flows according to claim 1, characterized in that, the calculation formula of the POI access probability is as follows: Among them, P r (P i ,(x,y),t) represents the probability that a taxi passenger gets off at the location (x,y) at time t and visits P i ; d((x,y),P i ) represents the distance between the passenger's drop-off location and the candidate point of interest P i ; β is the distance decay coefficient; P r (P i ,(x,y),t) ranges from 0 to 1, and the sum of the access probabilities of all POIs is equal to 1; P i (T s-e ) represents the total number of times that a POI of type i is visited during the time period T s-e ; P j (T s-e ) represents the total number of times that a POI of type j is visited during the time period T s-e .
4. The method for analyzing residents' travel characteristics based on spatio-temporal semantic clustering of OD flows according to claim 3, characterized in that, the construction of the corpus based on the POI access probability is specifically as follows: (a) Construct the set A of POIs within the target area, A = {P 1 , P 2 …, P n}; (b) If the getting-off time T of the taxi passenger i is not within i the business hours T s -T e of P, then construct the set B = {P i ,…P o}, the POI access probability in the set B is 0, let the set C = A - B, and eliminate the POIs in the set B from the set A; (c) Calculate the access probability of POIs in set C according to formula (2), and re - sort set C in descending order of the POI access probability. Set C is the corpus based on the dynamic POI access probability.
Citation Information
Patent Citations
Abnormal resident travel mode mining method based on taxi OD data
CN112836000A
Urban area function identification model and identification method based on space-time big data
CN113806419A