Space-time double-layer clustering stay point identification method based on Thiessen polygon network

By constructing a Tyson polygonal network and a space-time bilayer clustering algorithm, the accuracy and stability of mobile phone signaling data recognition in urban areas are solved, and a more efficient residence point recognition effect is achieved.

CN120547504APending Publication Date: 2025-08-26HUBEI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510190882.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

When using mobile phone signaling data to identify traffic stop points, there are inaccurate assumptions for base station coverage division, sensitive spatial threshold parameters, unstable identification accuracy, and "overfitting" and "underfitting", which leads to high probability of drifting data and ping-pong data, making it difficult to accurately identify stop points in urban areas.

Method used

The space-time bilayer clustering method based on Tyson polygon network is adopted. By constructing the Tyson polygon network, the spatial dynamic threshold of the base station is calculated, and the K nearest neighbor algorithm and DBSCAN clustering algorithm are combined to identify the stay points of mobile phone signaling data to improve the accuracy of traditional methods.

Benefits of technology

It improves the accuracy and stability of stop point recognition, reduces the "overfit" and "underfit" phenomena, improves the recognition accuracy, recall and F1 score, and performs better than traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120547504A_ABST
    Figure CN120547504A_ABST
Patent Text Reader

Abstract

The invention provides a space-time double-layer clustering stay point identification method based on a Thiessen polygon network, and the method comprises the following steps: S1, reading mobile phone signaling and base station positioning data, and cleaning the mobile phone signaling data; s2, constructing a Thiessen polygon network according to the base station positioning data; s3, according to the Thiessen polygon network, calculating a spatial dynamic threshold value of each base station by adopting a K nearest neighbor algorithm; and S4, according to the spatial dynamic threshold, identifying the stay point of the mobile phone signaling data through a space-time double-layer clustering algorithm. The method is based on the DBSCAN algorithm, combines the spatial dynamic threshold value calculated under the Thiessen polygon network hypothesis, fully considers the spatial and time characteristics in the travel track staying process, and effectively solves the problem that the clustering algorithm is sensitive to spatial parameters in combination with the spatial dynamic threshold value. According to the method, the recognition performance is better, and the recognition accuracy, the recall rate and the F1 score are respectively improved by 9.1%, 7.7% and 8.4% at most.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communications, and in particular to a spatiotemporal double-layer clustering stay point identification method based on a Thiessen polygon network. Background Art

[0002] With the advancement of mobile communication technology and the rapid adoption of smartphones, research on mining urban travel characteristics based on mobile phone signaling data has been increasing. Over the past two decades, many institutions and researchers have conducted resident travel surveys to understand travel patterns and patterns in urban planning and transportation development. Traditional resident travel surveys rely on questionnaires and electronic emails. These surveys are characterized by high costs and low sampling rates, making it difficult to obtain high-quality data to accurately extract travel patterns. With the continuous advancement of information technology, a growing number of methods are being used to capture resident travel patterns. Common data sources include floating vehicles, loops, radio frequency identification (RFID), and video cameras. These technologies suffer from limitations such as limited data coverage, sparse sampling points, and a single monitoring target, significantly limiting their application in mining urban travel characteristics. Later, GPS-assisted surveys, which record the continuous travel trajectories of respondents, significantly improved the challenges of missing travel information and low sampling rates. However, the high cost of the GPS equipment used in these survey methods makes them inadequate for large-scale urban travel surveys. Mobile phone signaling data can capture the real-time location of mobile phone users, dynamically assessing their travel patterns, and covering large populations and regions. Using mobile phone signaling data to analyze residents' travel characteristics can provide more valuable travel information to support traffic planning and management. It has many advantages, such as strong timeliness, large sample size, low collection cost, wide coverage, fast data update frequency, and strong objectivity and authenticity of data. It has quickly become a research hotspot in the field of transportation research.

[0003] Travel demand, or the OD chain, is a crucial component of travel characteristics. Identifying resident travel stops accurately characterizes the travel chains of individual mobile phone users and facilitates analysis of travel demand. Currently, most methods for identifying travel stops using mobile phone signaling data employ probability-based and clustering-based approaches. Probabilistic approaches primarily consider the frequency of trajectory points, ignoring their temporal continuity and directionality, potentially generating some meaningless stops. Furthermore, these approaches place high demands on data quality and are not robust to noise. With poor-quality data, data sparsity issues can easily arise. Clustering-based approaches utilize a designed distance measurement method integrated with a clustering algorithm to automatically discover clusters of spatiotemporally proximal continuous trajectory points, characterized by low spatial displacement and temporal order. These clusters are then labeled as stop points, while all other points are labeled as moving points. This approach automatically identifies spatiotemporally proximal continuous stop points without requiring manual rule-setting. However, these methods are highly sensitive to user-defined parameters, often leading to over- and under-identification of stop points, requiring the user to continuously adjust these parameters to find the appropriate ones. At the same time, different sources of trajectory data and different positioning accuracy also lead to certain difficulties and challenges in the selection of threshold parameters. For example, Chinese patent document CN110113718A records a method for identifying the population type of railway transportation hubs based on mobile phone signaling data, and CN111681421A records a method for analyzing the spatial distribution of external passenger transportation hubs based on mobile phone signaling data. That is, there is a large amount of drift data and ping-pong data, and these data are identified and merged. The problem is that since the location data of each base station have overlapping areas, the definition of the base station location is unclear, which aggravates the existence of drift data and ping-pong data, and after the merger, there is a disagreement on which specific base station location to set the stop point.

[0004] At the urban scale, due to the large differences in spatial resources and population distribution, mobile base stations are densely distributed in the city center and sparsely distributed in the suburbs, resulting in uneven base station distribution density in urban areas. Mobile phone signaling data is based on base station positioning, so there are shortcomings such as irregular sampling, inconsistent positioning accuracy, and uneven data quality. Therefore, using mobile phone signaling data to identify traffic stop points has the following problems: (1) The method of calculating spatial thresholds based on the assumption of traditional base station coverage range division (regular hexagon, fan, circle, etc.) at the urban scale is not accurate, the positioning is not clear, and the probability of drift data and ping-pong data is greatly increased; (2) The mobile phone signaling stop point identification algorithm based on clustering algorithm is sensitive to spatial threshold parameters at the urban scale, and the identification accuracy is unstable; (3) There are "over-identification" and "under-identification" phenomena of stop points in base station coverage areas with different degrees of sparsity. Summary of the Invention

[0005] The technical problem to be solved by this invention is to provide a spatiotemporal dual-layer clustering stay point identification method based on a Thiessen polygon network. This method can automatically partition the coverage areas of mobile phone base station towers within an urban area and calculate spatial dynamic thresholds through nearest neighbor analysis. This method improves the accuracy of spatial threshold calculations based on traditional base station coverage assumptions (e.g., regular hexagons, sectors, and circles). This method effectively avoids the overfitting and underfitting problems of traditional stay point identification methods using fixed thresholds, while also improving recognition accuracy.

[0006] To solve the above technical problems, the technical solution adopted by the present invention is: a spatiotemporal double-layer clustering stay point identification method based on Thiessen polygon network, comprising the following steps: S1, read mobile phone signaling and base station positioning data; S2. Construct a Thiessen polygon network based on base station positioning data; S3. Calculate the spatial dynamic threshold of each base station using a K-nearest neighbor algorithm based on a Thiessen polygon network; S4. Based on the spatial dynamic threshold, the stay points of the mobile phone signaling data are identified through the spatiotemporal double-layer clustering algorithm.

[0007] Preferably, step S1 includes the step of cleaning the mobile phone signaling data: S11, determining whether two adjacent records of the current data are from the same base station and the time difference is less than a preset threshold; if so, marking the current data record as ping-pong effect data and removing it; S12. Calculate the speed of two adjacent records. If the speed is greater than a preset speed threshold, mark the latter record as drift data and remove it.

[0008] Preferably, step S1 includes the step of cleaning the mobile phone signaling data: S101. Aggregate all mobile phone signaling data by user ID and sort by time to obtain a mobile phone signaling data set sorted by user ID and time, so as to facilitate subsequent processing of individual user data; S102, reading the mobile phone signaling data of a single user at one time; S103, read the i-th record (i starts at 1), i+1 record, and i+2 record one by one; if the data read reaches the last row or there are no last two rows of data, determine whether it is the last user, if not, return to step S102, otherwise end; S104: Determine whether the base stations recorded in the i-th and i+2-th rows are the same base station, and calculate the time difference ∆t between the i+1-th and i+2-th rows. i +1, judge Δt i +1 whether it is within the preset time threshold T; If both conditions are met, jump to step S105; otherwise, jump to step S106; S105, the i+1th record is ping-pong effect data, which is directly discarded and the process goes to S103; S106. Set i=i+1 and jump to S103.

[0009] Preferably, step S1 includes the step of cleaning the mobile phone signaling data: S111. Aggregate all mobile phone signaling data by user ID and sort by time to obtain a mobile phone signaling data set sorted by user ID and time, so as to facilitate subsequent processing of individual user data; S112, reading the mobile phone signaling data of a single user at one time; S113. Read the i-th record (i starts at 1) and i+1-th record one by one. If the data read reaches the last row or there is no next row of data, determine whether it is the last user. If not, return to step 2. Otherwise, end. S114, calculate the distance di and the interval time Δti between the i-th and i+1-th data, and then calculate the speed from point i to point i+1 vi, the distance is calculated as shown in formula (1) and (2), and the speed is calculated as shown in formula (3); (1); (2); (3); In the formula, R is the radius of the earth, which is approximately 6371 km; i ,lon i )、(lat i +1,lon i +1) are the spatial coordinates of the two records i and i+1, d i is the distance between two points, v i is the velocity of point i; S115, Comparison v i and the preset speed threshold v t The size of v i >v i , move to step S116; otherwise, move to step S117; S116, directly delete the (i+1)th record and jump to step S113; S117. Set i=i+1 and jump to step S113.

[0010] Preferably, in step S2, constructing the Thiessen polygon network includes the following steps: S21. Import base station positioning data and number the base station positioning data from left to right and from top to bottom according to spatial distribution; S22. Construct a triangulated network for the discrete base station positioning data, that is, connect each base station with its adjacent base stations to form triangles, number the formed triangles, and establish a mapping relationship between the triangles and the discrete base station numbers; S23, calculating the circumcenter of each triangle, and then connecting the circumcenters of the triangles adjacent to each discrete base station in S22, to obtain the Thiessen polygon corresponding to each discrete point; Preferably, for the Thiessen polygons at the boundary of the triangulated network, a closed polygon is formed by intersecting the perpendicular bisector with the boundary.

[0011] Preferably, in step S3, the current base station P1 is taken as the center, k nearest base stations of its adjacent M Thiessen polygons are selected, the distances between the current base station P1 and the k base stations are calculated, and the maximum value of the k distances is selected as the spatial threshold of the base station.

[0012] Preferably, calculating the spatial dynamic threshold of each base station includes the following steps: S31, set the set of all base stations P= { P 1 ,P 2 ,……,P n}, the Thiessen polygon set corresponding to each base station T= { T 1 ,T 2 ,…,T n}; S32, base station P 1 as an example, the Thiessen polygon where it is located T 1 has M adjacent Thiessen polygons, calculate the base station P 1 The Euclidean distance set to the adjacent Thiessen polygon control point base station D ={ D 1 ,D 2 ,…,D m}, and sort them from small to large; S33, judge according to the set k value, if K>M , then let K=M ; S34, then take the first k distance values ​​in the distance set D D 1 ,D 2 ,…,D k , then the base station P A spatial threshold of 1 D p1is the maximum value among these k distance values, and the calculation is shown in formula (4).

[0013] (4).

[0014] Preferably, in step S4, identifying the stop points of the mobile phone signaling data using the spatiotemporal double-layer clustering algorithm includes the following steps: Based on the DBSCAN clustering algorithm, stay point identification is performed according to spatial dynamic thresholds.

[0015] Preferably, the DBSCAN clustering algorithm comprises the following steps: S41. Import a user's one-day mobile phone signaling travel trajectory data and a base station spatial threshold dataset based on Thiessen polygons; S42, read the unprocessed points in the trajectory data in chronological order p , take out the point p Distance threshold in a spatial threshold dataset D p , calculated in order of time increments p Point and p i The distance between d ,until d > D p Stop and get the track point p Spatial Threshold D p A collection of points within the range P= { p 1 , p 2 ,…,p n}; S43, Judgment Point p arrive P The last point in the set p n The time difference ∆ t and time thresholds T Whether ∆ t > T If satisfied, execute step S44; otherwise, jump to step S42; S44. Create a new cluster C ,Will p Points and Sets P Add the points in C middle; S45, read clusters in sequence C Medium p Other trajectory points p i , calculate the trajectory point according to step S42 pi Spatial Threshold D pi A collection of points within the range P pi ={ p i1 ,p i2 ,…,p in}; S46, point p i arrive P pi Each point in the set p ij The time difference ∆ t and time thresholds T Whether ∆ t > T If satisfied, then point p ij Join cluster C; otherwise, jump to step S45; S47, determine whether all trajectory points have been processed, if so, end; Otherwise jump to step S42.

[0016] The present invention provides a method and construction method for identifying stop points using a spatiotemporal double-layer clustering method based on a Thiessen polygon network. Based on the DBSCAN algorithm and integrating the spatial dynamic threshold calculated under the Thiessen polygon network assumption, a method for identifying stop points using a spatiotemporal double-layer clustering method is designed. This method fully considers the spatial and temporal characteristics of the travel trajectory during the stop process, and effectively solves the problem of clustering algorithms being sensitive to spatial parameters in combination with the spatial dynamic threshold. In the test, based on the Ubicomp2021-SHL project dataset with stop point labels, the accuracy of stop point identification was compared between the method proposed in the present invention and other classic algorithms (TDBC, DJ-Cluster, CB-SMOT, etc.). The method proposed in the present invention performed better in recognition performance, with the highest recognition accuracy, recall rate, and F1 score increased by 9.1%, 7.7%, and 8.4%, respectively.

[0017] (1) Based on the Thiessen polygon theory, the present invention constructs a Thiessen polygon network for all base stations in the city, and integrates the KNN idea to construct a calculation method for spatial dynamic thresholds, which effectively eliminates the problem of spatial threshold sensitivity in the travel stop point identification method caused by the uneven positioning accuracy of mobile phone signaling data and irregular sampling points.

[0018] (2) This paper proposes a spatiotemporal double-layer clustering stop point identification method based on spatial dynamic thresholds. This method fully considers the spatial and temporal characteristics of the travel trajectory during the stop process, and combines the spatial dynamic threshold to effectively improve the accuracy of the traffic stop point identification method based on mobile phone signaling data. The method proposed in this paper is experimentally compared with other commonly used methods using the publicly available Ubicomp 2021-SHL project dataset, and it is verified that the accuracy of this method is higher than that of other methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The present invention will be further described below with reference to the accompanying drawings and examples: Figure 1 It is a schematic flow diagram of the present invention.

[0020] Figure 2 It is a flow chart of drift data processing of the present invention.

[0021] Figure 3 It is a flow chart of ping-pong data processing of the present invention.

[0022] Figure 4 It is a flow chart of the present invention for constructing a Thiessen polygon network.

[0023] Figure 5 This is a schematic diagram of the Thiessen polygon network constructed by the present invention.

[0024] Figure 6 It is a schematic diagram of the present invention for identifying stay points based on temporal and spatial dynamic thresholds. DETAILED DESCRIPTION

[0025] A method for identifying stay points by spatiotemporal double-layer clustering based on Thiessen polygon network, comprising the following steps: S1, read mobile phone signaling and base station positioning data; The mobile phone signaling data includes user ID, time, and location information; the base station location data includes base station ID and base station latitude and longitude coordinates; Based on the user ID and time fields in the mobile phone signaling data, all mobile phone signaling data is aggregated by user ID and sorted by time to obtain a mobile phone signaling dataset sorted by user ID and time. From this mobile phone signaling dataset sorted by user ID and time, all mobile phone signaling data for a single user is read. The time and location information fields in the mobile phone signaling data of a single user are obtained, and the location information is used to find the corresponding base station ID and base station latitude and longitude coordinates in the base station location data.

[0026] S2. Construct a Thiessen polygon network based on base station positioning data; S3. Calculate the spatial dynamic threshold of each base station using a K-nearest neighbor algorithm based on a Thiessen polygon network; S4. Based on the spatial dynamic threshold, the stay points of the mobile phone signaling data are identified through the spatiotemporal double-layer clustering algorithm.

[0027] Preferably, step S1 includes the step of cleaning the mobile phone signaling data: S11, determining whether two adjacent records of the current data are from the same base station and the time difference is less than a preset threshold; if so, marking the current data record as ping-pong effect data and removing it; S12. Calculate the speed of two adjacent records. If the speed is greater than a preset speed threshold, mark the latter record as drift data and remove it.

[0028] Determine whether the time difference between two adjacent mobile phone signaling data is less than a given time threshold. If so, calculate the distance between the two positioning points. If the calculated distance is greater than the given distance threshold, determine whether the speed of the second point relative to the first point exceeds the given speed threshold. If so, the second point data is marked as drift point data. If the base station ID of three adjacent mobile phone signaling data switches back and forth between two values, and the switching interval is less than the given time threshold, the middle data is marked as "ping-pong effect" data. Remove the mobile phone signaling data marked as drift point data and "ping-pong effect" data to obtain the denoised mobile phone signaling data for that user.

[0029] Determine whether the denoised mobile phone signaling data all falls within the polygonal area of ​​the corresponding base station. If not, mark the data as invalid. Remove the mobile phone signaling data marked as invalid, and finally obtain the user's one-day mobile phone signaling travel trajectory data after cleaning.

[0030] Preferably, step S1 includes the step of cleaning the mobile phone signaling data: S101. Aggregate all mobile phone signaling data by user ID and sort by time to obtain a mobile phone signaling data set sorted by user ID and time, so as to facilitate subsequent processing of individual user data; S102, reading the mobile phone signaling data of a single user at one time; S103, read the i-th record (i starts at 1), i+1 record, and i+2 record one by one; if the data read reaches the last row or there are no last two rows of data, determine whether it is the last user, if not, return to step S102, otherwise end; S104: Determine whether the base stations recorded in the i-th and i+2-th rows are the same base station, and calculate the time difference ∆t between the i+1-th and i+2-th rows. i +1, judge Δt i +1 whether it is within the preset time threshold T; If both conditions are met, jump to step S105; otherwise, jump to step S106; S105, the i+1th record is ping-pong effect data, which is directly discarded and the process goes to S103; S106: Set i=i+1 and jump to S103. Based on the above steps, for the drift data in the mobile phone signaling data, the distance, time and speed of two adjacent records are calculated, and the speed is compared with the preset speed threshold. If the speed is greater than the preset speed threshold, it is eliminated.

[0031] Preferably, step S1 includes the step of cleaning the mobile phone signaling data: S111. Aggregate all mobile phone signaling data by user ID and sort by time to obtain a mobile phone signaling data set sorted by user ID and time, so as to facilitate subsequent processing of individual user data; S112, reading the mobile phone signaling data of a single user at one time; S113. Read the i-th record (i starts at 1) and i+1-th record one by one. If the data read reaches the last row or there is no next row of data, determine whether it is the last user. If not, return to step 2. Otherwise, end. S114: Calculate the distance di and the interval time Δt between the i-th and i+1-th data. i , and then calculate the speed from point i to point i+1 vi, the distance is calculated as shown in formula (1) and (2), and the speed is calculated as shown in formula (3); (1); (2); (3); In the formula, R is the radius of the earth, which is approximately 6371 km; i ,lon i )、(lat i +1,lon i +1) are the spatial coordinates of the two records i and i+1, d i is the distance between two points, v i is the velocity of point i; S115, Comparison v i and the preset speed threshold v t The size of v i >v i , move to step S116; otherwise, move to step S117; S116, directly delete the (i+1)th record and jump to step S113; S117, set i=i+1, jump to step S113. According to the above steps, the ping-pong effect data in the mobile phone signaling data is identified and eliminated by determining whether the data between the two data are the same data and whether the time difference is less than a preset threshold.

[0032] Preferably, in step S2, constructing the Thiessen polygon network includes the following steps: S21. Import base station positioning data and number the base station positioning data from left to right and from top to bottom according to spatial distribution; S22. Construct a triangulated network for the discrete base station positioning data, that is, connect each base station with its adjacent base stations to form triangles, number the formed triangles, and establish a mapping relationship between the triangles and the discrete base station numbers; S23, calculating the circumcenter of each triangle, and then connecting the circumcenters of the triangles adjacent to each discrete base station in S22, to obtain the Thiessen polygon corresponding to each discrete point; Preferably, for the Thiessen polygons at the boundary of the triangulated network, a closed polygon is formed by intersecting the perpendicular bisector with the boundary.

[0033] Preferably, in step S3, the current base station P1 is taken as the center, k nearest base stations of its adjacent M Thiessen polygons are selected, the distances between the current base station P1 and the k base stations are calculated, and the maximum value of the k distances is selected as the spatial threshold of the base station.

[0034] Preferably, calculating the spatial dynamic threshold of each base station includes the following steps: S31, set the set of all base stations P= { P 1 ,P 2 ,……,P n}, the Thiessen polygon set corresponding to each base station T= { T 1 ,T 2 ,…,T n}; S32, base station P 1 as an example, the Thiessen polygon where it is located T 1 has M adjacent Thiessen polygons, calculate the base station P 1 The Euclidean distance set to the adjacent Thiessen polygon control point base station D ={ D 1 ,D 2 ,…,D m}, and sort them from small to large; S33, judge according to the set k value, if K>M, then let K=M ; S34, then take the first k distance values ​​in the distance set D D 1 ,D 2 ,…,D k , then the base station P A spatial threshold of 1 D p1 is the maximum value among these k distance values, and the calculation is shown in formula (4).

[0035] (4).

[0036] For each base station location, sort the distances from smallest to largest and select the top K distances. Based on these top K distances, calculate their mean or median to obtain the spatial dynamic threshold for that base station. Determine whether the spatial dynamic thresholds for all base stations have been calculated. If so, proceed to the next step; otherwise, return to step S22 to process the next base station. Summarize the spatial dynamic thresholds for all base stations to obtain the spatial dynamic threshold for the entire city area, which is then used by the subsequent stay point identification algorithm.

[0037] Preferably, in step S4, identifying the stop points of the mobile phone signaling data using the spatiotemporal double-layer clustering algorithm includes the following steps: Based on the DBSCAN clustering algorithm, stay point identification is performed according to spatial dynamic thresholds.

[0038] Preferably, the DBSCAN clustering algorithm comprises the following steps: S41. Import a user's one-day mobile phone signaling travel trajectory data and a base station spatial threshold dataset based on Thiessen polygons; S42, read the unprocessed points in the trajectory data in chronological order p , take out the point p Distance threshold in a spatial threshold dataset D p , calculated in order of time increments p Point and p i The distance between d ,until d > D p Stop and get the track point p Spatial Threshold D p A collection of points within the range P= { p 1 , p 2 ,…,p n}; S43, Judgment Point p arriveP The last point in the set p n The time difference ∆ t and time thresholds T Whether ∆ t > T If satisfied, execute step S44; otherwise, jump to step S42; S44. Create a new cluster C ,Will p Points and Sets P Add the points in C middle; S45, read clusters in sequence C Medium p Other trajectory points p i , calculate the trajectory point according to step S42 p i Spatial Threshold D pi A collection of points within the range P pi ={ p i1 ,p i2 ,…,p in}; S46, point p i arrive P pi Each point in the set p ij The time difference ∆ t and time thresholds T Whether ∆ t > T If satisfied, then point p ij Join cluster C; otherwise, jump to step S45; S47, determine whether all trajectory points have been processed, if so, end; Otherwise jump to step S42.

[0039] Example 2: The scheme of application embodiment 1 is used to sort the mobile phone signaling data in ascending order according to the user ID and timestamp, and then read the data records of each user in sequence. During the traversal process, if the current record is not the last one of the user, the next one is read. The Euclidean distance and time difference of the two records are calculated at the same time. Assuming that the distance is 500 meters and the time difference is 60 seconds, the speed is 500 / 60=8.33 meters per second. The calculated speed is compared with a set threshold value, such as 2 meters per second. If it is greater than the threshold, it is regarded as drift data and deleted. Then it is judged whether there is a situation where the same position exists in the three adjacent records. If so, if the base station IDs of the i-th and i+2-th items are consistent and the time difference is less than a given threshold value, such as 5 seconds, the middle i+1-th record is considered to be abnormal data generated by the ping-pong effect and is removed. After completing data cleaning, the base station network topology is established using the Thiessen polygon algorithm, and then the dynamic spatial threshold of each base station is calculated using the K nearest neighbor idea. For example, for a certain base station, the average distance of the eight nearest base stations around it is calculated as the spatial threshold of the base station. Finally, based on the above processing results, a spatiotemporal double-layer DBSCAN clustering algorithm is used to identify stay points, where the spatial threshold is an adaptive dynamic threshold and the time threshold is set to 10 minutes, for example.

[0040] For example, first import the original base station data, such as (116.3906, 39.9092) and (116.4512, 39.9194). These points are numbered from 1 to N according to their spatial distribution, from left to right and from top to bottom. Then, the Delaunay triangulation algorithm is used to construct a triangulated network (TMN) from these N discrete points. The resulting triangles are also numbered from 1 to M, and a mapping table is established between the triangle numbers and the numbers of their three vertices. Next, the coordinates of the center of each triangle's circumscribed circle are calculated. For example, the circumscribed center of the first triangle is (116.4032, 39.9157). Based on the mapping table, all triangles adjacent to each base station are found. By connecting the circumscribed centers of these triangles, the Thiessen polygon for that base station is obtained. Polygons located at the TMN boundary are then closed by calculating the intersection of the point with the perpendicular bisector of the boundary line. Furthermore, for each base station, the distances to all edges of the Thiessen polygon are calculated and sorted from smallest to largest. For example, for point 1, the distances to the six edges of the polygon are 223m, 302m, 317m, 396m, 472m, and 525m, respectively. The median of the first three distances, 302m, is taken as the spatial dynamic threshold for this point. This is repeated for all other base stations, ultimately resulting in a comprehensive list of spatial dynamic thresholds for the entire region, which is used for subsequent stop point identification analysis.

[0041] For any base station corresponding to a Thiessen polygon, the KD tree algorithm is used to quickly query its 10 adjacent Thiessen polygons. Next, within these 10 adjacent Thiessen polygons, the distance between the current base station and each adjacent base station is calculated using Euclidean distance. The five closest base stations are selected as the nearest neighbors. Subsequently, the distances between the current base station and these five nearest neighbors are calculated, resulting in a distance set of {1.2km, 1.5km, 1.8km, 2.1km, 2.3km}. The maximum value, 2.3km, from this distance set is selected as the spatial threshold for the current base station. The spatial threshold is calculated using the formula: Spatial Threshold = max{1.2, 1.5, 1.8, 2.1, 2.3} = 2.3km. If the current base station is the last base station, the calculation ends; otherwise, the next base station is selected and the above steps are repeated. By looping through all base stations, the spatial dynamic threshold corresponding to each base station is ultimately obtained. Based on the calculated spatial dynamic threshold, the mobile phone signaling data is used to identify the stay points. If a user stays within the coverage area of ​​a base station for more than 15 minutes and the moving distance is less than the spatial threshold of the base station, the area is determined to be the user's stay point, realizing dynamic identification of stay points.

[0042] The trajectory point dataset P and dynamic spatial threshold dataset D are obtained based on the mobile phone signaling data. The trajectory point data includes attributes such as location coordinates (such as longitude and latitude) and timestamps, and the spatial threshold data defines the distance thresholds at different locations to determine the spatial proximity of the trajectory points. The trajectory points are processed in chronological order. For each unprocessed point p, its corresponding distance threshold D is first obtained from the spatial threshold dataset D. p , then take p as the center and calculate the relationship between p and other trajectory points p in increasing time. i distance d, until d exceeds D p So far, we get the spatial neighborhood point set P of p. Next, we determine whether each point p in P i Whether the time difference Δt with p exceeds the preset time threshold T, for example 10 minutes, if Δt>T, then p i It is classified into the spatiotemporal cluster C with p as the core. In this way, trajectory points that meet the conditions are continuously added to C until all spatial neighborhood points of p are traversed, and finally a cluster C that meets the spatiotemporal constraints is obtained. If the scale of C |C| is greater than or equal to the set stop point scale threshold N, for example 50, then C is marked as a candidate stop point. All candidate stop points are merged by running a density clustering algorithm such as DBSCAN to obtain the final stop point set S. Traverse each stop point s in S i , determine whether it is included in the identified stay area, if so, merge it into the area, otherwise iAs a new stay area. Based on this, the stay area sequence of each user is extracted to construct an individual trip chain. Finally, by counting the number and time distribution of trips between different stay areas, a resident trip OD matrix is ​​generated, which is used to analyze traffic flow and travel demand characteristics between areas.

[0043] Although the present application has been described with reference to specific features and embodiments thereof, it is apparent that various modifications and combinations may be made thereto without departing from the spirit and scope of the present application. Accordingly, this specification and the drawings are merely illustrative of the present application as defined by the appended claims and are deemed to cover any and all modifications, variations, combinations or equivalents within the scope of the present application. Obviously, those skilled in the art may make various modifications and variations to the present application without departing from the scope of the present application. Thus, the present application is intended to include such modifications and variations if they fall within the scope of the claims of the present application and their equivalents.

Claims

1. A spatiotemporal double-layer clustering stay point identification method based on Thiessen polygon network, characterized by The following steps are involved: S1, read mobile phone signaling and base station positioning data; S2. Construct a Thiessen polygon network based on base station positioning data; S3. Calculate the spatial dynamic threshold of each base station using a K-nearest neighbor algorithm based on a Thiessen polygon network; S4. Based on the spatial dynamic threshold, the stay points of the mobile phone signaling data are identified through the spatiotemporal double-layer clustering algorithm.

2. The method for identifying stay points using a spatiotemporal double-layer clustering method based on a Thiessen polygon network according to claim 1, wherein: Step S1 includes the steps of cleaning the mobile phone signaling data: S11, determining whether two adjacent records of the current data are from the same base station and the time difference is less than a preset threshold; if so, marking the current data record as ping-pong effect data and discarding it; S12. Calculate the speed of two adjacent records. If the speed is greater than a preset speed threshold, mark the latter record as drift data and remove it.

3. The method for identifying stay points using a spatiotemporal double-layer clustering method based on a Thiessen polygon network according to claim 1, wherein: Step S1 includes the steps of cleaning the mobile phone signaling data: S101. Aggregate all mobile phone signaling data by user ID and sort by time to obtain a mobile phone signaling data set sorted by user ID and time, so as to facilitate subsequent processing of individual user data; S102, reading the mobile phone signaling data of a single user at one time; S103, read the i-th record (i starts at 1), i+1 record, and i+2 record one by one; if the data read reaches the last row or there are no last two rows of data, determine whether it is the last user, if not, return to step S102, otherwise end; S104: Determine whether the base stations recorded in the i-th and i+2-th rows are the same base station, and calculate the time difference ∆t between the i+1-th and i+2-th rows. i+1 , judge Δt i+1 Whether it is within the preset time threshold T; If both conditions are met, jump to step S105; Otherwise, jump to step S106; S105, the i+1th record is ping-pong effect data, which is directly discarded and the process goes to S103; S106. Set i=i+1 and jump to S103.

4. The method for identifying stay points by using a spatiotemporal double-layer clustering method based on a Thiessen polygon network according to claim 1, The method is characterized in that step S1 includes the step of cleaning the mobile phone signaling data: S111. Aggregate all mobile phone signaling data by user ID and sort by time to obtain a mobile phone signaling data set sorted by user ID and time, so as to facilitate subsequent processing of individual user data; S112, reading the mobile phone signaling data of a single user at one time; S113. Read the i-th record (i starts at 1) and i+1-th record one by one. If the data read reaches the last row or there is no next row of data, determine whether it is the last user. If not, return to step 2. Otherwise, end. S114, calculate the distance di and the interval time Δti between the i-th and i+1-th data, and then calculate the speed from point i to point i+1 vi, the distance is calculated as shown in formula (1) and (2), and the speed is calculated as shown in formula (3); (1); (2); (3); In the formula, R is the radius of the earth, which is approximately 6371 km; i ,lon i )、(lat i+1 ,lon i+1 ) are the spatial coordinates of the two records i and i+1, d i is the distance between two points, v i is the velocity of point i; S115, Comparison v i and the preset speed threshold v t The size of v i >v i , move to step S116; Otherwise, move to step S117; S116, directly delete the (i+1)th record and jump to step S113; S117. Set i=i+1 and jump to step S113.

5. The method for identifying stay points by using a spatiotemporal double-layer clustering method based on a Thiessen polygon network according to claim 1, wherein in step S2, constructing the Thiessen polygon network comprises the following steps: S21. Import base station positioning data and number the base station positioning data from left to right and from top to bottom according to spatial distribution; S22. Construct a triangulated network for the discrete base station positioning data, that is, connect each base station with its adjacent base stations to form triangles, number the formed triangles, and establish a mapping relationship between the triangles and the discrete base station numbers; S23. Calculate the circumcenter of each triangle, and then connect the circumcenters of these triangles according to the adjacent triangles of each discrete base station in S22 to obtain the Thiessen polygon corresponding to each discrete point.

6. The method for identifying stay points using a spatiotemporal double-layer clustering method based on a Thiessen polygon network according to claim 5, wherein: For the Thiessen polygons at the boundary of the triangulation network, a closed polygon is formed by intersecting the perpendicular bisector with the boundary.

7. The method for identifying stay points using a spatiotemporal double-layer clustering method based on a Thiessen polygon network according to claim 1 is characterized in that in step S3, the current base station P1 is taken as the center, k nearest base stations of its M adjacent Thiessen polygons are selected, the distances between the current base station P1 and the k base stations are calculated, and the maximum value of the k distances is selected as the spatial threshold of the base station.

8. The method for identifying stay points by using a spatiotemporal double-layer clustering method based on a Thiessen polygon network according to claim 7, wherein Calculating the spatial dynamic threshold of each base station includes the following steps: S31, set the set of all base stations P= { P 1 ,P 2 ,……,P n }, the Thiessen polygon set corresponding to each base station T= { T 1 , T 2 ,…,T n }; S32, base station P 1 as an example, the Thiessen polygon where it is located T 1 has M adjacent Thiessen polygons, calculate the base station P 1 The Euclidean distance set to the adjacent Thiessen polygon control point base station D ={ D 1 ,D 2 ,…,D m }, and sort them from small to large; S33, judge according to the set k value, if K>M , then let K=M ; S34, then take the first k distance values ​​in the distance set D D 1 ,D 2 ,…,D k , then the base station P A spatial threshold of 1 D p1 is the maximum value among these k distance values, and the calculation is shown in formula (4). (4)。 9. The method for identifying stay points based on a spatiotemporal double-layer clustering of Thiessen polygon networks according to claim 1, wherein in step S4, identifying stay points of mobile phone signaling data using a spatiotemporal double-layer clustering algorithm comprises the following steps: Based on the DBSCAN clustering algorithm, stay point identification is performed according to spatial dynamic thresholds.

10. The method for identifying stay points based on a spatiotemporal double-layer clustering of Thiessen polygon networks according to claim 1, wherein DB The SCAN clustering algorithm includes the following steps: S41. Import a user's one-day mobile phone signaling travel trajectory data and a base station spatial threshold dataset based on Thiessen polygons; S42, read the unprocessed points in the trajectory data in chronological order p , take out the point p Distance threshold in a spatial threshold dataset D p , calculated in order of time increments p Point and p i The distance between d ,until d > D p Stop and get the track point p Spatial Threshold D p A collection of points within the range P= { p 1 , p 2 ,…,p n }; S43, Judgment Point p arrive P The last point in the set p n The time difference ∆ t and time thresholds T Whether ∆ t > T If satisfied, execute step S44; otherwise, jump to step S42; S44. Create a new cluster C ,Will p Points and Sets P Add the points in C middle; S45, read clusters in sequence C Medium p Other trajectory points p i , calculate the trajectory point according to step S42 p i Spatial Threshold D pi A collection of points within the range P pi ={ p i1 ,p i2 ,…,p in }; S46, point p i arrive P pi Each point in the set p ij The time difference ∆ t and time thresholds T Whether ∆ t > T If satisfied, then point p ij Join cluster C; otherwise, jump to step S45; S47, determine whether all trajectory points have been processed, if so, end; Otherwise jump to step S42.

Citation Information

Patent Citations

  • Railway transportation junction population type identification method based on mobile phone signaling data

    CN110113718A

  • External passenger transport hub concentrated and sparse space distribution analysis method based on mobile phone signaling data

    CN111681421A