Bus service gap identification method based on mobile phone signaling and bus data coupling

By processing mobile signaling data and fusing multi-dimensional spatiotemporal features, gaps in public transportation services are identified, solving the problems of data lag and limited samples in traditional public transportation planning, and achieving precise optimization of the public transportation network.

CN121882479BActive Publication Date: 2026-05-22JILIN JIANZHU UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
JILIN JIANZHU UNIVERSITY
Filing Date
2026-03-20
Publication Date
2026-05-22

Smart Images

  • Figure CN121882479B_ABST
    Figure CN121882479B_ABST
Patent Text Reader

Abstract

The application discloses a bus service gap identification method based on mobile phone signaling and bus data coupling and belongs to the technical field of traffic supply condition supervision. The method specifically comprises the following steps: S1, data collection; S2, data processing; S3, identification of job and residence locations; S4, construction of a commuting OD matrix; S5, construction of a commuting demand index; S6, construction of a bus supply level index; and S7, analysis of a demand-supply matching degree index. First, job and residence locations of residents are identified by using mobile phone signaling data, and a commuting OD matrix is established; then, commuting demand intensity and bus supply level are quantified; finally, a demand-supply matching degree index model is constructed, and a bus service gap area is identified by using GIS spatial coupling analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of traffic supply monitoring technology, specifically relating to a method for identifying public transport service gaps based on the coupling of mobile phone signaling and public transport data. Background Technology

[0002] With urbanization, urban traffic congestion has become a prominent problem restricting sustainable development and affecting residents' quality of life. Developing transit-oriented urban transportation (TOD) systems is an important solution. However, traditional transit planning relies on periodic resident travel surveys, which suffers from high costs, slow updates, limited samples, and difficulty in reflecting dynamic spatiotemporal characteristics. This often leads to planning lagging behind urban development, resulting in a dilemma of "unclear demand and inaccurate supply."

[0003] In recent years, mobile signaling data has become an effective means of mining commuter origin-destination (OD) data and analyzing traffic demand due to its advantages such as wide coverage, large sample size, real-time continuity, and ability to track individual spatiotemporal trajectories. However, existing research mainly focuses on analyzing job-housing balance and commuting characteristics, while research on its application to precisely optimize public transportation networks is still relatively lacking. Summary of the Invention

[0004] To address the technical problem of a lack of research on precise optimization of public transport networks, this invention provides a public transport service gap identification method based on the coupling of mobile phone signaling and public transport data. This method identifies public transport service gaps in the research area, thereby providing technical support for optimizing the public transport network.

[0005] The method described in this invention is specifically as follows:

[0006] S1. Data Collection: Collect mobile phone signaling data of users in the study area, including user ID, user entry time at the stop point, user departure time at the stop point, latitude and longitude of the stop point;

[0007] S2. Data Processing: Encrypt user IDs, preprocess missing and duplicate data, preprocess ping-pong effect points and drift data, aggregate and filter preprocessed user dwell points, and identify valid dwell points.

[0008] S3. Residence and Employment Identification: Identify the residence and employment locations of users among the valid locations of stay, construct a confidence index, and select the locations that meet the confidence index from all the locations of stay corresponding to the residence and employment locations of users as the final residence and employment locations of stay.

[0009] S4. Construction of the commuting OD matrix: The study area is divided into N*N grids, and the grids to which the residential stop points and the employment stop points belong are numbered. The commuting OD flow between each grid is counted to form a gridded commuting OD matrix.

[0010] S5. Construction of Commuting Demand Index: Obtaining the commuting demand intensity index for the study area through the commuting OD matrix. ;

[0011] S6. Construction of Public Transport Supply Level Index: Obtaining the public transport supply level index of the study area through station coverage and route density. ;

[0012] S7. Demand-Supply Matching Index Analysis: Construct the demand-supply matching index MI=S / D, and analyze the gap in public transportation services through MI.

[0013] Furthermore, the preprocessing of missing and duplicate data specifically involves:

[0014] For missing value data, if the data has the following conditions: the ID field is empty or the format is abnormal, the entry time or departure time field is missing or the format is incorrect or there is a logical contradiction, or the longitude or latitude field of the stop point is empty, it will be marked as invalid data and deleted.

[0015] For duplicate data of the same user, if the spatial distance between two adjacent records is less than a set threshold and the difference between the departure time of the previous record and the entry time of the next record is less than the time tolerance, the two records will be merged, and the latitude and longitude of the stop point will be taken as a weighted average.

[0016] Furthermore, the ping-pong effect points are preprocessed by constructing a comprehensive ping-pong effect score. Proceed, if the first If the ping-pong effect score of a stop point is greater than 0.6, then this stop point is determined to be a ping-pong effect point, deleted, and adjacent recorded points are merged as new stop points. ,in, , and For coefficients, For the first A function for detecting the ping-pong effect round-trip pattern at each stop point. For the first Abnormal factors of dwell time at each stop point For the first Spatial oscillation index of a dwell point.

[0017] Furthermore, the drift data is preprocessed by constructing an identification method based on velocity discrimination and centroid outlier detection:

[0018] Calculate the first The implicit speed of movement between the current stop point and the next stop point ,in, and These represent the spatial distance and temporal distance between two adjacent stops, respectively. To prevent small constants from being divided by zero;

[0019] Calculate the weighted centroid coordinates of user dwell points ,in, For the first Duration of stay at each stop This indicates all user dwell times. Summation of latitudes This indicates all user dwell times. Summation of longitude, Represents all user dwell points The duration of stay is used as a weight;

[0020] Construct the outlier score function and calculate the first The centroid distance fraction of each rest point is shown in the formula: ,in, For the first The distance from each dwell point to the weighted centroid of the user's dwell point. The mean of the weighted centroid distances from all user dwell points to the user's dwell point. The standard deviation of the weighted centroid distance from all user dwell points to the user's dwell point;

[0021] Based on the above characteristics, the drift data is comprehensively determined as shown in the formula. As shown, data that conforms to the formula is identified as drift data and deleted.

[0022] in, The percentage of speed exceeding the limit, , Indicates the maximum permissible speed, according to Make dynamic adjustments.

[0023] Furthermore, the preprocessed user dwell points are aggregated and filtered to identify valid dwell points. Specifically, the preprocessed user dwell point data is clustered, with the neighborhood radius parameter set to... After obtaining the aggregated weighted dwell points m, calculate the total dwell time. If the total dwell time of the filtered dwell points is greater than the minimum effective dwell time threshold, then it is determined to be a valid dwell point.

[0024] Furthermore, the method for identifying the stop points corresponding to a user's residence is to select the stop points where the user spends the most time and has the highest frequency of stay during weekday nights as the stop points corresponding to their residence. The method for identifying the stop points corresponding to a user's workplace is to select the stop points where the user spends the most time and has the highest frequency of stay during weekday daytime hours as the stop points corresponding to their workplace. The confidence index is constructed as follows:

[0025] Construct residence score functions respectively and place of employment scoring function :

[0026] ,in, and These are the times for entering and leaving the stop, respectively. and These are weighting functions assigned based on the time period of residence and employment, respectively. This indicates the identified corresponding place of residence or place of employment.

[0027] Construct separate functions for determining place of residence and place of employment to determine the location of residence. and place of stay in employment :

[0028] ,in, For users All valid stops, The points of stay after excluding the place of residence from all of the user's valid points of stay;

[0029] Build users confidence index If the value is not lower than 0.5, it is considered a valid dwell point.

[0030] ;

[0031] The distance between the place of residence and the place of employment.

[0032] Furthermore, in the construction of the commuter OD matrix, the commuter OD flow between each grid is calculated as follows: ;

[0033] in, For grid functions, represented by user Place of residence and place of stay at employment Number them and map them to the corresponding grids; Table Grid To grid Commuter OD traffic, This indicates that for users Their place of residence is located at the address numbered Within the grid, all residential stops are filtered out. Users in

[0034] This indicates that for users Their work location is located at the address numbered Within the grid, all work location stops are filtered out. Users in This is an indicator function that takes the value 1 if the condition is met, and 0 otherwise.

[0035] Furthermore, the commuting demand intensity index of the study area was obtained through the commuting OD matrix. Specifically, this involves spatially overlaying the commuter OD matrix with the public transport network and station distribution, and then... To grid Commuter OD traffic is distributed to the connected grid From residential stop to grid Optimal path to workplace stop points On the various road segments traversed, each road segment Its commuting volume is recorded as Then the road segment Total commuting demand The sum of the allocated commuting volume: ; All road segments within a grid Total commuting demand Add them together to get the total commuting demand within the grid. Calculate the total commuting demand within all grids. The extreme value normalization method was then used to obtain the commuting demand intensity index for the study area. .

[0036] Furthermore, the study area's public transport supply level index was obtained through... Obtain, among which, and These are the weight functions, The site coverage rate after processing using the extreme value normalization method is... This refers to the line density.

[0037] Furthermore, when analyzing the public transport service gap using MI:

[0038] MI ≈ 1: This indicates that the level of public transport supply in the study area is basically equivalent to the intensity of commuting demand.

[0039] MI>1: This indicates that the public transport supply level in the study area is higher than the commuting demand intensity.

[0040] MI<1: This indicates that the public transportation supply is insufficient to meet commuting demand, and there is a service gap.

[0041] The beneficial effects of the method described in this invention are as follows:

[0042] First, mobile phone signaling data is used to identify residents' work and residence locations and establish a commuting OD matrix; then, the intensity of commuting demand and the level of public transport supply are quantified respectively; finally, a "demand-supply matching index" model is constructed, and GIS spatial coupling analysis is used to identify areas with public transport service gaps.

[0043] Compared to traditional single-rule, fixed-threshold, or pure clustering methods, this paper proposes a ping-pong effect and drift data identification method, and constructs a multi-dimensional spatiotemporal feature fusion system: In ping-pong effect detection, multimodal complementarity of triplet round-trip patterns, spatial oscillation index, and dwell time anomaly factors, combined with geometric modeling of direction change angle, is used to accurately characterize reciprocating oscillation behavior; In drift data identification, a displacement adaptive velocity threshold strategy is adopted, which dynamically adjusts the maximum allowable velocity according to the distance interval, and constructs a dual verification mechanism of velocity discrimination and centroid outlier, effectively balancing the identification needs of long-distance normal movement and short-distance high-speed abnormal jumps.

[0044] This method introduces a time-based weighted integral mechanism, replacing the traditional hard threshold division with a continuous weight function. This finely characterizes the differences in the contributions of residential and employment activities during core, transitional, and atypical periods, significantly improving the accuracy of dwell time behavior measurement. Simultaneously, it employs step-by-step identification and distance constraints to effectively prevent the same location from being repeatedly identified as a work-residence location and eliminates abnormal interference such as base station drift. Furthermore, a confidence index is constructed, automatically selecting reliable samples based on score proportions and work-residence distance, ensuring the representativeness of the output results. Compared to traditional methods based on frequency statistics, fixed time periods, or clustering, this method has significant advantages in semantic interpretability, data quality control, and practical application adaptability. Attached Figure Description

[0045] Figure 1 This is a flowchart of the method described in an embodiment of the present invention. Detailed Implementation

[0046] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.

[0047] This embodiment provides a method for identifying public transport service gaps based on the coupling of mobile phone signaling and public transport data, such as... Figure 1 The diagram shown is the overall flowchart of the method. The method will be explained in detail step by step below.

[0048] 1. Data Sources and Processing

[0049] 1.1 Data Source

[0050] The acquired mobile signaling data includes user ID, user entry time at the stop point, user exit time at the stop point, latitude of the stop point, and longitude of the stop point; some data samples are shown in Table 1.

[0051] Table 1. Sample Data

[0052]

[0053] Other auxiliary data included: (1) Administrative division data, road network data, POI data, land use data, and public transport network data were all uniformly converted to the WGS84 geographic coordinate system and were clipped, topology checked, and attribute supplemented and repaired.

[0054] 1.2 Data Processing

[0055] (1) Missing and duplicate

[0056] Data loss and duplicate records are the two most common quality defects in mobile signaling data.

[0057] For missing value data, if the data has the following conditions: the ID field is empty or the format is abnormal, the entry time or departure time field is missing or the format is incorrect or there is an obvious logical contradiction, or the longitude or latitude field is empty, it will be marked as invalid data and deleted.

[0058] For duplicate data of the same user, if the spatial distance between two adjacent records is less than the set threshold of 500 m and the difference between the departure time of the previous record and the entry time of the next record is less than the time tolerance of 5 min, then the two records will be merged and the coordinates will be taken as a weighted average.

[0059] (2) Ping-pong effect and drift data

[0060] The ping-pong effect is an inherent error phenomenon in cellular network positioning technology, which manifests as the user's location jumping back and forth between several adjacent points.

[0061] For the identification of ping-pong effect data, this embodiment proposes a multi-dimensional detection method based on spatiotemporal pattern recognition:

[0062] For users Construct a set of location sequences in chronological order. As shown in the formula:

[0063] ;

[0064] In the formula: For the first The coordinates of each stop point; For users to enter the The time spent at each stop; For users leaving the The duration of the stay; This represents the total number of times the same user has stayed.

[0065] Define a triplet round-trip pattern detection function to identify round-trip sequences, as shown in the equation:

[0066] ;

[0067] In the formula: For the first The ping-pong effect round-trip pattern detection function for each stop point has a value of 0 or 1. This is an indicator function; it returns 1 if the condition is true, and 0 otherwise. This is a logical operator, meaning AND; This is the Haversine distance formula, in meters. The spatial distance threshold is set to 500 m. The total time span of the three-point sequence is equal to the departure time of the last point minus the entry time of the first point, in seconds. The time window threshold is set to 300 seconds.

[0068] Calculate the angle of change of direction As shown in the formula:

[0069] ;

[0070] In the formula: From Click The displacement vector of a point; Let be the magnitude of the vector.

[0071] The Spatial Oscillation Index (SOI) is then constructed, as shown in the equation:

[0072] ;

[0073] In the formula: For the first The stopping point is within the radius of the sliding window. The spatial oscillation index within, ; The threshold for large-angle turnaround is set to 2.5 rad.

[0074] Define the stay duration anomaly factor as shown in the formula:

[0075] ;

[0076] In the formula: For the first Abnormal factors of dwell time at each stop point ; For the first Duration of stay at each stop , For reference dwell time, the value is 180 s.

[0077] Finally, the above features are combined to construct a comprehensive ping-pong effect score, as shown in the formula:

[0078] ;

[0079] In the formula: For the first The overall ping-pong effect score of each stop point ; , , These are weighting coefficients, summing to 1, and taking values ​​of 0.5, 0.2, and 0.3 respectively.

[0080] If the first If the ping-pong effect score of a stop point is greater than 0.6, then the point is determined to be a ping-pong effect point, the point is deleted and adjacent record points are merged, the coordinates are the weighted average of the stay duration, and the time is the earliest entry and the latest exit.

[0081] Drift data is characterized by large spatial jumps that occur within a very short time interval when a user is present.

[0082] For the detection of drift data, this embodiment constructs a recognition algorithm based on velocity discrimination and centroid outlier detection:

[0083] First, calculate the... The implicit speed of movement between each stop point and the next point is shown in the equation: ;

[0084] In the formula: and These represent the spatial distance and temporal distance between two adjacent points, respectively. To prevent small constants from being divided by zero, the value is set to 1 s.

[0085] The maximum allowable speed is dynamically adjusted based on the displacement distance, as shown in Table 2:

[0086] Table 2 Speed ​​Thresholds

[0087]

[0088] The user's weighted centroid coordinates are calculated as shown in the formula:

[0089] ;

[0090] In the formula: For the first Duration of stay at each stop This indicates all user dwell times. Summation of latitudes This indicates all user dwell times. Summation of longitude, Represents all user dwell points The duration of stay is used as a weight;

[0091] Then construct the outlier score function and calculate the first... The centroid distance fraction of each rest point is shown in the formula:

[0092] In the formula, For the first The distance from each dwell point to the weighted centroid of the user's dwell point. The mean of the weighted centroid distances from all user dwell points to the user's dwell point. The standard deviation of the weighted centroid distance from all user dwell points to the user's dwell point.

[0093] Finally, based on the above characteristics, drift data is comprehensively judged as shown in the formula. Data that meets the formula is determined to be drift data and deleted:

[0094] ;

[0095] in, The percentage of speed exceeding the limit, , Indicates the maximum permissible speed, according to Make dynamic adjustments.

[0096] Compared to traditional single-rule, fixed-threshold, or pure clustering methods, the ping-pong effect and drift data identification method proposed in this embodiment constructs a multi-dimensional spatiotemporal feature fusion system: In ping-pong effect detection, multimodal complementarity of triplet round-trip patterns, spatial oscillation index, and dwell time anomaly factors, combined with geometric modeling of direction change angle, achieves accurate characterization of reciprocating oscillation behavior; In drift data identification, a displacement adaptive velocity threshold strategy is adopted, dynamically adjusting the maximum allowable velocity according to the distance interval, and constructing a dual verification mechanism of velocity discrimination and centroid outlier, effectively balancing the identification needs of long-distance normal movement and short-distance high-speed abnormal jumps.

[0097] (3) Effective aggregation and screening of stop points

[0098] Based on the ST-DBSCAN spatial clustering method, the preprocessed stop point data is first clustered, with a neighborhood radius parameter of 500 m. After obtaining the aggregated weighted stop point attributes (latitude and longitude, and entry and exit times), the total stay duration is calculated. If the total stay duration of the filtered stop points is greater than the minimum effective stay duration threshold of 300 s, they are determined to be effective stop points.

[0099] After the above preprocessing steps, the retention rate of effective dwell point data after cleaning was 79.6%.

[0100] 2. Demand-Supply Coupling Analysis of Public Transportation Services

[0101] 2.1 Identification of Work and Residence

[0102] The following rules are used to extract work and residence data from mobile phone signaling data:

[0103] (1) Residence identification: The place where the user stays the longest and appears most frequently during weekday nights (20:00~24:00, 00:00-06:00) is selected as the residence.

[0104] (2) Identification of place of employment: Select the place of employment where the user spends the most time and appears most frequently during weekdays (10:00~16:00) (excluding the identified place of residence).

[0105] Weighting function based on work-residence time period and The weight of the core residential period is 1, the weight of the transition period is 0.7, and the weight of the others is 0.2, as shown in Tables 3 and 4.

[0106] Table 3 Weighting of Residential Period

[0107]

[0108] Table 4 Weighting of Employment Periods

[0109]

[0110] Construct score functions for place of residence and place of employment respectively:

[0111] ,in, To construct a residence score function, For the location of employment score function, and These are the times for entering and leaving the stop, respectively. and These are weighting functions assigned based on the time period of residence and employment, respectively. This indicates the identified corresponding place of residence or place of employment. In This indicates the identified corresponding residential location of stay. In This indicates the identified corresponding place of stay for employment.

[0112] Construct separate functions for determining place of residence and place of employment to determine the location of residence. and place of stay at employment :

[0113] ,in, For users All valid stops, The points of stay after excluding the place of residence from all of the user's valid points of stay;

[0114] Build users confidence index If the value is not lower than 0.5, it is considered a valid dwell point.

[0115] ;

[0116] The distance between the place of residence and the place of employment.

[0117] This method introduces a time-period weighted integral mechanism, replacing the traditional hard threshold division with a continuous weight function. This finely characterizes the differences in the contributions of residential and employment activities during core, transitional, and atypical periods, significantly improving the accuracy of dwelling behavior measurement. Simultaneously, it employs step-by-step identification and distance constraints to effectively prevent the same location from being repeatedly identified as a work-residence location and eliminates abnormal interference such as base station drift. Furthermore, it constructs a confidence index, automatically selecting reliable samples based on score proportion and work-residence distance, ensuring the representativeness of the output results. Compared to traditional methods based on frequency statistics, fixed time period total duration, or clustering, the method in this embodiment has significant advantages in semantic interpretability, data quality control, and practical application adaptability.

[0118] 2.2 Construction of Commuter OD Matrix

[0119] The study area was divided into 500 m × 500 m grids, and the grids to which the identified work and residence locations belonged were numbered. The commuting OD flow between each grid was counted to form a gridded commuting OD matrix.

[0120] In the construction of the commuter OD matrix, the commuter OD flow between each grid is calculated as follows: ;

[0121] in, For grid functions, represented by user Place of residence and place of stay at employment Number them and map them to the corresponding grids; Table Grid To grid Commuter OD traffic, This indicates that for users Their place of residence is located at the address numbered Within the grid, all residential stops are filtered out. Users in

[0122] This indicates that for users Their work location is located at the address numbered Within the grid, all work location stops are filtered out. Users in This is an indicator function that takes the value 1 if the condition is met, and 0 otherwise.

[0123] 2.3 Commuting Demand and Public Transportation Supply Matching Index

[0124] 2.3.1 Commuting Demand Index

[0125] Spatially overlay the commuter OD matrix with the public transport network and station distribution, and combine the grid... To grid Commuter OD traffic is distributed to the connected grid From residential stop to grid Optimal path to workplace stop points On the various road segments traversed, each road segment Its commuting volume is recorded as Then the road segment Total commuting demand The sum of the allocated commuting volume: ; All road segments within a grid Total commuting demand Add them together to get the total commuting demand within the grid. Calculate the total commuting demand within all grids. The extreme value normalization method was then used to obtain the commuting demand intensity index for the study area. :

[0126] ,in, This represents the total commuting demand across all grids. The minimum value in, This represents the total commuting demand across all grids. The maximum value in.

[0127] 2.3.2 Public transport supply level

[0128] The level of public transport supply is evaluated from two dimensions: "station coverage" and "route density".

[0129] (1) Station Coverage: A buffer zone with a certain radius is created centered on the bus stops, and the percentage of the buffer zone's area relative to the total area of ​​the study area is calculated. Simultaneously, for a more refined evaluation, the study area is divided into a 500 m × 500 m grid, and the bus stop coverage within each grid is calculated. A buffer zone with a radius of 500 m is created centered on the obtained bus stops, and their total area is calculated after merging. Total area of ​​the study area The ratio of the two values ​​yields the regional site coverage rate. As shown in the following formula:

[0130] ;

[0131] A more precise quantification method utilizes a grid to calculate the coverage of a single grid cell. The study area is divided into a 500 m × 500 m grid, and then the coverage area of ​​the bus stop buffer zone is calculated. With the total area of ​​the grid The ratio is given by the following formula:

[0132] ;

[0133] (2) Route density: Calculate the cumulative length of bus routes per unit area (km / km) 2 ):

[0134] ,in, For bus route density, This represents the total length of all bus routes within the study area.

[0135] The public transport supply level index is composed of both station coverage and route density. The station coverage rate C is obtained by normalizing the extreme values ​​of the two indicators. norm and line density D norm Then, the public transport supply level index S is synthesized using an equal-weighted method, as shown in the following formula:

[0136] ;

[0137] The higher the value, the higher the level of public transportation supply in the area.

[0138] 2.3.3 Demand-Supply Matching Index

[0139] To analyze the level of public transport service, a demand-supply matching index (MI) is constructed, as follows: MI = S / D;

[0140] The meaning of this index is:

[0141] MI ≈ 1: This indicates that the public transport supply level within the grid is basically equivalent to the commuting demand intensity, indicating a balance between supply and demand and a good service level.

[0142] MI>1: This indicates that the level of public transport supply is higher than the intensity of commuting demand, which may indicate a surplus of resources and that public transport capacity is not being fully utilized.

[0143] MI < 1: This indicates that the public transportation supply is insufficient to meet commuting demand, resulting in a service gap. The smaller the MI value, the more severe the gap.

[0144] Based on the relevant provisions of the "Classification Standard for Urban Public Transportation" and the "Planning Standard for Urban Comprehensive Transportation System", the following classification criteria have been established:

[0145] (1) Level 1 service gap (MI<0.5): This level of area has a serious supply and demand imbalance, which is manifested in the fact that the coverage rate of bus stops within 500 m is less than 50%, or the route density is less than 1.0 km / km², and the commuting demand intensity is high.

[0146] (2) Secondary service gap (0.5 ≤ MI<0.8): The public transport supply in this level area has a certain coverage, with a station coverage rate between 50% and 70%, or the departure frequency is low, or there are problems with inconvenient connections.

[0147] (3) Basic matching of supply and demand (0.8 ≤ MI<1.2): The supply of public transportation is relatively balanced with the commuting demand, and the service quality is good.

[0148] (4) Supply is relatively abundant (MI ≥ 1.2): Public transport resources are sufficient or even excessive, and appropriate adjustments can be made to areas with shortages.

[0149] Sensitivity analysis has verified that the two critical values ​​of MI, 0.5 and 0.8, can effectively distinguish different levels of service gaps and provide a clear quantitative basis for resource allocation decisions.

Claims

1. A method for identifying public transport service gaps based on the coupling of mobile phone signaling and public transport data, characterized in that, The method includes the following steps: S1. Data Collection: Collect mobile phone signaling data of users in the study area, including user ID, user entry time at the stop point, user departure time at the stop point, latitude and longitude of the stop point; S2. Data Processing: Encrypt user IDs, preprocess missing and duplicate data, preprocess ping-pong effect points and drift data, aggregate and filter preprocessed user dwell points, and identify valid dwell points. Preprocessing of ping-pong effect points to construct a comprehensive ping-pong effect score Proceed, if the first If the ping-pong effect score of a stop point is greater than 0.6, then this stop point is determined to be a ping-pong effect point, deleted, and adjacent recorded points are merged as new stop points. ,in, , and For coefficients, For the first A function for detecting the ping-pong effect round-trip pattern at each stop point. For the first Abnormal factors of dwell time at each stop point For the first The stopping point is within the radius of the sliding window. The spatial oscillation index within; S3. Residence and Employment Identification: Identify the residence and employment locations of users among the valid locations of stay, construct a confidence index, and select the locations that meet the confidence index from all the locations of stay corresponding to the residence and employment locations of users as the final residence and employment locations of stay. S4. Construction of the commuting OD matrix: The study area is divided into N*N grids, and the grids to which the residential stop points and the employment stop points belong are numbered. The commuting OD flow between each grid is counted to form a gridded commuting OD matrix. S5. Construction of Commuting Demand Index: Obtaining the commuting demand intensity index for the study area through the commuting OD matrix. ; S6. Construction of Public Transport Supply Level Index: Obtaining the public transport supply level index of the study area through station coverage and route density. ; S7. Demand-Supply Matching Index Analysis: Construct the demand-supply matching index MI=S / D, and analyze the gap in public transportation services through MI.

2. The method for identifying public transport service gaps based on the coupling of mobile phone signaling and public transport data according to claim 1, characterized in that, The preprocessing of missing and duplicate data is as follows: For missing value data, if the data has the following conditions: the ID field is empty or the format is abnormal, the entry time or departure time field is missing or the format is incorrect or there is a logical contradiction, or the longitude or latitude field of the stop point is empty, it will be marked as invalid data and deleted. For duplicate data of the same user, if the spatial distance between two adjacent records is less than a set threshold and the difference between the departure time of the previous record and the entry time of the next record is less than the time tolerance, the two records will be merged, and the latitude and longitude of the stop point will be taken as a weighted average.

3. The method for identifying public transport service gaps based on the coupling of mobile phone signaling and public transport data according to claim 2, characterized in that, Preprocessing of drift data is performed by constructing an identification method based on velocity discrimination and centroid outlier detection: Calculate the first The implicit speed of movement between the current stop point and the next stop point ,in, and These represent the spatial distance and temporal distance between two adjacent stops, respectively. To prevent small constants from being divided by zero; Calculate the weighted centroid coordinates of user dwell points ,in, For the first Duration of stay at each stop This indicates all user dwell times. Summation of latitudes This indicates all user dwell times. Summation of longitude, Represents all user dwell points The duration of stay is used as a weight; Construct the outlier score function and calculate the first The centroid distance fraction of each rest point is shown in the formula: ,in, For the first The distance from each dwell point to the weighted centroid of the user's dwell point. For the first The coordinates of each stop point The mean of the weighted centroid distances from all user dwell points to the user's dwell point. The standard deviation of the weighted centroid distance from all user dwell points to the user's dwell point; Based on the above characteristics, the drift data is comprehensively determined as shown in the formula. As shown, data that conforms to the formula is identified as drift data and deleted. in, The percentage of speed exceeding the limit, , Indicates the maximum permissible speed, according to Make dynamic adjustments.

4. The method for identifying public transport service gaps based on the coupling of mobile phone signaling and public transport data according to claim 3, characterized in that, The preprocessed user dwell points are aggregated and filtered to identify valid dwell points. Specifically, the preprocessed user dwell point data is clustered, and the neighborhood radius parameter is set to... After obtaining the aggregated weighted dwell points m, calculate the total dwell time. If the total dwell time of the filtered dwell points is greater than the minimum effective dwell time threshold, then it is determined to be a valid dwell point.

5. The public transport service gap identification method based on the coupling of mobile phone signaling and public transport data according to claim 4, characterized in that, The method for identifying the stop points corresponding to a user's residence is to select the stop points where the user spends the most time and has the highest frequency of stay during weekday nights as the stop points corresponding to their residence. The method for identifying the stop points corresponding to a user's workplace is to select the stop points where the user spends the most time and has the highest frequency of stay during weekday daytime hours as the stop points corresponding to their workplace. The confidence index is constructed as follows: Construct residence score functions respectively and place of employment scoring function : ,in, and These are the times for entering and leaving the stop, respectively. and These are weighting functions assigned based on the time period of residence and employment, respectively. This indicates the identified corresponding place of residence or place of employment. Construct separate functions for determining place of residence and place of employment to determine the location of residence. and place of stay at employment : ,in, For users All valid stops, The points of stay after excluding the place of residence from all of the user's valid points of stay; Build users confidence index If the value is not lower than 0.5, it is considered a valid dwell point. ; The distance between the place of residence and the place of employment.

6. The method for identifying public transport service gaps based on the coupling of mobile phone signaling and public transport data according to claim 5, characterized in that, In the construction of the commuter OD matrix, the commuter OD flow between each grid is calculated as follows: ; in, For grid functions, represented by user Place of residence and place of stay in employment Number them and map them to the corresponding grids; Table grid To grid Commuter OD traffic, This indicates that for users Their place of residence is located at the address numbered Within the grid, all residential stops are filtered out. Users in This indicates that for users Their work location is located at the address numbered Within the grid, all work location stops are filtered out. Users in This is an indicator function that takes the value 1 if the condition is met, and 0 otherwise.

7. The method for identifying public transport service gaps based on the coupling of mobile phone signaling and public transport data according to claim 6, characterized in that, The commuting demand intensity index of the study area was obtained using the commuting OD matrix. Specifically, this involves spatially overlaying the commuter OD matrix with the public transport network and station distribution, and then... To grid Commuter OD traffic is distributed to the connected mesh From residential location to grid Optimal path to workplace stop points On the various road segments traversed, each road segment Its commuting volume is recorded as Then the road segment Total commuting demand The sum of the allocated commuting volume: ; All road segments within a grid Total commuting demand Add them together to get the total commuting demand within the grid. Calculate the total commuting demand within all grids. The extreme value normalization method was then used to obtain the commuting demand intensity index for the study area. .

8. The method for identifying public transport service gaps based on the coupling of mobile phone signaling and public transport data according to claim 7, characterized in that, The study area's public transport supply level index was obtained through... Obtain, among which, and These are the weight functions, The site coverage rate after processing using the extreme value normalization method is... This refers to the line density.

9. The method for identifying public transport service gaps based on the coupling of mobile phone signaling and public transport data according to claim 8, characterized in that, When analyzing the public transport service gap using MI: MI = 1: This indicates that the level of public transport supply in the study area is comparable to the intensity of commuting demand. MI>1: This indicates that the public transport supply level in the study area is higher than the commuting demand intensity. MI<1: This indicates that the public transportation supply is insufficient to meet commuting demand, and there is a service gap.