Private car travel behavior identification method based on communication big data
By using spatiotemporal density clustering based on communication signaling data and multi-strategy trajectory matching, private car travel behavior is identified, solving the problems of limited coverage, high cost, and poor timeliness of traditional methods. This achieves accurate travel behavior identification across all scenarios at low cost, supporting traffic management decisions.
Patent Information
- Application Number
- CN202511722090.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-21
- Publication Date
- 2026-02-03
AI Technical Summary
Traditional traffic survey methods face problems such as limited coverage, high cost, poor timeliness, and insufficient data continuity and completeness in identifying private car travel behavior, making it difficult to meet the needs of modern urban traffic planning and management for accurate and real-time data.
By acquiring the communication signaling data of the target user, a single trip trajectory chain is constructed using a spatiotemporal density clustering algorithm. Combined with a multi-strategy trajectory matching method, the travel mode is identified, and private car travel behavior identification results are generated, including subway travel, low-speed travel, public transport travel, and private car travel. Deep packet inspection technology and a subway private network base station database are used for judgment, and a ride-hailing driver database is constructed for auxiliary analysis.
It enables full-scenario travel trajectory recognition without additional hardware, reducing costs, improving recognition accuracy and timeliness, providing comprehensive private car travel behavior analysis, and supporting traffic planning and management decisions.
Smart Images

Figure CN121456764A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of transportation planning technology, and in particular to a method for identifying private car travel behavior based on big data communication. Background Technology
[0002] With the acceleration of urbanization and the rapid increase in the number of private cars, problems such as traffic congestion and environmental pollution have become increasingly prominent. Accurately identifying private car travel behavior is crucial for optimizing urban traffic structure and formulating scientific traffic management policies. Traditional traffic survey methods, such as resident travel surveys and vehicle license plate recognition, while providing some traffic-related information, have significant limitations. These methods typically have limited coverage and cannot comprehensively reflect private car travel in all urban areas. Furthermore, these methods are costly and lack timeliness, failing to meet the real-time data requirements of urban traffic planning and management. In addition, traditional survey methods are insufficient in terms of data continuity and completeness, failing to provide complete information on the entire chain of private car travel, which poses a challenge to in-depth analysis of travel behavior and the formulation of precise traffic policies.
[0003] Therefore, traditional traffic survey methods face problems such as limited coverage, high cost, poor timeliness, and insufficient data continuity and integrity in the identification of private car travel behavior. These problems make it difficult for traditional traffic survey methods to meet the needs of modern urban traffic planning and management for accurate and real-time data. There is an urgent need for a new method that can overcome these shortcomings. Summary of the Invention
[0004] The main purpose of this application is to provide a method for identifying private car travel behavior based on big data communication, which aims to solve the problems of limited coverage, high cost, poor timeliness, and insufficient data continuity and integrity faced by traditional traffic survey methods in identifying private car travel behavior.
[0005] To achieve the above objectives, this application proposes a method for recognizing private car travel behavior based on big data communication, the method comprising: Acquire communication signaling data of the target user within a preset time period; Based on the communication signaling data, construct several single-trip trajectory chains for the target user; Identify the single travel mode corresponding to each of the single travel trajectory chains; Based on each of the aforementioned single travel modes, a private car travel behavior recognition result corresponding to the target user is generated.
[0006] In one possible implementation, acquiring the target user's communication signaling data within a preset time period further includes: Obtain the raw communication signaling data of the target user within a preset time period; Based on the timestamp and location information in the original communication signaling data, duplicate data in the original communication signaling data is removed to obtain the communication signaling processing data. According to the preset ping-pong effect elimination method, the ping-pong effect data in the communication signaling processing data is eliminated to obtain the communication signaling data.
[0007] In one possible implementation, constructing several single-trip trajectory chains for the target user based on the communication signaling data includes: A spatiotemporal density clustering algorithm is used to aggregate the spatiotemporally adjacent trajectory points in the communication signaling data into user dwell points, generating a user dwell point sequence that includes the dwell point location and dwell time. Based on the location of the dwell point and the dwell time, the time interval and spatial distance between two adjacent user dwell points are calculated. If the time interval is greater than a preset time threshold and the spatial distance is greater than a preset spatial threshold, then the two user dwell points are determined as a pair of travel origin and destination points. The city road network is acquired, and a multi-strategy trajectory matching method is used to match the trajectory points between each pair of origin and destination to the city road network, thus integrating them to form the trajectory chains of each single trip.
[0008] In one possible implementation, identifying the single travel mode corresponding to each of the single travel trajectory chains includes: Obtain the preset database of indoor distributed base stations in underground parking lots and the Internet access data in the communication signaling data; Based on the preset underground parking lot indoor distributed base station database, determine whether there are records of the single trip trajectory chain residing in the underground parking lot indoor distributed base station within a set time period before the start and a set time period before the end, and obtain the first judgment result; Deep packet inspection technology is used to parse the server network identification information in the Internet data to determine whether there are records of accessing navigation application services during the trip period, and a second judgment result is obtained. Based on the first judgment result and the second judgment result, the confidence level of private car travel is determined. If the confidence level of private car travel is medium or high, then the single travel mode is determined to be private car travel.
[0009] In one possible implementation, identifying the single travel mode corresponding to each of the single travel trajectory chains includes: Obtain the preset database of subway private network base stations; The base station identifiers corresponding to the trajectory points of the single trip trajectory chain are compared with the metro private network base station database. If at least one trajectory point has a base station identifier that belongs to the metro private network base station database, then the single trip mode is determined to be metro travel. If the determination result is not subway travel, then calculate the average travel speed of the single travel trajectory chain and the percentage of travel time in the single travel trajectory chain where the speed reaches or exceeds the preset high speed threshold. If the average driving speed is lower than a preset speed threshold and the percentage of driving time is lower than a preset percentage threshold, then the single trip mode is determined to be low-speed travel.
[0010] In one possible implementation, identifying the single travel mode corresponding to each of the single travel trajectory chains includes: Obtain the origin-destination pair corresponding to the single trip trajectory chain, and under the conditions of limiting the number of transfers, detour index and bus stop distance threshold, perform bus route navigation based on bus network data to generate several candidate bus routes; Calculate the similarity between the trajectory of each candidate bus route and the trajectory of the single trip trajectory chain; If the trajectory similarity of at least one candidate bus route exceeds a preset similarity threshold, then the single trip mode is determined to be public transportation.
[0011] In one possible implementation, generating the private car travel behavior recognition result corresponding to the target user based on each of the single travel modes includes: The number of private car trajectories in all single travel trajectory chains within the preset time period, where the single travel mode is private car travel, is counted, and the proportion of the number of private car trajectories to the total effective travel trajectories is calculated. If the percentage exceeds a preset percentage threshold, the target user is preliminarily determined to be a potential private car owner; Obtain and construct a ride-hailing driver database based on the daily mileage, daily driving time, and trajectory detour index from historical single-trip trajectory chains; The target users initially identified as potential private car owners are compared with the user identifiers and feature data in the ride-hailing driver database to generate private car travel behavior recognition results corresponding to the target users.
[0012] Furthermore, to achieve the above objectives, this application also proposes a private car travel behavior recognition device based on communication big data, the private car travel behavior recognition device based on communication big data comprising: The acquisition module is used to acquire communication signaling data of the target user within a preset time period; The construction module is used to construct several single-trip trajectory chains of the target user based on the communication signaling data; The identification module is used to identify the single travel mode corresponding to each of the single travel trajectory chains. The generation module is used to generate private car travel behavior recognition results for the target user based on each of the single travel modes.
[0013] Furthermore, to achieve the above objectives, this application also proposes a private car travel behavior recognition device based on communication big data. The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. The computer program is configured to implement the steps of the private car travel behavior recognition method based on communication big data as described above.
[0014] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the private car travel behavior recognition method based on communication big data as described above.
[0015] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the private car travel behavior recognition method based on communication big data as described above.
[0016] This application provides a method for identifying private car travel behavior based on communication big data. This method acquires communication signaling data of a target user within a preset time period, and then constructs several single-trip trajectory chains for the target user based on the communication signaling data. This identifies the single-trip mode corresponding to each single-trip trajectory chain, and then generates a private car travel behavior identification result for the target user based on each single-trip mode. Therefore, identification is performed based on communication signaling data, without the need for additional hardware deployment. It can cover the target user's full-scenario travel trajectory within a preset time period, avoiding the problems of limited coverage and high costs associated with traditional survey methods. Furthermore, constructing single-trip trajectory chains based on communication signaling data can reconstruct the user's actual travel path, reducing identification errors caused by trajectory ambiguity. Attached Figure Description
[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a flowchart illustrating an embodiment of the private car travel behavior recognition method based on big data communication provided in this application. Figure 2 A scenario implementation diagram of the single-trip trajectory chain provided for the private car travel behavior recognition method based on communication big data in this application; Figure 3 This diagram illustrates a scenario related to public transportation, providing a framework for the private car travel behavior recognition method based on big data communication, as described in this application. Figure 4 A simplified flowchart illustrating the private car travel behavior recognition method based on communication big data provided in this application; Figure 5 This is a schematic diagram of the module structure of the private car travel behavior recognition device based on communication big data according to an embodiment of this application; Figure 6 This is a schematic diagram of the hardware operating environment of the private car travel behavior recognition method based on communication big data in the embodiments of this application.
[0020] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0021] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0022] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0023] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone; or an electronic device, big data service platform, or private car travel behavior recognition system based on communication big data capable of realizing the above functions. The following description uses a private car travel behavior recognition system based on communication big data as an example to illustrate this embodiment and the subsequent embodiments.
[0024] Based on this, embodiments of this application provide a method for recognizing private car travel behavior based on big data communication, referring to... Figure 1 , Figure 1This is a flowchart illustrating an embodiment of the private car travel behavior recognition method based on communication big data provided in this application.
[0025] In this embodiment, the private car travel behavior recognition method based on communication big data includes steps S11 to S14: Step S11: Obtain the communication signaling data of the target user within a preset time period; It should be noted that the target user refers to the mobile communication terminal user who needs to be identified based on communication big data for private car travel behavior recognition. Their identity can be uniquely determined through user identifiers. The preset time refers to the pre-set time period for collecting communication signaling data. This period needs to be sufficient to cover multiple travel behaviors of the user to ensure the accuracy of recognition. The communication signaling data refers to standardized data that has been pre-processed to remove duplicate data and ping-pong effect data. It includes core information such as user identifiers, timestamps, base station identifiers, and base station locations. The sources include mobile phones, tablets, computers, watches, wristbands, and vehicle communication, etc., without any restrictions. In this way, relying on the wide coverage and real-time dynamic characteristics of communication signaling data, the problems of limited coverage and poor data timeliness of traditional recognition methods can be solved, ensuring that the travel trajectory information of the target user can be fully captured.
[0026] In one possible implementation, the preset time can be adjusted according to the actual application scenario. If it is used for daily traffic analysis, it can be set to 30 days; if it is used for short-term travel feature identification, it can be set to 7 days, so as to ensure that the data can cover a sufficient number of trips without data redundancy due to the excessively long period.
[0027] It should also be noted that, in this embodiment, the acquisition of communication signaling data must be based on legality and compliance, collected through authorized interfaces of communication operators, and only necessary information related to travel trajectories is extracted, without involving user privacy data, thus ensuring the security and compliance of data use. In one possible implementation, the system will de-identify the collected communication signaling data, replacing the user identifier with an anonymous code to avoid directly associating it with the user's real identity information.
[0028] Specifically, the system receives the target user's identification request and obtains the user's unique authorized identifier (such as an anonymized mobile phone number code). Then, the system submits the target user identifier and a preset time range through a legitimate data interface connected to the telecommunications operator, requesting to obtain the raw communication signaling data within that time period. Next, the system preprocesses the obtained raw communication signaling data, removing duplicate data based on timestamps and location information, and then eliminates invalid data caused by frequent base station handovers using a preset ping-pong effect elimination method (such as time interval method, spatial distance method, or comprehensive judgment method). Finally, the preprocessed standardized data is determined as the target user's communication signaling data within the preset time period.
[0029] Step S12: Based on the communication signaling data, construct several single-trip trajectory chains for the target user; It should be noted that a single trip trajectory chain refers to the complete road path record of a target user from the starting point of a trip to the end point, including information such as road segments, travel time, and travel speed.
[0030] Specifically, a spatiotemporal density clustering algorithm is used to aggregate spatiotemporally adjacent trajectory points in the communication signaling data into user dwell points, generating a sequence of user dwell points containing dwell point locations and dwell durations. Then, based on the dwell point locations and dwell durations, the time interval and spatial distance between two adjacent user dwell points are calculated. If the time interval is greater than a preset time threshold and the spatial distance is greater than a preset spatial threshold, the two user dwell points are determined as a pair of travel origin-destination points, thereby obtaining the urban road network. A trajectory matching method combining multiple strategies is used to match the trajectory points between each pair of travel origin-destination points to the urban road network, integrating them to form the single travel trajectory chains.
[0031] Step S13: Identify the single travel mode corresponding to each of the single travel trajectory chains; It should be noted that the mode of transportation for a single trip refers to the mode of transportation used by the target user to complete a single trip, including subway travel, low-speed travel, public transportation travel, private car travel, etc., without any restrictions. For details, please refer to steps S41-S63, which will not be elaborated here.
[0032] Step S14: Based on each of the single travel modes, generate the private car travel behavior recognition result corresponding to the target user.
[0033] It should be noted that the private car travel behavior recognition result refers to the final conclusion of whether the target user is a private car owner, as well as relevant auxiliary analysis data (such as private car travel frequency, average travel distance, etc.). The purpose of this step is to accurately determine whether the target user is a private car owner through statistical analysis and interference elimination, while outputting valuable auxiliary data to provide decision support for scenarios such as traffic planning, resource allocation, and commercial services, and to solve the problems of insufficient accuracy and lack of auxiliary analysis data in traditional recognition methods.
[0034] Specifically, the number of private car trajectories in all single-trip trajectory chains within the preset time period that involve private car travel is counted, and the proportion of the number of private car trajectories to the total valid travel trajectories is calculated. If the proportion exceeds a preset threshold, the target user is initially determined to be a potential private car owner. Based on the daily mileage, daily driving time, and trajectory detour index in the historical single-trip trajectory chains, a ride-hailing driver database is constructed. The target user initially determined to be a potential private car owner is then compared with the user identifiers and feature data in the ride-hailing driver database to generate a private car travel behavior identification result corresponding to the target user.
[0035] This embodiment acquires communication signaling data of a target user within a preset time period, and then constructs several single-trip trajectory chains for the target user based on the communication signaling data. This identifies the single-trip mode corresponding to each single-trip trajectory chain, and then generates a private car travel behavior recognition result for the target user based on each single-trip mode. Therefore, recognition is performed based on communication signaling data, without the need for additional hardware deployment, and can cover the target user's travel trajectory across all scenarios within a preset time period, avoiding the limited coverage and high costs of traditional survey methods. Furthermore, constructing single-trip trajectory chains based on communication signaling data can reconstruct the user's actual travel path, reducing recognition errors caused by trajectory ambiguity.
[0036] In one feasible implementation, the step of acquiring the target user's communication signaling data within a preset time period further includes: Step S21: Obtain the raw communication signaling data of the target user within a preset time period; It should be noted that raw communication signaling data refers to the unprocessed location and communication records generated by the target user during interaction with the communication network via a mobile terminal within a preset time period. This data includes core fields such as user identifier, timestamp, base station identifier, base station location, and signal strength. By collecting widely covered, real-time, and dynamic raw communication signaling data, the limitations of limited coverage and poor data timeliness in traditional traffic survey methods can be addressed. In one possible implementation, the preset time period must ensure that the data collection period covers different travel scenarios, such as weekdays and weekends.
[0037] Specifically, the system receives a request for identifying private car travel behavior based on big data from communications, and obtains anonymized target user identifiers authorized by the user (such as encrypted mobile phone number codes) to ensure the compliance and privacy security of data collection. Then, the system submits the target user identifier and a preset time range through a standardized data interface that interfaces with telecommunications operators, and initiates a request for collecting raw communication signaling data. The interface uses an encrypted transmission protocol to ensure data transmission security. Next, the telecommunications operator filters all signaling records of the target user within the corresponding time range according to the request, including the user's base station access information and location update records at different time points, forming a raw communication signaling dataset. Finally, the system receives the raw communication signaling dataset returned by the operator, verifies the data integrity, and confirms whether it contains continuous signaling records within the preset time period. If data is missing, a supplementary collection request is initiated to the operator. After the verification is passed, the data is stored in the local raw database for subsequent preprocessing.
[0038] Step S22: Based on the timestamp and location information in the original communication signaling data, duplicate data in the original communication signaling data is removed to obtain communication signaling processing data; It should be noted that the timestamp refers to the specific time of signaling generation recorded in the original communication signaling data, accurate to the second or millisecond level, used to identify the time attribute of the signaling; location information refers to the corresponding base station location information in the signaling record, usually represented by latitude and longitude coordinates, used to identify the geographical location of the user when the signaling was generated; duplicate data refers to redundant signaling records in the original communication signaling data where both the timestamp and location information meet preset similarity conditions. This type of data is generated by repeated transmission and signal retransmission in the communication network, and has no practical analytical value. Therefore, it is necessary to remove redundant information in the dataset to reduce the interference of invalid data on subsequent processing, reduce the consumption of computing resources, and at the same time ensure the uniqueness and accuracy of the data, so as to provide a clean data source for subsequent trajectory chain construction.
[0039] In one possible implementation, the preset timestamp similarity condition can be set to a time difference of less than 30 seconds, and the location information similarity condition can be set to a latitude and longitude coordinate deviation of less than 50 meters. If either condition is met, the data is determined to be duplicated.
[0040] Specifically, first, the system reads the raw communication signaling data from the original database, sorts it in ascending order by user identifier and timestamp to ensure that the data is processed in chronological order. Then, the system iterates through each signaling record, extracts the timestamp and location information (latitude and longitude coordinates) of the current record, and compares it with the previous retained signaling record. Next, it determines whether the difference in timestamps between the two records is less than a preset time threshold and whether the latitude and longitude deviation of the location information is less than a preset spatial threshold. If both conditions are met, the current record is determined to be duplicate data and is removed. If neither condition is met, the current record is retained and used as the comparison benchmark for the next record. Finally, after the iteration is completed, the communication signaling processing dataset after removing duplicate data is obtained, stored in the preprocessing database, and a duplicate data removal report is generated, recording the number of removed data and the reasons for removal, for users to view.
[0041] Step S23: According to the preset ping-pong effect elimination method, eliminate the ping-pong effect data in the communication signaling processing data to obtain the communication signaling data.
[0042] It should be noted that ping-pong effect data refers to high-frequency handover records in communication signaling data that have no actual movement significance due to users frequently switching between adjacent base stations within a short period of time because of overlapping coverage areas. This type of data can interfere with the judgment of the user's actual movement trajectory. The preset ping-pong effect elimination method refers to pre-set rules or algorithms used to identify and remove ping-pong effect data, including time interval methods, spatial distance methods, and comprehensive judgment methods, among which: 1. Time Interval Method: Set a minimum dwell time. If a user's dwell time at a base station is less than this time, it is considered a ping-pong effect. 2. Spatial Distance Method: Set a minimum movement distance. If a user moves a distance less than this distance in a short period, it is considered a ping-pong effect. 3. Comprehensive Judgment Method: Combine time and spatial information for a comprehensive judgment.
[0043] This eliminates interference from base station handover, restores the user's true movement trajectory, ensures the accuracy of subsequent trajectory chain construction and travel mode identification, and solves the trajectory distortion problem caused by the ping-pong effect.
[0044] In one possible implementation, the preset ping-pong effect elimination method can adopt a comprehensive judgment method, combining a time interval threshold (e.g., 60 seconds) and a spatial distance threshold (e.g., 100 meters) to double-verify whether the signaling record is ping-pong effect data.
[0045] Furthermore, the ping-pong effect manifests differently depending on the base station coverage density in different scenarios. Therefore, the preset ping-pong effect elimination method needs to support dynamic parameter adjustment to adapt to the communication network environment of different cities and regions. In one possible implementation, the system automatically adjusts the time interval threshold and spatial distance threshold based on the base station density of the target user's area. The threshold can be appropriately reduced in areas with high base station density (such as urban core areas) and appropriately increased in areas with low base station density (such as suburbs), thereby improving the adaptability of ping-pong effect elimination.
[0046] Specifically, the system reads the communication signaling processing dataset after removing duplicate data from the preprocessing database, sorts it in ascending order by timestamp, and extracts the timestamp, base station identifier, and corresponding location information for each record. Then, the system uses a preset ping-pong effect elimination method (such as a comprehensive judgment method), sets time interval thresholds and spatial distance thresholds, and traverses each signaling record, comparing the current record with multiple adjacent signaling records. Next, it determines whether the time interval between the current record and adjacent records is less than a preset time threshold, and whether the spatial distance between the corresponding base station locations is less than a preset spatial distance threshold. If these conditions are met, the current record is determined to be ping-pong effect data, marked, and temporarily stored. After traversal, the system counts the consecutively marked ping-pong effect data segments. If the number of signaling records for a segment exceeds a preset number (such as 5), all data in that segment is removed to avoid accidentally deleting real movement trajectory data. Finally, the dataset after removing ping-pong effect data is determined as communication signaling data and stored in a standardized database to provide high-quality data support for subsequent trajectory chain construction.
[0047] This embodiment reduces invalid interference information by eliminating duplicate and ping-pong effect data, ensuring that communication signaling data accurately reflects the actual movement trajectory of the target user, improving data quality, reducing the amount of processing redundant and invalid data, avoiding wasting computing power on invalid data in subsequent algorithms, improving the processing efficiency of trajectory construction, travel mode identification, and other stages, shortening the overall identification cycle, and reducing computational costs. Among these, duplicate data can easily lead to redundant trajectory points, and ping-pong effect data can interfere with the judgment of the actual movement path. Eliminating both can avoid the misjudgment of the origin and destination of the trip and trajectory matching deviation caused by these problems, thereby improving the accuracy of private car travel behavior recognition based on communication big data.
[0048] In one feasible implementation, constructing a plurality of single-trip trajectory chains for the target user based on the communication signaling data includes: Step S31: Use a spatiotemporal density clustering algorithm to aggregate the spatiotemporally adjacent trajectory points in the communication signaling data into user dwell points, and generate a user dwell point sequence containing dwell point location and dwell time. It should be noted that the spatiotemporal density clustering algorithm refers to a clustering algorithm that considers both spatial location and timestamp dimensions. It can aggregate discrete trajectory points that meet preset spatiotemporal distance conditions into a meaningful set. A trajectory point refers to the user's geographical location point corresponding to each record in the communication signaling data, which is determined by the base station location or positioning information. A user dwell point refers to the point where a user stays at a certain geographical location for a certain period of time, and it is the core element constituting the origin and destination of a trip. The dwell point location refers to the geographical coordinates (usually latitude and longitude) of the user dwell point. The dwell time refers to the time the user stays at the dwell point. The user dwell point sequence refers to the set of all user dwell points arranged in chronological order, reflecting the distribution of user stays within a preset time period.
[0049] The purpose of this step is to extract the user's actual location from discrete communication signaling trajectory points, solving the problem of difficulty in determining the origin and destination of travel caused by trajectory point fragmentation, and ensuring the accuracy of subsequent trajectory chain construction. In one possible implementation, the parameters of the spatiotemporal density clustering algorithm can be dynamically adjusted, with a spatial radius threshold set to 50-100 meters and a time threshold set to 10-15 minutes, to adapt to the differences in base station coverage density in different areas (such as urban core areas and suburbs).
[0050] Specifically, first, the system reads communication signaling data from a standardized database, extracts the timestamp and trajectory point location (latitude and longitude) for each data point, and sorts them in ascending order by timestamp. Then, it initializes the parameters of the spatiotemporal density clustering algorithm, including the spatial radius threshold (e.g., 80 meters) and time threshold (e.g., 12 minutes), and sets the minimum number of trajectory points for clustering (e.g., 3). Next, the system traverses all trajectory points, using the current trajectory point as the core, and searches for other trajectory points within the spatial radius threshold and whose timestamps are within the time threshold. If the minimum number of trajectory points is met, a cluster is formed. Then, the system calculates the center position of each cluster (the average of the latitude and longitude of all trajectory points) as the dwell point position, and calculates the time difference between the first and last trajectory points in the cluster as the dwell duration. Finally, it arranges the dwell points corresponding to all clusters in timestamp order, generates a user dwell point sequence containing dwell point location, dwell duration, and dwell start time, and stores it in the trajectory analysis database.
[0051] Step S32: Based on the location of the dwelling point and the dwelling duration, calculate the time interval and spatial distance between two adjacent user dwelling points. If the time interval is greater than a preset time threshold and the spatial distance is greater than a preset spatial threshold, then determine the two user dwelling points as a pair of travel origin and destination points. It should be noted that two adjacent user dwell points refer to two dwell points that are sequentially adjacent in the user dwell point sequence; the time interval refers to the difference between the end time of the previous dwell point and the start time of the next dwell point; the spatial distance refers to the straight-line distance between the locations (latitude and longitude coordinates) of two dwell points; the preset time threshold refers to the minimum time interval set in advance to determine whether it constitutes an independent trip, which must be greater than the user's usual short stay duration; the preset spatial threshold refers to the minimum spatial distance set in advance to determine whether it constitutes an independent trip, which must be greater than the spatial radius threshold of dwell point clustering; the trip origin-end point pair refers to the combination of the origin dwell point and the destination dwell point of an independent trip, which is the core foundation for constructing a single trip trajectory chain.
[0052] The purpose of this step is to accurately classify a user's independent trips using both time and space criteria, avoiding misclassifying temporary stops within the same trip as multiple trips or merging two independent trips into one, and ensuring that subsequent trajectory chains correspond one-to-one with actual travel behavior. In one possible implementation, the preset time threshold can be set to 30 minutes, and the preset space threshold can be set to 500 meters. The thresholds can also be dynamically adjusted based on user travel characteristics (such as commuters or leisure users) to improve classification accuracy.
[0053] Specifically, the system reads the user's stop point sequence from the trajectory analysis database, extracts the stop location, stop start time, and stop duration of two adjacent stop points in chronological order, and then calculates the time interval: subtracting the sum of the stop start time and stop duration of the previous stop point from the stop start time of the latter stop point. Next, the system uses the Haversine formula to calculate the straight-line distance (latitude and longitude) between the two stop points to obtain the spatial distance. Then, the calculated time interval is compared with a preset time threshold (e.g., 30 minutes), and the spatial distance is compared with a preset spatial threshold (e.g., 500 meters). If both the time interval and the spatial distance are greater than the preset time threshold and the spatial distance is greater than the preset spatial threshold, it is determined that the two stop points correspond to an independent trip, and the former stop point is determined as the trip start point and the latter stop point is determined as the trip end point, forming a trip start-end point pair. Finally, all adjacent stop points are traversed to generate all trip start-end point pairs that meet the conditions, and each trip start-end point pair is associated with a corresponding time range (from the end time of the start point stop to the start time of the end point stop), which is stored in the trajectory analysis database.
[0054] Step S33: Obtain the urban road network and use a multi-strategy trajectory matching method to match the trajectory points between each pair of travel origin and destination to the urban road network, and integrate them to form each single travel trajectory chain.
[0055] It should be noted that the urban road network refers to a dataset containing information such as the location, topological relationship, road grade, and number of lanes of all road segments within the city, which serves as the geographical reference basis for trajectory point matching; the multi-strategy trajectory matching method refers to a composite matching strategy that integrates projection method, nearest neighbor method, and road network topology constraints, taking into account both matching speed and accuracy; the trajectory points between the origin and destination pairs refer to all trajectory points in the communication signaling data that are located within the time range corresponding to a certain pair of origin and destination pairs, which are the original data constituting the travel path; a single travel trajectory chain refers to the continuous trajectory record formed after matching the trajectory points between the origin and destination pairs to the urban road network, which includes the complete travel path, travel time of each road segment, and travel speed.
[0056] The purpose of this step is to transform discrete communication signaling trajectory points into continuous trajectories that closely match actual roads, thus resolving the problem of ambiguous travel paths caused by the disconnect between trajectory points and roads. In one possible implementation, urban road network data can be obtained through a third-party geographic information service interface and updated regularly to ensure the timeliness of road information. In multi-strategy matching, a projection method can be used for rapid preliminary matching first, and then the nearest neighbor method and road network topology constraints can be used to correct deviations and improve matching accuracy.
[0057] It should also be noted that different cities have different road network densities and road topology complexities. The trajectory matching method that combines multiple strategies supports dynamic adjustment of parameters. For example, in densely roaded areas in the city center, the projection distance threshold can be reduced to improve the accuracy of matching; in sparsely roaded areas in the suburbs, the threshold can be appropriately increased to avoid matching failure due to sparse trajectory points.
[0058] Specifically, the system obtains road network data of the target city through an authorized geographic information interface, including the latitude and longitude coordinate sequence of road segments, road topology, road grade, and other information, and constructs a local road network index. Then, for each pair of origin and destination points, it extracts all trajectory points within the corresponding time range from the communication signaling data and sorts them in ascending order by timestamp. Next, it processes the group of trajectory points using a trajectory matching method that combines multiple strategies: The first step involves projecting each trajectory point onto the nearest road segment using a projection method. If the projection distance is less than a preset projection threshold (e.g., 50 meters), a preliminary match is made to that road segment. The second step involves using the nearest neighbor method to find the nearest road segment for trajectory points whose projection distance exceeds the threshold, and correcting the matching results by considering road network topology relationships (e.g., the connection relationships between adjacent road segments). The third step integrates all matched road segments in chronological order, eliminating discontinuous road segments caused by matching errors to ensure the path conforms to the road network topology. Then, the travel time (time difference between adjacent trajectory points on that road segment) and travel speed (the ratio of road segment length to travel time) for each road segment are calculated. Finally, the matched road segment sequence, travel time and speed for each segment, and origin / destination information are integrated to form a single-trip trajectory chain. Each "origin / destination pair" corresponds to one single-trip trajectory chain, which is stored in the trajectory analysis database for reference. Figure 2 .
[0059] This embodiment aggregates spatiotemporally adjacent trajectory points through a spatiotemporal density clustering algorithm, which can accurately extract the user's actual location and duration of stay, avoiding misjudging short stays as fixed locations. Combined with urban road networks and multi-strategy trajectory matching methods, it accurately matches discrete trajectory points to real roads, restores the user's actual driving path, and solves the trajectory ambiguity problem caused by discrete communication signaling data.
[0060] In one feasible implementation, identifying the single travel mode corresponding to each of the single travel trajectory chains includes: Step S41: Obtain the preset underground parking lot indoor distributed base station database and the Internet access data in the communication signaling data; It should be noted that the pre-built underground parking lot indoor distributed base station database refers to a pre-constructed and stored database containing information such as the identifiers, locations, and coverage areas of indoor distributed base stations corresponding to all underground parking lots in the city. This base station database is generated by compiling network planning data from telecommunications operators and is regularly updated to ensure information timeliness. Internet access data in the communication signaling data refers to network access records generated by the target user using the mobile network during their travels, including core fields such as server network identification information, access timestamps, and data transmission volume. Server network identification information refers to the server-related identifiers recorded in the internet access data, including IP addresses, port numbers, and Uniform Resource Locators (URLs). In one possible implementation, the underground parking lot indoor distributed base station database can be stored by city administrative region, supporting rapid retrieval by geographical location. Internet access data can be filtered by timestamp, extracting only data within the travel time period corresponding to a single travel trajectory chain, reducing the amount of invalid data processed.
[0061] Specifically, the system accesses a pre-defined database of indoor distributed base stations in underground parking lots via a dedicated interface that interfaces with telecommunications operators. This interface employs an encrypted transmission protocol to ensure data security. Upon receiving the data, the system indexes it by base station identifier for easy and rapid subsequent retrieval. Next, the system reads the target user's communication signaling data from a standardized database and filters internet access records within the specified travel time period (from the start time to the end time) corresponding to a single travel trajectory chain. Then, the filtered internet access data undergoes format standardization, extracting key fields such as server network identifier information and access timestamps from each record, while removing invalid fields and redundant data. Finally, the processed internet access data is associated with and stored along with the corresponding single travel trajectory chain, while the database of indoor distributed base stations in the underground parking lot is cached locally for rapid retrieval in subsequent decision-making steps.
[0062] Step S42: Based on the preset underground parking lot indoor distributed base station database, determine whether there are records of the single trip trajectory chain residing in the underground parking lot indoor distributed base station within the set time period before the start and the set time period before the end, and obtain the first judgment result; It should be noted that the "pre-start time period" refers to the preset time length before the start time (origin and end time of stay) of a single trip trajectory chain, used to detect whether the user departs from the underground parking lot before traveling; the "pre-end time period" refers to the preset time length before the end time (end and start time of stay) of a single trip trajectory chain, used to detect whether the user enters the underground parking lot after traveling; the record of staying at the underground parking lot indoor distributed base station refers to the signaling record in the communication signaling data of the user accessing the underground parking lot indoor distributed base station within the set time period; the first judgment result refers to the combined judgment result including the two sub-results "staying at the underground parking lot before travel" and "staying at the underground parking lot after travel", each sub-result being "yes" or "no". The purpose of this step is to explore one of the core characteristics of private car travel - the behavior of entering and exiting underground parking lots, and to use the exclusive coverage attribute of the underground parking lot indoor distributed base station to distinguish between private cars and ride-hailing vehicles (ride-hailing vehicles rarely enter and exit underground parking lots), providing a key basis for determining the confidence of private car travel.
[0063] In one possible implementation, the time period set before the start and the time period set before the end can be uniformly set to 10 minutes. This ensures that the behavior of entering and exiting the underground parking lot can be captured, while avoiding confusion with other travel behaviors due to excessively long time periods. It also supports dynamic adjustment according to the actual scenario.
[0064] Specifically, firstly, the system determines the start and end times of a single trip's trajectory chain, and calculates the time ranges for a pre-trip time period (e.g., 10 minutes) (trip start time minus the pre-trip time period) and the end time period (e.g., 10 minutes) (trip end time minus the pre-trip time period). Then, from the target user's communication signaling data, signaling records within these two time ranges are selected, and the base station identifier is extracted from each record. Next, the selected base station identifiers are compared one by one with a pre-defined database of indoor distributed base stations in underground parking lots. If a base station identifier belonging to this database exists in the signaling records within the pre-trip time period, the sub-result "staying in the underground parking lot before trip" is determined to be "yes"; otherwise, it is "no". Similarly, the base station identifiers within the pre-trip time period are compared to determine the sub-result "staying in the underground parking lot after trip" to be "yes" or "no". Finally, the two sub-results are combined to form the first judgment result, which is then associated with and stored with the corresponding single trip's trajectory chain.
[0065] Step S43: Use deep packet inspection technology to parse the server network identification information in the Internet access data, determine whether there is a record of accessing navigation application services during the trip period, and obtain a second judgment result; It should be noted that Deep Packet Inspection (DPI) is a technology that can deeply analyze application-layer information in network data packets, extracting server-side network identification information and matching it with corresponding application service types. Internet access data refers to network access records within the duration of a single trip's trajectory chain. Server-side network identification information refers to information such as IP address, port number, and Uniform Resource Locator (URL) included in the internet access data to identify the server. The trip duration refers to the time from the start of the trip to the start of the destination within the trajectory chain. Navigation application services refer to network services provided by mobile applications that offer route planning and real-time navigation, such as Gaode Maps and Baidu Maps. The second judgment result refers to the result indicating whether there are records of accessing navigation application services within the trip duration, categorized as "yes" or "no." The purpose of this step is to capture another core characteristic of private car travel—navigation application usage behavior—and to accurately identify navigation service access records using DPI, further distinguishing between private cars and ride-hailing vehicles (where passengers use navigation applications less frequently), thus providing crucial evidence for determining the confidence level of private car travel. In one possible implementation, the system can pre-build a feature library for navigation application services, which includes information such as the official domain names, associated IP address ranges, and service ports of mainstream navigation applications. The feature library can be updated regularly to adapt to changes in application services.
[0066] Specifically, first, the system calls a pre-built feature library for navigation applications, which contains feature information such as the official domain names, associated IP address ranges, and commonly used service ports of mainstream navigation applications. Then, the system performs packet-by-packet parsing of the internet access data corresponding to a single trip's trajectory chain, extracting server network identification information such as the server IP address, port number, and Uniform Resource Locator (URL) for each piece of internet access data using deep packet inspection technology. Next, the extracted server network identification information is matched against the navigation application service feature library in multiple dimensions: if the URL contains the official domain name of the navigation application, or the server IP address belongs to the associated IP range of the navigation application, or the port number is a commonly used service port of the navigation application, then it is determined that the internet access data corresponds to accessing a navigation application service. Finally, if at least one record in the internet access data within the trip period meets the above matching conditions, the second judgment result is "yes"; otherwise, it is "no," and the result is associated with and stored with the corresponding single trip's trajectory chain.
[0067] Step S44: Based on the first judgment result and the second judgment result, determine the confidence level of private car travel. If the confidence level of private car travel is medium or high, then determine that the single travel mode is private car travel.
[0068] It should be noted that the first judgment result refers to the combined result of the two sub-results: "staying in the underground parking lot before the trip" and "staying in the underground parking lot after the trip"; the second judgment result refers to whether navigation application services were accessed during the trip; the confidence level of private car travel refers to the degree of credibility in determining whether a single trip belongs to a private car based on private car-specific characteristics (entry and exit from the underground parking lot, use of navigation applications), and is divided into three levels: high confidence, medium confidence, and low confidence; private car travel refers to a trip completed by the target user driving a private car. The core purpose of this step is to accurately distinguish private car travel from other modes of travel (especially ride-hailing travel) by determining confidence through the combination of multiple features, thus solving the industry problem of difficulty in distinguishing between the two due to similar single trajectories. In one possible implementation, the confidence determination rule can be flexibly configured. If it is necessary to improve the recognition accuracy, the feature requirements for high confidence can be adjusted; if it is necessary to expand the recognition coverage, the feature threshold for medium confidence can be appropriately lowered to adapt to the needs of different application scenarios.
[0069] Specifically, firstly, the system clarifies the rules for determining the confidence level of private car travel: if both sub-results in the first judgment result are "yes," or one sub-result is "yes" and the second judgment result is "yes," then it is determined to be of high confidence; if any sub-result in the first judgment result is "yes," or the second judgment result is "yes," then it is determined to be of medium confidence; if both sub-results in the first judgment result are "no" and the second judgment result is "no," then it is determined to be of low confidence. Next, the system extracts the first and second judgment results corresponding to a single travel trajectory chain, matches them according to the above rules, and determines the corresponding confidence level of private car travel. Then, it determines the confidence level: if it is high or medium confidence, then the single travel mode is determined to be private car travel; if it is low confidence, then the single travel mode is determined to be unknown and excluded from subsequent private car owner statistics. Finally, the travel mode determination result is associated with and stored with the corresponding single travel trajectory chain to provide data support for subsequent steps.
[0070] This embodiment focuses on the unique behavioral characteristics of private cars (entry and exit from underground parking lots, navigation usage) to solve the core problem of difficulty in distinguishing between the two based on similar single trajectories, thereby improving the accuracy of travel mode identification. Relying on the underground parking lot base station database and deep packet inspection technology, it transforms abstract behavioral characteristics into objective data that can be directly judged, without the need for complex manual intervention, and is adapted to automated identification processes.
[0071] In one feasible implementation, identifying the single travel mode corresponding to each of the single travel trajectory chains includes: Step S51: Obtain the preset metro private network base station database; It should be noted that the pre-built subway dedicated network base station database refers to a pre-constructed and stored database containing information on dedicated network base stations covering all subway lines and stations within the city. Core information includes base station identifiers, location coordinates, coverage areas, and their association with the corresponding subway lines and stations. This base station database is generated by telecommunications operators in conjunction with subway network planning data and is regularly updated based on the opening of new subway lines and base station equipment upgrades to ensure the timeliness and completeness of the information. The purpose of this step is to provide a dedicated and accurate reference for subway travel identification. Leveraging the dedicated coverage characteristics of subway dedicated network base stations, it quickly distinguishes subway travel from other ground travel modes, laying the foundation for subsequent tiered screening of non-private car travel modes. In one possible implementation, the subway dedicated network base station database can be stored in partitions according to subway lines, and an association index can be established between base station identifiers and subway lines and stations. This supports rapid matching of base station information within a given area based on trajectory point location, improving comparison efficiency.
[0072] Specifically, the system initiates a request to access the metro private network base station database through an encrypted data interface connected to the telecommunications operator. The interface transmission process employs a secure encryption protocol to prevent data leakage. Then, the operator returns complete metro private network base station database data based on the request. Upon receiving the data, the system performs format verification to confirm that core fields such as base station identifiers and location coordinates are complete and without errors. Next, the system categorizes and organizes the base station database data according to metro lines, establishing a fast retrieval index for base station identifiers. Simultaneously, it converts the base station location coordinates into a unified latitude and longitude format for easy association with trajectory point locations. Finally, the processed metro private network base station database is cached in the local data module for rapid access in subsequent trajectory point base station identifier comparison steps, reducing the time consumed by repetitive data transmission.
[0073] Step S52: Compare the base station identifiers corresponding to the trajectory points of the single trip trajectory chain with the metro private network base station database. If at least one trajectory point's base station identifier belongs to the metro private network base station database, then determine that the single trip mode is metro travel. It should be noted that the base station identifier corresponding to the trajectory point refers to the unique identifier of the communication base station accessed by the user's mobile terminal when each trajectory point is generated in a single travel trajectory chain; subway travel refers to the travel mode completed by the target user by taking the subway. The purpose of this step is to utilize the exclusive coverage attributes of subway dedicated network base stations to quickly and accurately identify subway travel as the first priority in the travel mode stratification determination, excluding travel modes such as subway travel that have significantly different characteristics from private car travel, thereby reducing invalid calculations in subsequent identification processes. In one possible implementation, the comparison process can adopt a "batch retrieval + precise matching" approach, first filtering the base station database data within the corresponding time range according to the travel time of the trajectory point, and then performing identifier comparison to further improve the determination efficiency.
[0074] Specifically, firstly, the system extracts the single-trip trajectory chain to be determined from the trajectory analysis database, obtaining the base station identifiers corresponding to all trajectory points included in the trajectory chain, as well as the timestamps of each trajectory point. Then, the system extracts base station data that intersects with the travel time period of the trajectory chain from the locally cached metro dedicated network base station database (if the trajectory chain's travel time period is 07:30-08:15, then base stations in the base station database without time restrictions or covering this time period are filtered). Next, the system compares all base station identifiers of the trajectory chain with the base station identifiers in the filtered metro dedicated network base station database one by one, or searches for matching items through batch retrieval. If at least one trajectory point's base station identifier is completely identical to a base station identifier in the metro dedicated network base station database, then the single-trip mode is directly determined to be metro travel. If no matching item is found in the metro dedicated network base station database for any of the trajectory point's base station identifiers, the determination result is not metro travel. Finally, the determination result is associated with and stored in relation to the corresponding single-trip trajectory chain.
[0075] Step S53: If the determination result is not subway travel, calculate the average travel speed of the single travel trajectory chain and the proportion of travel time in the single travel trajectory chain that reaches a preset high speed threshold. It should be noted that average driving speed refers to the ratio of the total driving distance to the total driving time of a single trip trajectory chain, reflecting the overall speed level of the trip; total driving distance refers to the sum of the driving distances of all adjacent trajectory points in a single trip trajectory chain after being matched to the urban road network; total driving time refers to the duration of the trip in a single trip trajectory chain (i.e., the difference between the end time of the starting point and the start time of the ending point); the preset high-speed threshold refers to a pre-set speed threshold used to distinguish between high-speed and low-speed driving, which is set in conjunction with the maximum driving speed of common non-motorized vehicles (walking, bicycles, motorcycles); the percentage of driving time at speeds exceeding the preset high-speed threshold refers to the ratio of the sum of the driving times of road segments in a single trip trajectory chain where the driving speed exceeds the preset high-speed threshold to the total driving time, reflecting the proportion of high-speed driving in the trip. The purpose of this step is to provide objective data support for subsequent low-speed travel determination by quantifying speed-related indicators, and to further exclude non-private car travel modes by utilizing the significant differences in speed characteristics between low-speed travel and motorized travel. In one possible implementation, the preset high-speed threshold can be set to 40 km / h, which can effectively distinguish between non-motorized vehicles and motorized vehicles, adapt to the normal driving speed scenario of urban roads, and support dynamic adjustment according to the traffic conditions of different cities.
[0076] Specifically, the system extracts single-trip trajectory chains from the trajectory analysis database that are not determined to be subway trips. It then obtains information on each road segment after matching the trajectory chain to the urban road network, including the travel distance, travel time, and average speed for each segment. Next, it calculates the total travel distance by summing the travel distances of all road segments. It also calculates the total travel time by extracting the duration of the travel period for the trajectory chain or by summing the travel times of each road segment. Then, it calculates the average speed by dividing the total travel distance by the total travel time (in kilometers per hour). Afterward, it sets a preset high-speed threshold (e.g., 40 km / h) and filters out road segments with speeds exceeding this threshold. The system calculates the sum of the travel times for these road segments as the speed-compliant travel time. Finally, it calculates the percentage of travel time exceeding the speed-compliant travel time by dividing the total travel time, obtaining the percentage of travel time exceeding the preset high-speed threshold (rounded to two decimal places). The average speed and percentage results are then associated and stored with the corresponding single-trip trajectory chain.
[0077] Step S54: If the average driving speed is lower than a preset speed threshold and the percentage of driving time is lower than a preset percentage threshold, then the single trip mode is determined to be low-speed travel.
[0078] It should be noted that the preset speed threshold refers to a pre-set critical value for the average driving speed used to determine low-speed travel. This value is set based on the average driving speed range of common low-speed travel modes (walking, cycling, and motorcycles). The preset percentage threshold refers to a pre-set critical value for the percentage of time spent traveling at the required speed to determine low-speed travel, used to assist in verifying whether the travel mode is low-speed. Low-speed travel refers to travel completed by the target user using non-motorized vehicles such as walking, cycling, and motorcycles. Its core characteristics are low speed and an extremely low percentage of high-speed travel. The purpose of this step is to further exclude non-private car travel modes based on speed characteristics. By using the dual thresholds of average speed and high-speed percentage, the accuracy of low-speed travel identification is improved, avoiding misclassification of low-speed motor vehicles as low-speed travel, and simultaneously narrowing the scope and focusing the target for subsequent private car travel identification. In one possible implementation, the preset speed threshold can be set to 40 km / h, and the preset percentage threshold can be set to 50%. Only when both conditions are met simultaneously is it determined to be low-speed travel, which ensures the accuracy of identification and can adapt to speed fluctuations in different low-speed travel scenarios.
[0079] Specifically, firstly, the system extracts the average driving speed, the percentage of driving time exceeding a preset high-speed threshold, preset speed thresholds, and preset percentage thresholds corresponding to a single trip trajectory chain from the trajectory analysis database. Then, it compares the average driving speed with the preset speed thresholds and the percentage of driving time with the preset percentage thresholds. If the average driving speed is lower than the preset speed threshold and the percentage of driving time is lower than the preset percentage threshold, the single trip is determined to be a low-speed trip. If the average driving speed is higher than or equal to the preset speed threshold, or the percentage of driving time is higher than or equal to the preset percentage threshold, or neither condition is met, the single trip is determined not to be a low-speed trip. Finally, the determination results are associated with and stored in relation to the corresponding single trip trajectory chain.
[0080] This embodiment first accurately identifies subway travel through the subway dedicated network base station database, and then determines low-speed travel based on speed characteristics. This efficiently eliminates two core non-private car travel modes, narrows the identification range of private car travel, and improves subsequent identification efficiency. It relies on easily implemented technical means such as base station identifier comparison, speed and proportion calculation, and does not require complex algorithms to quickly complete the determination of the two types of travel modes, reducing the computational cost of the overall identification process.
[0081] In one feasible implementation, identifying the single travel mode corresponding to each of the single travel trajectory chains includes: Step S61: Obtain the origin and destination pairs corresponding to the single trip trajectory chain. Under the conditions of limiting the number of transfers, detour index and bus stop distance threshold, perform bus route navigation based on bus network data to generate several candidate bus routes. It should be noted that the origin-destination pair refers to the combination of the starting and ending points of a single trip's trajectory chain, including the coordinates of the origin and destination locations and the travel time; the limited number of transfers refers to the pre-set maximum allowed number of transfers for a bus route (e.g., 0, 1, 2 times), used to filter bus routes that match the user's actual travel habits; the detour index is the ratio of the total length of the bus route to the straight-line distance between the origin and destination, limiting the detour index can avoid generating unreasonable routes with excessive detours; the bus stop distance threshold refers to the maximum allowed walking distance between the origin / destination and the bus stop (e.g., 500 meters), used to filter bus stops that the user can actually reach; the bus network data refers to a dataset containing information such as the stops, routes, operating hours, and station latitude and longitude of all bus routes in the city; and the candidate bus routes refer to possible bus route schemes from the origin to the destination that meet the above-mentioned constraints, including information such as route name, stops along the way, and travel trajectory.
[0082] The purpose of this step is to simulate the route selection logic of users' actual public transportation trips, generate public transportation route plans that may match the trajectory chain of a single trip, and ensure the targeting and accuracy of public transportation trip identification. In one possible implementation, the number of transfers can be limited to 0-2 (covering more than 90% of public transportation trip scenarios), the detour index can be limited to ≤2.5 (to avoid excessive route detours), and the bus stop distance threshold can be set to 500 meters (adapting to the distance of a typical walk to a bus stop), and can be adjusted according to the size of the city.
[0083] Specifically, the system extracts the origin-destination pairs corresponding to a single trip trajectory chain from the trajectory analysis database, obtaining the latitude and longitude coordinates of the origin and destination. Then, it calls the public transport network data to filter out all bus stops within 500 meters of the origin as candidate origin stations and all bus stops within 500 meters of the destination as candidate destination stations. Next, it sets constraints: the number of transfers ≤ 2 and the detour index ≤ 2.5. Based on the route direction and station association relationships in the public transport network data, it calls a path planning algorithm (such as Dijkstra's algorithm) to calculate all eligible bus routes from the candidate origin stations to the candidate destination stations. After that, the generated bus routes are deduplicated and optimized, eliminating duplicate routes and obviously unreasonable routes (such as routes whose total travel time far exceeds the trajectory chain length). Finally, the optimized routes are retained as candidate bus routes. Each route includes the latitude and longitude sequence of the stations it passes through (forming a route trajectory), the number of transfers, the detour index, and other information, which are stored in association with the corresponding single trip trajectory chain.
[0084] Step S62: Calculate the similarity between the trajectory of each candidate bus route and the trajectory of the single trip trajectory chain; It should be noted that the trajectory of a candidate bus route refers to a continuous path composed of the latitude and longitude coordinates of the stops along the route in the order of travel; the trajectory of a single trip trajectory chain refers to a continuous path after the trajectory chain is matched to the urban road network, composed of a sequence of road segments' latitude and longitude coordinates; trajectory similarity refers to the degree of overlap between two trajectories in spatial paths, with a value ranging from 0 to 1, where a value closer to 1 indicates a higher degree of overlap. The purpose of this step is to improve the accuracy of public transport travel identification by quantitatively analyzing the matching degree between candidate bus routes and actual travel trajectories, thus addressing the potential for misjudgment that may result from matching only the origin and destination points. In one possible implementation, trajectory similarity can be calculated using the Dynamic Time Warping (DTW) algorithm, which can effectively handle differences in time and spatial scales between two trajectories, and is particularly suitable for scenarios where the intervals between bus stops and signaling trajectory points are inconsistent.
[0085] Specifically, the system extracts trajectory data (latitude and longitude sequences of stops along the route) and trajectory data (latitude and longitude sequences matched to the road network) of candidate bus routes, unifying the coordinate format. Then, it preprocesses the two trajectories, sampling uniformly at time or spatial intervals to reduce data volume and ensure uniformity of comparison. Next, it uses the Dynamic Time Warping (DTW) algorithm to calculate the similarity between the two trajectories: by constructing a distance matrix, it finds the optimal matching path between the two trajectories, calculates the cumulative distance along the path, and then normalizes the cumulative distance to obtain a similarity value (0-1); or it uses a point-pair matching method: it calculates the shortest distance between each point in the trajectory chain and the bus route trajectory, and counts the proportion of points with a distance less than a preset threshold (e.g., 100 meters) out of the total number of points, using this as the similarity value. Finally, it records the trajectory similarity between each candidate bus route and the single-trip trajectory chain, and stores it in association with the corresponding candidate route.
[0086] Step S63: If the trajectory similarity of at least one candidate bus route exceeds a preset similarity threshold, then the single trip mode is determined to be public transportation.
[0087] It should be noted that the preset similarity threshold refers to a pre-set critical value used to determine the degree of trajectory matching, ranging from 0 to 1. A value exceeding this threshold indicates that the two trajectories are highly similar. Public transport travel refers to the travel mode completed by the target user using public buses (including BRT, regular buses, etc.). The purpose of this step is to ultimately determine whether a single trip is a public transport trip based on the quantified results of trajectory similarity. By setting a reasonable threshold, the accuracy and coverage of identification are balanced, avoiding misclassification of similar trajectories as public transport trips while ensuring that genuine public transport trips are not missed. In one possible implementation, the preset similarity threshold can be set to 0.6. This value has been verified through a large amount of sample data and can achieve a recognition accuracy of over 85% in most urban public transport scenarios. It also supports dynamic adjustment based on the overlap between bus routes and road networks (e.g., the threshold can be appropriately lowered for BRT dedicated lanes).
[0088] Specifically, the process of determining public transportation travel based on trajectory similarity is as follows: First, the system extracts the trajectory similarity of all candidate bus routes corresponding to a single trip trajectory chain from the associated stored information, as well as a preset similarity threshold (e.g., 0.6). Then, it iterates through the similarity values of all candidate bus routes to determine whether at least one route has a similarity exceeding the preset threshold. If so, the single trip is determined to be public transportation travel. If the similarity of all candidate bus routes does not exceed the preset threshold, the single trip is determined not to be public transportation travel. Finally, the determination result is associated with and stored with the corresponding single trip trajectory chain to provide a basis for subsequent comprehensive identification of travel modes.
[0089] In one embodiment, reference may be made to Figure 3 , Figure 3 The black trajectory represents the candidate bus route, and the red trajectory represents the user's actual trajectory. If the origin-destination (OD) of two trajectories is the same, a similarity score is needed to confirm whether the user traveled via that bus route. The trajectory similarity calculation methods mainly include: 1. Distance Similarity: Calculates the distance difference between two trajectories. 2. Angle Similarity: Calculates the angle difference between two trajectories. 3. Shape Similarity: Calculates the shape difference between two trajectories. 4. Comprehensive Similarity: Combines multiple similarity indicators for comprehensive calculation.
[0090] In practical applications, multiple similarity metrics can be combined to calculate trajectory similarity, thereby improving the accuracy of the calculation. For example, the DTW (Dynamic Time Warping) algorithm can be used to calculate the time series similarity of trajectories, and Frechet distance can be used to calculate the spatial similarity of trajectories. Then, the overall similarity of the trajectories can be judged. This will not be elaborated on in detail here.
[0091] Furthermore, a spatiotemporal weighted similarity calculation model for public transport routes can be used. The core of this model is to simultaneously consider both spatial path overlap and temporal synchronization, accurately quantifying the degree of matching between a user's actual travel trajectory and candidate public transport routes. This avoids misjudgments caused by traditional methods that only calculate spatial overlap (such as situations where private cars and buses travel the same route but travel at different times). The specific calculation formula is as follows:
[0092] in, The spatiotemporal weighted similarity is used (the value ranges from [0,1], and ≥0.7 is used to determine public transportation travel). The normalized position coordinates (x-axis) of the user trajectory at time t; Here are the normalized position coordinates (x-axis) of the candidate bus route at time t. , The average location coordinates of the user's trajectory and bus route; T: Time steps (dividing the trip into 1-minute intervals); The time weighting coefficient (highlights the importance of "matching locations at the same time point" to avoid misjudgments due to spatial overlap but temporal misalignment) This embodiment generates candidate bus routes by limiting the number of transfers and detour index, and combines trajectory similarity comparison to effectively distinguish between public transportation and other motor vehicle travel such as private cars, thereby improving the accuracy of travel mode classification. Furthermore, through multiple limiting conditions (transfer, detour, and distance to bus stops) + trajectory similarity double verification, it avoids misjudging other travel that accidentally overlaps with the bus trajectory as public transportation, reducing the bias of a single judgment dimension.
[0093] In one feasible implementation, generating the private car travel behavior recognition result corresponding to the target user based on each of the single travel modes includes: Step S71: Count the number of private car trajectories in all single travel trajectory chains within the preset time period, where the single travel mode is private car travel, and calculate the proportion of the number of private car trajectories to the total effective travel trajectories. It should be noted that the number of private car trajectories refers to the total number of trajectory chains for a single trip identified as a private car trip within a preset time period; the total number of valid trip trajectories refers to the total number of single trip trajectory chains for which the trip mode has been determined (not "unknown") within the preset time period; the percentage refers to the ratio of the number of private car trajectories to the total number of valid trip trajectories (rounded to two decimal places), used to quantify the degree to which users rely on private car travel. The purpose of this step is to avoid misjudgments due to the randomness of a single trip by statistically analyzing the percentage of users' private car travel within a certain period, ensuring that the judgment results are based on users' stable travel habits. In one possible implementation, the preset time period can be set to 30 days (covering a full month of travel behavior), and invalid data (such as trajectory chains with a trajectory point missing rate exceeding 30%) must be removed from the total number of valid trip trajectories to ensure the validity of the statistical base.
[0094] Specifically, the system extracts all single-trip trajectory chains of the target user within a preset time period (e.g., 30 days) from the trajectory analysis database, filters out the trajectory chains whose travel mode determination result is "private car travel", and counts the number as the number of private car trajectories; then, it filters out all trajectory chains whose travel mode determination result is not "unknown" within the preset time period (including private car, subway, bus, low-speed travel, etc.), and counts the number as the total number of valid travel trajectories; next, it calculates the percentage: the number of private car trajectories is divided by the total number of valid travel trajectories. If the total number of valid travel trajectories is 0, it is marked as "no valid data"; finally, the number of private car trajectories, the total number of valid travel trajectories, and the calculated percentage results are associated with the target user and stored in the user behavior analysis database.
[0095] Step S72: If the percentage exceeds a preset percentage threshold, the target user is preliminarily determined to be a potential private car owner. It should be noted that the preset percentage threshold refers to a pre-set critical value (e.g., 50%) used to determine whether a user relies on private cars for travel. Exceeding this value indicates that the user primarily uses private cars for travel. Potential private car owners refer to users who are initially judged to be likely to own and frequently use private cars based on the proportion of their travel modes, and further exclusion of interfering groups such as ride-hailing drivers is necessary. The purpose of this step is to filter out potential individuals with a high proportion of private car travel from the user group by setting a reasonable percentage threshold, narrowing the scope for subsequent accurate identification of private car owners and ensuring that the identification process focuses on core target users. In one possible implementation, the preset percentage threshold can be set to 50% (i.e., private car travel accounts for more than half). This value has been verified by a large number of samples and can effectively distinguish users who primarily use private cars for travel from other users. It also supports adjustments based on the city's traffic structure (e.g., the threshold can be appropriately increased in cities with high car ownership).
[0096] Specifically, first, the system extracts the percentage of private car trajectories of target users and a preset percentage threshold (e.g., 50%) from the user behavior analysis database; then, it compares the percentage of target users with the preset percentage threshold. If the percentage is greater than the preset percentage threshold, the user is preliminarily determined to be a potential private car owner; if the percentage is less than or equal to the preset percentage threshold, the user is determined to be a non-potential private car owner; finally, the preliminary determination result is associated with the target user and stored in the user behavior analysis database to provide a basis for subsequent steps to exclude ride-hailing drivers.
[0097] Step S73: Obtain and construct a ride-hailing driver database based on the daily mileage, daily driving time, and trajectory detour index in the historical single trip trajectory chain; It should be noted that the historical single-trip trajectory chain refers to the travel trajectory data of multiple ride-hailing users before a preset time (such as the past 3 months), used to mine long-term stable travel characteristics; daily mileage refers to the sum of the total travel distance of all user trips in a single day; daily travel time refers to the sum of the total travel time of all user trips in a single day; trajectory detour index refers to the ratio of the actual travel distance of a single trip trajectory to the straight-line distance from the origin to the destination, reflecting the degree of trajectory detour; the ride-hailing driver database refers to a database storing user identifiers and feature data with typical ride-hailing driver travel characteristics, used to distinguish between private car owners and ride-hailing drivers. The purpose of this step is to extract the unique behavioral characteristics of ride-hailing drivers (such as high mileage, high travel time, and high detour index) to build a feature database to exclude potential private car owners who are ride-hailing drivers, thus solving the identification confusion problem caused by the similarity of their travel modes.
[0098] Specifically, firstly, the system extracts a large number of historical single-trip trajectory chains of ride-hailing users from the historical trajectory database, calculates each user's daily mileage (total distance per day), daily driving time (total time per day), and detour index for each trajectory, and then calculates the average daily mileage, average daily driving time, and average trajectory detour index for each user. Next, it sets characteristic thresholds for ride-hailing drivers (e.g., average daily mileage > 150 km, average daily driving time > 6 hours, average detour index > 1.8; multiple feature combinations ensure the accuracy of users in the database, and the characteristic thresholds can be dynamically adjusted according to the operational characteristics of ride-hailing services in different cities), and filters out users who simultaneously meet all three thresholds. Then, it stores the selected user identifiers and their characteristic data (average daily mileage, average daily driving time, and average detour index) into the ride-hailing driver database and establishes a user identifier index to improve subsequent comparison efficiency. Finally, it updates the ride-hailing driver database regularly (e.g., monthly), adding new users who meet the characteristics and removing users who no longer meet the characteristics to ensure the timeliness of the database data.
[0099] Alternatively, in one possible implementation, ride-hailing drivers are selected using a dynamic threshold adjustment combined with a comprehensive feature scoring formula to adapt to different city scenarios. The specific formula is as follows:
[0100] in, Ride-hailing drivers are scored based on their characteristics (≥1.0 indicates they are ride-hailing drivers, <0.8 excludes them, and 0.8~1.0 indicates they are suspected and require secondary verification). This refers to the user's average daily mileage. The dynamic mileage threshold (calibrated based on urban ride-hailing operation data) is calculated as follows:
[0101] in, 150km is the baseline mileage threshold; P is the population density of the target city (unit: people / square kilometer). The baseline density is 5,000 people per square kilometer. This is the calibration coefficient.
[0102] Average daily driving time for users (unit: hours); Dynamic duration threshold (set based on different cities); The detour index threshold is 1.5. The average detour index of the user's trajectory is calculated as follows:
[0103] Let be the actual trajectory length of the i-th trip. (The straight-line distance between the start and end points).
[0104] Step S74: The target users initially identified as potential private car owners are compared with the user identifiers and feature data in the ride-hailing driver database to generate the private car travel behavior recognition results corresponding to the target users.
[0105] It should be noted that "target user" refers to users initially identified as potential private car owners; "user identifier" refers to information used to uniquely identify users (such as anonymized phone number hash values); "feature data" refers to the average daily mileage, average daily driving time, and average trajectory detour index stored in the ride-hailing driver database; and "private car travel behavior identification result" refers to the final result determining whether the target user is a genuine private car owner. The core purpose of this step is to exclude ride-hailing drivers among potential private car owners by comparing them with the ride-hailing driver database, thus resolving the identification confusion caused by the frequent use of motor vehicles by both, ensuring the accuracy of the final identification result, and providing a reliable user classification basis for scenarios such as traffic planning and car owner services. In one possible implementation, the comparison process adopts an "identifier priority + feature assistance" strategy: if the target user's identifier is in the ride-hailing driver database, they are directly identified as not being a private car owner; if the identifier is not in the database, but the feature data is close to the ride-hailing driver threshold (such as an average daily mileage of 140-150 kilometers), they are marked as "suspected ride-hailing driver" and require further verification.
[0106] Specifically, firstly, the system extracts the identifiers and characteristic data (average daily mileage, average daily driving time, and average detour index) of target users initially identified as potential private car owners from the user behavior analysis database. Then, it compares the target user identifiers with user identifiers in the ride-hailing driver database. If a match is found, the user is identified as "non-private car owner (ride-hailing driver)". If no match is found, the target user's characteristic data is further compared with characteristic thresholds in the ride-hailing driver database. If none of the thresholds are met, the user is identified as "private car owner". If the characteristic data is close to the threshold (e.g., meeting two thresholds), the user is marked as "suspected ride-hailing driver", requiring further data verification. Alternatively, based on the extracted characteristics, if a target user has had a T-type traffic violation within the last 30 days... taxi If a user's profile shows a travel history of 15 days that matches the characteristics of a ride-hailing driver, the user can be identified as a ride-hailing driver. Finally, the final private car travel behavior identification result is associated with the target user and stored in the user behavior analysis database, completing the entire identification process. (See reference below.) Figure 4 .
[0107] In another embodiment, the determination of whether a user is confirmed as a private car traveler for a given trajectory chain can also be made using a comprehensive calculation formula for the confidence level of private car travel. The formula is as follows:
[0108] Wherein, C: confidence level of private car travel (value range [0,1], ≥0.6 is high confidence level, 0.4~0.6 is medium confidence level, <0.4 is low confidence level); Contribution value to parking lot dwell time (1 for dwell time before / after travel, 0 for no dwell time); Contribution value for navigation applications (1 for accessing navigation applications during the trip, 0 for not accessing them); Contribution value to trajectory feature matching (calculated based on the matching degree between the trajectory and the main urban road), = Matching the length of the main road's trajectory / Total trajectory length, with a value range of [0,1]); These are weighting coefficients to highlight the priority of core features.
[0109] Furthermore, a formula for determining private car owners is designed by integrating multi-dimensional features, as follows:
[0110] Where F is the private car owner identification index (≥0.65 indicates a private car owner, <0.65 excludes them); This represents the number of private car travel routes within a preset time period. This represents the total number of valid travel trajectories within a preset time period. Scoring ride-hailing drivers based on their characteristics; The average confidence level of C for all private car travel trajectories; , , The weighting coefficients are used to prioritize core statistical features while also considering interference exclusion and confidence levels.
[0111] This embodiment transforms user travel behavior into quantifiable indicators by statistically analyzing the proportion of private car trips, avoiding subjective judgment bias and making the initial assessment of potential private car owners more objective. Furthermore, by constructing a database of ride-hailing drivers and comparing their characteristics, users who match the behavioral characteristics of ride-hailing drivers are eliminated, thus solving the problem of confusion between the travel characteristics of private cars and ride-hailing drivers and improving the accuracy of the final identification results.
[0112] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0113] This application also provides a private car travel behavior recognition device based on communication big data, please refer to... Figure 5 The private car travel behavior recognition device based on communication big data includes: The acquisition module 51 is used to acquire the communication signaling data of the target user within a preset time period; Construction module 52 is used to construct several single-trip trajectory chains of the target user based on the communication signaling data; The identification module 53 is used to identify the single travel mode corresponding to each of the single travel trajectory chains; The generation module 54 is used to generate the private car travel behavior recognition result corresponding to the target user based on each of the single travel modes.
[0114] The private car travel behavior recognition device based on communication big data provided in this application adopts the private car travel behavior recognition method based on communication big data in the above embodiments, which can solve the technical problems in the background art. Compared with the prior art, the beneficial effects of the private car travel behavior recognition device based on communication big data provided in this application are the same as the beneficial effects of the private car travel behavior recognition method based on communication big data provided in the above embodiments, and other technical features in the private car travel behavior recognition device based on communication big data are the same as the features disclosed in the methods of the above embodiments, and will not be repeated here.
[0115] This application provides a private car travel behavior recognition device based on communication big data. The private car travel behavior recognition device based on communication big data includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the private car travel behavior recognition method based on communication big data in the above embodiment 1.
[0116] The following is for reference. Figure 6 This document illustrates a structural schematic diagram of a private car travel behavior recognition device based on communication big data, suitable for implementing embodiments of this application. The private car travel behavior recognition device based on communication big data in this application embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The private car travel behavior recognition device based on big data communication shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0117] like Figure 6As shown, a private car travel behavior recognition device based on communication big data may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory 1002 or a program loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the private car travel behavior recognition device based on communication big data. The processing unit 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. Communication device 1009 allows the private car travel behavior recognition device based on communication big data to wirelessly or wiredly communicate with other devices to exchange data. While the figure shows private car travel behavior recognition devices based on communication big data with various systems, it should be understood that implementing or having all of the systems shown is not required. More or fewer systems may be implemented alternatively.
[0118] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0119] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0120] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the private car travel behavior recognition method based on communication big data provided by the above methods.
[0121] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0122] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0123] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.
Claims
1. A method for recognizing private car travel behavior based on big data communication, characterized in that, include: Acquire communication signaling data of the target user within a preset time period; Based on the communication signaling data, construct several single-trip trajectory chains for the target user; Identify the single travel mode corresponding to each of the single travel trajectory chains; Based on each of the aforementioned single travel modes, a private car travel behavior recognition result corresponding to the target user is generated.
2. The method for recognizing private car travel behavior based on big data communication as described in claim 1, characterized in that, The acquisition of communication signaling data of the target user within a preset time period includes: Obtain the raw communication signaling data of the target user within a preset time period; Based on the timestamp and location information in the original communication signaling data, duplicate data in the original communication signaling data is removed to obtain the communication signaling processing data. According to the preset ping-pong effect elimination method, the ping-pong effect data in the communication signaling processing data is eliminated to obtain the communication signaling data.
3. The method for recognizing private car travel behavior based on big data communication as described in claim 1, characterized in that, The step of constructing several single-trip trajectory chains for the target user based on the communication signaling data includes: A spatiotemporal density clustering algorithm is used to aggregate the spatiotemporally adjacent trajectory points in the communication signaling data into user dwell points, generating a user dwell point sequence that includes the dwell point location and dwell time. Based on the location of the dwell point and the dwell time, the time interval and spatial distance between two adjacent user dwell points are calculated. If the time interval is greater than a preset time threshold and the spatial distance is greater than a preset spatial threshold, then the two user dwell points are determined as a pair of travel origin and destination points. The city road network is acquired, and a multi-strategy trajectory matching method is used to match the trajectory points between each pair of origin and destination to the city road network, thus integrating them to form the trajectory chains of each single trip.
4. The method for recognizing private car travel behavior based on big data communication as described in claim 1, characterized in that, The identification of the single travel mode corresponding to each of the single travel trajectory chains includes: Obtain the preset database of indoor distributed base stations in underground parking lots and the Internet access data in the communication signaling data; Based on the preset underground parking lot indoor distributed base station database, determine whether there are records of the single trip trajectory chain residing in the underground parking lot indoor distributed base station within a set time period before the start and a set time period before the end, and obtain the first judgment result; Deep packet inspection technology is used to parse the server network identification information in the Internet data to determine whether there are records of accessing navigation application services during the trip period, and a second judgment result is obtained. Based on the first judgment result and the second judgment result, the confidence level of private car travel is determined. If the confidence level of private car travel is medium or high, then the single travel mode is determined to be private car travel.
5. The method for recognizing private car travel behavior based on big data communication as described in claim 1, characterized in that, The identification of the single travel mode corresponding to each of the single travel trajectory chains includes: Obtain the preset database of subway private network base stations; The base station identifiers corresponding to the trajectory points of the single trip trajectory chain are compared with the metro private network base station database. If at least one trajectory point has a base station identifier that belongs to the metro private network base station database, then the single trip mode is determined to be metro travel. If the determination result is not subway travel, then calculate the average travel speed of the single travel trajectory chain and the percentage of travel time in the single travel trajectory chain where the speed reaches or exceeds the preset high speed threshold. If the average driving speed is lower than a preset speed threshold and the percentage of driving time is lower than a preset percentage threshold, then the single trip mode is determined to be low-speed travel.
6. The method for recognizing private car travel behavior based on big data communication as described in claim 1, characterized in that, The identification of the single travel mode corresponding to each of the single travel trajectory chains includes: Obtain the origin-destination pair corresponding to the single trip trajectory chain, and under the conditions of limiting the number of transfers, detour index and bus stop distance threshold, perform bus route navigation based on bus network data to generate several candidate bus routes; Calculate the similarity between the trajectory of each candidate bus route and the trajectory of the single trip trajectory chain; If the trajectory similarity of at least one candidate bus route exceeds a preset similarity threshold, then the single trip mode is determined to be public transportation.
7. The method for recognizing private car travel behavior based on big data communication as described in claim 1, characterized in that, The step of generating private car travel behavior recognition results for the target user based on each of the aforementioned single travel modes includes: The number of private car trajectories in all single travel trajectory chains within the preset time period, where the single travel mode is private car travel, is counted, and the proportion of the number of private car trajectories to the total effective travel trajectories is calculated. If the percentage exceeds a preset percentage threshold, the target user is preliminarily determined to be a potential private car owner; Obtain and construct a ride-hailing driver database based on the daily mileage, daily driving time, and trajectory detour index from historical single-trip trajectory chains; The target users initially identified as potential private car owners are compared with the user identifiers and feature data in the ride-hailing driver database to generate private car travel behavior recognition results corresponding to the target users.
8. A private car travel behavior recognition device based on communication big data, characterized in that, include: The acquisition module is used to acquire communication signaling data of the target user within a preset time period; The construction module is used to construct several single-trip trajectory chains of the target user based on the communication signaling data; The identification module is used to identify the single travel mode corresponding to each of the single travel trajectory chains. The generation module is used to generate private car travel behavior recognition results for the target user based on each of the single travel modes.
9. A private car travel behavior recognition device based on communication big data, characterized in that, The private car travel behavior recognition device based on communication big data includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the private car travel behavior recognition method based on communication big data as described in any one of claims 1 to 7.
10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the method for recognizing private car travel behavior based on big data communication as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Mobile phone user trip mode identification method based on mobile phone signaling data and navigation route data
CN106197458A
Mobile phone signaling data travel mode identification method based on interest points and navigation data
CN111341135A
Urban travel mode comprehensive identification method based on mobile phone signaling data
CN111653093A
Multi-section travel mode identification method based on mobile phone signaling data
CN117858024A
Travel mode identification method and device, storage medium and program product
CN118820833A