Urban road traffic state identification method based on taxi track big data
By reconstructing HMM and using FCM/DBSCAN clustering methods based on taxi trajectory big data, the traffic status and congestion areas of urban roads are identified, solving the problem of inaccurate identification in existing technologies, achieving more precise traffic management and planning, and improving the traffic efficiency of urban road networks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGXI TEACHERS EDUCATION UNIV
- Filing Date
- 2025-12-30
- Publication Date
- 2026-04-17
AI Technical Summary
Existing technologies cannot accurately identify road traffic conditions or accurately extract congestion hotspots in urban road networks. Traditional traffic data collection equipment is costly, difficult to maintain, and has limited coverage. Floating car trajectory data is still insufficient in congestion identification.
Based on taxi trajectory big data, the algorithm reconstructs the HMM for map matching, combines FCM clustering and DBSCAN 3D clustering to identify road traffic status boundaries and extract congested areas, and calculates traffic operation index through traffic flow parameters to accurately identify morning and evening peak hours.
It enables accurate identification of urban road traffic conditions and precise extraction of congested areas, helping traffic management departments to implement more effective traffic management, reduce the possibility of congestion and accidents, and improve road network efficiency.
Smart Images

Figure CN121884584A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of smart city technology, and in particular to a method for identifying urban road traffic status based on taxi trajectory big data. Background Technology
[0002] With the increasing severity of urban traffic congestion, people have begun to focus on how to effectively solve this problem. Identifying congestion hotspots in cities is key to optimizing road network structure. All analysis and research presupposes an accurate data acquisition system. Traditional traffic data acquisition methods mainly rely on fixed-point detection equipment, such as geomagnetic coils, piezoelectric detectors, video monitoring, and microwave, ultrasonic, and infrared radar. However, these devices all have limitations. Geomagnetic coils are easily damaged, difficult to maintain, and require installation on closed road sections, making them difficult to apply to existing roads with high traffic volume. The accuracy of video monitors is greatly affected by weather and other environmental factors. Microwave, ultrasonic, and infrared radar are affected not only by weather conditions but also by road infrastructure. Most cities in my country have limited funds for fixed detection equipment, and its maintenance costs are high. Installation also requires consideration of the environment, including tunnels, bridges, toll stations, and other road facilities, making it difficult to establish a comprehensive traffic data acquisition system and achieve full coverage of urban road network data. With the widespread adoption of Global Navigation Satellite System (GNSS) devices, floating car trajectory data acquisition technology has emerged. This technology uses GNSS devices carried by drivers or vehicles to provide information such as vehicle location, timestamps, speed, and direction at regular intervals. Due to its low cost, it is widely used on floating cars and has been vigorously promoted in major Chinese cities in recent years, accumulating massive amounts of positioning data and making full data acquisition coverage possible. This acquisition equipment has advantages such as low cost, wide coverage, rich variety, high precision, and high timeliness, providing rich data support for research on road traffic status identification and urban road network congestion area identification. Secondly, taxis, as an important component of the urban transportation system, can provide real-time dynamic information on urban traffic through their operational trajectory data. By mining and analyzing this data, the potential patterns and distribution characteristics of urban traffic can be effectively identified, providing reliable decision-making basis for traffic management departments. In practical applications, internet companies such as Gaode and Didi have provided relatively accurate results for real-time road network status evaluation through extensive mobile user data. Therefore, using GNSS floating car trajectory positioning data to analyze urban roads has a solid theoretical basis, data foundation, practical feasibility, and practical significance.
[0003] In recent years, with the continuous development of geographic information technology and transportation information technology in my country, an Intelligent Transportation System (ITS) has emerged. This system primarily improves urban traffic conditions by real-time monitoring of road traffic status, optimizing signal control, providing real-time traffic information and intelligent navigation. It offers travelers and traffic managers real-time, accurate, and comprehensive urban traffic information, achieving safe, efficient, environmentally friendly, and sustainable development of the transportation system, and effectively alleviating the increasingly serious traffic congestion problem. Supported by big data technology, ITS has collected and stored a large amount of historical traffic information from urban road networks. Researchers can use effective data analysis methods to find historical patterns in traffic flow from historical data, uncover the underlying patterns, and capture and identify signs of traffic flow changes during congestion, thereby achieving monitoring and effective intelligent control of urban road network traffic conditions. However, while existing research has achieved some results in commuting and congestion identification using floating car trajectory data, it still cannot accurately identify road traffic conditions or accurately extract congestion hotspots in urban road networks.
[0004] Therefore, a method for identifying urban road traffic conditions based on big data of taxi trajectories is needed. Summary of the Invention
[0005] To address the limitations of existing technologies in accurately identifying road traffic conditions and extracting urban road network congestion hotspots, this invention provides a method for identifying urban road traffic conditions based on taxi trajectory big data. This method utilizes taxi trajectory positioning big data to match urban road networks, comprehensively considers various traffic flow factors at different times, and identifies the boundaries of five traffic conditions for roads of different levels and peak hours based on the FCM clustering model. Furthermore, it merges complex road segments of the same type to identify urban congestion areas. This method is of great significance for supplementing and improving the system and methods for identifying road segment traffic conditions and urban road congestion hotspots. The specific technical solution is as follows: A method for identifying urban road traffic conditions based on taxi trajectory big data includes the following steps: The first step is to preprocess taxi trajectory data and road network data, calculate the due north azimuth of each road segment, select the weight of directional factors by judging road type, construct a new observation probability model, and use the reconstructed HMM algorithm to perform large-scale map matching of the road network in the study area after selecting the radius threshold. The second step is to statistically analyze the traffic flow parameters of each level of road segment after equal duration, select candidate parameters that are strongly correlated with the key parameters, perform FCM clustering to identify the boundaries between the traffic states of each road, extract the road segments in the road network that are in a state of severe congestion, and calculate the traffic operation index. The third step involves preprocessing the road network and extracting midpoints through equidistant segmentation. The traffic status boundary is used as the neighborhood speed threshold, and the shortest path is used as the neighborhood path threshold to reconstruct the DBSCAN 3D clustering algorithm. The highest level of traffic status is used as the seed point for expansion and spread. The resulting clustering results map to the loop network, and similar road segments are connected to form congested areas.
[0006] Preferably, the process of using the reconstructed HMM algorithm to perform large-scale map matching of the road network in the study area is as follows: Constructing a buffer zone: Construct a buffer zone with the positioning point as the center and the maximum error of the positioning device as the radius. The road segments within the buffer zone are candidate road segments. Mapping candidate points: Find a point on the candidate road segment that minimizes the distance between two points; this point is called a candidate point. Obtaining the shortest distance between candidate points: The relationship between the location point and the candidate point is many-to-many (N:M). It is not possible to directly calculate the shortest path between the candidate points of two adjacent location points. It is necessary to obtain the distance between the upstream and downstream endpoints of the road segment where the candidate point is located to calculate the shortest path distance between the two adjacent points. Constructing the HMM model: The joint distributed model of HMM is shown in the following equation: In the formula, It is a joint distributed HMM; It is the initial state probability; It is the probability of observation; It is the transition probability; Construct the observation probability model as shown in the following equation: In the formula, Distance factor for observation probability; For positioning points; Candidate road sections; for and Great circle distance between them; The standard deviation of the measured values at the positioning point; pi (π) Natural constant ; Construct the state transition probability model as follows: In the formula, for Great circle distance between them; for The shortest path distance between; for and The absolute value of the difference is taken as the median; natural constant ; The Viterbi algorithm is used to solve the HMM model to obtain the matching points.
[0007] Preferably, the method for obtaining the shortest distance between candidate points is as follows: Step 1: The two points are candidate points at two different times. for , Distance from the point to the upstream node for Distance from downstream node, via = judge Dot and If the points are on the same road segment, proceed to step two; otherwise, proceed to step three. Step Two: If The path length is 0; otherwise, output... Path output ,Finish; Step 3: Judgment , To determine if a path is directly connected, check the topology table to see if the road segments are adjacent. If so, output the path length. Path output If the above steps are not completed, proceed to step four. Step 4: Calculation , Shortest path to a point and length Path length output Path output The steps are now complete.
[0008] Preferably, the traffic flow parameters include at least average speed, speed standard deviation, number of matching points, average stopping time, and taxi traffic volume.
[0009] Preferably, the calculation process for the traffic operation index is as follows: Step 1: Set statistical intervals, calculate the average speed of each road segment within the statistical intervals, classify the traffic conditions of each road segment according to the road traffic condition level classification boundaries, filter out the road segments in a severely congested state, and calculate the severely congested mileage ratio of each road level using the following formula: In the formula, For the first The ratio of severely congested mileage on graded roads; For the first The length of road segments in graded roads that are severely congested; For the first The total length of graded road sections.
[0010] Step 2: Calculate the vehicle-kilometers on roads of each grade, the total vehicle-kilometers of the overall road network, and the weighting factors for each road's operating grade, as shown in the following formula: In the formula, For the first Vehicle mileage on graded roads; For the first Number of taxis on graded roads; For the first The total length of graded road sections; The total vehicle kilometers of the road network; For the first Weighting factors for graded roads; Step 3: Using the proportion of vehicle-kilometers of roads of each level as weights, calculate the overall road network's severely congested mileage ratio, as shown in the following formula.
[0011] In the formula, The ratio of severely congested road network mileage; For the first Weighting factors for graded roads; For the first The ratio of severely congested mileage on graded roads.
[0012] Step 4: Based on the pre-defined conversion relationship between the ratio of severely congested mileage in the road network and the traffic operation index, convert the ratio of severely congested mileage in the road network into the traffic operation index.
[0013] Preferably, the specific process for reconstructing the DBSCAN 3D clustering algorithm is as follows: Step 1: Initialization: Sample point dataset All points within the cluster are marked as unvisited; cluster list Used to store clusters Sample points; cluster labels ; Speed threshold =0; Road operation level With neighborhood velocity threshold dictionary Based on the road operation level of the seed point Decision, when hour, ,when hour, ,when hour, ,when hour, Input dataset Road network Neighborhood radius threshold Neighborhood path threshold The neighborhood density threshold for the minimum number of neighboring points required to form a core point. ; Step 2: Determine seed points: Filter out specific road traffic status levels The sample points are used as candidate seed point datasets ,from Randomly select an unvisited, unclassified sample point And mark it as visited, with a dot Threshold for the radius of the circle's center neighborhood Establish a buffer zone for the radius, and calculate the distance from the sample points within the buffer zone to the center point. shortest path Filter out those that meet the neighborhood path threshold The sample points are used as candidate neighborhoods ,like Then the marker point For noisy points, return to the previous step from the candidate seed point dataset. Randomly select an unvisited, unclassified point ;like Then the point Add to this cluster class In the middle, that is, the list of clusters. Add points According to the point Road operation level From the dictionary Obtain the neighborhood velocity threshold of this cluster class. ,mark For untraversed regions, start from the candidate neighborhood. Choose one that is not yet visited and has the shortest path distance to the center of the circle. Minimum sample point Mark it as traversed and add it to the cluster list. In the middle, calculate the list of cluster classes. inner field mean ,like Then the point From the list of clusters and candidate neighborhood Remove from the middle, if after removal Then the point Mark as noise points and remove from cluster list Remove and return to the candidate point dataset. Randomly select an unvisited point If removed Then return to the previous step from the candidate neighborhood. Select the untraversed circle and the circle center. shortest path distance The smallest point ;like Then return directly to the previous step from the candidate neighborhood. Select untraversed points and their intersection with the center of the circle. shortest path distance smallest Repeat the above operation until candidate neighborhoods are reached. The neighborhood of all sample points within the range has been traversed. Exit the loop and set the point Mark as seed point and seed point neighborhood (Besides yourself) join the spread queue middle; Step 3: Spread and expand. Similar to Step 1, from the spread list... Randomly select unvisited and unclassified sample points And mark it as visited, with a dot Threshold for the radius of the circle's center neighborhood Establish a buffer zone for the radius, and calculate the distance from the sample points within the buffer zone to the center point. shortest path Filter out those that meet the neighborhood path threshold The sample points are used as candidate neighborhoods ,like Then the marker point For the boundary point, return to the previous step from the spread list. Randomly select unvisited and unclassified sample points ;like Then mark For untraversed regions, start from the candidate neighborhood. Choose an untraversed circle that is adjacent to the center. shortest path distance Minimum sample point Mark it as traversed and add it to the cluster list. In the middle, calculate the list of cluster classes. inner field mean ,like Then the point From the list of clusters and candidate neighborhood Remove from the middle, if after removal Then the point Mark as boundary point, return to the spread list Randomly select unvisited and unclassified sample points If removed Then return to the previous step from the candidate neighborhood. Select the untraversed circle and the circle center. shortest path distance The smallest point .like Then return directly to the previous step from the candidate neighborhood. Select the untraversed circle and the circle center. shortest path distance The smallest point Repeat the above operation until candidate neighborhoods are reached. The neighborhood of all sample points within the range has been traversed. Exit the loop and set the point Mark it as the core point, and put the core point neighborhood (Besides yourself) join the spread queue Middle; Repeat step three until the spread queue continues. All sample points in the loop have been visited, so exit the loop.
[0014] Step 4: Iterate through the loop: as the queue spreads When all sample points in the cluster have been visited, the cluster has finished spreading, and a new cluster is created: If the candidate seed point dataset If not all visits are made, return to step two: From Randomly select an unvisited, unclassified sample point Candidate seed point dataset All have been visited. ,like Then return to step two: filter out the specific road traffic status level. The sample points are used as candidate seed point datasets ;like Then output the cluster classes. Cluster list gather.
[0015] Preferably, the process for obtaining congested areas is as follows: sample set The DBSCAN 3D clustering algorithm is used to reconstruct the clusters for each traffic status level. These clusters are then connected to road segments via cluster sample point IDs to obtain the clustered road network data. Seed points connected to road segments form seed road segments, core points form core road segments, and boundary points form boundary road segments. Following a traffic status level from highest to lowest, for road segments of the same category, the shortest path between the seed point and its corresponding core point is searched based on the seed road segment's location. If a road segment on the shortest path is not classified, it is assigned to that cluster; otherwise, it is assigned to the higher-level cluster based on its traffic status. All core road segments are traversed, and the shortest path from each core segment to a boundary road segment is searched. The core road segments are then clustered again using the same method as the seed road segments. This process is repeated until all core road segments have been traversed, and the resulting list of road segment clusters is output.
[0016] A computer-readable storage medium includes a stored program, wherein, when the program is executed, it controls the device on which the computer-readable storage medium is located to perform the urban road traffic status recognition method based on taxi trajectory big data as described above.
[0017] A processor for running a program, wherein the program executes the urban road traffic status recognition method based on taxi trajectory big data as described above.
[0018] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention first preprocesses taxi trajectory data and road network data, calculates the due north azimuth of each road segment, selects appropriate directional factor weights by determining road type, constructs a new observation probability model, and uses a reconstructed HMM algorithm to perform large-scale map matching on the road network of the study area after selecting an appropriate radius threshold. Then, traffic flow parameters of each level of road segment are statistically analyzed according to equal time intervals, and candidate parameters strongly correlated with key parameters are selected for FCM clustering to identify the boundaries between the traffic states of each road, extracting road segments in a severely congested state in the road network and calculating traffic operation indexes to accurately identify morning and evening peak hours. Finally, the road network is preprocessed and equidistantly segmented to extract midpoints. The traffic state boundaries are used as neighborhood speed thresholds and the shortest path is used as neighborhood path thresholds to reconstruct the DBSCAN three-dimensional clustering algorithm. The highest level of traffic state is used as the seed point for expansion and spread. The obtained clustering results map to a loop network, and road segments of the same type are connected to form congested areas. In summary, this invention utilizes taxi trajectory big data to focus on the identification of road traffic conditions and congested areas in the urban road network, forming a full-chain urban traffic analysis technology system of "precise matching - status identification - index evaluation - regional clustering". This helps traffic management departments understand more accurate urban road traffic conditions and the distribution of congestion hotspots, enabling them to implement more accurate and effective traffic management and urban planning measures to reduce the likelihood of traffic congestion and accidents, and improve the traffic efficiency of the urban road network. Attached Figure Description
[0019] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the accompanying drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn to scale.
[0020] Figure 1 This is a flowchart of the method of the present invention; Figure 2 Here is a flowchart of the HMM-based map matching algorithm; Figure 3 A schematic diagram illustrating the process of finding candidate points for constructing a buffer zone; Figure 4 Here is a flowchart of the fuzzy C-means clustering algorithm; Figure 5 Here is a flowchart of the FCM-based road traffic status recognition model; Figure 6 To reconstruct the DBSCAN clustering algorithm flowchart; Figure 7 This is a schematic diagram of a road network congestion area identification model based on DBSCAN. Detailed Implementation
[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0023] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0024] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0025] In one embodiment of the present invention, a method for identifying urban road traffic conditions based on taxi trajectory big data is provided, such as... Figure 1 As shown, it includes the following steps: The first step involves preprocessing taxi trajectory data and road network data, calculating the due north azimuth angle of each road segment, selecting appropriate directional factor weights based on road type, constructing a new observation probability model, and using a reconstructed HMM algorithm to perform large-scale map matching of the road network in the study area after selecting an appropriate radius threshold. The second step involves statistically analyzing traffic flow parameters for each level of road segment after matching at equal time intervals, selecting candidate parameters strongly correlated with key parameters, and using FCM clustering to identify the boundaries between road traffic states. This process extracts road segments in a severely congested state and calculates traffic operation indices to accurately identify morning and evening peak hours. The third step involves preprocessing the road network and extracting midpoints through equidistant segmentation. Using traffic state boundaries as neighborhood speed thresholds and the shortest path as a neighborhood path threshold, a DBSCAN 3D clustering algorithm is reconstructed. The highest-level traffic state is used as a seed point for expansion and sprawl. The resulting clustering results map to a loop network, with similar road segments connecting to form congested areas.
[0026] The following is a detailed explanation of each step: I. Data Preprocessing: Taxi location data is collected by a receiving device at certain sampling intervals. During this process, interference from various external factors such as positioning equipment, atmospheric environment, and building facilities can cause errors in transmission or storage, resulting in abnormal data. This data cannot reasonably reflect the actual taxi operating status, and direct use will seriously affect subsequent identification of road traffic conditions. Therefore, this embodiment formulates six types of data filtering methods for the following situations: ① Data filtering. The filtering study area has longitudes of 108.128265~108.612418 and latitudes of 22.638756~22.977737 (decimal system). Based on the "Design Specifications for Highway Speed Limit Signs" issued by the Ministry of Transport...
[58] The speed limit on highways is 120 km / h; therefore, data with speeds between 0 and 120 km / h are retained. ② Data redundancy: For data entries with multiple identical entries, retain the first one and delete the rest. ③ Missing data: If several fields in a row are NAN or a field is continuously missing over time, delete it. If a row is missing but adjacent rows are normal, use average interpolation. ④ Data anomalies: Such as garbled characters or illogical data (vehicle speed is 0 but position changes), delete them. ⑤ Rest periods: For drivers resting by the roadside, filter data where the speed is 0 and the position remains unchanged for half an hour; these are considered rest periods and deleted. ⑥ Data drift: The distance between the current position and the position at adjacent times is too large, significantly exceeding the distance a vehicle can travel within a normal timeframe. Define data outlier factors (Equations 2.3-1 to 2.3-3) to identify data drift and determine the validity of trajectory data.
[0027] Equation 2.3-1 Equation 2.3-2 Equation 2.3-3 In the formula, It is the first The location point at any given moment; for and The Euclidean distance; and for and The recorded moment; for arrive The average speed; The maximum permissible speed; Outlier factor; For vehicles Cumulative outlier factor of the trajectory. If the velocity is much greater than the normal value, then... If the value is 1, the cumulative value of the outlier factor for the entire trajectory is calculated. A validity threshold is set. Data below the threshold is retained, and drift data is processed by average interpolation. Otherwise, the data is deleted.
[0028] Trajectory segment division: Since the device is not always working, the trajectory segment needs to be divided to prepare for the subsequent map matching algorithm. The data within a day is sorted by license plate number and time. The time interval between two adjacent points is calculated. If it is greater than the threshold, it is determined to be a new trajectory and marked with a sequence number.
[0029] Coordinate system transformation: The trajectory data is in the Mars coordinate system GCJ-02, a non-linear topographic map processing technique for data security. However, the road network is in the geographic coordinate system WGS-84. A unified coordinate system is required. Due to the massive amount of data, GIS software struggles to process it efficiently. This embodiment uses the Python third-party library PyProj to perform coordinate system transformation on hundreds of millions of rows of data. Example: x2, y2, z2 = transform(p1, p2, x1, y1, z1, radians=False), where p1 is the original coordinate system, p2 is the target coordinate system, x1, y1, and z1 are the latitude, longitude, and elevation in the p1 coordinate system (default z1=none), and the radians parameter indicates whether to return values in radians (default False). The transformation result is shown below.
[0030] II. Road Network Data Preprocessing: The OSM road network includes not only motor vehicle lanes but also non-motorized roads such as bicycle lanes, sidewalks, and pedestrian walkways. Roads designated for motor vehicles and those accessible by motor vehicles are identified. Roads categorized as maintenance roads, unclassified, or unknown are then further classified. This is based on the "Urban Road Traffic Operation Evaluation Index System" of multiple regions. [59,60] Roads are classified into four operating levels based on the road type field fclass (Table 1).
[0031] Table 1. Specific Classification of Road Operation Levels in OSM In the OSM road network, intersections of low-level non-same-named roads do not generate intersection points. To create subsequent topology relationships, it is necessary to break the intersections. Then, the network dataset is generated to extract the intersections.shp data, assign identification codes and latitude and longitude to the intersections, and spatially connect the road network with the intersections. Roads with more than 2 connections are broken, and roads with 0 or 1 connections are deleted to ensure that each road has only 2 intersections. A topology table is generated based on the spatial connection results.
[0032] In this embodiment, a new directional factor is added to the map matching. Roads need to be broken into segments by nodes. The north azimuth (decimal degree) of each segment is calculated using the "Create COGO Field" function in ArcCatalog. When the direction of the segment is from west to east, the azimuth is 90.00, and when it is from west to east, it is 270.00.
[0033] III. Map matching based on taxi trajectory big data (projecting coordinates onto the road network, this process is called map matching (MM)). First, let's further explain the following possible terms: Road Network: A vector road dataset of a road network model, consisting of several roads. Road: Two intersections are considered as one road, composed of several road segments. Road Segment: A road segment is formed by connecting two endpoints with a straight line; it is the smallest unit in the road network and has direction. Location Point: A point observed using a positioning device. The attribute structure is shown in Table 2. Observation Sequence: Also known as a trajectory sequence or trajectory. It consists of a set of adjacent location points. Composition. Candidate road segment: Location point The set of road segments to be matched. Candidate points: location points. A set of possible corresponding actual locations.
[0034] Table 2. Structure of Road Network Attributes osm_id: Unique identifier for OSM road network data. name: Road name. fclass: Road type, such as motorway, trunk, primary, etc. ref: Road number, according to the "Highway Route Identification Rules and National Highway Numbering".
[57] The road numbering system uses uppercase letters and numbers. Oneway: Whether it is a one-way street. F: Double-lane one-way street, T: Single-lane one-way street, B: Single-lane double-lane street. Bridge: Whether it is a bridge. F: Yes, T: No. Tunnel: Whether it is a tunnel. F: Yes, T: No.
[0035] This embodiment uses a Hidden Markov Model (HMM)-based map matching algorithm to reconstruct observation probabilities and establish a directional factor model to improve the algorithm's accuracy. The HMM-based map matching process can be viewed as a decoding problem of the HMM model. This algorithm considers not only the positional relationship between locations and roads, the connectivity and reachability between roads, but also the correlation between two adjacent locations at two different time points. The main steps are as follows: Figure 2 As shown: (1) Constructing a buffer zone: Under normal circumstances, the error of GNSS positioning equipment will not exceed a certain range. Therefore, a buffer zone is constructed with the positioning point as the center and the maximum error of the positioning equipment as the radius. The road segments within the buffer zone are candidate road segments.
[0036] (2) Mapping candidate points: Find a point on the candidate road segment that minimizes the distance between two points; this point is called a candidate point. Steps (1) and (2) are as follows. Figure 3 As shown.
[0037] (3) Obtaining the shortest distance between candidate points: The relationship between the location points and candidate points is many-to-many (N:M). It is not possible to directly calculate the shortest path between candidate points of two adjacent location points. Instead, the shortest path distance between adjacent points needs to be calculated by obtaining the distance between the upstream and downstream endpoints of the road segment where the candidate point is located. The specific steps are as follows: Step 1: The two points are candidate points at two different times. for , Distance from the point to the upstream node for Distance from downstream node, via = judge Dot and If the points are on the same road segment, proceed to step two; otherwise, proceed to step three.
[0038] Step Two: If The path length is 0; otherwise, output... Path output ,Finish.
[0039] Step 3: Judgment , To determine if a path is directly connected, check the topology table to see if the road segments are adjacent. If so, output the path length. Path output If the above steps are not completed, proceed to step four. Step 4: Calculation , Shortest path to a point and length Path length output Path output The steps are now complete.
[0040] (4) Constructing the HMM model: The HMM model is a Markov process for solving unobserved states, and its structure has a set of lengths of Observation state sequence (trajectory point sequence) ), and a set of candidate point sequences The observed state corresponds to a location point; the unobserved state corresponds to a matching point; there is a set of observation probabilities between observed and unobserved states. , as the positioning point Projected onto candidate road segments The probability of the observed states; there is a set of transition probabilities between observed states. To locate the candidate road segment Drive to another candidate road section The probability of; the joint distribution of HMM is shown in Equation 3.3-1: Equation 3.3-1 In the formula, It is a joint distributed HMM; It is the initial state probability; It is the probability of observation; It is the transition probability.
[0041] (5) Constructing an observation probability model: observation probability This refers to the location point being matched with the candidate road segment. The probability. In the map matching process, the actual location of a point can be considered as being distributed in the surrounding area with a certain probability. For road section, For positioning points On the road section Candidate points on, and for The shortest path distance between them. for The great circle distance between them, according to Newson et al., is based on the observation that the positioning points follow a normal distribution, as shown in Equation 3.3-2: Equation 3.3-2 In the formula, Distance factor for observation probability; For positioning points; Candidate road sections; for and Great circle distance between them; The standard deviation of the measured values at the positioning point; pi (π) Natural constant .
[0042] Equation 3.3-2 only considers the positional relationship between the positioning point and the road, and its model parameters are only related to distance. Since the study area contains a large number of two-way single-lane roads, and the road length accounts for 32.13% of the total length of the study road network, this embodiment, based on the oneway field in the road data, considers the direction of single-lane and double-lane roads and the direction of the positioning point, and establishes different direction weight models. Equations 3.3-3~4 are used to calculate the angle between the direction of the positioning point and the direction of the candidate road segment. When the angle is larger, the weight is smaller; when When the included angle is 90°, the weight is 0; when the direction coincides with the road direction, the weight is 1.
[0043] Equation 3.3-3 Equation 3.3-4 In the formula, Directional factors for the probability of observation; Whether it is a one-way street; It is the angle between the location point and the road direction; The angle between the location point and due north; This is the angle between the road segment and due north. Taking into account both distance and direction factors, the expression for the observation probability model is shown in Equation 3.3-5: Equation 3.3-5 (6) Constructing a state transition probability model: state transition probability Designated sites from candidate road segments Drive to The probability of state transitions. Newson et al. found that the state transition probabilities exhibit an exponential distribution.
[62] As shown in equations 3.3-6~7.
[0044] Equation 3.10 Equation 3.11 In the formula, for Great circle distance between them; for The shortest path distance between; for and The absolute value of the difference is taken as the median; natural constant .
[0045] (7) Solving the HMM model using the Viterbi algorithm: The most commonly used algorithm for decoding the HMM model is the Viterbi algorithm, which is a general dynamic programming algorithm for finding the shortest path in a sequence. The algorithm has three loops: the outermost loop drives the Viterbi cells to move, the middle loop updates the reached nodes, and the innermost loop updates the starting node and calculates the shortest path and two local variables. and .
[0046] The HMM map matching algorithm based on the reconstructed observation probability was used to perform map matching on the road network within the Nanning Expressway Ring Road. Some matching results are shown in Table 3.
[0047] Table 3: Attribute Table of Partial Map Matching Results VID: License plate number. TID: Track segment number. LON, LAT: Location point latitude and longitude. GPSDATETIME: Recording time. MATCHLON, MATCHLAT: Matching point latitude and longitude. LINKID: Location point matching road sign.
[0048] III. Identifying Urban Road Traffic Status Based on FCM Model This embodiment uses the FCM model to cluster roads at three levels and divide them into five categories of road traffic status boundaries. Based on the boundaries, severely congested road sections are extracted, and the Traffic Performance Index (TPI) based on the severe congestion mileage ratio is calculated. This provides a general quantitative evaluation of the overall traffic operation of the road network in the study area and identifies peak congestion periods.
[0049] Road traffic status refers to classifying roads into states such as smooth traffic, basically smooth traffic, light congestion, moderate congestion, and severe congestion based on certain indicators. Road operation level is a method of classifying road operation status based on indicators such as traffic flow, speed, and density. The specific classification in this embodiment is shown in Figure 1.
[0050] There is no uniform numerical standard for the range of traffic conditions from smooth to severe congestion. Different conditions are generally classified based on local road conditions. For example, in the United States, traffic conditions are mainly divided into six levels based on parameters such as average travel speed, road load factor, and road saturation.
[64] Germany classifies road traffic conditions into five levels based on traffic flow density as a parameter.
[65] In 2002, the Ministry of Public Security of my country issued the "Urban Traffic Management Evaluation Index System" manual, which stipulated the use of average speed of road sections to classify traffic into four levels. In 2011, the "Urban Road Traffic Operation Evaluation Index System" was released.
[67] The speed ranges for each road grade are further specified in Table 4; this is from the "Urban Traffic Operation Status Evaluation Standard" released in 2016.
[68] The standard classifies road traffic conditions into five categories based on the relationship between average speed and free velocity, as shown in Table 5. For road section exist The average speed of vehicles at any given time; For road section The free flow velocity.
[0051] Table 4. Classification Standards for the Evaluation Index System of Urban Road Traffic Operation Table 5 Classification Standards for the "Urban Traffic Operation Status Evaluation Standard" Selecting appropriate traffic flow parameters is crucial for identifying traffic conditions. These parameters, through quantitative description of road traffic characteristics, transform traffic condition research from qualitative to quantitative. This embodiment uses the average speed of the road segment as the key parameter; and uses the speed mean square error of the road segment, the number of matching points, the stopping time, and taxi traffic volume as candidate parameters.
[0052] This embodiment selects representative road segments from each road class as the study segments. Based on the map matching results, the average speed, speed variance, number of matching points, average parking time, and taxi traffic volume on the study segments are calculated sequentially at 10-minute statistical intervals. Details are as follows: (1) Average speed is a key indicator for identifying traffic conditions and can reflect the current congestion situation of the road segment very intuitively. Calculate the average speed of each taxi trajectory on the road segment, and then average the average speed of each taxi trajectory to obtain the average speed of the road segment (Equations 4.4-1~2).
[0053] Equation 4.4-1 Equation 4.4-2 In the formula, For road section The average speed at; The number of taxis passing through the section of road; For the first The average speed of a vehicle passing through a road segment; For the first vehicle number The instantaneous velocity of each matching point; For the first The number of matching points left by a vehicle on a road segment.
[0054] (2) The speed mean square deviation can reflect the overall speed fluctuation. When the road segment is unobstructed, the speed fluctuation is small and the speed mean square deviation is small; when the road segment is congested, the fluctuation of vehicle speed and queuing speed is large and the road segment speed mean square deviation is large. The road segment speed mean square deviation is the average of the speed variances of all vehicles, as shown in Equations 4.4-3~4.
[0055] Equation 4.4-3 Equation 4.4-4 In the formula, The mean squared error of the velocity; The number of taxis that passed through; For the first Variance of vehicle speed; For the first vehicle number The instantaneous velocity of each matching point; For the first Average speed of vehicles; For the first The number of vehicle matching points.
[0056] (3) The number of matching points can reflect the traffic status of a road segment to a certain extent. Since the data collection device records information such as the location of vehicles at certain time intervals, when the road segment is congested, the number of location points collected on the road segment will be greater. The total number of matching points for a road segment is shown in Equation 4.4-5.
[0057] Equation 4.4-5 In the formula, This represents the total number of matching points for the road segment. The number of taxis passing through the section of road; For the first The number of location points a taxi passes through on a given road segment. This parameter can become unstable due to driver rest stops. This embodiment has cleaned the data for rest stops, but it cannot clean short-term rest stops of less than half an hour.
[0058] (4) The average parking time is the average of the total parking time of each taxi on the road segment. The longer the queuing time, the more congested the road segment. As shown in Equations 4.4-6~7.
[0059] Equation 4.4-6 Equation 4.4-7 In the formula, This represents the average parking time. For the first Total parking time of the vehicle; For the first The first taxi The duration of each parking session. In addition to being affected by rest behavior, the parking duration is also affected by the sampling interval of the positioning and data collection equipment. The larger the sampling interval, the longer the time between adjacent points, making it difficult to ensure that the vehicle's movement status remains unchanged during the period.
[0060] (5) Taxi traffic volume is the number of taxis passing through the road segment. The actual traffic volume is greater than the taxi traffic volume. This parameter can indirectly reflect the load situation of the road segment.
[0061] To find more suitable model parameters, it is necessary to explore the correlation between candidate parameters and key parameters. Since both key parameters and candidate parameters are linearly correlated, this embodiment uses the Pearson coefficient to explore the linear relationship between key parameters and candidate parameters.
[0062] Pearson correlation, also known as product-moment correlation, is usually used to measure the linear relationship between two random variables. It can effectively avoid using redundant feature parameters when clustering. It is defined as the quotient of the product of the covariance and standard deviation of the two variables, as shown in Equation 4.4-8.
[0063] Equation 4.4-8 In the formula, For parameters Pearson coefficient; For the first parameter in the parameter set Number; for The average value of the Pearson coefficient is 1. The closer the absolute value of the Pearson coefficient is to 1, the stronger the linear correlation between the two parameters, and vice versa, as shown in Table 6 below.
[0064] Table 6. Correlation Classification of Pearson Coefficient like Figure 4 As shown, this embodiment uses FCM clustering, which places greater emphasis on membership degree, allowing data points to belong to multiple categories simultaneously in the form of membership degree. Membership degree refers to the probability that a data point belongs to a certain cluster center. The higher the membership degree of a data point to a cluster center, the more likely it is to belong to that cluster, and the sum of the membership degrees of a data point to all cluster centers is 1. The objective function of FCM is shown in Equations 4.5-1 and 4.5-2.
[0065] Equation 4.5-1 Equation 4.5-2 In the formula, This is the objective function; the smaller the value, the better the clustering effect. This represents the number of data points. For the number of clusters ; for belong Membership degree; As a fuzzy factor, it controls the degree of fuzziness in clustering. The larger the value, the higher the degree of fuzziness, and the more clustering sessions are typically performed. Clustering yields the best results; For the first One data point; For the first The center of a cluster; for and The Euclidean distance between them. The fuzzy C-means algorithm is an iterative convergence process that minimizes the objective function, which can be solved iteratively using the Lagrange multiplier method. The minimum value is determined by the following process: Step 1: Initialization. Set the number of clusters. Input dataset Fuzzy factors Convergence threshold Iteration threshold Initialize the iteration count and membership matrix Randomly selected from the dataset Each sample data point was used as the initial cluster center.
[0066] Step 2: Update the membership matrix Calculate the membership degree of each data point in the dataset to each cluster center according to Equation 4.5-3.
[0067] Equation 4.5-3 In the formula, for belong Membership degree; The number of clusters; For the first in the dataset One data point; and For the first Cluster centers; is the fuzzy coefficient.
[0068] Step 3: Update cluster centers. Calculate each cluster center according to Equation 4.5-4.
[0069] Equation 4.5-4 In the formula, As cluster center; for belong Membership degree; For the first in the dataset One data point; is the fuzzy coefficient.
[0070] Step 4: Determine convergence. If and The difference is less than the convergence threshold or reaching the loop limit Stop the iteration if necessary; otherwise, return to step two.
[0071] The Traffic Performance Index (TPI) is a comprehensive indicator used to quantify the overall operational status of an urban road network. It integrates multi-dimensional traffic data (such as vehicle speed, congestion distance, and travel time) to reflect the operational efficiency of the transportation system in a straightforward numerical format. The TPI typically ranges from 0 to 10, with higher values indicating worse traffic conditions and more severe overall road network congestion, as shown in Table 7 below. Table 7 Relationship between Traffic Operation Index and Overall Road Network Traffic Status The calculation process for the traffic operation index based on the severe congestion mileage ratio mainly consists of the following four steps: Step 1: Using a statistical interval of no more than 15 minutes, calculate the average speed of each road segment within the statistical interval. Divide each road segment into traffic states according to the classification boundaries of road traffic state levels, screen out the road segments in a severely congested state, and calculate the severely congested mileage ratio of each road level according to Formula 4.6-1.
[0072] Equation 4.6-1 In the formula, For the first The ratio of severely congested mileage on graded roads; For the first The length of road segments in graded roads that are severely congested; For the first The total length of graded road sections.
[0073] Step 2: Calculate the vehicle kilometers on each road level, the total vehicle kilometers of the overall road network, and the weighting factors for each road operation level, as shown in equations 4.6-2 to 4 below.
[0074] Equation 4.6-2 Equation 4.6-3 Equation 4.6-4 In the formula, For the first Vehicle mileage on graded roads; For the first Number of taxis on graded roads; For the first The total length of graded road sections; The total vehicle kilometers of the road network; For the first Weighting factors for graded roads.
[0075] Step 3: Using the proportion of vehicle-kilometers of roads of each level as weights, calculate the overall road network's severely congested mileage ratio, as shown in Formula 4.6-5 below.
[0076] Equation 4.6-5 In the formula, The ratio of severely congested road network mileage; For the first Weighting factors for graded roads; For the first The ratio of severely congested mileage on graded roads.
[0077] Step 4: Based on the conversion relationship between the severe congestion mileage ratio of the road network and the Traffic Performance Index (TPI), convert it into TPI, as shown in Table 8 below.
[0078] Table 8. Relationship between the ratio of severely congested road network mileage and traffic operation index Based on the characteristics of taxi trajectory data, a road traffic status recognition model based on FCM is established using fields such as speed and recording time within the data. The basic idea of this model is as follows: Figure 5 As shown: First, representative road segments from each level were selected as study segments. Traffic flow parameters were statistically analyzed at 10-minute intervals. The parameters with the highest correlation to key parameters were selected using the Pearson coefficient to form a two-parameter system for identifying road traffic conditions. Maximum and minimum value normalization was performed using Equation 4.7-1 to map the parameter values to the [0,1] interval, and the number of clusters in the model was set. Fuzzy factors Convergence threshold Iteration threshold This embodiment categorizes road traffic conditions into five types: smooth traffic, mostly smooth traffic, light congestion, moderate congestion, and severe congestion. The objective function is calculated iteratively using the FCM clustering algorithm. Membership degree of each data point ,until The cluster centers are output and inverse normalization is performed. Based on the results, the road traffic status boundaries are defined and the road operation index is calculated to obtain the morning and evening peak time periods.
[0079] Equation 4.7-1 In the formula, The result is the data after normalization, and its range is between [0,1]. This is the original data; The minimum value in the dataset; This represents the maximum value in the dataset.
[0080] V. Identifying Road Network Congestion Areas Based on 3D DBSCAN Model This embodiment clusters road segments based on speed rather than density. However, traditional clustering methods require pre-determining the number of clusters, which may not accurately reflect the actual situation, and they can only converge to a local optimum, with performance heavily influenced by the initial centroids. Therefore, this embodiment, considering road connectivity, utilizes the DBSCAN extension concept to simulate the congestion spread in a real road network for algorithm reconstruction, achieving 3D clustering. The core idea is to merge scattered congested road segments through a spreading mechanism to achieve dimensionality reduction, thereby better identifying congestion hotspots. The specific algorithm flowchart is as follows. Figure 6As shown.
[0081] Step 1: Initialization. Sample point dataset. All points within the cluster are marked as unvisited; cluster list Used to store clusters Sample points; cluster labels ; Speed threshold =0; Road operation level With neighborhood velocity threshold dictionary Based on the road operation level of the seed point Decision, when hour, ,when hour, ,when hour, ,when hour, (Refer to Table 4.8-5). Input dataset Road network Neighborhood radius threshold Neighborhood path threshold The neighborhood density threshold for the minimum number of neighboring points required to form a core point. (Based on the needs of the spread, this point itself serves as a bridge connecting the upper and lower points.)
[0082] Step 2: Determine seed points. Filter out specific road traffic status levels. The sample points are used as candidate seed point datasets ,from Randomly select an unvisited, unclassified sample point And mark it as visited, with a dot Threshold for the radius of the circle's center neighborhood Establish a buffer zone for the radius, and calculate the distance from the sample points within the buffer zone to the center point. shortest path Filter out those that meet the neighborhood path threshold The sample points are used as candidate neighborhoods ,like Then the marker point For noisy points, return to the previous step from the candidate seed point dataset. Randomly select an unvisited, unclassified point ;like Then the point Add to this cluster class In the middle, that is, the list of clusters. Add points According to the point Road operation level From the dictionary Obtain the neighborhood velocity threshold of this cluster class. ,mark For untraversed regions, start from the candidate neighborhood. Choose one that is not yet visited and has the shortest path distance to the center of the circle. Minimum sample point Mark it as traversed and add it to the cluster list. In the middle, calculate the list of cluster classes. inner field mean ,like Then the point From the list of clusters and candidate neighborhood Remove from the middle, if after removal Then the point Mark as noise points and remove from cluster list Remove and return to the candidate point dataset. Randomly select an unvisited point If removed Then return to the previous step from the candidate neighborhood. Select the untraversed circle and the circle center. shortest path distance The smallest point .like Then return directly to the previous step from the candidate neighborhood. Select untraversed points and their intersection with the center of the circle. shortest path distance smallest Repeat the above operation until candidate neighborhoods are reached. The neighborhood of all sample points within the range has been traversed. Exit the loop and set the point Mark as seed point and seed point neighborhood (Besides yourself) join the spread queue middle.
[0083] Step 3: Spread and expand. Similar to Step 1, from the spread list... Randomly select unvisited and unclassified sample points And mark it as visited, with a dot Threshold for the radius of the circle's center neighborhood Establish a buffer zone for the radius, and calculate the distance from the sample points within the buffer zone to the center point. shortest path Filter out those that meet the neighborhood path threshold The sample points are used as candidate neighborhoods ,like Then the marker point For the boundary point, return to the previous step from the spread list. Randomly select unvisited and unclassified sample points ;like Then mark For untraversed regions, start from the candidate neighborhood. Choose an untraversed circle that is adjacent to the center. shortest path distance Minimum sample point Mark it as traversed and add it to the cluster list. In the middle, calculate the list of cluster classes. inner field mean ,like Then the point From the list of clusters and candidate neighborhood Remove from the middle, if after removal Then the point Mark as boundary point, return to the spread list Randomly select unvisited and unclassified sample points If removed Then return to the previous step from the candidate neighborhood. Select the untraversed circle and the circle center. shortest path distance The smallest point .like Then return directly to the previous step from the candidate neighborhood. Select the untraversed circle and the circle center. shortest path distance The smallest point Repeat the above operation until candidate neighborhoods are reached. The neighborhood of all sample points within the range has been traversed. Exit the loop and set the point Mark it as the core point, and put the core point neighborhood (Besides yourself) join the spread queue In the middle. Repeat step three until the queue spreads. All sample points in the loop have been visited, so exit the loop.
[0084] Step 4: Iterate through the loop. As the queue spreads... When all sample points in the cluster have been visited, the cluster has finished spreading, and a new cluster is created: If the candidate seed point dataset If not all visits are made, return to step two: From Randomly select an unvisited, unclassified sample point Candidate seed point dataset All have been visited. ,like Then return to step two: filter out the specific road traffic status level. The sample points are used as candidate seed point datasets ;like Then output the cluster classes. Cluster list gather.
[0085] The propagation condition is that the sample point has not been visited. When a sample point belongs to two classes with the same traffic status level, the partitioning is based on a priority principle. If a sample point belongs to two clusters with different traffic status levels, the higher-level cluster is prioritized, so the sample point will be preferentially included in the higher-level cluster. In this algorithm, cluster members are dynamically adjusted to avoid interference from noise points. The road operation level is bound to the neighborhood speed threshold to achieve dynamic parameter adjustment. At the same time, road network connectivity and traffic status are considered to improve the practical significance of the clustering results.
[0086] VI. Based on the reconstructed DBSCAN 3D clustering recognition model Based on the reconstructed DBSCAN 3D clustering algorithm, the road network data containing the traffic status levels of each road in the previous part needs to be preprocessed. The road network data is then projected to obtain the projected coordinate system: WGS_1984_UTM. The length of each road segment is calculated, and a segment threshold is set. All road segments with a length greater than the road segment threshold within the road network are divided at equal step intervals according to the threshold. The segmented road network data are then reassigned with ID and length. The midpoints of road segments whose road traffic status is not Nan are extracted as the sample set. The attribute structure of the sample set is shown in Table 9 below: Table 9. Attribute Structure of Sample Set ID: Unique identifier for the midpoint, inherited from the segment ID after segmentation. This field can be used to quickly project the point onto the segment; Run_LV: Road operation level; Opera_LV: Road traffic status level; Avg_Speed: Average speed of the segment; LON, LAT: Latitude and longitude of the midpoint (unit: m).
[0087] The algorithm outputs clustered sample points, but it cannot directly identify congested areas of the road network. It is still necessary to cluster road segments using a point-to-line mapping method. The sample set... The DBSCAN 3D clustering algorithm is used to reconstruct the clusters for each traffic status level. These clusters are then connected to road segments via cluster sample point IDs to obtain the clustered road network data. Seed points connected to road segments form seed road segments, core points form core road segments, and boundary points form boundary road segments. Following a traffic status level from highest to lowest, for road segments of the same category, the shortest path between the seed point and its corresponding core point is searched based on the seed road segment's location. If a road segment is not classified within the shortest path, it is assigned to that cluster; otherwise, it is assigned to the higher-level cluster based on its traffic status. This process is repeated for all core road segments, searching for the shortest path from each core segment to a boundary road segment, and then clustering the core segments again using the same method as the seed road segments. This process is repeated until all core road segments have been traversed, and the resulting list of road segment clusters is output. The model for identifying congested areas in the road network based on the reconstructed DBSCAN 3D clustering is as follows: Figure 7 As shown.
[0088] The reconstructed DBSCAN 3D clustering algorithm can cluster road segments at or above the current traffic condition level based on traffic status restrictions. After clustering, sample points are mapped to road segments. Road segments grouped into the same category are connected by the shortest path. Road segments on the shortest path that are not classified or have a lower traffic condition than the current category are also grouped into that category, with severe congestion having the highest priority. The traffic condition level of the cluster is determined by the traffic condition of the seed point. This model can not only merge congested road segments concentrated within a certain range, but also identify specific congestion areas and quantities that are consistent with the actual situation.
[0089] DBSCAN clustering algorithm uses a neighborhood radius threshold. and neighborhood density threshold The values of these parameters are extremely sensitive, and the clustering results are significantly affected by two key parameters. The model in this embodiment effectively avoids the influence of these parameters. In the preprocessing stage, road segments are divided with equal step sizes according to the road segment threshold. The neighborhood radius threshold is set slightly larger than the road segment threshold. Road segments are obtained by connecting two endpoints with straight lines, so this condition will definitely propagate congested road segments. The model is updated in real time through its propagation mechanism. The value-based approach can effectively simulate the buffering effect of congested traffic entering unobstructed road sections. When a long unobstructed road section is divided equally, the longer the unobstructed road, the more midpoints there are. The calculation is performed after adding clusters. The value drops to the neighborhood speed threshold, thus simulating the buffering situation of congested traffic entering unobstructed road sections. Set the neighborhood path threshold. ,in, A value greater than or equal to 1 indicates that the shortest path from a sample point within the neighborhood to the center of the circle is always greater than or equal to the neighborhood radius threshold. ,when If the condition is such that the road is not connected, then it can be determined that the neighborhood density threshold is not sufficient; the set neighborhood density threshold only needs to satisfy the purpose of spreading congested road sections, that is... This point acts as a bridge connecting the two points above and below, extending the congested section of the road outwards.
[0090] In summary, this embodiment, based on the HMM map matching algorithm, considers the actual situation of the road network in the study area, adds a directional factor, reconstructs the observation probability model, and performs large-scale road network matching. Introducing a map matching algorithm evaluation system, comparative analysis reveals that when the buffer radius is 50m and the sampling interval is 5s, the average matching accuracy of this embodiment's algorithm reaches 98.37%, which is 4.13% higher than the HMM algorithm, but the total time consumption of this embodiment's algorithm is the longest. When the sampling interval is 15s, the accuracy of this embodiment's algorithm is close to 94%, and the time consumption is less. When the sampling interval exceeds 15s, the matching accuracy of this embodiment drops to around 90%, and the time gain decreases. Taking all factors into consideration, this embodiment selected a sampling interval of 15 seconds and a buffer radius of 50 meters to conduct large-scale map matching based on the HMM model before and after reconstruction. Forty randomly selected matching trajectories were compared, revealing that the average accuracy of the reconstructed HMM matching algorithm was improved by 7.50% compared to the algorithm before reconstruction, the average proportion of extra matched road segments decreased by 1.44%, and the proportion of unmatched road segments decreased by 7.50%, with a significant improvement in stability compared to the original HMM algorithm. The algorithm in this embodiment has a similar average matching accuracy curve to DeepMM, and after updating the direction weight model, it shows a more significant optimization compared to IVMM, with the effect becoming more pronounced with a larger sampling interval. The algorithm in this embodiment exhibits more stable matching efficiency when encountering more complex roads. When the positioning point matches a single-line road with bidirectional traffic, the HMM map matching algorithm shows a large number of incorrect matching segments indicating reverse traffic, while the reconstructed HMM algorithm has stronger direction recognition capabilities, enabling it to maintain higher matching accuracy even in more complex road network environments.
[0091] Those skilled in the art will recognize that the units of the various examples described in connection with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components of each example have been generally described in terms of function in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0092] In the embodiments provided by the present invention, it should be understood that the division of units is only a logical functional division. In actual implementation, there may be other division methods, such as multiple units can be combined into one unit, one unit can be split into multiple units, or some features can be ignored.
[0093] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0094] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0095] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be covered within the scope of the claims and specification of the present invention.
Claims
1. A method for identifying urban road traffic status based on taxi trajectory big data, characterized in that, Includes the following steps: The first step is to preprocess taxi trajectory data and road network data, calculate the due north azimuth of each road segment, select the weight of directional factors by judging road type, construct a new observation probability model, and after selecting the radius threshold, use the reconstructed HMM algorithm to perform large-scale map matching on the road network of the study area. The second step is to statistically analyze the traffic flow parameters of each level of road segment after equal duration, select candidate parameters that are strongly correlated with the key parameters, perform FCM clustering to identify the boundaries between the traffic states of each road, extract the road segments in the road network that are in a state of severe congestion, and calculate the traffic operation index. The third step involves preprocessing the road network and extracting midpoints through equidistant segmentation. The DBSCAN 3D clustering algorithm is reconstructed using the traffic status boundary as the neighborhood speed threshold and the shortest path as the neighborhood path threshold. The highest-level traffic status is used as the seed point for expansion and spread. The resulting clustering results map to the loop network, and similar road segments are connected to form congested areas.
2. The method for identifying urban road traffic status based on taxi trajectory big data according to claim 1, characterized in that, The process of using the reconstructed HMM algorithm to perform large-scale map matching of the road network in the study area is as follows: Constructing a buffer zone: Construct a buffer zone with the positioning point as the center and the maximum error of the positioning device as the radius. The road segments within the buffer zone are candidate road segments. Mapping candidate points: Find a point on the candidate road segment that minimizes the distance between two points; this point is called a candidate point. Obtaining the shortest distance between candidate points: The relationship between the location point and the candidate point is many-to-many (N:M). It is not possible to directly calculate the shortest path between the candidate points of two adjacent location points. It is necessary to obtain the distance between the upstream and downstream endpoints of the road segment where the candidate point is located to calculate the shortest path distance between the two adjacent points. Constructing the HMM model: The joint distributed model of HMM is shown in the following equation: In the formula, It is a joint distributed HMM; It is the initial state probability; It is the probability of observation; It is the transition probability; Construct the observation probability model as shown in the following equation: In the formula, Distance factor for observation probability; For positioning points; Candidate road sections; for and Great circle distance between them; The standard deviation of the measured values at the positioning point; pi (π) Natural constant ; Construct the state transition probability model as follows: In the formula, for Great circle distance between them; for The shortest path distance between; for and The absolute value of the difference is taken as the median; natural constant ; The Viterbi algorithm is used to solve the HMM model to obtain the matching points.
3. The method for identifying urban road traffic status based on taxi trajectory big data according to claim 2, characterized in that, The specific steps to obtain the shortest distance between candidate points are as follows: Step 1: The two points are candidate points at two different times. for , Distance from the point to the upstream node for Distance from downstream node, via = judge Dot and If the points are on the same road segment, proceed to step two; otherwise, proceed to step three. Step Two: If The path length is 0; otherwise, output... Path output ,Finish; Step 3: Judgment , To determine if a path is directly connected, check the topology table to see if the road segments are adjacent. If so, output the path length. Path output If the above steps are not completed, proceed to step four. Step 4: Calculation , Shortest path to a point and length Path length output Path output The steps are now complete.
4. The method for identifying urban road traffic status based on taxi trajectory big data according to claim 1, characterized in that, Traffic flow parameters include at least average speed, speed standard deviation, number of matching points, average stopping time, and taxi traffic volume.
5. The method for identifying urban road traffic status based on taxi trajectory big data according to claim 1, characterized in that, The calculation process for the traffic operation index is as follows: Step 1: Set statistical intervals, calculate the average speed of each road segment within the statistical intervals, classify the traffic conditions of each road segment according to the road traffic condition level classification boundaries, filter out the road segments in a severely congested state, and calculate the severely congested mileage ratio of each road level using the following formula: In the formula, For the first The ratio of severely congested mileage on graded roads; For the first The length of road segments in graded roads that are severely congested; For the first The total length of graded road sections; Step 2: Calculate the vehicle-kilometers on roads of each grade, the total vehicle-kilometers of the overall road network, and the weighting factors for each road's operating grade, as shown in the following formula: In the formula, For the first Vehicle mileage on graded roads; For the first Number of taxis on graded roads; For the first The total length of graded road sections; The total vehicle kilometers of the road network; For the first Weighting factors for graded roads; Step 3: Using the proportion of vehicle-kilometers for each road class as weights, calculate the overall road network's severely congested mileage ratio, as shown in the following formula: In the formula, The ratio of severely congested road network mileage; For the first Weighting factors for graded roads; For the first The ratio of severely congested mileage on graded roads; Step 4: Based on the pre-defined conversion relationship between the ratio of severely congested mileage in the road network and the traffic operation index, convert the ratio of severely congested mileage in the road network into the traffic operation index.
6. The method for identifying urban road traffic status based on taxi trajectory big data according to claim 1, characterized in that, The specific process for reconstructing the DBSCAN 3D clustering algorithm is as follows: Step 1: Initialization: Sample point dataset All points within the cluster are marked as unvisited; cluster list Used to store clusters Sample points; cluster labels ; Speed threshold =0; Road operation level With neighborhood velocity threshold dictionary Based on the road operation level of the seed point Decision, when hour, ,when hour, ,when hour, ,when hour, Input dataset Road network Neighborhood radius threshold Neighborhood path threshold The neighborhood density threshold for the minimum number of neighboring points required to form a core point. ; Step 2: Determine seed points: Filter out specific road traffic status levels The sample points are used as candidate seed point datasets ,from Randomly select an unvisited, unclassified sample point And mark it as visited, with a dot Threshold for the radius of the circle's center neighborhood Establish a buffer zone for the radius, and calculate the distance from the sample points within the buffer zone to the center point. shortest path Filter out those that meet the neighborhood path threshold The sample points are used as candidate neighborhoods ,like Then the marker point For noisy points, return to the previous step from the candidate seed point dataset. Randomly select an unvisited, unclassified point ;like Then the point Add to this cluster class In the middle, that is, the list of clusters. Add points According to the point Road operation level From the dictionary Obtain the neighborhood velocity threshold of this cluster class. ,mark For untraversed regions, start from the candidate neighborhood. Choose one that is not yet visited and has the shortest path distance to the center of the circle. Minimum sample point Mark it as traversed and add it to the cluster list. In the middle, calculate the list of cluster classes. inner field mean ,like Then the point From the list of clusters and candidate neighborhood Remove from the middle, if after removal Then the point Mark as noise points and remove from cluster list Remove and return to the candidate point dataset. Randomly select an unvisited point If removed Then return to the previous step from the candidate neighborhood. Select the untraversed circle and the circle center. shortest path distance The smallest point ;like Then return directly to the previous step from the candidate neighborhood. Select untraversed points and their intersection with the center of the circle. shortest path distance smallest Repeat the operation until candidate neighborhoods are reached. The neighborhood of all sample points within the range has been traversed. Exit the loop and set the point Mark as seed point and seed point neighborhood (Besides yourself) join the spread queue middle; Step 3: Spread and expand: From the spread list Randomly select unvisited and unclassified sample points And mark it as visited, with a dot Threshold for the radius of the circle's center neighborhood Establish a buffer zone for the radius, and calculate the distance from the sample points within the buffer zone to the center point. shortest path Filter out those that meet the neighborhood path threshold The sample points are used as candidate neighborhoods ,like Then the marker point For the boundary point, return to the previous step from the spread list. Randomly select unvisited and unclassified sample points ;like Then mark For untraversed regions, start from the candidate neighborhood. Choose an untraversed circle that is adjacent to the center. shortest path distance Minimum sample point Mark it as traversed and add it to the cluster list. In the middle, calculate the list of cluster classes. inner field mean ,like Then the point From the list of clusters and candidate neighborhood Remove from the middle, if after removal Then the point Mark as boundary point, return to the spread list Randomly select unvisited and unclassified sample points If removed Then return to the previous step from the candidate neighborhood. Select the untraversed circle and the circle center. shortest path distance The smallest point ;like Then return directly to the previous step from the candidate neighborhood. Select the untraversed circle and the circle center. shortest path distance The smallest point Repeat the operation until candidate neighborhoods are reached. The neighborhood of all sample points within the range has been traversed. Exit the loop and set the point Mark it as the core point, and put the core point neighborhood (Besides yourself) join the spread queue Middle; Repeat step three until the spread queue continues. All sample points in the loop have been visited; exit the loop. Step 4: Iterate through the loop: as the queue spreads When all sample points in the cluster have been visited, the cluster has finished spreading, and a new cluster is created: If the candidate seed point dataset If not all visits are made, return to step two: From Randomly select an unvisited, unclassified sample point Candidate seed point dataset All have been visited. ,like Then return to step two: filter out the specific road traffic status level. The sample points are used as candidate seed point datasets ;like Then output the cluster classes. Cluster list gather.
7. The method for identifying urban road traffic status based on taxi trajectory big data according to claim 1, characterized in that, The process of obtaining congested areas is as follows: sample set The DBSCAN 3D clustering algorithm is used to reconstruct the clusters for each traffic status level. These clusters are then connected to road segments via cluster sample point IDs to obtain the clustered road network data. Seed points connected to road segments form seed road segments, core points form core road segments, and boundary points form boundary road segments. Following a traffic status level from highest to lowest, for road segments of the same category, the shortest path between the seed point and its corresponding core point is searched based on the seed road segment's location. If a road segment is not classified within the shortest path, it is assigned to that cluster; otherwise, it is assigned to the higher-level cluster based on its traffic status. All core road segments are traversed, and the shortest path from each core segment to a boundary road segment is searched. The core road segments are then clustered again using the same method as the seed road segments. This process is repeated until all core road segments have been traversed, and the resulting list of road segment clusters is output.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein, when the program is executed, it controls the device containing the computer-readable storage medium to perform the urban road traffic status recognition method based on taxi trajectory big data as described in any one of claims 1 to 7.
9. A processor, characterized in that, The processor is used to run a program, wherein the program executes the urban road traffic status recognition method based on taxi trajectory big data as described in any one of claims 1 to 7.