Scenic spot vehicle behavior identification method based on trip chain state machine and unsupervised clustering

By constructing a travel chain state machine and using unsupervised clustering, the problems of low vehicle recognition accuracy and insufficient semantic expression ability in scenic areas are solved. This enables continuous characterization and high-precision recognition of vehicle behavior, and is applicable to scenic areas with different geographical structures and traffic scales, providing reliable traffic management data support.

CN121963494AActive Publication Date: 2026-05-01GUANGZHOU UNIVERSITY
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGZHOU UNIVERSITY
Filing Date
2026-04-03
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing vehicle recognition methods for scenic areas are ill-suited to the traffic management needs of smart scenic areas. They suffer from low accuracy in vehicle behavior recognition, weak semantic expression capabilities, and insufficient versatility. They are unable to depict the continuous traffic behavior of vehicles within scenic areas and are difficult to adapt to the analysis needs of different geographical structures, traffic scales, and data scales.

Method used

We employ a method based on travel chain state machine and unsupervised clustering. By constructing a vehicle travel chain, combining the travel chain state machine model and unsupervised clustering algorithm, we build vehicle behavior feature vectors, perform feature optimization and adaptive weighting, execute anomaly detection and multi-round clustering analysis, and obtain stable vehicle behavior clustering results.

Benefits of technology

It achieves a complete characterization of continuous vehicle behavior, improves recognition accuracy and temporal consistency, reduces reliance on manual rules, enhances the adaptability of the method and the stability of recognition results, and can be flexibly applied in different scenic areas, providing reliable traffic management data support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963494A_ABST
    Figure CN121963494A_ABST
Patent Text Reader

Abstract

The invention specifically discloses a scenic spot vehicle behavior identification method based on a trip chain state machine and unsupervised clustering, and relates to the technical field of intelligent traffic. The method comprises the following steps: firstly, judging vehicle access states according to passing data of scenic spot checkpoints, grouping and sequencing, and constructing three types of trip chains; dividing the vehicles into large vehicles, medium vehicles and small vehicles, aggregating and analyzing travel chain data of each vehicle type, and constructing vehicle behavior feature vectors; obtaining a weighted vehicle trip chain feature vector after feature optimization, standardization and adaptive weighting; after the abnormal trip chain is detected and isolated, an initial clustering result of vehicle behaviors is obtained, and the optimal clustering number is determined; obtaining a stable vehicle behavior clustering result through stability evaluation parameter adjustment; and finally, outputting a final recognition result according to the distribution difference of the vehicle type clustering centers. The method improves the recognition accuracy and adaptability, provides data support for fine management of scenic spot traffic, and can also be applied to traffic scenes with clear access boundaries.
Need to check novelty before this filing date? Find Prior Art

Description

A Scenic Area Vehicle Behavior Recognition Method Based on Trip Chain State Machine and Unsupervised Clustering Technical Field

[0001] This invention relates to the field of intelligent transportation technology, and in particular to a method for recognizing vehicle behavior in scenic areas based on travel chain state machines and unsupervised clustering. Background Technology

[0002] With the deep integration of smart scenic area construction and intelligent transportation systems, scenic area traffic management is gradually transforming from the traditional model of manual statistics and experience-based judgment to a model of refined analysis of multi-source data relying on equipment such as checkpoints and video monitoring. Various traffic sensing devices can continuously collect core information such as vehicle travel time, location, and frequency, providing a solid data foundation for analyzing vehicle behavior and assessing traffic conditions in scenic areas, and laying the data support for achieving intelligent and refined management and control of scenic area traffic.

[0003] Current methods for vehicle identification and classification in scenic areas are still insufficient to meet the traffic management needs of smart scenic areas, exhibiting significant technical shortcomings. Existing methods often rely solely on the time, location, or frequency of a vehicle's single passage through a checkpoint as a classification criterion, neglecting the complete travel chain characteristics formed by vehicles across multiple time periods and nodes. This fails to characterize the continuous travel behavior of vehicles within the scenic area, making it difficult to reflect real traffic behavior patterns and thus limiting the accuracy of vehicle behavior identification.

[0004] Meanwhile, existing technologies suffer from insufficient semantic representation of vehicle behavior and poor method versatility. On the one hand, rule-based discrimination methods, based on human experience, struggle to effectively distinguish between various behavior types such as tourist vehicles, passing vehicles, resident vehicles, and park work vehicles, exhibiting weak adaptability to complex scenarios. On the other hand, key steps such as travel chain construction, feature extraction, and behavior classification are fragmented, lacking a structured and reusable unified analysis process. This makes it difficult to apply these methods across scenic areas with different geographical structures and traffic volumes, and also fails to meet the analysis needs of varying data scales. Therefore, there is an urgent need for an automatic vehicle behavior recognition method in scenic areas that integrates temporal features of vehicle travel chains with unsupervised learning to address the pain points of existing technologies. Summary of the Invention

[0005] The purpose of this invention is to propose a scenic area vehicle behavior recognition method based on travel chain state machine and unsupervised clustering, which solves the problems of low accuracy, weak semantic expression ability and insufficient versatility of existing scenic area vehicle behavior recognition methods.

[0006] To achieve the above objectives, this invention proposes a method for identifying vehicle behavior in scenic areas based on a travel chain state machine and unsupervised clustering. The steps are as follows: Step S1: Obtain vehicle passage record data collected from entrance and exit checkpoints within the scenic area during the statistical period, including checkpoint name, direction, license plate number, vehicle passage time, and vehicle type; perform deduplication, missing value removal, and outlier cleaning on the vehicle passage record data to form structured vehicle passage data; Step S2: Based on the checkpoint name and direction data, determine whether the vehicle behavior status is entering or leaving the scenic area, and classify the structured vehicle passage data according to the license plate number. The data is grouped and sorted according to the order of vehicle passage time; Step S3: Based on vehicle entry and exit status and temporal continuity rules, three types of travel chain structures are constructed for each vehicle, including complete travel chain, entry-only chain, and exit-only chain; Step S4: Based on vehicle type label data, vehicles are initially classified into small cars, medium-sized cars, and large cars; Based on the travel chain state machine model, the travel chain data of small cars, medium-sized cars, and large cars are aggregated and analyzed to construct nine behavioral feature vectors, forming vehicle behavioral feature vectors; Step S5: The travel chain data constructed for each vehicle type is further analyzed. The feature vectors are optimized, and the optimized behavioral feature vectors are standardized to transform the features of each dimension to a uniform scale space. Step S6: The contribution of each behavioral feature in the standardized behavioral feature vectors is evaluated, and each feature is adaptively weighted to obtain a weighted vehicle travel chain feature vector. Step S7: Before performing unsupervised clustering analysis, anomaly detection and isolation are performed on the weighted vehicle travel chain feature vectors. Step S8: After isolating abnormal travel chains, K-M clustering is applied to small car samples, medium-sized car samples, and large car samples respectively. The EANS clustering algorithm performs vehicle-type clustering analysis to obtain initial clustering results for vehicle behavior, and determines the optimal number of clusters for each vehicle type based on the silhouette coefficient, CH index, or intra-cluster sum of squares. Step S9: To address the reliability issue of the initial clustering results, the clustering results are evaluated for stability and parameters are adjusted until stable vehicle behavior clustering results are obtained. Step S10: Based on the distribution differences of cluster centers for each vehicle type in terms of dwell characteristics, frequency characteristics, and time period characteristics, the clusters are interpreted semantically to map different clustering results to the corresponding vehicle behavior types and output the final recognition results.

[0007] Preferably, in step S3, each vehicle travel chain is represented as follows: ;in, Let be the statistical values ​​of the i-th vehicle travel chain over k time intervals. Let j be the passing position for the j-th pass. Let be the passage time for the j-th passage. , The total number of passes; the passage time for the j-th pass. satisfy: ;in, A preset threshold is set for the time interval; when the time interval between two adjacent passage records exceeds the preset threshold... At that time, the travel process is disconnected, and a new vehicle travel chain is generated.

[0008] Preferably, the state representation corresponding to the j-th passage behavior in the vehicle travel chain is as follows: ; ;in, Let j be the state corresponding to the j-th travel behavior in the vehicle travel chain. For state functions, To get into the state, In a stationary state, In the "out of state" state, It is in a cross-day state.

[0009] Preferably, in step S4, the nine behavioral characteristics include: stay characteristics: average stay duration and standard deviation of stay duration; frequency characteristics: number of travel chains, number of travel days, and whether there is a cross-day chain; time period characteristics: average entry time and average departure time; behavioral pattern characteristics of the origin and destination: total number of snapshots; the vehicle behavior feature vector is represented as follows: ;in, Let be the multi-dimensional behavioral feature vector of the i-th vehicle travel chain. Let m be the value of the travel chain of the i-th vehicle on the m-th behavioral feature dimension. Multidimensional behavioral characteristics include vehicle dwell time characteristics, travel frequency characteristics, entry and exit time distribution characteristics, cross-day behavior characteristics, and checkpoint traffic intensity characteristics.

[0010] Preferably, in step S5, the behavioral feature vector is subjected to feature optimization processing, specifically: first, Spearman correlation analysis is performed to remove highly correlated features, and then feature ablation experiments are conducted to evaluate the influence of features on the clustering result structure, optimize the feature dimension combination, and obtain the optimized behavioral feature vector.

[0011] Preferably, in step S6, the weighted vehicle travel chain feature vector is represented as follows: ;in, Let be the feature vector of the i-th vehicle travel chain after adaptive weighting. The weight of the m-th feature; feature weight satisfy: ;in, For the feature contribution evaluation function, This represents the contribution value of the m-th feature.

[0012] Preferably, in step S7, the characteristic deviation of the vehicle travel chain is defined as: ;in, Let be the deviation of the i-th vehicle travel chain in the feature space. Let be the reference center vector in the feature space. For vector norm; when the deviation satisfy: The vehicle's travel chain was determined to be an abnormal travel chain; among which, This is the threshold for anomaly detection.

[0013] Preferably, in step S8, R clustering analyses are performed on the same feature dataset, and the consistency of the clustering results is calculated using a consistency formula; the consistency formula is as follows: ;in, As a consistency indicator, For the first Clustering results under each round For the first Clustering results under each round The similarity function between the two clustering results. , R is the round number of the clustering repeat experiment, where R is an integer; when the consistency index Less than the preset threshold If necessary, the clustering parameters are adjusted and the clustering analysis is re-executed until a stable clustering result of vehicle behavior is obtained.

[0014] Preferably, in step S10, the vehicle types include tourist private cars, taxis, resident vehicles, staff vehicles, transit vehicles, tourist buses, and scenic area shuttle buses.

[0015] Therefore, this invention proposes a method for identifying vehicle behavior in scenic areas based on travel chain state machine and unsupervised clustering. Its beneficial effects are as follows: (1) This invention constructs a vehicle travel chain and combines it with a travel chain state machine to complete the state representation of the travel process, which fully depicts the continuous behavior process of vehicles in the scenic area from entering, staying, leaving to traveling across days. It avoids the problem of behavior fragmentation caused by relying solely on single-passage information and greatly improves the accuracy and temporal consistency of vehicle behavior identification.

[0016] (2) This invention relies on unsupervised clustering analysis to mine the inherent patterns of vehicle behavior. It does not require preset vehicle category labels and manual judgment rules. At the same time, through a unified behavior modeling and recognition process, it breaks through the limitations of scenic area geographical structure, traffic scale and data scale. It can be flexibly applied in different scenic areas and can also be applied to traffic scenarios with clear entry and exit boundaries such as airports and railway hubs.

[0017] (3) This invention uses feature adaptive weighting, abnormal travel chain isolation and multi-round clustering stability assessment to ensure that the vehicle behavior recognition results remain highly consistent across multiple time scales and batches of data updates. The output vehicle behavior types and related feature analysis results can provide reliable data support for the judgment of traffic congestion during peak hours in scenic areas, the assessment of operational status and the construction of congestion prediction models, thus helping to achieve refined and intelligent management of traffic in scenic areas.

[0018] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0019] Figure 1 is a flowchart of a scenic area vehicle behavior recognition method based on travel chain state machine and unsupervised clustering according to the present invention; Figure 2 is a schematic diagram of the Spearman correlation matrix of small cars in an embodiment of the present invention; Figure 3 is a schematic diagram of the Spearman correlation matrix of medium-sized cars in an embodiment of the present invention; Figure 4 is a schematic diagram of the Spearman correlation matrix of large cars in an embodiment of the present invention; Figure 5 is a schematic diagram of the relationship between the contour coefficient and K value of small cars in an embodiment of the present invention; Figure 6 is a schematic diagram of the relationship between the contour coefficient and K value of medium-sized cars in an embodiment of the present invention; Figure 7 is a schematic diagram of the relationship between the contour coefficient and K value of large cars in an embodiment of the present invention; Figure 8 is a PCA scatter plot of small cars in an embodiment of the present invention; wherein, principal component 1 is a feature reflecting the differences in the number of active days, stay characteristics, travel frequency and entry and exit time of small cars. The feature combination, with principal component 2 being a feature combination reflecting the time-period characteristics, dwell dispersion, and cross-day characteristics of small vehicles; Figure 9 is a PCA scatter plot of medium-sized vehicles in an embodiment of the present invention; wherein, principal component 1 is a feature combination reflecting the number of active days, dwell characteristics, travel frequency, and differences in entry and exit times of medium-sized vehicles, and principal component 2 is a feature combination reflecting the time-period characteristics, dwell dispersion, and cross-day characteristics of medium-sized vehicles; Figure 10 is a PCA scatter plot of large vehicles in an embodiment of the present invention; wherein, principal component 1 is a feature combination reflecting the number of active days, dwell characteristics, travel frequency, and differences in entry and exit times of large vehicles, and principal component 2 is a feature combination reflecting the time-period characteristics, dwell dispersion, and cross-day characteristics of large vehicles; Figure 11 is a schematic diagram of the behavior composition of different vehicle types in an embodiment of the present invention; Figure 12 is a schematic diagram of the overall vehicle behavior distribution in an embodiment of the present invention. Detailed Implementation

[0020] To make the technical solutions, advantages, and objectives of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below. The described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the described embodiments of the present invention without creative effort are within the protection scope of the present invention.

[0021] Unless otherwise defined, the technical or scientific terms used in this invention shall have the ordinary meaning as understood by one of ordinary skill in the art to which this invention pertains.

[0022] Example 1 illustrates the specific implementation process and effects of the scenic area vehicle behavior recognition method based on travel chain state machine and unsupervised clustering proposed in this invention, using actual traffic operation data from a scenic area.

[0023] As shown in Figure 1, this invention provides a method for identifying vehicle behavior in scenic areas based on travel chain state machines and unsupervised clustering. The specific steps are as follows: Step S1: Obtain vehicle passage record data collected from entrance and exit checkpoints within the scenic area during the statistical period, including checkpoint name, direction, license plate number, vehicle passage time, and vehicle type; perform deduplication, missing value removal, and outlier cleaning on the vehicle passage record data to form structured vehicle passage data; in this embodiment, multiple fixed checkpoints deployed at the main road entrances and exits of a scenic area are selected as the source of vehicle passage data collection; the checkpoint equipment automatically generates passage records when vehicles pass through, and each record includes at least the checkpoint name, direction, license plate number, vehicle passage time, and vehicle type.

[0024] This type of checkpoint data features continuous collection, stable coverage, and high temporal accuracy, providing a relatively complete picture of vehicle entry and exit behavior within the scenic area. To ensure the accuracy of subsequent analysis, the original traffic record data underwent deduplication, missing value removal, and cleaning of obvious anomalies to create structured vehicle traffic data.

[0025] In this embodiment, a total of 48,234 original vehicle passage records were captured within the time range of 00:00:20 on August 28, 2023 to 23:55:30 on September 1, 2023. The original data fields include: capture location, lane number, license plate number, license plate color, vehicle type, and passage time. The original vehicle passage record data is preprocessed to form structured vehicle passage data, specifically including: 1. Field integrity check.

[0026] The original record's capture location, license plate number, vehicle type, and passage time fields undergo integrity checks. In this embodiment, none of the above key fields have missing values, therefore the record is not deleted due to missing values. If any of the capture location, license plate number, vehicle type, or passage time fields is empty, the record can be determined as invalid and deleted.

[0027] 2. Remove duplicate records.

[0028] Multiple records with identical "capture location, lane number, license plate number, and passage time" are identified as duplicate records, and only one of them is retained. In this embodiment, a total of 35 completely duplicate records were identified and deleted after deduplication.

[0029] 3. Standardize the time field.

[0030] The "Ever Time" field is uniformly converted to a standard timestamp format and parsed according to "year-month-day hour:minute:second"; records whose time formats cannot be parsed or are outside the statistical period range are removed. In this embodiment, the ever time of all records can be successfully parsed and all fall within the statistical period.

[0031] 4. Standardization of capture location and semantic mapping of direction.

[0032] The "Capture Location" field undergoes text standardization processing, mapping different expressions to standard checkpoint names and directional semantics. In this embodiment, capture locations include the following four categories: Checkpoint A towards the scenic area; Tunnel B towards the scenic area; Checkpoint A towards location C; Tunnel B towards location D. Records containing "towards the scenic area" or "in the direction of the scenic area" are marked as entering the scenic area; records containing "towards location C" or "towards location D" are marked as leaving the scenic area.

[0033] 5. Standardize the processing of license plate numbers.

[0034] The "License Plate Number" field is cleaned by removing leading and trailing spaces and abnormal separators, and the character format is standardized. In this embodiment, license plate numbers may be abbreviated or anonymized, and some license plates are 3, 4, 5, 6, or 7 digits long. Therefore, only records with empty license plate numbers, records consisting entirely of meaningless characters, or records that cannot be used as unique vehicle identifiers are deleted.

[0035] 6. Vehicle type cleaning and screening.

[0036] In this embodiment, vehicle types include: small cars, medium-sized cars, large cars, unknown, two-wheeled vehicles, and three-wheeled vehicles. Among them, small cars, medium-sized cars, and large cars are retained as valid motor vehicle samples for subsequent travel chain construction and behavior recognition; "unknown," "two-wheeled vehicles," and "three-wheeled vehicles" are excluded in this embodiment due to their significant differences from the behavior patterns of motor vehicles in scenic areas and will not participate in subsequent motor vehicle clustering analysis.

[0037] After the above data deduplication, missing value removal, and cleaning of obvious abnormal data, a total of 37,256 valid data records were retained, forming structured vehicle traffic data, and finally a structured vehicle traffic data table was obtained. The fields of the structured vehicle traffic data table include: standardized license plate number, standardized capture location (checkpoint name), traffic direction status (entering / leaving the scenic area), vehicle type, vehicle passage time, and vehicle passage date.

[0038] Step S2: Based on the checkpoint name and direction data, determine whether the vehicle behavior status is entering or leaving the scenic area, and group the structured vehicle passage data according to the license plate number, and sort them according to the order of vehicle passage time.

[0039] Step S3: Based on vehicle entry / exit status and temporal continuity rules, construct three types of travel chain structures for each vehicle, including complete travel chains, entry-only chains, and exit-only chains; the representation of each vehicle travel chain is as follows: ;in, Let be the statistical values ​​of the i-th vehicle travel chain over k time intervals. Let j be the passing position for the j-th pass. Let be the passage time for the j-th passage. , The total number of passes; the passage time for the j-th pass. satisfy: ;in, A preset threshold is set for the time interval; when the time interval between two adjacent passage records exceeds the preset threshold... At that time, the travel process is disconnected, and a new vehicle travel chain is generated.

[0040] The state representation corresponding to the j-th passage behavior in the vehicle travel chain is as follows: ; ;in, Let j be the state corresponding to the j-th travel behavior in the vehicle travel chain. For state functions, To get into the state, In a stationary state, In the "out of state" state, It is in a cross-day state.

[0041] Step S4: Based on vehicle type label data, vehicles are initially classified into small cars, medium cars, and large cars. Based on the travel chain state machine model, the travel chain data for small cars, medium cars, and large cars are aggregated and analyzed to construct nine behavioral feature vectors, forming the vehicle behavior feature vector. The nine behavioral features include: Dwelling characteristics: average dwell time, standard deviation of dwell time; Frequency characteristics: number of travel chains, number of travel days, existence of cross-day chains; Time period characteristics: average entry time, average departure time; Behavioral pattern characteristics of origin and destination: total number of snapshots. The vehicle behavior feature vector is represented as follows: ;in, Let be the multi-dimensional behavioral feature vector of the i-th vehicle travel chain. Let m be the value of the travel chain of the i-th vehicle on the m-th behavioral feature dimension. Step S5: Perform feature optimization processing on the behavioral feature vectors constructed for each vehicle model, and standardize the optimized behavioral feature vectors to transform the features of each dimension to a unified scale space. Specifically, the feature optimization processing of the behavioral feature vectors involves: first performing Spearman correlation analysis to remove highly correlated features, and then evaluating the influence of features on the clustering result structure through feature ablation experiments to optimize the feature dimension combination and obtain the optimized behavioral feature vectors.

[0042] As shown in Figure 2, a moderate positive correlation exists among some frequency-related indicators for small vehicles (such as the number of trip chains, the number of snapshots, and the average number of trips per day), indicating that these indicators collectively reflect the intensity of vehicle activity to some extent. Meanwhile, some time-related features and stay-related features show relatively low correlation, suggesting that they characterize vehicle behavior patterns from different dimensions. These results indicate that the overall data features do not suffer from severe collinearity, but some information overlap still exists, providing a statistical basis for subsequent feature selection or feature ablation, thus helping to improve the stability and interpretability of the clustering model.

[0043] As shown in Figure 3, some activity frequency indicators exhibit strong positive correlations, indicating that vehicle activity level is an important dimension describing the behavior of medium-sized vehicles. Conversely, the correlations between dwell time, time, and frequency characteristics are relatively weak, suggesting that these variables can supplement vehicle behavior from different perspectives. Overall, the features show some correlation while maintaining good independence, providing a reasonable feature basis for subsequent cluster analysis.

[0044] As shown in Figure 4, some activity intensity indicators exhibit strong positive correlations, while the correlation between time features and dwell time features is relatively moderate. This indicates that the behavior patterns of large vehicles are influenced by both activity frequency and time structure and dwell time behavior. Overall, the features maintain a certain correlation while also exhibiting good information complementarity. This demonstrates that the constructed behavioral feature system can effectively characterize the travel behavior of large vehicles from multiple dimensions, providing a solid data foundation for subsequent cluster analysis and behavioral pattern recognition.

[0045] Step S6: Evaluate the contribution of each behavioral feature in the standardized behavioral feature vector, and perform adaptive weighting on each feature to obtain the weighted vehicle trip chain feature vector; the weighted vehicle trip chain feature vector is represented as follows: ;in, Let be the feature vector of the i-th vehicle travel chain after adaptive weighting. The weight of the m-th feature; feature weight satisfy: ;in, For the feature contribution evaluation function, This represents the contribution value of the m-th feature.

[0046] Step S7: Before performing unsupervised clustering analysis, perform anomaly detection and isolation on the weighted vehicle travel chain feature vector; define the feature deviation of the vehicle travel chain as: ;in, Let be the deviation of the i-th vehicle travel chain in the feature space. Let be the reference center vector in the feature space. For vector norm; when the deviation satisfy: The vehicle's travel chain was determined to be an abnormal travel chain; among which, This is the threshold for anomaly detection; in this embodiment, .

[0047] Step S8: After isolating the abnormal travel chains, K-Means clustering algorithm is used to perform vehicle-type clustering analysis on small car samples, medium car samples, and large car samples respectively to obtain the initial clustering results of vehicle behavior. The optimal number of clusters for each vehicle type is determined based on the silhouette coefficient, CH index, or intra-cluster sum of squares. As shown in Figure 5, the silhouette coefficient reaches its highest value when K=3, indicating that the intra-cluster sample similarity is the highest, the inter-cluster difference is the greatest, and the clustering structure is the clearest. When K>3, the silhouette coefficient generally shows a downward trend, indicating that as the number of categories increases, the originally clear behavior patterns are over-subdivided, leading to the gradual blurring of inter-cluster boundaries. Therefore, Figure 5 proves that the behavior patterns of small cars have a relatively stable clustering structure as a whole. It also shows that using the silhouette coefficient for data-driven K value selection can effectively avoid the problem of subjectively setting the number of clusters, thereby improving the objectivity and reliability of the clustering results.

[0048] As shown in Figure 6, the silhouette coefficient reaches its maximum value when K=2, indicating that the mid-sized vehicle group mainly exhibits two relatively obvious structural divisions in terms of behavioral patterns. As K increases, the silhouette coefficient decreases significantly, indicating that further subdivision weakens inter-class differences and reduces cluster stability. This result shows that the behavioral structure of mid-sized vehicles is relatively simple and concentrated, and a small number of categories can well characterize their main behavioral patterns. It also shows that the constructed behavioral features have good cluster separability.

[0049] As shown in Figure 7, the silhouette coefficient reaches its relative maximum when K=3, indicating that the large vehicle group can be well divided into three main structural categories in terms of behavior patterns. As K continues to increase, the silhouette coefficient gradually decreases, indicating that too many category divisions will reduce the clarity of the cluster structure. This result shows that large vehicles have a relatively stable behavior type structure in the scenic area transportation system, and also proves that parameter selection based on the silhouette coefficient can effectively improve the rationality of the clustering model.

[0050] Perform R clustering analyses on the same feature dataset and calculate the consistency of the clustering results using a consistency formula; the consistency formula is as follows: ;in, As a consistency indicator, For the first Clustering results under each round For the first Clustering results under each round The similarity function between the two clustering results. , R is the round number of the clustering repeat experiment, where R is an integer; when the consistency index Less than the preset threshold If necessary, the clustering parameters are adjusted and the clustering analysis is re-executed until a stable clustering result of vehicle behavior is obtained.

[0051] As shown in Figure 8, different clusters form multiple relatively independent point cloud regions in the principal component space. Some clusters exhibit significant spatial separation, indicating that the constructed behavioral features can effectively distinguish different types of small vehicle travel patterns. Meanwhile, some clusters still have certain proximity or transitional regions, suggesting a degree of behavioral continuity or mixed characteristics within the small vehicle group. Overall, Figure 8 visually verifies the rationality of the clustering results and demonstrates that the selected features have good discriminative ability in characterizing differences in vehicle behavior.

[0052] As shown in Figure 9, the three main point clusters exhibit clear spatial separation in the principal component space, and the points within each cluster are relatively concentrated, indicating that the behavioral patterns of medium-sized vehicles have relatively clear structural differences in the feature space. The large spatial distance between different clusters indicates that the constructed behavioral features can effectively capture the differences of medium-sized vehicles in terms of dwell characteristics, travel frequency, and temporal structure, thus forming stable and interpretable clustering results.

[0053] As shown in Figure 10, different categories form three distinct regions in the principal component space, and the points within each cluster are relatively concentrated, indicating strong differences in the behavioral patterns of large vehicles. Compared to small vehicles, medium and large vehicles tend to have more clearly defined functional attributes and usage scenarios, making it easier for their behavioral characteristics to form clear structural divisions in the feature space. This further verifies that the clustering model has good stability and interpretability in identifying the behavioral patterns of large vehicles.

[0054] Step S9: To address the reliability issue of the initial clustering results, the clustering results are evaluated for stability and parameters are adjusted until stable vehicle behavior clustering results are obtained. Step S10: Based on the distribution differences of cluster centers of each vehicle type in terms of dwell characteristics, frequency characteristics, and time period characteristics, the clusters are interpreted semantically to map different clustering results to the corresponding vehicle behavior types and output the final recognition results.

[0055] Vehicle types include private cars for tourists, taxis, resident vehicles, staff vehicles, transit vehicles, tourist buses, and scenic area shuttle buses.

[0056] As shown in Figure 11, different vehicle types exhibit significant differences in their behavioral composition. Large vehicles are almost entirely composed of tour buses, accounting for nearly 100%, indicating that large vehicles primarily serve as group tourist transport vehicles within the scenic area's transportation system, exhibiting a relatively simple and stable behavioral pattern. Medium-sized vehicles are almost entirely composed of scenic area shuttle buses, also showing a highly concentrated behavioral structure. This suggests that medium-sized vehicles mainly provide shuttle and connection services within the scenic area's internal transportation system, with a relatively fixed operating mode and distinct operational characteristics. In contrast, the behavioral composition of small vehicles is significantly more complex and diverse. Small vehicles mainly consist of tourist vehicles (approximately 87%), while also including a certain proportion of resident vehicles, staff vehicles, transit vehicles, and taxis. This structure indicates that small vehicles not only encompass tourist travel but also include local resident commuting, scenic area staff daily travel, and non-scenic area-related transit traffic. Therefore, the behavioral patterns of the small vehicle group exhibit significant diversity and complexity.

[0057] As shown in Figure 11, the behavioral patterns of large and medium-sized vehicles are highly concentrated, indicating that their traffic functions are relatively clear; while small vehicles exhibit mixed characteristics of multi-source behaviors, making them the most complex vehicle group in the scenic area's transportation system. Therefore, in subsequent behavior pattern recognition and cluster analysis, it is necessary to focus on the behavior classification of small vehicles to improve the accuracy and interpretability of the vehicle behavior recognition model.

[0058] As shown in Figure 12, tourist vehicles dominate, accounting for approximately 80.1%, far exceeding other types of vehicles. This result indicates that the main source of traffic flow in the scenic area's transportation system is tourist vehicles, and their travel demands constitute the core of the traffic flow. Therefore, tourist vehicle travel behavior should be a key research focus in the process of optimizing traffic management and organization in scenic areas.

[0059] Besides tourist vehicles, shuttle buses account for approximately 6.0%, resident vehicles for approximately 5.1%, staff vehicles for approximately 3.0%, transit vehicles for approximately 2.6%, tour buses for approximately 1.8%, and taxis for approximately 1.4%. Although these vehicle types represent a relatively small proportion, they still play an important role in the scenic area's transportation system. For example, shuttle buses serve as internal connections for tourists, resident and staff vehicles reflect local transportation needs around the scenic area, while transit vehicles indicate that some traffic flow is not related to travel within the scenic area but rather traverses the scenic area's road network.

[0060] By statistically analyzing the proportions of different vehicle behavior patterns, it can be clearly seen that the scenic area's transportation system is dominated by tourist travel, while also including a certain proportion of local traffic and commercial vehicles. This result provides an important data foundation for subsequent vehicle behavior identification, traffic demand analysis, and the formulation of scenic area traffic management strategies.

[0061] Practical application of the above embodiments demonstrates that the method proposed in this invention can effectively identify various behavior types, including tourist vehicles, passing vehicles, vehicles belonging to scenic area residents, and vehicles belonging to scenic area staff, relying solely on vehicle traffic data. Compared to existing methods that classify vehicles based on single passage records or simple statistical rules, this invention avoids fragmentation of cross-day behavior through travel chains; improves the stability of clustering results through adaptive feature weighting and isolation of abnormal travel chains; and maintains good consistency of vehicle behavior identification results across different time scales and data sizes through a multi-round clustering consistency evaluation mechanism, demonstrating higher engineering applicability and application value.

[0062] It is worth noting that all contents not described in detail in this invention are existing technologies and are well known to those skilled in the art.

[0063] Therefore, this invention provides a method for identifying vehicle behavior in scenic areas based on travel chain state machines and unsupervised clustering. By combining travel chain state machine modeling with unsupervised clustering, the method achieves vehicle behavior identification in scenic areas. This not only fully characterizes the continuous travel behavior of vehicles and improves the temporal consistency and accuracy of identification, but also reduces reliance on manual rules, enhances the adaptability of the method to different scenarios, and ensures the stability and reusability of the identification results. It provides reliable data support for the refined management of traffic in scenic areas and can be applied to various traffic scenarios with clear entry and exit boundaries, such as airports and railway hubs.

[0064] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for recognizing vehicle behavior in scenic areas based on travel chain state machine and unsupervised clustering, characterized in that, The steps are as follows: Step S1: Obtain vehicle passage record data collected from entrance and exit checkpoints within the scenic area during the statistical period, including checkpoint name, direction, license plate number, vehicle passage time, and vehicle type; perform deduplication, missing value removal, and outlier cleaning on the vehicle passage record data to form structured vehicle passage data; Step S2: Based on the checkpoint name and direction data, determine whether the vehicle behavior status is entering or leaving the scenic area, and group the structured vehicle passage data according to the license plate number, and sort them according to the order of vehicle passage time; Step S3: Based on the vehicle entry and exit status and temporal continuity rules, construct three types of travel chain structures for each vehicle, including complete travel chain, entry-only chain, and exit-only chain; Step S4: Based on the vehicle type label data, perform preliminary stratification and classification of vehicles into small cars, medium-sized cars, and large cars; Based on the travel chain state machine model, travel chain data for small cars, medium-sized cars, and large cars are aggregated and analyzed separately to construct nine behavioral feature vectors, forming vehicle behavioral feature vectors. Step S5: Feature optimization processing is performed on the behavioral feature vectors constructed for each car type, and the optimized behavioral feature vectors are standardized to transform each dimension of features to a unified scale space. Step S6: The contribution of each behavioral feature in the standardized behavioral feature vector is evaluated, and each feature is adaptively weighted to obtain a weighted vehicle travel chain feature vector. Step S7: Before performing unsupervised clustering analysis, anomaly detection and isolation are performed on the weighted vehicle travel chain feature vector. Step S8: After isolating abnormal travel chains, K-Means clustering algorithm is used to perform vehicle-type clustering analysis on small car samples, medium car samples, and large car samples to obtain initial clustering results of vehicle behavior. The optimal number of clusters for each vehicle type is determined based on the silhouette coefficient, CH index, or intra-cluster sum of squares. Step S9: To address the reliability issue of the initial clustering results, the stability of the clustering results is evaluated and parameters are adjusted until stable vehicle behavior clustering results are obtained. Step S10: Based on the distribution differences of cluster centers of each vehicle type in terms of dwell characteristics, frequency characteristics, and time period characteristics, behavioral semantic interpretation is performed on the clusters, mapping different clustering results to the corresponding vehicle behavior types and outputting the final recognition results.

2. The scenic area vehicle behavior recognition method based on travel chain state machine and unsupervised clustering according to claim 1, characterized in that, In step S3, each vehicle travel chain is represented as follows: ;in, Let be the statistical values ​​of the i-th vehicle travel chain over k time intervals. Let j be the passing position for the j-th pass. Let be the passage time for the j-th passage. , The total number of passes; the passage time for the j-th pass. satisfy: ;in, A preset threshold is set for the time interval; when the time interval between two adjacent passage records exceeds the preset threshold... At that time, the travel process is disconnected, and a new vehicle travel chain is generated.

3. The scenic area vehicle behavior recognition method based on travel chain state machine and unsupervised clustering according to claim 2, characterized in that, The state representation corresponding to the j-th passage behavior in the vehicle travel chain is as follows: ; ;in, Let j be the state corresponding to the j-th travel behavior in the vehicle travel chain. For state functions, To get into the state, In a stationary state, In the "out of state" state, It is in a cross-day state.

4. The scenic area vehicle behavior recognition method based on travel chain state machine and unsupervised clustering according to claim 1, characterized in that, In step S4, the nine behavioral characteristics include: Dwelling characteristics: average dwell time and standard deviation of dwell time; Frequency characteristics: number of travel chains, number of travel days, and whether there is a cross-day chain; Time period characteristics: average entry time and average departure time; Behavioral pattern characteristics of the origin and destination: total number of snapshots; The vehicle behavior feature vector is represented as follows: ;in, Let be the multi-dimensional behavioral feature vector of the i-th vehicle travel chain. Let m be the value of the travel chain of the i-th vehicle on the m-th behavioral feature dimension. Multidimensional behavioral characteristics include vehicle dwell time characteristics, travel frequency characteristics, entry and exit time distribution characteristics, cross-day behavior characteristics, and checkpoint traffic intensity characteristics.

5. The scenic area vehicle behavior recognition method based on travel chain state machine and unsupervised clustering according to claim 1, characterized in that, In step S5, the behavioral feature vector is subjected to feature optimization processing. Specifically, Spearman correlation analysis is first performed to remove highly correlated features. Then, feature ablation experiments are conducted to evaluate the influence of features on the clustering result structure, optimize the feature dimension combination, and obtain the optimized behavioral feature vector.

6. The scenic area vehicle behavior recognition method based on travel chain state machine and unsupervised clustering according to claim 4, characterized in that, In step S6, the weighted vehicle travel chain feature vector is represented as follows: ;in, Let be the feature vector of the i-th vehicle travel chain after adaptive weighting. The weight of the m-th feature; feature weight satisfy: ;in, For the feature contribution evaluation function, This represents the contribution value of the m-th feature.

7. The scenic area vehicle behavior recognition method based on travel chain state machine and unsupervised clustering according to claim 6, characterized in that, In step S7, the characteristic deviation of the vehicle travel chain is defined as: ;in, Let be the deviation of the i-th vehicle travel chain in the feature space. Let be the reference center vector in the feature space. For vector norm; when the deviation satisfy: The vehicle's travel chain was determined to be an abnormal travel chain; among which, This is the threshold for anomaly detection.

8. The scenic area vehicle behavior recognition method based on travel chain state machine and unsupervised clustering according to claim 1, characterized in that, In step S8, R clustering analyses are performed on the same feature dataset, and the consistency of the clustering results is calculated using the consistency formula; the consistency formula is as follows: ;in, As a consistency indicator, For the first Clustering results under each round For the first Clustering results under each round The similarity function between the two clustering results. 、 R is the round number of the clustering repeat experiment, where R is an integer; when the consistency index Less than the preset threshold If necessary, the clustering parameters are adjusted and the clustering analysis is re-executed until a stable clustering result of vehicle behavior is obtained.

9. The scenic area vehicle behavior recognition method based on travel chain state machine and unsupervised clustering according to claim 1, characterized in that, In step S10, the vehicle types include tourist private cars, taxis, resident vehicles, staff vehicles, transit vehicles, tourist buses, and scenic area shuttle buses.

Citation Information

Patent Citations

  • Method for evaluating congestion risk of each travel mode by applying license plate recognition data

    CN118430237A

  • Expressway abnormal state detection method and system based on support vector machine

    CN119720023A

  • A safety management system for highway maintenance

    CN119784120A

  • Truck trip chain identification method based on trajectory data

    CN121075131A

  • Intersection abnormal vehicle trajectory identification and analysis method based on hierarchical clustering

    WO2021077761A1