A user segmentation method based on high-speed ETC charging data

By preprocessing ETC toll data and using the SOM clustering algorithm, the problem of identifying and classifying highway users was solved, enabling fast and accurate user segmentation, improving data utilization efficiency and accuracy, and supporting highway management decisions.

CN114519388BActive Publication Date: 2026-04-07SHANDONG HI SPEED COMPANY +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-30
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing technologies cannot effectively utilize highway ETC toll data for user identification and classification. Traditional methods suffer from low data quality, high cost, and low sampling rate, making it difficult to meet the needs of highway operation and management.

Method used

A user segmentation method based on ETC toll data is adopted. Through preprocessing, data cleaning and SOM clustering algorithm, the time, space and personal attribute indicators of users are extracted. The SOM clustering algorithm is used to classify highway users and identify commuting, operation, business and sporadic travel.

Benefits of technology

It enables rapid and accurate classification of highway users, providing a basis for highway planning and management, improving the amount and accuracy of data information, and supporting operational and congestion management decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114519388B_ABST
    Figure CN114519388B_ABST
Patent Text Reader

Abstract

A kind of user segmentation method based on high-speed ETC charging data, the pre-processing of highway toll data is carried out, the field information required for highway user classification is extracted, and the basic information is stored with the license plate number of highway user as key field, to form the travel basic data of highway user;The high-speed toll records of each highway user are sorted according to time, and the data is cleaned according to the abnormal state of time and space, to obtain the high-speed toll data after data cleaning;According to the cleaned data, the information of three dimensions of highway user time index, space index and personal attribute index is extracted respectively, to form the user classification evaluation index system, and the classification of highway user is completed;According to the time index and space index of highway user travel, the classification is carried out in month cycle, to identify various travels such as commuting travel, operation travel, sporadic travel and business travel.The present application has complete information and high precision, and provides basis for highway planning and construction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for identifying and classifying highway users. In particular, it relates to a user segmentation method based on highway ETC toll data. Background Technology

[0002] Highways are an integral part of urban transportation, and understanding the travel needs of highway users is crucial for highway planning and management. The "Outline for Building a Powerful Transportation Nation" sets higher requirements for highway operation management and travel services. However, traditional MTC (Manual Toll Collection) systems have limited data fields related to users, making continuous analysis of highway users impossible. Furthermore, manual surveys such as traffic surveys and questionnaires suffer from disadvantages such as long cycles, low sampling rates, and high costs, and the low data quality often fails to achieve the desired results.

[0003] With the development of information technology and infrastructure, the ETC system has been widely adopted, generating massive amounts of ETC toll data as highways operate. ETC toll data uniquely identifies users, enabling one person, one vehicle, one unique tag, and providing the possibility to identify highway users' commuting, commercial, business, and casual travel. In October 2020, the usage rate of the ETC non-stop toll collection system approached 70%, covering the majority of highway users. By mining users' travel characteristics, an opportunity has been provided for more in-depth identification and classification of highway users.

[0004] SOM is a representative semi-supervised machine learning algorithm. Unlike traditional k-means clustering and fuzzy clustering methods, the SOM algorithm does not require setting an initial value for the number of clusters, making it easier to operate. It can not only automatically find the intrinsic relationships between sample attributes, but also reduce the dimensionality and complexity of the data. A typical SOM model is hierarchical, generally with only an input layer and a competition layer, which has great advantages for processing large-scale complex data.

[0005] There are currently no relevant literature reports. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to overcome the shortcomings of the prior art and provide a user segmentation method based on highway ETC toll data that can quickly and accurately identify and classify highway users.

[0007] The technical solution adopted in this invention is: a user segmentation method based on highway ETC toll data, which identifies the travel purpose of highway users, including commuting, commercial travel, business travel, and sporadic travel, and includes the following steps:

[0008] 1) Preprocess the highway toll data within the set period, extract the field information required for highway user classification, and store the basic information with the highway user's license plate number as the key field to form the basic travel data of highway users.

[0009] 2) Sort the toll records of each highway user within the set period by time, and clean the data according to the abnormal conditions of time and space to obtain the cleaned highway toll data.

[0010] 3) Based on the data cleaned in step 2), extract information from three dimensions of highway users within a set period: time indicators, spatial indicators, and personal attribute indicators, to form a user classification evaluation index system, and use the SOM clustering algorithm to classify highway users.

[0011] 4) Classify highway users’ travel based on time and space indicators on a monthly basis, and identify various types of travel such as commuting, operational travel, sporadic travel, and business travel.

[0012] Step 1) includes: sorting highway toll records within a set period according to the user's license plate number, removing abnormal data records with missing fields or incorrect license plate numbers, and forming the following basic travel data storage format.

[0013] [License plate number, entry time, entry location, exit time, exit location, billing distance, final charge].

[0014] Step 2) The data cleaning based on time anomalies is as follows: Read the exit time and entry time of a highway user's trip record within a set period, and calculate the travel time in the record. If the travel time is negative, that is, the exit time is less than the entry time, or the travel time exceeds 24 hours, then the consumption record is determined to be abnormal time data of the highway user and is removed.

[0015] Step 2) involves cleaning the data based on the abnormal spatial conditions: reading the exit time, entry time, and billing distance of a single trip record of a highway user within a set period, calculating the travel speed of this trip, and determining that the speed is greater than 120km / h or the billing distance is greater than 1000km, thus identifying this consumption record as abnormal spatial data of the highway user and removing it.

[0016] Step 3) The method for extracting highway user time indicators is as follows: count the number of days each highway user travels on weekdays and non-weekdays within a set period, count the number of days during peak and off-peak periods, wherein the peak period is the morning peak from 7:00 to 9:00 and the evening peak from 17:00 to 19:00, and the rest of the time is the off-peak period.

[0017] Step 3) describes the method for extracting highway user spatial indicators as follows: Extract all toll station origins and destinations for each user's trips within a set period and assign them a number 'a'. Then, based on the number, calculate the travel frequency of each user at each origin and destination within the set period. Finally, calculate the travel percentage of each user at each origin and destination within the set period. The calculation formula is as follows:

[0018]

[0019]

[0020] Where 'a' represents the origin and destination numbers of the toll station within a set period, 'C' represents the total travel frequency of each user on the highway within the set period, 'A' represents the set of all origins and destinations traversed by each user within the set period, and 'C' represents the total travel frequency of each user on the highway within the set period. a To define the travel frequency of each user at origin and destination a within a given period, Q a To define the percentage of trips taken by each user at origin and destination a within a given period.

[0021] Step 3) describes the method for extracting personal attribute indicators for highway users: The total tollable distance for each highway user within a set period is calculated using an aggregation function, as shown in the following formula:

[0022]

[0023] Where 'a' represents the origin and destination numbers within the set period of the toll station, A represents the set of all origin and destination 'a' traversed by each user within the set period, and S represents the total tollable distance for each user on the highway. a The single billing distance is the distance from origin to destination a.

[0024] Step 3) describes classifying highway users using the SOM clustering algorithm. This involves using the extracted temporal and spatial travel indicators of highway users as input, and setting the size of the adaptive neural network's competitive layer to N*N, where N is the number of neurons, obtained from the following formula:

[0025] Where sample is the number of highway users.

[0026] Cluster analysis was performed using the python-minisom tool within the SOM clustering algorithm. Based on the cluster analysis results, the average temporal and spatial metrics of highway users in each cluster were calculated, resulting in the following storage format.

[0027]

[0028] Step 4) describes the method for identifying commuter and commercial travel as follows: Select cluster IDs of highway users who travel an average of more than 3 days per week. Then, calculate the total number of days that highway users in the cluster travel during peak and off-peak hours (7:00-9:00 and 17:00-19:00), specifically selecting the k-th cluster ID.

[0029]

[0030]

[0031] Among them, W k Let M be the total number of days of travel for highway users during peak hours in month k; k The total number of days of highway travel during off-peak hours in month k;

[0032] If, W k >M k If the cluster ID is "highway user", then the highway users included in the cluster ID are defined as commuter users; otherwise, the highway users in the cluster ID are defined as daily operating users.

[0033] Step 4) describes a method for identifying sporadic and business trips: Select cluster IDs for highway users whose average weekday travel is less than 3 days, and then calculate the travel frequency for all origins and destinations for each highway user in month k.

[0034]

[0035]

[0036] Among them, P kj P represents the frequency of highway users traveling at the j-th origin and destination in the k-th month; k Let q be the total travel frequency of highway users in month k; q is the total number of origin and destination points.

[0037] Calculate the percentage of each origin and destination for highway users in this cluster ID relative to all origins and destinations. If the percentage of the largest origin and destination exceeds 40%, then the highway users in this cluster ID are defined as business travelers; otherwise, the highway users in this cluster ID are defined as casual travelers.

[0038] The user segmentation method based on highway ETC toll data of the present invention has the following advantages:

[0039] (1) This invention makes full use of ETC toll data on highways, which can quickly and accurately classify commuter, daily operation, morning and sporadic travel users, providing a basis for highway planning and construction.

[0040] (2) The basic data of this invention comes from the highway travel records of ETC users with unique identifiers, which has the characteristics of complete information and high accuracy compared with traditional traffic sampling surveys and other methods.

[0041] (3) The SOM classification method used in this invention is flexible and easy to use, and has significant advantages in processing large-scale ETC toll data, and can quickly obtain classification results.

[0042] (4) The highway user classification results of the present invention can more accurately reflect the differences in the spatial and temporal distribution of highway users, and can provide support for highway operation and congestion management decisions. Attached Figure Description

[0043] Figure 1 This is a flowchart of a user segmentation method based on high-speed ETC toll data according to the present invention;

[0044] Figure 2 This is a schematic diagram of SOM clustering in the invention;

[0045] Figure 3 This is a schematic diagram illustrating the user allocation for highways in the invention. Detailed Implementation

[0046] The following describes in detail a user segmentation method based on high-speed ETC toll data according to the present invention, with reference to embodiments and accompanying drawings.

[0047] This invention provides a user segmentation method based on highway ETC toll data, which identifies the travel purpose of highway users, such as commuting, commercial travel, business travel, and occasional travel. Figure 1 As shown, it includes the following steps:

[0048] 1) Preprocess highway toll data within a set period, extract the fields required for highway user classification, and store basic information using highway user license plate numbers as the key field to form basic travel data for highway users; including:

[0049] Based on the user's license plate number, highway toll records within a set period are sorted, and abnormal data records with missing fields or incorrect license plate numbers are removed, resulting in the following basic travel data storage format.

[0050] [License plate number, entry time, entry location, exit time, exit location, billing distance, final charge];

[0051] 2) For each highway user's toll records within a set period, sort them by time, and perform data cleaning based on temporal and spatial anomalies to obtain cleaned highway toll data; among which,

[0052] The data cleaning based on time anomalies is as follows: read the exit time and entry time of a highway user's trip record within a set period, and calculate the travel time in the record. If the travel time is negative, i.e., the exit time is less than the entry time, or the travel time exceeds 24 hours, then the consumption record is determined to be abnormal time data of the highway user and is removed.

[0053] The aforementioned data cleaning based on spatial anomalies involves: reading the exit time, entry time, and toll distance of a single trip record of a highway user within a set period; calculating the travel speed for this trip; and determining that if the speed exceeds 120 km / h or the toll distance exceeds 1000 km, the consumption record is considered spatial anomaly data for the highway user and is removed.

[0054] 3) Based on the data cleaned in step 2), information on three dimensions of highway users—time indicators, spatial indicators, and personal attribute indicators—is extracted within a set period to form a user classification and evaluation index system. The SOM clustering algorithm is then used to classify highway users.

[0055] The method for extracting highway user time indicators is as follows: count the number of days each highway user travels on weekdays and non-weekdays within a set period, and count the number of days during peak and off-peak periods. The peak period refers to the morning peak from 7:00 to 9:00 and the evening peak from 17:00 to 19:00, and the rest of the time is the off-peak period.

[0056] The method for extracting highway user spatial indicators is as follows: Extract all toll station origins and destinations for each user's trips within a set period and assign them a number 'a'. Then, based on these numbers, calculate the travel frequency of each user at each origin and destination within the set period. Finally, calculate the travel percentage of each user at each origin and destination within the set period. The calculation formula is as follows:

[0057]

[0058]

[0059] Where 'a' represents the origin and destination numbers of the toll station within a set period, 'C' represents the total travel frequency of each user on the highway within the set period, 'A' represents the set of all origins and destinations traversed by each user within the set period, and 'C' represents the total travel frequency of each user on the highway within the set period. a To define the travel frequency of each user at origin and destination a within a given period, Q a To define the percentage of trips taken by each user at origin and destination a within a given period.

[0060] The method for extracting personal attribute indicators of highway users is as follows: The total toll distance for each highway user within a set period is calculated using an aggregation function, and the calculation formula is as follows:

[0061]

[0062] Where 'a' represents the origin and destination numbers within the set period of the toll station, A represents the set of all origin and destination 'a' traversed by each user within the set period, and S represents the total tollable distance for each user on the highway. a The single billing distance is the distance from origin to destination a.

[0063] The aforementioned use of the SOM clustering algorithm to classify highway users utilizes, for example... Figure 2 The SOM clustering algorithm shown takes extracted highway user time and space travel indicators as input and sets the size of the adaptive neural network competitive layer to N*N, where N is the number of neurons, obtained by the following formula:

[0064] Where sample is the number of highway users.

[0065] Cluster analysis was performed using the python-minisom tool within the SOM clustering algorithm. Based on the cluster analysis results, the average temporal and spatial metrics of highway users in each cluster were calculated, resulting in the following storage format.

[0066]

[0067] 4) such as Figure 3 As shown, based on monthly time and spatial indicators of highway user travel, various types of travel are categorized, including commuting, commercial travel, sporadic travel, and business travel; among them,

[0068] The method for identifying commuter and commercial travel is as follows: Select cluster IDs of highway users who travel an average of more than 3 days per week. Then, calculate the total number of days that highway users in the cluster travel during peak hours (7:00-9:00, 17:00-19:00) and off-peak hours, specifically selecting the k-th ID for calculation.

[0069]

[0070]

[0071] Among them, W k Let M be the total number of days of travel for highway users during peak hours in month k; k The total number of days of highway travel during off-peak hours in month k;

[0072] If, W k >M k If the cluster ID is "highway user", then the highway users included in the cluster ID are defined as commuter users; otherwise, the highway users in the cluster ID are defined as daily operating users.

[0073] The method for identifying sporadic and business trips is as follows: select cluster IDs for highway users whose average number of trips per week is less than 3 days, and then calculate the trip frequency of all origins and destinations for each highway user in month k:

[0074]

[0075]

[0076] Among them, P kj P represents the frequency of highway users traveling at the j-th origin and destination in the k-th month; k Let q be the total travel frequency of highway users in month k; q is the total number of origin and destination points.

[0077] Calculate the percentage of each origin and destination for highway users in this cluster ID relative to all origins and destinations. If the percentage of the largest origin and destination exceeds 40%, then the highway users in this cluster ID are defined as business travelers; otherwise, the highway users in this cluster ID are defined as casual travelers.

[0078] The following are specific examples:

[0079] According to the method of this invention, the ETC toll data of a specific lane of a highway in July 2019 was used to classify users into commuters, operators, business travelers, and casual travelers based on the highway ETC toll data, as shown in the flowchart.

[0080] Step 101: Preprocess the high-speed ETC data.

[0081] The amount of ETC toll data on highways is enormous, exceeding 100GB. To improve storage efficiency, key fields were extracted from the raw data according to time and space characteristics. The highway toll records were sorted, and abnormal data records such as missing fields and incorrect license plate numbers were removed. The resulting basic data storage format contains 20 million records and more than 1.4 million users.

[0082] [License plate number, entry time, entry location, exit time, exit location, billing distance, final fare]

[0083] Step 102: Clean up the user's travel records based on time and space anomalies.

[0084] Because there are errors in the system input and recognition of highway ETC data, data cleaning is necessary before data processing. First, the travel records of each user are sorted by time, and then the following steps are performed:

[0085] Step 1021: Record abnormal data during cleaning time.

[0086] Read the exit time and entry time of a highway user's trip record, and calculate the travel time for that record. If the travel time is negative (exit time is less than entry time) or the travel time exceeds 24 hours, then the current consumption record is determined to be abnormal time data of the highway user.

[0087] Step 1022: Clean up abnormal data records in the cleaning space.

[0088] The system reads the exit time, entry time, and toll distance of a single trip from a highway user, and calculates the travel speed. If the speed exceeds 120 km / h, or the toll distance exceeds 1000 km, the trip is considered spatial anomaly data for the highway user. After data cleaning, approximately 1.35 million highway users remain.

[0089] Step 1023: Extract user travel metrics based on time, location, and personal attributes.

[0090] The number of weekday and non-weekday trips within the statistical period was calculated, with 7:00-9:00 as the morning peak and 17:00-19:00 as the evening peak. The trip frequency of each origin and destination for highway users was calculated, and the proportion of each origin and destination in all trips was also calculated. Aggregate functions were used to calculate the total trip frequency and total toll distance for each highway user within the study period, thus obtaining the travel indicators for all highway users. The travel indicators for a specific highway user are shown in Table 1.

[0091] Table 1

[0092]

[0093] Step 103: Use SOM clustering to complete the highway user clustering.

[0094] The python-minisom tool in the SOM clustering method was used to perform cluster analysis on the above-mentioned highway user time, space and personal attribute indicators. The input parameters of the SOM clustering algorithm include the number of travel days of highway users on weekdays and non-weekdays, the number of travel days of monthly users during peak and non-peak periods, the proportion of the most frequently used origin and destination in all trips, and the size of the adaptive neural network competition layer is set to N×N=76×76.

[0095] After SOM clustering, 6 categories were finally obtained. Then, the average value of all user travel indicators in this cluster was calculated according to the cluster number ID, and the data format shown in Table 2 was formed for each cluster.

[0096] Table 2

[0097]

[0098] Step 104: Based on the highway user identification principles, classify users into commuter, commercial, business, and casual travel users.

[0099] Highway users in clusters 1 and 4 both travel more than three times per week on average on weekdays. However, users in cluster 1 travel more concentrated during peak hours, while those in cluster 4 travel more dispersedly. Therefore, cluster 1 is defined as commuter users, and cluster 4 as commercial users. Highway users in the remaining clusters 2, 3, 5, and 6 travel less frequently, averaging less than three times per weekday. However, in cluster 3, the most frequently used origin-destination trips account for over 40%, and the travel routes are more concentrated. Therefore, cluster 3 is defined as business travel users, while the remaining clusters 2, 5, and 6 are defined as sporadic travel users.

[0100] The above embodiments are provided merely for the purpose of describing the present invention and are not intended to limit the scope of the invention. The scope of the invention is defined by the appended claims. Various equivalent substitutions and modifications made without departing from the spirit and principles of the invention should be covered within the scope of the invention.

Claims

1. A user segmentation method based on highway ETC toll data, characterized in that, It identifies the travel purpose of highway users, including commuting, commercial travel, business travel, and occasional travel, and includes the following steps: 1) Preprocess highway toll data within a set period, extract the fields required for highway user classification, and store basic information using highway user license plate numbers as the key field to form basic travel data for highway users; including: Based on the user's license plate number, highway toll records within a set period are sorted, and abnormal data records with missing fields or incorrect license plate numbers are removed, resulting in the following basic travel data storage format. ; 2) Sort the toll records of each highway user within the set period by time, and clean the data according to the abnormal conditions of time and space to obtain the cleaned highway toll data; The data cleaning based on time anomalies is as follows: read the exit time and entry time of a highway user's trip record within a set period, and calculate the travel time in the record. If the travel time is negative, i.e., the exit time is less than the entry time, or the travel time exceeds 24 hours, then the consumption record is determined to be abnormal time data of the highway user and is removed. The aforementioned data cleaning based on spatial anomalies involves: reading the exit time, entry time, and billing distance of a single trip record of a highway user within a set period, calculating the travel speed of this trip, and determining that the speed is greater than 120km / h or the billing distance is greater than 1000km, thus identifying this consumption record as spatial anomaly data of the highway user and removing it. 3) Based on the data cleaned in step 2), extract information from three dimensions of highway users within a set period: time indicators, spatial indicators, and personal attribute indicators, to form a user classification evaluation index system, and use the SOM clustering algorithm to classify highway users. The method for extracting highway user spatial indicators within a set period is as follows: extract all toll station origins and destinations for each user's trips within the set period and assign them a number 'a'. Then, based on the number, calculate the travel frequency of each user at each origin and destination within the set period. Finally, calculate the travel percentage of each user at each origin and destination within the set period. The calculation formula is as follows: ; ; Where 'a' represents the origin and destination number of the toll station within a set period. To determine the total travel frequency for each user on the highway within a given period, This is the set of all start and end points traversed by each user within a given period. To set the travel frequency of each user at origin and destination a within a set period, To determine the percentage of trips made by each user at origin and destination a within a given period; The method for extracting personal attribute indicators of highway users within a set period is as follows: The total toll distance for each highway user within the set period is calculated using an aggregation function, and the calculation formula is as follows: ; Where 'a' represents the origin and destination number of the toll station within a set period. Let be the set of all origin and destination points 'a' traversed by each user within a given period. The total toll distance for each user on the highway. The single-trip billing distance for origin and destination a; The aforementioned classification of highway users using the SOM clustering algorithm involves using the extracted temporal and spatial travel indicators of highway users as input, and setting the size of the competitive layer of the adaptive neural network to N*N, where N is the number of neurons, obtained from the following formula: ; Cluster analysis was performed using the python-minisom tool in the SOM clustering algorithm. Based on the cluster analysis results, the average values ​​of highway users in time and space were calculated for each cluster, resulting in the following storage format. ; 4) Classify highway users’ travel based on time and space indicators on a monthly basis, and identify various types of travel such as commuting, operational travel, sporadic travel, and business travel.

2. The user segmentation method based on highway ETC toll data according to claim 1, characterized in that, Step 3) The method for extracting highway user time indicators within a set period is as follows: count the number of days each highway user travels on weekdays and non-weekdays within the set period, count the number of days during peak and non-peak periods, wherein the peak period is the morning peak from 7:00 to 9:00 and the evening peak from 17:00 to 19:00, and the rest of the time is the non-peak period.

3. The user segmentation method based on highway ETC toll data according to claim 1, characterized in that, Step 4) describes the method for identifying commuter and commercial travel as follows: Select cluster IDs where highway users average more than 3 days of travel per week. Then, calculate the total number of days within each cluster ID that highway users traveled during peak and off-peak hours (7:00-9:00 and 17:00-19:00). Specifically, select the k-th cluster ID for calculation. ; ; in, This represents the total number of days spent by highway users during peak hours in month k. This represents the total number of days that highway users travel during off-peak hours in month k. if, If the cluster ID is "highway user", then the highway users included in the cluster ID are defined as commuter users; otherwise, the highway users in the cluster ID are defined as daily operating users.

4. The user segmentation method based on highway ETC toll data according to claim 1, characterized in that, Step 4) describes a method for identifying sporadic and business trips: Select cluster IDs for highway users whose average weekday travel is less than 3 days, and then calculate the travel frequency for all origins and destinations for each highway user in month k. ; ; in, The frequency of highway users traveling at the j-th origin and destination in the k-th month; Let q be the total travel frequency of highway users in month k; q is the total number of origin and destination points. Calculate the percentage of each origin and destination for highway users in this cluster ID relative to all origins and destinations. If the percentage of the largest origin and destination exceeds 40%, then the highway users in this cluster ID are defined as business travelers; otherwise, the highway users in this cluster ID are defined as casual travelers.

Citation Information

Patent Citations

  • Vehicle traveling analysis method based on gate plate recognition data

    CN108717790A