Comprehensive traffic transportation channel identification method based on clustering analysis
By using K-means++ to select the initial clustering center in the clustering algorithm, the clustering structure missed or merged problems caused by initial selection sensitivity is solved, and a more accurate comprehensive transportation channel identification is achieved.
Patent Information
- Application Number
- CN202510676908.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-06-24
AI Technical Summary
In the prior art, when the clustering algorithm recognizes a comprehensive transportation channel, the selection of the initial clustering center is sensitive, which may lead to the clustering structure being missed or the different channels being erroneously merged, and the traffic channel cannot be accurately identified.
Using K-means++-based cluster analysis method, the initial cluster center is selected by determining the number of clusters K and using K-means++ to reduce the randomness of the initial selection and improve the accuracy of the clustering results.
Through the K-means++ algorithm, more representative comprehensive transportation channels can be identified, and the cluster center reflects the core characteristics of the channels, improving the accuracy and consistency of channel identification.
Smart Images

Figure CN120199079A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of transportation, and in particular to a method for identifying an integrated transportation corridor based on cluster analysis. Background Art
[0002] With the acceleration of globalization and the advancement of regional integration, the demand for transportation in social and economic development is increasing day by day. As an important infrastructure connecting different regions and promoting the flow of materials and people, the identification and construction of an integrated transportation corridor are of great significance for promoting regional economic development and optimizing resource allocation. Especially in the context of current global economic integration, the smoothness of the integrated transportation corridor is directly related to the competitiveness and sustainable development ability of the regional economy.
[0003] In the prior art, in the scenario of identifying an integrated transportation corridor, the clustering algorithm is also somewhat sensitive to the selection of the initial clustering center. When there are multiple clustering centers, if the initial clustering center is not properly selected, some true clustering structures may be missed, or different clusters may be wrongly merged into one, resulting in the inability to accurately identify all transportation corridors, or wrongly regarding different transportation corridors as the same corridor. Therefore, a method for identifying an integrated transportation corridor based on cluster analysis is proposed. Summary of the Invention
[0004] The purpose of the present invention is to solve the disadvantages existing in the prior art that when there are multiple clustering centers, if the initial clustering center is not properly selected, some true clustering structures may be missed, or different clusters may be wrongly merged into one, resulting in the inability to accurately identify all transportation corridors, or wrongly regarding different transportation corridors as the same corridor, and to propose a method for identifying an integrated transportation corridor based on cluster analysis.
[0005] In order to achieve the above purpose, the present invention adopts the following technical solutions: A method for identifying an integrated transportation corridor based on cluster analysis, comprising the following steps: S1: Data preparation: Collect traffic data including multiple transportation modes, such as traffic flow, speed, occupancy rate, etc., and process the collected traffic data; S2: Application of K-means++: Determine the number of clusters K, use K-means++ to select the initial clustering center, and assign each data point to the nearest clustering center according to the distance from each data point to the clustering center, and recalculate the center point of each cluster as the new clustering center. Repeat this process until the clustering center no longer changes or reaches the maximum number of iterations to obtain a stable clustering result; S3: Cluster result analysis: Evaluate the clustering results, use the silhouette coefficient and Rand index to evaluate the quality of the clustering results, and adjust and optimize K-means++ according to the evaluation results. Identify the comprehensive transportation corridors based on the clustering results. Each cluster represents a comprehensive transportation corridor, and the cluster center reflects the core features of the corridor. S4: Processing and optimization: Further analyze and optimize the identified comprehensive transportation corridors, plan and adjust the transportation network according to the clustering results, and optimize the allocation of transportation resources.
[0006] The above further includes: Further, in S1, the traffic data of multiple transportation modes include roads, railways, waterways, and aviation. The road data is obtained by collecting traffic flow, speed, occupancy, etc. of major roads such as highways, national roads, and provincial roads. The road data is acquired through devices such as traffic monitoring systems and vehicle detectors. The railway data is obtained by collecting train operation data of railway lines, including the number of trains, running speed, running time, etc. The railway data is retrieved from the database of the railway department. The waterway data is obtained by collecting freight volume, passenger volume, number of ships, etc. of waterway transportation such as ports and waterways. The waterway data is retrieved from the database of the maritime department or relevant enterprises. The aviation data is obtained by collecting flight information of airports, including the number of flights, takeoff and landing times, passenger throughput, etc. The aviation data is retrieved from the database of the civil aviation department.
[0007] Further, in S1, the processing of the collected traffic data includes data preprocessing, data transformation and feature extraction, and data verification and storage. The data preprocessing includes data cleaning, missing value handling, and normalization. The purpose of normalization is to transform the data to the same scale for subsequent calculations and analyses. In the K-means++ algorithm, normalization can improve the convergence speed of the algorithm and the accuracy of the clustering results. Suppose there is a dataset containing traffic flow and speed. The range of traffic flow is [100, 1000], and the range of speed is [30, 120]. Use the normalization formula to transform these two attributes to the range of [0, 1].
[0008] Further, the specific steps of the data transformation: Discretize the data of continuous attributes (such as traffic flow) into different intervals, and use one-hot encoding or label encoding to transform the data of categorical attributes (such as transportation modes) into numerical data. The feature extraction applies PCA, and the data after transformation is set as M samples { }, each sample has N-dimensional features , each feature Each has its own eigenvalue; First, all features are decentralized, that is, the mean value is removed. The average value of each feature is calculated, and then for all samples, each feature is subtracted from its own mean value. The respective mean values are ; After decentralization, the covariance matrix is calculated , where the diagonal elements are the variances of the features and respectively, and the non-diagonal elements are the covariances. The calculation formula of is From this, the covariance matrix C of M samples under these N-dimensional features is obtained; After obtaining the covariance matrix, according to the characteristic equation
[0009] Furthermore, in S2, the number of clusters K represents the number of integrated transportation channels expected to be identified.
[0010] In one embodiment, for the above S2, the specific steps of using K-means++ to select the initial cluster centers in S2 are as follows: Randomly select the first initial cluster center: K-means++ randomly selects a data point from the dataset as the first initial cluster center. This selection is random, but once selected, subsequent selections will be based on this initial point; Calculate the distance from each data point to the nearest cluster center: After selecting the first initial cluster center, K-means++ calculates the distance from each data point in the dataset to this nearest cluster center. The distance represents the relative position relationship between the data point and the selected cluster center; Select the next initial cluster center according to the distance: K-means++ selects a data point as the next initial cluster center with a predetermined probability according to the distance from each data point to the nearest cluster center, that is, the probability of a data point being selected is proportional to the square of its distance to the nearest cluster center. The farther a data point is from the selected cluster center, the greater the possibility of being selected as the next initial cluster center. The purpose of this selection strategy is to ensure that the initial cluster centers are more evenly distributed in the dataset, thus avoiding the clustering result falling into a local optimal solution; Repeat the selection process until K initial cluster centers are reached: K-means++ repeats the above selection process until the preset number of clusters K is reached. In each selection, based on the currently selected cluster centers and the distances from each data point to the nearest cluster center, the next initial cluster center is selected with a predetermined probability.
[0011] Further, in S3, the silhouette coefficient and the Rand index are used to evaluate the quality of the clustering result; The silhouette coefficient is an index used to evaluate the quality of clustering. Its value ranges from -1 to 1. The larger the silhouette coefficient value, the better the clustering effect. The formula for calculating the silhouette coefficient is , where a is the average distance from data point i to other data points within its affiliated cluster (i.e., the within-cluster dissimilarity), b is the average minimum distance from data point i to all data points in other clusters (i.e., the between-cluster dissimilarity), and the silhouette coefficient of the entire data set is the average of the silhouette coefficients of all data points.
[0012] Further, the Rand index is an index used to evaluate the consistency between the clustering result and the true labels. Its value ranges from 0 to 1. The larger the value of the Rand index, the higher the consistency between the clustering result and the true labels. The formula for calculating the Rand index is , where c represents the number of data point pairs that are correctly classified in both the clustering result and the true labels, d represents the number of data point pairs that are misclassified in both the clustering result and the true labels (i.e., the situation where a data point belongs to one cluster in the clustering result and another cluster in the true labels), represents the total number of data point pairs formed by randomly selecting two data points from n data points.
[0013] The present invention has the following beneficial effects: In the present invention, K-means++ is used to select the initial cluster centers, making the distances between the initial cluster centers as far as possible, thereby avoiding the influence of the randomness of the initial selection on the clustering result, reducing randomness and improving the accuracy of clustering. By clustering through the K-means++ algorithm, more representative comprehensive transportation channels can be identified. Each cluster represents a comprehensive transportation channel, and the cluster center reflects the core characteristics of the channel. This clustering method not only considers the complexity of the transportation network but also combines the actual needs, making the identified channels more in line with the actual situation and contributing to subsequent analysis and optimization. The silhouette coefficient and the Rand index help us more intuitively understand the quality of the clustering result, so as to adjust parameters or optimize the algorithm to obtain better clustering results. Description of the Drawings
[0014] Figure 1It is a step diagram of a comprehensive transportation corridor identification method based on cluster analysis proposed by the present invention. Detailed implementation manners
[0015] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0016] Please refer to Figure 1 As shown, the present invention is a comprehensive transportation corridor identification method based on cluster analysis, including the following steps: S1: Data preparation: Collect traffic data including various transportation modes, such as traffic flow, speed, occupancy rate, etc., and process the collected traffic data; S2: Application of K-means++: Determine the number of clusters K, use K-means++ to select the initial cluster centers, and assign each data point to the nearest cluster center according to the distance from each data point to the cluster centers. Recalculate the center points of each cluster as the new cluster centers, and repeat this process until the cluster centers no longer change or reach the maximum number of iterations to obtain a stable clustering result; S3: Cluster result analysis: Evaluate the clustering results, use the silhouette coefficient and the Rand index to evaluate the quality of the clustering results. According to the evaluation results, adjust and optimize K-means++. According to the clustering results, identify the comprehensive transportation corridors. Each cluster represents a comprehensive transportation corridor, and the cluster center reflects the core features of the corridor; S4: Processing and optimization: Further analyze and optimize the identified comprehensive transportation corridors. According to the clustering results, plan and adjust the transportation network and optimize the allocation of transportation resources.
[0017] In one embodiment, for the above S1, in S1, the traffic data of multiple transportation modes includes highways, railways, waterways, and aviation. The highway data is obtained by collecting traffic flow, speed, occupancy, etc. of major roads such as expressways, national highways, and provincial highways. The highway data is acquired through devices such as traffic monitoring systems and vehicle detectors. The railway data is obtained by collecting train operation data of railway lines, including the number of trains, running speed, running time, etc. The railway data is retrieved from the database of the railway department. The waterway data is obtained by collecting freight volume, passenger volume, number of ships, etc. of waterway transportation such as ports and waterways. The waterway data is retrieved from the database of the maritime department or relevant enterprises. The aviation data is obtained by collecting flight information of airports, including the number of flights, takeoff and landing times, passenger throughput, etc. The aviation data is retrieved from the database of the civil aviation department.
[0018] In one embodiment, for the above S1, in S1, the processing of the collected traffic data includes data preprocessing, data conversion and feature extraction, and data verification and storage. The data preprocessing includes data cleaning, missing value processing, and normalization. The purpose of normalization is to convert the data to the same scale for subsequent calculations and analyses. In the K-means++ algorithm, normalization can improve the convergence speed of the algorithm and the accuracy of the clustering results. Suppose there is a dataset containing traffic flow and speed. The range of traffic flow is [100, 1000], and the range of speed is [30, 120]. Use the normalization formula to convert these two attributes to the range of [0, 1].
[0019] In one embodiment, for the above data conversion, the specific steps of the data conversion are as follows: For the data of continuous attributes (such as traffic flow), discretize it into different intervals. For the data of categorical attributes (such as transportation modes), use one-hot encoding or label encoding to convert it into numerical data. The feature extraction applies PCA, and the data after being converted is set as M samples { }, each sample having N-dimensional features , and each feature has its own feature value; First, centralize all features, that is, remove the mean value, calculate the mean value of each feature, and then for all samples, each feature subtracts its own mean value, where the respective mean values are ; After centralization, calculate the covariance matrix , where the diagonal elements are the variances of features and respectively, and the non-diagonal elements are the covariances. The calculation formula is , from which the covariance matrix C of M samples under these N-dimensional features is obtained; After obtaining the covariance matrix, according to the characteristic equation its eigenvalues and their corresponding eigenvectors are obtained, where λ is the eigenvalue and μ is its corresponding eigenvector. The largest first k eigenvalues and the corresponding eigenvectors are selected for projection. The projection is the process of dimensionality reduction, which reduces the original features from high dimensions to low dimensions. After dimensionality reduction, a large amount of redundant information is removed, and at least more than 85% of the original information is retained.
[0020] In one embodiment, for the above S2, in S2, the clustering number K represents the number of integrated transportation channels expected to be recognized.
[0021] In one embodiment, for the above S2, in S2, the specific steps of using K-means++ to select the initial clustering centers are as follows: Randomly select the first initial clustering center: K-means++ randomly selects a data point from the dataset as the first initial clustering center. This selection is random, but once selected, subsequent selections will be based on this initial point; Calculate the distance from each data point to the nearest clustering center: After selecting the first initial clustering center, K-means++ calculates the distance from each data point in the dataset to this nearest clustering center. The distance represents the relative position relationship between the data point and the selected clustering center; Select the next initial clustering center according to the distance: K-means++ selects a data point as the next initial clustering center with a predetermined probability according to the distance from each data point to the nearest clustering center, that is, the probability of a data point being selected is proportional to the square of its distance to the nearest clustering center. The farther the data point is from the selected clustering center, the greater the possibility of being selected as the next initial clustering center. The purpose of this selection strategy is to ensure that the initial clustering centers are more evenly distributed in the dataset, thus avoiding the clustering result falling into a local optimal solution; Repeat the selection process until K initial clustering centers are reached: K-means++ repeats the above selection process until the preset clustering number K is reached. In each selection, according to the currently selected clustering centers and the distance from each data point to the nearest clustering center, the next initial clustering center is selected with a predetermined probability. Analysis of the clustering result: Evaluate the clustering result, and use the silhouette coefficient and the Rand index to evaluate the quality of the clustering result. According to the clustering result, identify the integrated transportation channels. Each clustering cluster represents an integrated transportation channel, and the clustering center reflects the core features of the channel.
[0022] In one embodiment, for the above S3, in S3, the silhouette coefficient and the Rand index are used to evaluate the quality of the clustering result; The silhouette coefficient is an index used to evaluate the quality of clustering. Its value ranges from -1 to 1. The larger the silhouette coefficient value, the better the clustering effect. The formula for calculating the silhouette coefficient is , where a is the average distance from data point i to other data points within its affiliated cluster (i.e., the within-cluster dissimilarity), b is the average minimum distance from data point i to all data points in other clusters (i.e., the between-cluster dissimilarity), and the silhouette coefficient of the entire data set is the average of the silhouette coefficients of all data points; Suppose there is a data set, and after clustering, 3 clusters are obtained. For each data point, calculate the average distance ai from it to other data points within its affiliated cluster, and the average minimum distance bi to all data points in other clusters. Then, calculate the silhouette coefficient of each data point according to the formula and take the average to obtain the silhouette coefficient of the entire data set. If the silhouette coefficient is close to 1, it indicates that the clustering effect is very good; if the silhouette coefficient is close to -1, it indicates that the clustering effect is very poor; if the silhouette coefficient is close to 0, it indicates that the clustering effect is average.
[0023] In one embodiment, for the above Rand index, the Rand index is an index used to evaluate the consistency between the clustering result and the true labels. Its value ranges from 0 to 1. The larger the Rand index value, the higher the consistency between the clustering result and the true labels. The formula for calculating the Rand index is , where c represents the number of data point pairs that are correctly classified in both the clustering result and the true labels, d represents the number of data point pairs that are misclassified in both the clustering result and the true labels (i.e., the situation where a data point belongs to one cluster in the clustering result and another cluster in the true labels), represents the total number of data point pairs formed by randomly selecting two data points from n data points.
[0024] Suppose there is a data set and its true labels are known. After clustering, compare the clustering result with the true labels and count the values of c and d. Then, calculate the Rand index according to the formula. If the Rand index is close to 1, it indicates that the consistency between the clustering result and the true labels is very high; if the Rand index is close to 0, it indicates that the consistency between the clustering result and the true labels is very low.
[0025] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A comprehensive transportation corridor identification method based on cluster analysis, characterized in that Including the following steps: S1: Data preparation: Collect traffic data including multiple transportation modes, and process the collected traffic data; S2: Application of K-means++: Determine the number of clusters K, use K-means++ to select the initial cluster centers, assign data points to the nearest cluster center according to the distance from each data point to the cluster center, recalculate the center point of each cluster as the new cluster center, and repeat this process until the cluster centers no longer change or reach the maximum number of iterations to obtain a stable clustering result; S3: Analysis of clustering results: Evaluate the clustering results, use the silhouette coefficient and Rand index to evaluate the quality of the clustering results, adjust and optimize K-means++ according to the evaluation results, identify the comprehensive transportation corridors according to the clustering results, each cluster represents a comprehensive transportation corridor, and the cluster center reflects the core features of the corridor; S4: Processing and optimization: Further analyze and optimize the identified comprehensive transportation corridors, plan and adjust the traffic network according to the clustering results, and optimize the allocation of traffic resources.
2. The integrated transportation corridor identification method based on cluster analysis according to claim 1, characterized in that In S1, the traffic data of multiple transportation modes include highways, railways, waterways, and aviation. The highway data is collected by collecting data of main roads, the railway data is collected by collecting train operation data of railway lines, the waterway data is collected by collecting waterway traffic data, and the aviation data is collected by collecting flight information of airports.
3. A method for identifying an integrated transportation corridor based on cluster analysis according to claim 1, characterized in that, In S1, processing the collected traffic data includes data preprocessing, data transformation and feature extraction, and data verification and storage. The data preprocessing includes data cleaning, missing value processing, and normalization. The purpose of normalization is to transform the data to the same scale.
4. The comprehensive transportation corridor identification method based on cluster analysis according to claim 3, characterized in that, The specific steps of the data transformation are as follows: Discretize the data of continuous attributes into different intervals, and use one-hot encoding or label encoding to transform the data of categorical attributes into numerical data; The feature extraction applies PCA, and the data after being digitalized is set as M samples { }, and each sample has N-dimensional features . For each feature , there is its own eigenvalue; First, all features are decentralized, that is, the mean is removed. The average value of each feature is calculated, and then for all samples, each feature is subtracted by its own mean, where the respective means are ; After decentralization, the covariance matrix is calculated , where the variances of features and are on the diagonal respectively, and the covariance is on the non-diagonal. The calculation formula of is, and thus the covariance matrix C of M samples under these N-dimensional features is obtained; After obtaining the covariance matrix, according to the characteristic equation calculate its eigenvalues and the corresponding eigenvectors, where λ is the eigenvalue and μ is the corresponding eigenvector. Select the top k largest eigenvalues and the corresponding eigenvectors for projection. The projection is the process of dimensionality reduction, which reduces the original features from high dimensions to low dimensions. After dimensionality reduction, a large amount of redundant information is removed, and at least 85% of the original information is retained.
5. The comprehensive transportation corridor identification method based on cluster analysis according to claim 1, wherein In S2, the number of clusters K represents the number of comprehensive transportation corridors expected to be identified.
6. The integrated transportation corridor identification method based on cluster analysis according to claim 1, wherein In S2, the specific steps of using K-means++ to select the initial cluster centers are as follows: Randomly select the first initial cluster center: K-means++ randomly selects a data point from the dataset as the first initial cluster center. This selection is random, but once selected, subsequent selections will be based on this initial point; Calculate the distance from each data point to the nearest cluster center: After selecting the first initial cluster center, K-means++ calculates the distance from each data point in the dataset to this nearest cluster center. The distance represents the relative position relationship between the data point and the selected cluster center; Select the next initial cluster center according to the distance: K-means++ selects a data point as the next initial cluster center with a predetermined probability according to the distance from each data point to the nearest cluster center, that is, the probability of a data point being selected is proportional to the square of its distance to the nearest cluster center. The farther the data point is from the selected cluster center, the greater the possibility of being selected as the next initial cluster center; Repeat the selection process until K initial cluster centers are reached: K-means++ repeats the above selection process until the preset number of clusters K is reached. In each selection, based on the currently selected cluster centers and the distances from each data point to the nearest cluster center, the next initial cluster center is selected with a predetermined probability.
7. The comprehensive transportation corridor identification method based on cluster analysis according to claim 1, wherein In S3, the silhouette coefficient and the Rand index are used to evaluate the quality of the clustering result; The silhouette coefficient is an index used to evaluate the quality of clustering results. Its value ranges from -1 to 1. The larger the silhouette coefficient value, the better the clustering effect. The calculation formula of the silhouette coefficient is , where a is the average distance from data point i to other data points within its affiliated cluster, b is the average minimum distance from data point i to all data points in other clusters, and the silhouette coefficient of the entire data set is the average value of the silhouette coefficients of all data points.
8. A comprehensive transportation corridor identification method based on clustering analysis according to claim 7, characterized in that, The Rand Index is a metric used to evaluate the consistency between clustering results and true labels. Its value ranges from 0 to 1. The larger the value of the Rand Index, the higher the consistency between the clustering results and the true labels. The formula for calculating the Rand Index is , where c represents the number of data point pairs that are correctly classified in both the clustering results and the true labels, d represents the number of data point pairs that are misclassified in both the clustering results and the true labels, represents the total number of data point pairs formed by randomly selecting two data points from n data points.
Citation Information
Patent Citations
Road network traffic state discrimination method based on clustering and graph convolutional network
CN113450562A
Steel pipe surface defect detection method based on machine vision
CN116735610A
Image recognition system based on machine learning
CN117152583A
Intelligent traffic management system based on big data
CN117251722A
Comprehensive traffic transportation channel identification method based on clustering analysis
CN118568526A