A cross-region road construction priority sorting method, system, terminal and storage medium
Patent Information
- Application Number
- CN202311195413.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-17
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-09-17
AI Technical Summary
[0033]1)本发明一种基于大数据K均值聚类的跨区道路建设优先度排序方法,以及相应的排序系统、终端和存储介质,,第一步利用面积最小的行政区的特征长度来确定城市内部跨行政区地带的范围,该步骤的有益效果一方面是避免出现缓冲区过小而导致各指标信息提取不充分的问题,另一方面是为了避免缓冲区过大将整个行政区包含,从而影响跨行政区地带指标的精确性问题。
Smart Images

Figure CN118071044B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of urban cross-regional integration feature technology, and in particular to a method, system, terminal and storage medium for prioritizing cross-regional road construction. Background Technology
[0002] Against the backdrop of high-quality, integrated development, regions are abandoning competition and moving towards a new landscape of cross-administrative regional integration and development characterized by "resource sharing, complementary strengths, and win-win development." Simultaneously, with the advancement of urbanization, urban structures are shifting from single-center to multi-center network models, and urban development is expanding from the center outwards, leading to a trend of cross-regional integration and development in areas bordering administrative districts. Urban development is an ongoing process, and the integration and development of different urban areas vary significantly, thus placing different demands on the development of cross-regional transportation. Due to limited funds for urban infrastructure construction, it is necessary to determine the scale of transportation infrastructure supply and phased implementation strategies based on the characteristics of cross-regional integration. For example, it is necessary to calculate the priority of cross-regional roads in different integration zones based on the degree of integration to guide road construction decisions.
[0003] Therefore, this invention aims to eliminate the current barriers to research on cross-regional integrated transportation, scientifically and quantitatively identify the characteristics of cross-regional integration in cities, and provide a scientific basis for precise policy implementation in road construction and urban transportation planning in cross-regional areas. Summary of the Invention
[0004] To address the aforementioned technical issues, the present invention aims to provide a method for prioritizing cross-regional road construction based on big data K-means clustering, along with a corresponding ranking system, terminal, and storage medium. This method aims to eliminate current barriers to research on cross-regional integrated transportation, scientifically and quantitatively identify the characteristics of urban cross-regional integration, perform relatively accurate quantitative calculations and ranking of cross-regional road construction priorities, and provide a scientific basis for precise policy implementation in road construction and urban transportation planning in cross-regional areas.
[0005] The technical solution of the present invention is as follows:
[0006] A method for prioritizing cross-regional road construction based on big data K-means clustering includes the following steps:
[0007] S1 uses the characteristic length of the smallest administrative region to determine the extent of the cross-administrative region zone within the city;
[0008] S2 combines big data to extract the integration characteristics of cross-administrative regions within the city, resulting in a sample dataset;
[0009] S3 preprocesses the sample data in the sample dataset obtained in step S2, then selects a certain number of principal components as input, and performs K-means clustering with the number of clusters as a variable. Finally, the optimal number of clusters k is determined based on the silhouette coefficient and the sum of squared errors, representing the reasonable number of types of cross-administrative regions within the city.
[0010] S3 specifically involves: First, normalizing the sample data using the max-min method; then, selecting p principal components using principal component analysis, ensuring that the cumulative variance contribution rate of the p principal components is not less than 85%; finally, using the principal components as input and the number of clusters q as a variable, performing K-means clustering, and taking the k value corresponding to the inflection point of the sum of squared errors and the maximum silhouette coefficient as the optimal number of clusters, representing the reasonable number of types of cross-administrative-region zones within the city;
[0011] S4 obtains the result of the K-means algorithm when the optimal number of clusters k is obtained in step S3, and obtains the final clustering result. It also calculates the average value of each index of each cluster, obtains the fusion characteristics of different types of cross-administrative regions, calculates the cross-regional road priority of different fusion regions based on the fusion characteristics, and sorts the cross-regional road priority.
[0012] Preferably, S1 specifically involves determining the extent of the cross-administrative-region zone within the city using the characteristic length of the smallest administrative region. The specific calculation method is as follows: Where r represents the radius of the buffer zone across administrative regions, and S is the area of the administrative region corresponding to the smallest administrative region among all administrative regions.
[0013] Preferably, S2 specifically involves selecting 12 indicators to characterize the fusion features: urban construction land ratio, inter-regional corridor spacing, inter-regional highway ratio, inter-regional corridor completion rate, population density, employment density, population density ratio, employment density ratio, road modal share, passenger station ratio, passenger flow density, and average saturation. Through big data analysis, the actual values of each indicator across administrative regions are obtained using surveys and statistics. The nth indicator of the m-th inter-administrative region is designated as x. nm The above 12 indicators, spanning administrative regions, are extracted to form an n×m dimensional sample dataset X: X=[x1,x2,...,x m ], where x m This represents the index vector of the m-th cross-administrative region.
[0014] Preferably, the population density ratio is the ratio of the population density within the buffer zone to the average population density of the centers of two adjacent zones, and the employment density ratio is the ratio of the employment density within the buffer zone to the average employment density of the centers of two adjacent zones.
[0015] Preferably, the detailed steps of S3 are as follows:
[0016] S31: First, normalize the sample data using the max-min method;
[0017] Subtract the minimum value of each indicator from the sample dataset and divide by the range of that indicator to shrink the values of each indicator to the [0,1] interval. The specific calculation method is as follows: Where x is an indicator vector spanning a certain administrative region, x * This is the normalized index vector for a specific cross-administrative region.
[0018] S32: Then, use principal component analysis to select p principal components, such that the cumulative variance contribution rate of the p principal components is not less than 85%.
[0019] The index vectors spanning administrative regions after normalization are decentralized, and the covariance matrix is calculated to obtain n eigenvalues {λ}. i |i=1,2,...,n}(λ1>λ2>...>λ n ) and the corresponding standard eigenvectors {ω i Given |i=1,2,...,n}, select the first p largest eigenvalues {λ} i The standard eigenvectors corresponding to |i=1,2,...,p} form the projection transformation matrix W=[ω1,ω2,...,ω p ] T We obtain a p×m dimensional principal component matrix Y, such that the cumulative variance contribution rate α of the p principal components is... p Not less than 85%, that is Its calculation method is Y = WX, Y = [y1, y2, ..., y m ], where y m Let x be the index vector of the m-th cross-administrative region. m The principal component vector obtained after dimensionality reduction to p dimensions;
[0020] Step S33: Finally, using the selected principal components as input and the number of clusters q as a variable, K-means clustering is performed respectively, and the k value corresponding to the inflection point of the sum of squared errors and the maximum silhouette coefficient is taken as the optimal number of clusters, representing the reasonable number of types of cross-administrative regions within the city.
[0021] The sum of squared errors refers to the sum of the squares of the distances between each data point in q clusters and the centroid of its respective cluster. The silhouette coefficient is the difference between the average distance of the data point to other data points in the same cluster and the average distance to its nearest neighbor in a different cluster, divided by the maximum of the two. The specific calculation methods for the two indicators are as follows: Where SSE is the sum of squared errors, SC is the silhouette coefficient, q is the cluster index after clustering, and D is the sum of squared errors.q Let y be the set of principal components of samples belonging to the q-th cluster, and let y be the index vector x of the m-th cross-administrative region. m The principal component vector obtained after dimensionality reduction to p dimensions, μ i Let be the mean of all principal component vectors in the i-th cluster, a(i) be the mean distance between the principal component vector corresponding to sample i and the principal component vectors of all sample points in its cluster, and b(i) be the mean distance between the principal component vector corresponding to sample i and the principal component vectors corresponding to all sample points in its nearest cluster.
[0022] Preferably, S4 specifically involves: using the K-means clustering result when the optimal number of clusters is k as the final result, calculating the average value of each indicator for each cluster, obtaining the mean vector of each indicator for different k types of cross-administrative region areas, classifying the fusion areas by analyzing the mean vectors of different types of cross-administrative region areas, calculating the cross-regional road priority of different fusion areas, and ranking the cross-regional road priority. This step is detailed as using the K-means clustering result when the optimal number of clusters is k as the final result, and calculating the average value of each indicator for each cluster. Where x is an indicator vector spanning a certain administrative region, and D k Let K be the set of cross-administrative region zones belonging to the k-th cluster. The mean vectors of various indicators for different types of cross-administrative region zones are obtained. By analyzing the mean vectors, based on the degree of integration, the integration zones are divided into four types: highly integrated zone, basically integrated zone, initially integrated zone, and low-degree integrated zone. Then, the priority of cross-regional roads in different integration zones is calculated. The specific calculation method for the priority of cross-regional roads is as follows: I ij =μ ij (β1F1+β2F2+β3F3), where μ ij β is the regional integration coefficient corresponding to the cross-regional road belonging to the i-th integration zone and of level j, obtained through expert scoring and analytic hierarchy process. i Let F1 be the weight coefficient of the i-th indicator, which is determined by the entropy method. F2 is the project benefit factor, F3 is the construction cost factor, and F4 is the construction difficulty factor.
[0023] Where the regional fusion coefficient μ ij The possible values are as follows:
[0024] i = 1, 2, 3, 4 represent highly integrated zones, basically integrated zones, initially integrated zones, and low-level integrated zones, respectively; j = 1, 2, 3, 4, 5 represent expressways, rapid transit roads, arterial roads, secondary arterial roads, and local roads, respectively.
[0025] A system for prioritizing cross-regional road construction based on big data K-means clustering, used to implement the aforementioned method for prioritizing cross-regional road construction based on big data K-means clustering, characterized in that the system includes:
[0026] 1) Module for determining the scope of cross-administrative-region areas within the city: used to implement step S1;
[0027] 2) A cross-administrative region feature extraction module within the city is used to implement step S2;
[0028] 3) A module for determining the reasonable types and quantities of cross-administrative-region areas within the city, used to implement step S3;
[0029] 4) A priority ranking module for cross-regional road construction, used to implement step S4;
[0030] A terminal for a cross-regional road construction priority ranking system based on big data K-means clustering includes a memory, a processor, and at least one instruction or at least one computer program stored in the memory and loadable and run on the processor. The processor loads and runs at least one instruction or at least one computer program to implement the steps of the cross-regional road construction priority ranking method based on big data K-means clustering.
[0031] A non-transitory computer-readable storage medium storing a computer program thereon, characterized in that, when the computer program is executed by a processor, it implements the steps of a method for prioritizing cross-regional road construction based on big data K-means clustering.
[0032] The beneficial effects of this invention are:
[0033] 1) This invention provides a method for prioritizing cross-regional road construction based on big data K-means clustering, as well as a corresponding ranking system, terminal, and storage medium. The first step utilizes the feature length of the smallest administrative region to determine the scope of cross-administrative region areas within the city. The beneficial effects of this step are twofold: firstly, it avoids the problem of insufficient extraction of various indicator information due to an excessively small buffer; secondly, it avoids the problem of an excessively large buffer encompassing the entire administrative region, thereby affecting the accuracy of cross-administrative region indicators.
[0034] 2) This invention provides a method for prioritizing cross-regional road construction based on big data K-means clustering, as well as a corresponding ranking system, terminal, and storage medium. The second step combines big data to extract the fusion characteristics of cross-administrative regions within the city to obtain a sample dataset. The beneficial effect of this step is that it comprehensively and systematically extracts various indicators of cross-administrative regions, which can objectively and accurately reflect the cross-regional fusion characteristics of cross-administrative regions.
[0035] 3) This invention provides a method for prioritizing cross-regional road construction based on big data K-means clustering, as well as a corresponding sorting system, terminal, and storage medium. The third step is to determine the optimal number of clusters k based on the principal components, contour coefficients, and sum of squared errors. The beneficial effect of this step is to scientifically and rationally determine the number of cross-administrative regions, save a lot of computing resources, and improve the clustering effect of the algorithm.
[0036] 4) This invention provides a method for prioritizing cross-regional road construction based on big data K-means clustering, along with a corresponding ranking system, terminal, and storage medium. The fourth step involves calculating the average value of each indicator for each cluster based on the results of the K-means algorithm when the optimal number of clusters is k, obtaining the fusion characteristics of different k types of cross-administrative region areas, and ranking the priority of cross-regional road construction based on these fusion characteristics. The beneficial effect of this step is the effective classification of cross-regional urban areas. By analyzing the characteristics of each classification and considering factors such as the social benefits, construction costs, and implementation difficulties of cross-regional road construction projects, the priority of cross-regional road construction is ranked. This provides a more accurate quantitative calculation and ranking of cross-regional road construction priorities, offering a scientific basis for precise policy implementation in cross-regional road construction and urban transportation planning. Attached Figure Description
[0037] Figure 1 This is a flowchart of the steps of a cross-regional road construction priority ranking method based on big data K-means clustering according to the present invention;
[0038] Figure 2 This is a graph showing the variation of the k value and the contour coefficient in one embodiment;
[0039] Figure 3 This is a graph showing the change of the k value versus the sum of squared errors in one embodiment;
[0040] Figure 4 This is a heatmap of the mean vectors of each index of the cluster under each optimal k value in one embodiment; Detailed Implementation
[0041] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments. The step numbers in the following embodiments are merely for ease of explanation and do not imply any limitation on the order of the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art. Furthermore, it should be noted that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0042] Figure 1This is a flowchart illustrating the steps of a cross-regional road construction priority ranking method based on big data K-means clustering, as described in this invention. Referring to the flowchart, the cross-regional road construction priority ranking method based on big data K-means clustering includes the following steps:
[0043] S1 uses the characteristic length of the smallest administrative region to determine the extent of the cross-administrative region zone within the city;
[0044] S1 specifically refers to determining the extent of the cross-administrative-region zone within a city using the characteristic length of the smallest administrative region. The specific calculation method is as follows: Where r represents the radius of the buffer zone across administrative regions, and S is the area of the administrative region corresponding to the smallest administrative region among all administrative regions;
[0045] S2 combines big data to extract the integration characteristics of cross-administrative regions within the city;
[0046] S2 specifically involves selecting 12 indicators from four dimensions—urban land use, urban population, urban roads, and traffic operation—across administrative regions within the city. These indicators include the proportion of urban construction land, the distance between inter-regional passages, the proportion of inter-regional highways, the completion rate of inter-regional passages, population density, employment density, population density ratio, employment density ratio, road modal share, passenger station ratio, passenger flow density, and average saturation. These indicators are used to present the characteristics of cross-regional integration. The actual values of each integration indicator in different cross-administrative regions are obtained through big data such as mobile phone signaling and checkpoints, using survey and statistical methods to form a sample dataset.
[0047] S3 preprocesses the sample data in the sample dataset obtained in step S2, then selects a certain number of principal components as input, and performs K-means clustering with the number of clusters as a variable. Finally, it determines the optimal number of clusters k based on the calculated silhouette coefficient and the sum of squared errors, representing the reasonable number of types of cross-administrative regions within the city.
[0048] S3 specifically involves: First, normalizing the sample data using the max-min method to eliminate the influence of dimensions on the algorithm; then, selecting p principal components using principal component analysis, ensuring that the cumulative variance contribution rate of the p principal components is not less than 85%, thereby achieving dimensionality reduction. Finally, using the principal components as input and the number of clusters q as a variable, K-means clustering is performed, and the k value corresponding to the inflection point of the sum of squared errors and the maximum silhouette coefficient is taken as the optimal number of clusters, representing the reasonable number of cross-administrative region types within the city.
[0049] S4 calculates the result of the K-means algorithm when the optimal number of clusters k is obtained from step S3, obtains the final clustering result, calculates the average value of each indicator of each cluster, obtains the fusion characteristics of each indicator of different k types of cross-administrative regions, and formulates strategies for transportation development based on the fusion characteristics.
[0050] S4 specifically involves: taking the K-means clustering result when the optimal number of clusters k is used as the final result, calculating the average value of each indicator for each cluster, obtaining the mean vector of each indicator for different types of cross-administrative regions, analyzing the mean vector of different types of cross-administrative regions, calculating the priority of cross-regional roads in different integration areas based on integration characteristics, and ranking the priority of cross-regional road construction as described in the transportation development strategy.
[0051] A system for prioritizing cross-regional road construction based on big data K-means clustering, used to implement the aforementioned method for prioritizing cross-regional road construction based on big data K-means clustering, characterized in that the system includes:
[0052] 1) Module for determining the scope of cross-administrative-region areas within the city: used to implement step S1;
[0053] 2) A cross-administrative region feature extraction module within the city is used to implement step S2;
[0054] 3) A module for determining the reasonable types and quantities of cross-administrative-region areas within the city, used to implement step S3;
[0055] 4) A priority ranking module for cross-regional road construction, used to implement step S4;
[0056] A terminal for a cross-regional road construction priority ranking system based on big data K-means clustering includes a memory, a processor, and at least one instruction or at least one computer program stored in the memory and loadable and run on the processor. The processor loads and runs at least one instruction or at least one computer program to implement the steps of the cross-regional road construction priority ranking method based on big data K-means clustering.
[0057] A non-transitory computer-readable storage medium storing a computer program thereon, characterized in that, when the computer program is executed by a processor, it implements the steps of a method for prioritizing cross-regional road construction based on big data K-means clustering.
[0058] Taking Guangzhou as an example, this paper obtains integrated characteristic data across the 19 administrative districts of Guangzhou through mobile phone signaling, checkpoint data, and other data, using survey and statistical methods, and conducts case analysis. This embodiment provides a method for prioritizing cross-regional road construction based on big data K-means clustering, including the following steps:
[0059] S1 uses the characteristic length of the smallest administrative region to determine the extent of the cross-administrative region zone within the city;
[0060] S2 combines big data to extract the integration characteristics of cross-administrative regions within the city;
[0061] S3 preprocesses the sample data in the sample dataset obtained in step S2, then selects a certain number of principal components as inputs, and performs K-means clustering with the number of clusters as variables. Finally, the optimal number of clusters k is determined based on the calculated silhouette coefficient and the sum of squared errors, representing the reasonable number of types of urban cross-administrative regions.
[0062] S4 calculates the result of the K-means algorithm when the optimal number of clusters k is obtained from step S3, obtains the final clustering result, calculates the average value of each index of each cluster, obtains the fusion characteristics of different cross-administrative regions, calculates the cross-regional road priority of different fusion regions based on the fusion characteristics, and sorts the cross-regional road construction priority mentioned in the transportation development strategy.
[0063] S1 specifically refers to determining the extent of the cross-administrative-region zone within a city using the characteristic length of the smallest administrative region. The specific calculation method is as follows: Where r represents the radius of the buffer zone across administrative regions, and S is the area of the administrative region corresponding to the smallest administrative region among all administrative regions.
[0064] S2 specifically involves selecting 12 indicators from four dimensions—urban land use, urban population, urban roads, and traffic operation—across administrative regions within the city. These indicators include the proportion of urban construction land, the distance between inter-regional passages, the proportion of inter-regional highways, the completion rate of inter-regional passages, population density, employment density, population density ratio, employment density ratio, road modal share, passenger station ratio, passenger flow density, and average saturation. These indicators are used to present the characteristics of cross-regional integration. The actual values of each integration indicator in different cross-administrative regions are obtained through big data such as mobile phone signaling and checkpoints, using survey and statistical methods to form a sample dataset.
[0065] The population density ratio is the ratio of the population density within the buffer zone to the average population density of the centers of the two adjacent zones, and the employment density ratio is the ratio of the employment density within the buffer zone to the average employment density of the centers of the two adjacent zones.
[0066] S3 specifically involves: First, normalizing the sample data using the max-min method; then, selecting p principal components using principal component analysis, ensuring that the cumulative variance contribution rate of the p principal components is not less than 85%; finally, using the principal components as input and the number of clusters q as a variable, performing K-means clustering, and taking the k value corresponding to the inflection point of the sum of squared errors and the maximum of the silhouette coefficient as the optimal number of clusters, representing the reasonable number of types of cross-administrative-region areas within the city.
[0067] S31: First, the sample data is normalized using the max-min method to eliminate the influence of units on the algorithm;
[0068] S31: First, normalize the sample data using the max-min method;
[0069] Subtract the minimum value of each indicator from the sample dataset and divide by the range of that indicator to shrink the values of each indicator to the [0,1] interval. The specific calculation method is as follows: Where x is an indicator vector spanning a certain administrative region, x * This is the normalized index vector for a specific cross-administrative region.
[0070] S32: Then, use principal component analysis to select p principal components, such that the cumulative variance contribution rate of the p principal components is not less than 85%.
[0071] The index vectors spanning administrative regions after normalization are decentralized, and the covariance matrix is calculated to obtain n eigenvalues {λ}. i |i=1,2,...,n}(λ1>λ2>...>λ n ) and the corresponding standard eigenvectors {ω i Given |i=1,2,...,n}, select the first p largest eigenvalues {λ} i The standard eigenvectors corresponding to |i=1,2,...,p} form the projection transformation matrix W=[ω1,ω2,...,ω p ] T We obtain a p×m dimensional principal component matrix Y, such that the cumulative variance contribution rate α of the p principal components is... p Not less than 85%, that is Its calculation method is Y = WX, Y = [y1, y2, ..., y m ], where y m Let x be the index vector of the m-th cross-administrative region. m The principal component vector obtained after dimensionality reduction to p dimensions;
[0072] Step S33: Finally, using the selected principal components as input and the number of clusters q as a variable, K-means clustering is performed respectively, and the k value corresponding to the inflection point of the sum of squared errors and the maximum silhouette coefficient is taken as the optimal number of clusters, representing the reasonable number of types of cross-administrative regions within the city.
[0073] The sum of squared errors refers to the sum of the squares of the distances between each data point in q clusters and the centroid of its respective cluster. The silhouette coefficient is the difference between the average distance of the data point to other data points in the same cluster and the average distance to its nearest neighbor in a different cluster, divided by the maximum of the two. The specific calculation methods for the two indicators are as follows: Where SSE is the sum of squared errors, SC is the silhouette coefficient, q is the cluster index after clustering, and D is the sum of squared errors. q Let y be the set of principal components of samples belonging to the q-th cluster, and let y be the index vector x of the m-th cross-administrative region. m The principal component vector obtained after dimensionality reduction to p dimensions, μ i Let be the mean of all principal component vectors in the i-th cluster, a(i) be the mean distance between the principal component vector corresponding to sample i and the principal component vectors of all sample points in its cluster, and b(i) be the mean distance between the principal component vector corresponding to sample i and the principal component vectors corresponding to all sample points in its nearest cluster.
[0074] S4 specifically involves: using the K-means clustering result when the optimal number of clusters is k as the final result, calculating the average value of each indicator for each cluster, obtaining the mean vector of each indicator for different k types of cross-administrative regions, classifying the fusion areas by analyzing the feature vectors formed by each fusion feature, calculating the cross-regional road priority of different fusion areas, and ranking the cross-regional road priority. This step is detailed as follows:
[0075] The K-means clustering result when the optimal number of clusters is k is taken as the final result, and the average value of each index for each cluster is calculated. Where x is an indicator vector spanning a certain administrative region, and D k Let K be the set of cross-administrative region zones belonging to the k-th cluster. We obtain the mean vectors of various indicators for different types of cross-administrative region zones. By analyzing the mean vectors, based on the degree of integration, the integration zones are divided into four types: highly integrated, basically integrated, initially integrated, and low-integration zones. Then, we calculate the importance of cross-regional roads in different integration zones. The specific calculation method for the importance of cross-regional roads is as follows: I ij =μ ij (β1F1+β2F2+β3F3), where μ ij The regional integration coefficient, β, is the cross-regional road belonging to the i-th integration zone and of level j. It can be obtained through expert scoring and the analytic hierarchy process.i F1 is the weighting coefficient of the i-th indicator, which can be determined by the entropy method. F2 is the project benefit factor, which is represented by the predicted traffic volume after the project is completed and running stably. F3 is the construction cost factor, which is the construction cost of the cross-regional road. F4 is the construction difficulty factor, which is the area of unfavorable factors such as demolition and basic farmland involved in the project.
[0076] Where the regional fusion coefficient μ ij The values can be referenced. i = 1, 2, 3, 4 represent highly integrated zones, basically integrated zones, initially integrated zones, and low-level integrated zones, respectively; j = 1, 2, 3, 4, 5 represent expressways, rapid transit roads, arterial roads, secondary arterial roads, and local roads, respectively.
[0077] To visually demonstrate the effects of the present invention, Figures 2-4 The results of three experiments demonstrate that this invention can accurately identify and classify areas with different degrees of integration, mainly dividing them into four types and obtaining 12 characteristics of the four integration zones. This intuitively shows the different characteristics of the four integration zones, eliminating the current barriers to cross-regional integrated transportation research. It scientifically and quantitatively identifies the characteristics of cross-regional integration in cities, and can perform relatively accurate quantitative calculation and ranking of the priority of cross-regional road construction. This can provide a scientific basis for the precise implementation of road construction and urban transportation planning in cross-regional areas.
[0078] The above is a detailed description of the preferred embodiments of the present invention. However, the present invention is not limited to the embodiments described. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present invention. All such equivalent modifications or substitutions are included within the scope defined by the claims of this application.
Claims
1. A method for prioritizing cross-regional road construction based on big data K-means clustering, characterized in that, Includes the following steps: S1 uses the characteristic length of the smallest administrative region to determine the extent of the cross-administrative region zone within the city; S1 specifically refers to determining the extent of the cross-administrative-region zone within a city using the characteristic length of the smallest administrative region. The specific calculation method is as follows: ,in Indicates the radius of the buffer zone spanning administrative regions. The area of the administrative region corresponding to the smallest administrative region among all administrative regions; S2 combines big data to extract the integration characteristics of cross-administrative regions within the city, resulting in a sample dataset; S2 specifically involves selecting 12 indicators to represent the integrated characteristics: urban construction land ratio, inter-regional corridor spacing, inter-regional highway ratio, inter-regional corridor completion rate, population density, employment density, population density ratio, employment density ratio, road modal share, passenger station ratio, passenger flow density, and average saturation. Through big data analysis, actual values of each indicator across administrative regions are obtained using surveys and statistics. The first cross-administrative region Individual markers are The above 12 indicators, which span across administrative regions, are extracted to form... Dimensional sample dataset : ,in, Indicates the first m A cross-administrative region indicator vector; The population density ratio is the ratio of the population density within the buffer zone to the average population density of the centers of two adjacent zones, and the employment density ratio is the ratio of the employment density within the buffer zone to the average employment density of the centers of two adjacent zones. S3 preprocesses the sample data in the sample dataset obtained in step S2, then selects a certain number of principal components as inputs, and performs K-means clustering with the number of clusters as a variable. Finally, the optimal number of clusters is determined based on the silhouette coefficient and the sum of squared errors. k This indicates the number of reasonable types of areas that cross administrative districts within a city; S3 specifically involves: First, normalizing the sample data using the max-min method; then, selecting p principal components using principal component analysis, ensuring that the cumulative variance contribution rate of the p principal components is not less than 85%; finally, using the principal components as input and the number of clusters q as a variable, performing K-means clustering, and taking the k value corresponding to the inflection point of the sum of squared errors and the maximum of the silhouette coefficient as the optimal number of clusters, representing the reasonable number of types of cross-administrative-region areas within the city. S4 is based on the optimal cluster number obtained in step S3. k The optimal number of clusters is obtained as follows: k The results of the K-means algorithm are used to obtain the final clustering results. The average value of each index of each cluster is calculated to obtain the fusion characteristics of different types of cross-administrative regions. The cross-regional road priority of different fusion regions is calculated based on the fusion characteristics, and the cross-regional road priority is sorted. S4 specifically refers to: In detail, determining the optimal number of clusters... k The K-means clustering results were used as the final result, and the average values of each index for each cluster were calculated. ,in x For a certain cross-administrative region, For belonging to the first k The set of cross-administrative region zones composed of various clusters yields the mean vector of each indicator for different types of cross-administrative region zones. By analyzing the mean vector, based on the degree of integration, the integration zones are divided into four types: highly integrated zone, basically integrated zone, preliminary integrated zone, and low-degree integrated zone. Then, the priority of cross-regional roads in different integration zones is calculated. The specific calculation method for the priority of cross-regional roads is as follows: ,in The project belongs to the first i One integration zone, level 1 j The regional integration coefficient corresponding to the cross-regional roads was obtained through expert scoring and the analytic hierarchy process. For the first i The weighting coefficients of each indicator are determined using the entropy method. As a project benefit factor, As a construction cost factor, To determine the difficulty factor in construction; Among them, the regional integration coefficient The possible values are as follows: = , =1, 2, 3, 4 represent the highly integrated zone, the basically integrated zone, the preliminary integrated zone, and the low-level integrated zone, respectively; =1, 2, 3, 4, 5 represent highways, expressways, arterial roads, secondary arterial roads, and local roads, respectively.
2. The method for prioritizing cross-regional road construction based on big data K-means clustering as described in claim 1, characterized in that, The detailed steps for S3 are as follows: S31: First, normalize the sample data using the max-min method; Subtract the minimum value of each indicator from the sample dataset and divide by the range of that indicator to shrink the values of each indicator to the [0,1] interval. The specific calculation method is as follows: ,in, For a certain cross-administrative region, This is the normalized index vector for a specific cross-administrative region. S32: Then, principal component analysis is used to select... p Each principal component makes p The cumulative variance contribution rate of each principal component is no less than 85%; The index vectors spanning administrative regions after normalization are decentralized, and the covariance matrix is calculated to obtain... eigenvalues and the corresponding standard feature vectors Before selection p The largest eigenvalue The corresponding standard eigenvectors form the projection transformation matrix. ,get Principal component matrix of dimension , making p Cumulative variance contribution rate of each principal component Not less than 85%, that is The calculation method is as follows , ,in For the first m Indicator vectors spanning multiple administrative regions Dimensional reduction p The principal component vector obtained after dimensioning; Step S33: Finally, using the selected principal components as input, and the number of clusters... q As variables, K-means clustering was performed separately, and the inflection point of the sum of squared errors and the location of the maximum silhouette coefficient were identified. k The value, representing the optimal cluster number, indicates the reasonable number of types across administrative regions within the city. Among them, the sum of squared errors refers to q The silhouette coefficient is the sum of the squares of the distances between each data point in a cluster and the centroid of its cluster. It is the difference between the average distance of the data point to other data points in the same cluster and the average distance to its nearest neighbor in a different cluster, divided by the maximum of these two values. The specific calculation methods for these two indicators are as follows: , ,in, For the sum of squared errors, For the profile coefficient, q This refers to the cluster index after clustering. For belonging to the first q The set of principal components of samples from each cluster. y For the first m Indicator vectors spanning multiple administrative regions Dimensional reduction p The principal component vector obtained after dimensioning, For the first i The mean of all principal component vectors in each cluster. For the sample i The mean distance between the corresponding principal component vector and the principal component vectors of all sample points in its cluster. For the sample i The mean distance between the corresponding principal component vector and the principal component vectors corresponding to all sample points in its nearest cluster.
3. A cross-regional road construction priority ranking system based on big data K-means clustering, used to implement the cross-regional road construction priority ranking method based on big data K-means clustering as described in any one of claims 1-2, characterized in that, The system includes: 1) Module for determining the scope of cross-administrative regions within the city: used to implement step S1; 2) A cross-administrative region feature extraction module within the city, used to implement step S2; 3) A module for determining the reasonable types and quantities of cross-administrative-region areas within the city, used to implement step S3; 4) The priority ranking module for cross-regional road construction is used to implement step S4.
4. A terminal for a cross-regional road construction priority ranking system based on big data K-means clustering, comprising a memory, a processor, and at least one instruction or at least one computer program stored in the memory and loadable and executable on the processor, characterized in that, The processor loads and runs at least one instruction or at least one computer program to implement the steps of the cross-regional road construction priority ranking method based on big data K-means clustering as described in any one of claims 1 to 2.
5. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the cross-regional road construction priority ranking method based on big data K-means clustering as described in any one of claims 1 to 2.
Citation Information
Patent Citations
Road classification method based on adaptive K-means clustering algorithm
CN110598747A
City edge region extraction method based on multi-source data fusion
CN111651545A