Metro station planning method, device and equipment based on multi-source data and medium

Through the combination of multi-source data and clustering algorithms, the problem of traditional subway site planning methods relying on a single data source and failing to fully consider geographical factors is solved, and more reasonable site planning is achieved, and operational efficiency and service quality are improved.

CN120146509APending Publication Date: 2025-06-13CHINA HYDROELECTRIC ENGINEERING CONSULTING GROUP CHENGDU RESEARCH HYDROELECTRIC INVESTIGATION DESIGN AND INSTITUTE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510262777.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The traditional subway site planning method relies on a single or limited data source and fails to fully consider the geographical topography and surrounding land use types, resulting in unreasonable site planning, wasted resources and high construction costs.

Method used

The subway station planning method based on multi-source data is adopted, and the traffic flow data, population data and geographical information data of each sampling area of ​​the city is obtained, and the city is divided into different regions in combination with the clustering algorithm, and the comprehensive score of each potential subway station is calculated to determine the target subway station.

Benefits of technology

By combining multi-source data and geographical factors, more reasonable subway station planning is achieved, operational efficiency and service quality are improved, and resource waste and construction costs are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146509A_ABST
    Figure CN120146509A_ABST
Patent Text Reader

Abstract

The invention discloses a subway station planning method, device and equipment based on multi-source data, and a medium. The planning method comprises the following steps: acquiring traffic flow data, population data and geographic information data of each sampling area of a city; based on the traffic flow data and the population data, a clustering algorithm is adopted to divide a city into different clustering areas; calculating the central position of each clustering area, and determining each central position as a potential subway station; calculating a comprehensive score of each potential subway station based on the traffic flow data, the population data and the geographic information data of the sampling area where each potential subway station is located; and determining the potential subway station with the comprehensive score higher than a first preset value as a target subway station. According to the method, the subway stations are planned by combining the multi-source data and considering the geographic factors, so that the planning of the subway stations is more reasonable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of urban traffic planning, and particularly to a subway station planning method, device, equipment and medium based on multi-source data. Background Art

[0002] With the acceleration of the urbanization process, the urban population has grown rapidly, the traffic demand has increased sharply, and the urban traffic is facing huge pressure. As a large-capacity, fast and punctual public transportation mode, the subway plays a key role in alleviating urban traffic congestion, and the rationality of subway station planning directly affects the operation efficiency and service quality of the subway system. There are some unreasonable aspects in the current subway station planning methods.

[0003] On the one hand, traditional subway station planning methods often rely on single or limited data sources. Some planning is only based on population data, and stations are set in densely populated areas, but the actual traffic situation is ignored. This may lead to stations being set in areas with a large population but relatively smooth traffic, unable to effectively share traffic pressure and causing waste of resources.

[0004] On the other hand, the current planning does not comprehensively consider factors such as geographical terrain and surrounding land use types, which may lead to a substantial increase in construction costs and even construction difficulties. Moreover, there is a lack of overall consideration in the connection with surrounding buildings and public facilities, and it is not possible to achieve seamless transfer with other transportation modes such as buses and taxis, and effective connection with crowded areas such as commercial areas and residential areas, reducing the convenience and attractiveness of the subway. Summary of the Invention

[0005] The purpose of the present invention is to provide a subway station planning method, device, equipment and medium based on multi-source data to solve the problems that traditional subway station planning methods rely on single or limited data sources and do not comprehensively consider factors such as geographical terrain and surrounding land use types. By combining multi-source data and considering geographical factors for subway station planning, the planning of subway stations is made more reasonable.

[0006] The present invention is achieved by the following technical solutions:

[0007] In the first aspect, the present invention provides a subway station planning method based on multi-source data, including:

[0008] Obtaining traffic flow data, population data and geographical information data of each sampling area in the city;

[0009] Based on the traffic flow data and population data, using a clustering algorithm to divide the city into different clustering regions;

[0010] Calculate the central position of each clustering area and determine each central position as a potential subway station;

[0011] Based on the traffic flow data, population data, and geographic information data of the sampling areas where each potential subway station is located, calculate the comprehensive score of each potential subway station;

[0012] Determine the potential subway stations with comprehensive scores higher than the first preset value as target subway stations.

[0013] As a further solution of the present invention, based on the traffic flow data and population data, use the clustering algorithm to divide the city into different clustering areas, including the following steps:

[0014] Step 1: Based on the population data of n sampling areas, process to obtain the population density p i1 , the peak-hour population inflow p i2 and the peak-hour population outflow p i3 corresponding to the n sampling areas; based on the traffic flow data of the n sampling areas, process to obtain the vehicle flow p i4 , vehicle speed p i5 , the number of bus transfers p i6 and the number of taxi pick-ups and drop-offs p i7 ; where i = 1, 2,..., n, representing the serial number of the sampling area; p i1 , p i2 , p i3 , p i4 , p i5 , p i6 and p i7 are all represented by n-dimensional vectors;

[0015] Step 2: Take p i1 , p i2 , p i3 , p i4 , p i5 , p i6 and p i7 as feature data, and each sampling area corresponds to a feature data group P i =(p i1 , p i2 ,..., p ij ); where j = (1, 2,..., 7), representing the feature serial number of the feature data;

[0016] Step 3: Based on each feature data group P i , use the K-means clustering algorithm to obtain multiple clustering areas.

[0017] As a further solution of the present invention, based on each feature data group P i, using the K-means clustering algorithm to obtain multiple clustering regions, including the following steps:

[0018] Step 1: Determine the number of clusters K; where the number of clusters K is used to divide each feature data group P i into K clusters, K = (1, 2,..., n); among the K clusters, the set of feature data groups P i contained in the k-th cluster is C k , and the cluster center of the k-th cluster is E k = (μ k1 , μ k2 , …, μ kj ), k = (1, 2,..., K);

[0019] Step 2: Randomly select K from the n feature data groups P i as the cluster centers; where the initial cluster centers are E 1 , E 2 , …, E K ;

[0020] Step 3: Normalize each feature data group P i = (p i1 , p i2 , …, p ij ) to Z i = (z i1 , z i2 , …, z ij ), and calculate the distance d(Z i ) from each feature data group P i to each cluster center E k , and the calculation formula is i , E k )

[0021] Step 4: Assign each feature data group P i to the cluster C i , E k ) closest to it; where k = argmin k d(Z k ; where k = argmin l=1,2,...,K d(Z i , E l );

[0022] Step 5: For each cluster C k , recalculate and update its cluster center E k = (μ k1 , μ k2 , …, μ kj ); where if the cluster C kThere are n k feature data groups P i . For the j-th feature, represents the mean value of the feature data groups P k contained in the cluster C i on the j-th feature;

[0023] Step 6. Repeat Step 3 to Step 5 until for all k = (1, 2,..., K), K different clusters are obtained; where is the updated cluster center, is the cluster center before update, and ∈ is the threshold;

[0024] Step 7. Take the total area composed of the sampling areas corresponding to all the feature data groups P i contained in each cluster as the cluster area of each cluster, and K different cluster areas are obtained.

[0025] As a further solution of the present invention, calculating the central position of each cluster area and determining each central position as a potential subway station includes the following steps:

[0026] Obtain the central coordinates (x i , y i , y i ) in the geographic coordinate system of the sampling area corresponding to each feature data group P

[0027] contained in each of the K different clusters; i , y i ) to calculate the central positions of the K different cluster areas The calculation formula is: where n k represents the number of feature data groups P i contained in each of the K different clusters;

[0028] Take each central position as a potential subway station.

[0029] As a further solution of the present invention, calculating the comprehensive score of each potential subway station based on the traffic flow data, population data and geographic information data of the sampling area where each potential subway station is located includes the following steps:

[0030] Based on the population data, traffic flow data and geographic information data of the sampling area to which the central position corresponding to each potential subway station belongs, obtain the potential station population density ρ k , the potential station peak-hour population inflow Ik 、Population outflow volume O during peak hours at potential sites k 、Traffic flow data acquisition vehicle flow Q at potential sites k 、Vehicle speed V at potential sites k 、Number of bus transfer passengers B at potential sites k 、Number of taxi pick-ups and drop-offs D at potential sites k 、Area proportion of different land use types at potential sites Average slope θ of potential sites k 、Average elevation h of potential sites k ; Among them, m represents the serial number of different land use types;

[0031] According to the formula S k =α×R k +β×T k +γ×L k +δ×δ k Calculate the comprehensive score S of each potential subway site k ; Among them, R k represents the population score of the sampling area to which the central position corresponding to each potential subway site belongs, T represents the traffic flow score of the sampling area to which the central position corresponding to each potential subway site belongs, L k represents the land use type score of the sampling area to which the central position corresponding to each potential subway site belongs, G represents the geographical terrain score of the sampling area to which the central position corresponding to each potential subway site belongs; α is the weight coefficient of R k is the weight coefficient of T is the weight coefficient of L k is the weight coefficient of G is the weight coefficient of G k is the weight coefficient of G k is the weight coefficient of G k is the weight coefficient of G k .

[0032] As a further solution of the present invention, before calculating the comprehensive score of each potential subway site based on the traffic flow data, population data, and geographical information data of the sampling area where each potential subway site is located, the following steps are further included:

[0033] Calculate the feasibility score of each potential subway site based on geographical information data;

[0034] Remove potential subway sites with feasibility scores lower than the second preset value.

[0035] As a further solution of the present invention, calculating the feasibility score of each potential subway site based on geographical information data includes the following steps:

[0036] Based on the central location corresponding to each potential subway station and the geographical information data of the sampling area to which it belongs, obtain the geological condition priority W of the potential station k , the density H of underground pipelines at the potential station k , and the protection level A of surrounding buildings at the potential station k ;

[0037] According to the formula F k = λ 1 ×W k - λ 2 ×H k - λ 3 ×A k , calculate the feasibility score of each potential subway station; where, λ 1 is the weight coefficient of W k , λ 2 is the weight coefficient of H k , and λ 3 is the weight coefficient of A k .

[0038] Second, the present invention also provides a subway station planning device based on multi-source data, including:

[0039] A data acquisition module, configured to acquire traffic flow data, population data, and geographical information data of each sampling area in the city;

[0040] A region division module, configured to divide the city into different clustering regions by using a clustering algorithm based on the traffic flow data and the population data;

[0041] A potential station determination module, configured to calculate the central location of each clustering region and determine each central location as a potential subway station;

[0042] A potential station scoring module, configured to calculate the comprehensive score of each potential subway station based on the traffic flow data, population data, and geographical information data of the sampling area where each potential subway station is located;

[0043] A target station confirmation module, configured to determine the potential subway stations with the comprehensive score higher than a first preset value as target subway stations.

[0044] Third, the present invention also provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, it implements the subway station planning method based on multi-source data as in the first aspect.

[0045] Fourthly, the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the subway station planning method based on multi-source data as in the first aspect.

[0046] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0047] In the present invention, traffic flow data, population data, and geographical information data of each sampling area in the city are obtained; then, based on the traffic flow data and population data, a clustering algorithm is used to divide the city into different clustering regions; then, the central position of each clustering region is calculated, and each central position is determined as a potential subway station; subsequently, based on the traffic flow data, population data, and geographical information data of the sampling area where each potential subway station is located, the comprehensive score of each potential subway station is calculated; finally, the potential subway stations with comprehensive scores higher than the first preset value are determined as target subway stations; the present invention plans subway stations by combining multi-source data and considering geographical factors, making the planning of subway stations more reasonable. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts. In the drawings:

[0049] Figure 1 is a schematic flow chart of the subway station planning method based on multi-source data in the present invention;

[0050] Figure 2 is a schematic structural diagram of the subway station planning device based on multi-source data in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] To make the objectives, technical solutions, and advantages of the present invention more clearly understood, the present invention will be further described in detail below in conjunction with the embodiments and the drawings. The illustrative embodiments and descriptions thereof of the present invention are only used to explain the present invention and are not intended to limit the present invention.

[0052] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs; the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above drawings are intended to cover non-exclusive inclusion.

[0053] As used herein, the mention of "embodiments" means that the specific features, structures, or characteristics described in connection with the embodiments may be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein may be combined with other embodiments.

[0054] Please refer to Figure 1 , a subway station planning method based on multi-source data provided in an embodiment of the present application includes the following steps:

[0055] S101. Obtain traffic flow data, population data, and geographic information data of each sampling area in the city;

[0056] S102. Based on the traffic flow data and population data, use a clustering algorithm to divide the city into different clustering regions;

[0057] S103. Calculate the central position of each clustering region and determine each central position as a potential subway station;

[0058] S104. Based on the traffic flow data, population data, and geographic information data of the sampling areas where each potential subway station is located, calculate the comprehensive score of each potential subway station;

[0059] S105. Determine the potential subway stations with comprehensive scores higher than the first preset value as target subway stations.

[0060] Specifically, for step S101:

[0061] For traffic flow data, it needs to be collected through various channels. Traffic monitoring devices such as induction coils and video monitoring cameras can be set at main roads, intersections, etc. in the city. These devices can record information such as the passing volume and driving speed of motor vehicles and non-motor vehicles at different times in real time. At the same time, relevant data can also be obtained from the public transportation system, including the passenger capacity, operation frequency of buses, and the station information of passengers getting on and off, so as to comprehensively evaluate the traffic flow conditions of each sampling area.

[0062] For population data, the census data of government departments can be used, which details information such as the population quantity, age distribution, and family structure of each region. The data of mobile operators can also be used. By analyzing the distribution density and activity time of mobile phone users in different regions, the population flow and resident population quantity of each sampling area can be roughly estimated.

[0063] For geographic information data, satellite remote sensing technology can be used to obtain topographic and geomorphic information, such as the height of the terrain and whether there are natural obstacles such as rivers and lakes. At the same time, it also includes the existing building distribution data in the city, such as the specific locations and scales of commercial areas, residential areas, and industrial areas.

[0064] For step S102:

[0065] Based on the obtained traffic flow data and population data, using a clustering algorithm to divide the city into different clustering regions is a complex and crucial step.

[0066] Select a suitable clustering algorithm, such as the K-Means clustering algorithm. When using this algorithm, the traffic flow data and population data are used as input features. Each parameter in the traffic flow data (such as the traffic volume of different roads, the passenger volume of public transportation, etc.) and the population data (population quantity, population density, etc.) together constitute a high-dimensional data space.

[0067] Still taking the K-Means algorithm as an example, in the initialization stage, the number of clusters needs to be preset in advance or determined by a certain method. Then, randomly select the initial cluster centers. Next, calculate the distances from the data points in each sampling area to these cluster centers. Here, the distance metric can select a suitable method according to the characteristics of the data, such as the Euclidean distance, etc. According to the distance, divide the data points in each sampling area into the class represented by the nearest cluster center.

[0068] After each division, recalculate the center points of each cluster until the cluster centers no longer change significantly or reach the preset number of iterations. Through such a clustering process, areas with similar traffic flow and population characteristics can be divided into the same class, thus forming different clustering regions.

[0069] For step S103:

[0070] For each clustering region, since it is composed of multiple sampling areas, these sampling areas have their own coordinates in the geographical space. The central position can be calculated using the weighted average method. Using the traffic flow and population data as weights, higher weights are assigned to the sampling areas with large traffic flow and high population density.

[0071] Through this weighted average calculation method, the location of potential subway stations can be more biased towards areas with more concentrated traffic flow and population. The potential subway stations determined in this way can better serve the surrounding residents and relieve traffic pressure.

[0072] For step S104:

[0073] For traffic flow data, it is necessary to analyze the traffic volume of roads around potential subway stations, the transfer volume of public transportation, etc. If the traffic volume around a potential subway station is large and public transportation transfers are frequent, it indicates that the station is of great significance for alleviating traffic congestion and should be given a higher score. Corresponding scoring criteria can be set according to the magnitude of traffic flow. For example, stations with traffic volume within a certain high threshold range receive higher scores, those in the medium traffic volume range receive moderate scores, and those with lower traffic volume receive lower scores.

[0074] For population data, it is necessary to consider the population density, population flow patterns, etc. around potential subway stations. In areas with high population density, the service demand for the station is large, and the score should be increased accordingly. The score can be determined based on the classification of population density. At the same time, for areas with frequent population flow (such as near commercial areas and office areas), appropriate additional points should also be given because the demand for the subway in these areas varies significantly at different times.

[0075] For geographical information data. If a potential subway station is located in an area with flat terrain, no large obstacles, and a surrounding building layout conducive to subway construction and passenger entry and exit, a higher score should be given. For example, situations where there are no large rivers, mountains, etc. around the station that impede traffic evacuation, and there are convenient connection channels with surrounding commercial areas and residential areas can all increase the score.

[0076] By quantifying and weighting these factors above, the comprehensive score of each potential subway station is finally calculated.

[0077] Regarding step S105:

[0078] The setting of the first preset value needs to comprehensively consider factors such as the overall urban planning, construction budget, and traffic development strategy. This preset value can be obtained through scientific analysis and summary of practical experience.

[0079] After calculating the comprehensive score of each potential subway station, the score is compared with the first preset value. If the comprehensive score is higher than the first preset value, it indicates that the potential subway station meets relatively high standards in terms of traffic flow service capacity, population coverage, and geographical conditions. The locations of these identified target subway stations will be the focus of subway station planning, and more detailed engineering design, line planning, etc. can be carried out around these target stations in the future to ensure that the construction of subway stations can meet the needs of urban traffic and residents' travel to the greatest extent.

[0080] In an alternative embodiment, based on traffic flow data and population data, a clustering algorithm is used to divide the city into different clustering regions, including the following steps:

[0081] Step 1: Based on the population data of n sampling areas, process to obtain the population density p corresponding to the n sampling areas i1 , the peak-hour population inflow p i2 and the peak-hour population outflow p i3 ; Based on the traffic flow data of n sampling areas, process to obtain the vehicle flow p i4 , vehicle speed p i5 , the number of bus transfers p i6 and the number of taxi pick-ups and drop-offs p i7 ; where i = 1, 2, …, n, representing the serial number of the sampling area; p i1 , p i2 , p i3 , p i4 , p i5 , p i6 and p i7 are all represented by n-dimensional vectors;

[0082] Step 2: Take p i1 , p i2 , p i3 , p i4 , p i5 , p i9 and p i7 as feature data, and each sampling area corresponds to a feature data group P i = (p i1 , p i2 , …, p ij ); where j = (1, 2, …, 7), representing the feature serial number of the feature data;

[0083] Step 3: Based on each feature data group P i , adopt the K-means clustering algorithm to obtain multiple clustering areas.

[0084] Specifically, for Step 1:

[0085] Regarding the population data, the population density, peak-hour population inflow, and peak-hour population outflow are obtained through analysis. The population density reflects the concentration degree of the population in a specific area and can be calculated by dividing the resident population quantity of the sampling area by the area of the region. The peak-hour population inflow reflects the number of people flowing into the sampling area from other areas during the traffic peak period. This can be determined by counting the number of people in the main transportation hubs (such as subway stations, bus stops, etc.) entering the area during the peak hour. Similarly, the peak-hour population outflow refers to the number of people flowing from the sampling area to other areas during the peak hour and can be obtained by monitoring the number of people at the main exits of the area. These three indicators are represented by n-dimensional vectors, providing specific population characteristic parameters for subsequent analysis.

[0086] Regarding traffic flow data, vehicle flow, vehicle speed, the number of bus transfers, and the number of taxi pick-ups and drop-offs are processed. Vehicle flow can be counted by traffic monitoring devices (such as induction coils, cameras, etc.) set on the road to count the number of vehicles passing through the sampling area within a specific time period. Vehicle speed can be calculated by the passing time and driving distance of vehicles recorded by the monitoring devices. The number of bus transfers can be obtained from the monitoring data or ticketing system of bus stops, reflecting the activity of public transportation in this area. The number of taxi pick-ups and drop-offs can be counted through the operation data of taxis, reflecting the traffic demand and mobility in this area. These indicators are also represented as n-dimensional vectors, comprehensively reflecting the traffic conditions of the sampling area.

[0087] For Step 2:

[0088] Take the population density p i1 , the peak-hour population inflow p i2 , the peak-hour population outflow p i3 , the vehicle flow p i4 , the vehicle speed p i5 , the number of bus transfers p i6 and the number of taxi pick-ups and drop-offs p i7 as feature data. Each sampling area corresponds to a feature data group p i = (p i1 , p i2 , …, p ij ). Such a feature data group combines information on both population and traffic, and can more comprehensively describe the characteristics of each sampling area. Among them, j = (1, 2, …, 7) represents the feature serial number of the feature data, which is convenient for distinguishing and processing different features in subsequent analysis.

[0089] By integrating these feature data together, it provides a specific data basis for dividing the city using clustering algorithms.

[0090] For Step 3:

[0091] First, select an appropriate number of clusters K, which can be determined by methods such as empirical judgment and the elbow method. Then, randomly initialize K cluster centers. For each feature data group P i , calculate its distance from each cluster center, and metric methods such as Euclidean distance can be used. According to the principle of the closest distance, assign this feature data group to the corresponding cluster.

[0092] After the allocation is completed, recalculate the center of each cluster, which is the mean of each feature. Then, perform the processes of allocation and center update again until the cluster centers no longer change significantly or reach the preset number of iterations. Through such an iterative process, the sampling areas with similar population and traffic characteristics are divided into the same cluster area.

[0093] The multiple finally obtained cluster areas have high internal similarity in terms of population and traffic characteristics, while there are obvious differences between different cluster areas.

[0094] In an alternative embodiment, based on each feature data group P i , using the K-means clustering algorithm to obtain multiple cluster areas, including the following steps:

[0095] Step 1: Determine the number of clusters K; where the number of clusters K is used to divide each feature data group P i into K clusters, K = (1, 2,..., n); in the K clusters, the set of feature data groups P i contained in the k-th cluster is C k , and the cluster center of the k-th cluster is E k = (μ k1 , μ k2 ,..., μ kj ), k = (1, 2,..., K);

[0096] Step 2: Randomly select K from the n feature data groups P i as the cluster centers; where the initial cluster centers are E 1 , E 2 ,..., E K ;

[0097] Step 3: Normalize each feature data group P i = (p i1 , p i2 ,..., p ij ) to Z i = (z i1 , z i2 ,..., z ij ), and calculate the distance d(Z i , E i ) from each feature data group P k to each cluster center E i , and the calculation formula is k )

[0098] Step 4: Allocate each feature data group P i to the cluster center E i to which its distance d(Z k)The nearest cluster center E k The cluster C to which it belongs k ; where k = argmin l=1,2,…,K d(Z i , E l );

[0099] Step Five: For each cluster C k , recalculate and update its cluster center E k =(μ k1 , μ k2 ,…, μ kj ); where, if there are n k feature data groups P k in the cluster C i , for the j-th feature, represents the mean value of the feature data group P k contained in the cluster C i on the j-th feature;

[0100] Step Six: Repeat Steps Three to Five until for all k=(1, 2,…, K), K different clusters are obtained; where, is the updated cluster center, is the cluster center before update, and ∈ is the threshold;

[0101] Step Seven: Take the total area composed of the sampling areas corresponding to all the feature data groups P i contained in each cluster as the cluster area of each cluster, and obtain K different cluster areas.

[0102] Specifically, for Step One:

[0103] The number of clusters K is used to divide each feature data group P i into K clusters, and its value range is K=(1, 2,…, n). Among these K clusters, the set of feature data groups P i contained in the k-th cluster is C k , and the cluster center of the k-th cluster is k k =(μ k1 , μ k2 ,…, μ kj ). An appropriate number of clusters K can make the clustering result better reflect the internal structure of the data and divide the sampling areas with similar features into the same cluster.

[0104] For Step Two:

[0105] These initial cluster centers are denoted as E 1 , E 2 ,…, E K. Randomly selecting the cluster centers is an initial step of the K-means algorithm, which provides initial reference points for the subsequent iterative process. Through random selection, the algorithm can gradually adjust the positions of the cluster centers in the subsequent iterations to more accurately represent the characteristics of each cluster.

[0106] For Step 3:

[0107] First, normalize each feature data group P i =(p i1 , p i2 , …, p ij ) to Z i =(z i1 , z i2 , …, z ij ). The normalization formula can be represents the feature serial number; represents the mean of the j-th feature, represents the standard deviation of the j-th feature. The purpose of normalization is to eliminate the influence caused by factors such as measurement units among different features, so that each feature has the same weight when calculating distances.

[0108] Then, calculate the distance d(Z i ) from each feature data group P i to each cluster center E k . The calculation formula is i , E k ) as This distance formula is a form of Euclidean distance. It measures the distance between the feature data group and the cluster center by calculating the square root of the sum of the squares of the differences in each feature dimension. In this way, the proximity of each feature data group to each cluster center can be determined, providing a basis for the subsequent assignment operation.

[0109] For Step 4:

[0110] This step is to assign each feature data group P i to the cluster C i , E k ) that is closest to it. Among them, k = argmin k d(Z k ), which means selecting the one that makes the distance d(Z l=1,2,…,K ), E i , E l ) the smallest. This indicates choosing the one that makes the distance d(Z i ), E k) The cluster corresponding to the smallest cluster center. Through this assignment method, each feature data group is divided into the cluster closest to its features, so that the feature data groups in the same cluster have more similar features, thus initially forming different clusters.

[0111] For step five:

[0112] For each cluster C k , it is necessary to recalculate and update its cluster center E k =(μ k1 ,μ k2 ,…,μ kj ). If there are n k feature data groups P k in cluster C i , for the j-th feature, which means the mean value of the feature data group P k included in cluster C i on the j-th feature. By recalculating the cluster center, it can more accurately reflect the average features of the feature data groups within the current cluster. This is an important iterative step in the K-means algorithm. As the iteration progresses, the cluster center will gradually stabilize and the clustering result will be more accurate.

[0113] For step six:

[0114] Repeat steps three to five until for all k=(1,2,…,K), K different clusters are obtained.

[0115] For The Euclidean norm (L2 norm) can be used to calculate The calculation formula is:

[0116]

[0117] Where, is the updated cluster center, is the cluster center before update, and ∈ is the threshold. This threshold ∈ determines when the algorithm stops iterating. When the change in the cluster center is less than this threshold, it is considered that the clustering has reached a stable state and the algorithm stops iterating. By continuously repeating the iterative process, the cluster center and the clustering division will be continuously optimized, and finally a more ideal clustering result will be obtained.

[0118] For step seven:

[0119] All the feature data groups P included in each cluster iThe total area composed of the corresponding sampling areas is used as the clustering area for each cluster, and K different clustering areas are obtained. In this way, through the previous series of steps, the clustering operation at the data level is transformed into the division of clustering areas at the geographical area level.

[0120] In an alternative embodiment, calculating the central position of each clustering area and determining each central position as a potential subway station includes the following steps:

[0121] Obtain each feature data group P included in each of the K different clusters i The central coordinates (x i , y i ) of the corresponding sampling area in the geographical coordinate system;

[0122] Calculate the central position of the K different clustering areas according to the coordinates (x i , y i ) The calculation formula is: Where n k represents the number of feature data groups P included in each of the K different clusters i ;

[0123] Take each central position as a potential subway station.

[0124] Specifically, since each feature data group P i corresponds to a sampling area, and these sampling areas have their specific positions in the geographical space, it is necessary to determine their central coordinates. This may require the use of a Geographic Information System (GIS) tool or relevant geographical data to obtain accurate coordinate information.

[0125] Calculate the central position of the K different clustering areas according to the obtained coordinates (x i , y i ) The calculation formula is and and Where n k represents the number of feature data groups P included in each of the K different clusters i . This calculation process is actually a weighted average of the central coordinates of all sampling areas within each cluster (here the weights are the same, all being ). In this way, a coordinate point that can represent the central position of the entire clustering area can be obtained, and this coordinate point comprehensively considers the position information of each sampling area within the clustering area.

[0126] Finally, take the calculated central position As a potential subway station. The determination of the location of the potential subway station is considered based on the overall characteristics and geographical distribution of the clustering area. Setting the station near the center of the clustering area is conducive to better serving the population within the clustering area and can, to a certain extent, evenly cover the traffic demand of the entire clustering area.

[0127] In an alternative embodiment, based on the traffic flow data, population data, and geographical information data of the sampling areas where each potential subway station is located, calculate the comprehensive score of each potential subway station, including the following steps:

[0128] Based on the central location corresponding to each potential subway station Obtain the population density ρ of the potential station from the population data, traffic flow data, and geographical information data of the sampling area to which it belongs k 、The population inflow I during the peak period of the potential station k 、The population outflow O during the peak period of the potential station k 、Obtain the traffic flow Q from the traffic flow data of the potential station k 、The vehicle speed V of the potential station k 、The number of bus transfer passengers B at the potential station k 、The number of taxi pick-ups and drop-offs D at the potential station k 、The area proportion of different land use types at the potential station The average slope θ of the potential station k 、The average elevation h of the potential station k ; where, The m in represents the serial number of different land use types;

[0129] According to the formula S k =α×R k +β×T k +γ×L k +δ×G k , calculate the comprehensive score S of each potential subway station k ; where, R k Represents the population score of the sampling area to which the central location corresponding to each potential subway station belongs T k Represents the traffic flow score of the sampling area to which the central location corresponding to each potential subway station belongs L k Represents the land use type score of the sampling area to which the central location corresponding to each potential subway station belongs G k Represents the geographical terrain score of the sampling area to which the central location corresponding to each potential subway station belongs α is the weight coefficient of R k β is the weight coefficient of T kThe weight coefficient, γ is L k The weight coefficient, δ is G k The weight coefficient.

[0130] Specifically, for R k , Represents the central position corresponding to each potential subway station The population density score of the sampling area to which it belongs, Represents the central position corresponding to each potential subway station The population flow score of the sampling area to which it belongs, ω 1 Is The weight coefficient of, ω 2 Is The weight coefficient of, ω 1 +ω 2 = 1, ρ max Represents the maximum population density of the city, I max Is the maximum population inflow of the city, O max Is the maximum population outflow of the city;

[0131] For T k , Is the central position corresponding to each potential subway station The road traffic flow score of the sampling area to which it belongs, Is the central position corresponding to each potential subway station The public transportation transfer score of the sampling area to which it belongs, ω 3 Is The weight coefficient of, ω 4 Is The weight coefficient of, ω 3 +ω 4 = 1, Q max Is the maximum traffic flow of the city, V max Is the maximum vehicle speed of the city, B max Is the maximum number of people transferring by bus in the city, D max Is the maximum number of passengers getting on and off taxis in the city;

[0132] For L k , σ m Is the central position corresponding to each potential subway station The corresponding scores of different land use types of the sampling area to which it belongs;

[0133] For G k , Is the central position corresponding to each potential subway station The terrain slope score of the sampling area to which it belongs is the central position corresponding to each potential subway station The elevation score of the sampling area to which it belongs, ω 5 is the weight coefficient of, ω 6 is the weight coefficient of, ω 5 + ω 6 = 1, θ max is the maximum slope of the city, h ref is the reference elevation of the city, Δh max is the maximum elevation difference of the city.

[0134] In an alternative embodiment, before calculating the comprehensive score of each potential subway station based on the traffic flow data, population data, and geographic information data of the sampling area where each potential subway station is located, the following steps are further included:

[0135] Based on the geographic information data, calculate the feasibility score of each potential subway station;

[0136] Remove the potential subway stations whose feasibility scores are lower than the second preset value.

[0137] Specifically, first, extract multiple key factors from the geographic information data to evaluate feasibility. For example, considering the terrain factor, if a potential subway station is located in a flat area, compared with an area with large terrain undulations, the engineering difficulty faced during construction will be relatively small, so a higher score can be given. For areas with complex terrains such as mountains and steep slopes, the score will be correspondingly reduced.

[0138] Secondly, analyze the geological conditions. A stable geological structure is crucial for the construction and long-term operation safety of subway stations. If the geological exploration data shows that the geological conditions of the area where a potential subway station is located are good, such as stable rock layers and strong soil bearing capacity, a higher feasibility score should be given. Conversely, if there are potential geological disaster hazards, such as seismic fault zones and karst development areas, the score will be very low.

[0139] Furthermore, consider the distribution of surrounding buildings and facilities. If there are no large immovable buildings (such as historical and cultural relics, important industrial facilities, etc.) around a potential subway station, and the connection with the existing urban infrastructure (such as roads, water supply and drainage systems, etc.) is relatively convenient, this will be beneficial to the construction and subsequent operation of the station, and the score can be appropriately increased.

[0140] Finally, comprehensively consider the above various geographic information factors, and by setting reasonable weights, score each potential subway station to obtain the feasibility score of each potential subway station.

[0141] After calculating the feasibility scores of each potential subway station, compare them with the second preset value. If the feasibility score of a potential subway station is lower than the second preset value, it means that there are significant obstacles or unfavorable factors in the geographical conditions of this station, which may lead to problems such as excessive construction costs, excessive engineering difficulties, or serious impacts on the surrounding environment.

[0142] Therefore, removing these potential subway stations with lower feasibility scores can effectively avoid unnecessary planning and investment in areas where the geographical conditions are not feasible, thereby improving the overall efficiency and scientific nature of subway station planning and ensuring that subsequent planning and construction work can proceed more smoothly.

[0143] In an alternative embodiment, based on the geographical information data, calculating the feasibility scores of each potential subway station includes the following steps:

[0144] Based on the geographical information data of the sampling area to which the central position corresponding to each potential subway station belongs obtain the geological condition priority W k of the potential station, the density H k of underground pipelines at the potential station, and the protection level A k of surrounding buildings of the potential station;

[0145] According to the formula F k = λ 1 ×W k - λ 2 ×H k - λ 3 ×A k , calculate the feasibility scores of each potential subway station; where λ 1 is the weight coefficient of W k , λ 2 is the weight coefficient of H k , and λ 3 is the weight coefficient of A k .

[0146] Specifically, for the geological condition priority W k of the potential station, it can be determined through geological exploration data. If the geological conditions are good, such as stable strata, firm rocks, and no obvious hidden geological hazards, then the value of W k will be relatively high. This is because good geological conditions are conducive to the construction and long-term stable operation of subway stations, reducing engineering risks and costs.

[0147] For the density H k, it can be obtained from the distribution data of urban underground pipelines. If the underground pipelines are dense, a large amount of pipeline relocation or protection work may be required during the construction of subway stations, which will increase the complexity and cost of the project. Therefore, the value of H k will be determined accordingly according to the density. The higher the density, the k larger the value of H.

[0148] The protection level A of the buildings around the potential station k also needs to be considered. Some historical and cultural buildings, important public facilities or buildings for special purposes may have a high protection level. If such buildings exist around the potential subway station, then the value of A k will be relatively high. This is because special measures need to be taken during the construction to protect these buildings and avoid damaging them.

[0149] Formula F k = λ 1 ×W k - λ 2 ×H k - λ 3 ×A k is used to comprehensively consider the above geographical information factors to calculate the feasibility score of the potential subway station.

[0150] Among them, λ 1 is the weight coefficient of W k , λ 2 is the weight coefficient of H k , and λ 3 is the weight coefficient of A k . These weight coefficients can be determined according to the importance of each factor in the feasibility of subway station construction.

[0151] For example, if the impact of geological conditions on station construction is considered to be the most critical, then the value of λ 1 may be relatively large to highlight the priority of geological conditions and the importance of W k in calculating the feasibility score.

[0152] By multiplying the values of each factor by the corresponding weight coefficients and performing operations, a feasibility score F k can be obtained, which is used to evaluate the feasibility of the potential subway station in terms of geographical information.

[0153] In the embodiments of the present application, a subway station planning method based on multi-source data is provided. This method plans the subway station by combining multi-source data and considering geographical factors, making the planning of the subway station more reasonable.

[0154] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this document, the execution of these steps has no strict order limitation, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0155] Based on the same inventive concept, an embodiment of the present application also provides a subway station planning device based on multi-source data for implementing the subway station planning method based on multi-source data described above. The implementation solutions provided by this device to solve problems are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more embodiments of the subway station planning device based on multi-source data provided below can refer to the limitations on the subway station planning method based on multi-source data in the above text, and will not be repeated here.

[0156] In an exemplary embodiment, as Figure 2 shown, a subway station planning device 200 based on multi-source data is provided, including:

[0157] A data acquisition module 201, configured to acquire traffic flow data, population data, and geographic information data of each sampling area in the city.

[0158] A region division module 202, configured to divide the city into different clustering regions based on the traffic flow data and population data by using a clustering algorithm.

[0159] A potential site determination module 203, configured to calculate the central position of each clustering region and determine each central position as a potential subway station.

[0160] A potential site scoring module 204, configured to calculate the comprehensive score of each potential subway station based on the traffic flow data, population data, and geographic information data of the sampling area where each potential subway station is located.

[0161] A target site confirmation module 205, configured to determine the potential subway stations with a comprehensive score higher than a first preset value as target subway stations.

[0162] Optionally, the region division module 202 includes:

[0163] A data processing unit for processing population data of n sampling regions to obtain population densities p corresponding to the n sampling regions i1 , population inflow p during peak hours i2 and population outflow p during peak hours i3 ; processing traffic flow data of n sampling regions to obtain traffic volumes p corresponding to the n sampling regions i4 , vehicle speeds p i5 , bus transfer numbers p i6 and taxi pick-up and drop-off numbers p i7 ; where i = 1, 2, …, n represents the serial number of the sampling region; p i1 , p i2 , p i3 , p i4 , p i5 , p i6 and p i7 are all represented by n-dimensional vectors.

[0164] A feature data group generation unit for using p i1 , p i2 , p i3 , p i4 , p i5 , p i6 and p i7 as feature data, with each sampling region corresponding to a feature data group p i = (p i1 , p i2 , …, p ij ); where j = (1, 2, …, 7) represents the feature serial number of the feature data.

[0165] A clustering unit for obtaining multiple clustering regions based on each feature data group P i using the K-means clustering algorithm.

[0166] Optionally, the clustering unit is specifically used to perform the following operations:

[0167] Step 1: Determine the number of clusters K; where the number of clusters K is used to divide each feature data group P i into K clusters, K = (1, 2, …, n); among the K clusters, the set of feature data groups P i contained in the k-th cluster is C k , and the cluster center of the k-th cluster is E k = (μ k1 , μ k2 , …, μ kj ), k = (1, 2, …, K);

[0168] Step 2: Select from the n feature data groups Pi Randomly select K from them as the clustering centers; among them, the initial clustering centers are E 1 , E 2 , …, E K ;

[0169] Step 3. Standardize each feature data group P i =(p i1 , p i2 , …, p ik ) to Z i =(z i1 , z i2 , …, z ij ), and calculate the distance d(Z i ) from each feature data group P i to each clustering center E k . The calculation formula is i , E k )

[0170] Step 4. Assign each feature data group P i to the cluster C i , E k ) to which the nearest clustering center E k belongs; where k = argmin l=1,2,…,K d(Z i , E l ); k )

[0171] Step 5. For each cluster C k , recalculate and update its clustering center E k1 =(μ k2 , μ kj , …, μ k ); where, if there are n k feature data groups P i in the cluster C k , for the j-th feature, represents the mean of the feature data group P i included in the cluster C i on the j-th feature;

[0172] Step 6. Repeat Step 3 to Step 5 until for all k=(1, 2, …, K), K different clusters are obtained; where is the updated clustering center, is the clustering center before update, and ∈ is the threshold;

[0173] Step 7. All feature data groups P included in each clusteri The total area composed of the corresponding sampling areas is used as the clustering area for each cluster, and K different clustering areas are obtained.

[0174] Optionally, the potential site determination module 203 includes:

[0175] A geographic coordinate acquisition unit for acquiring the central coordinates (x i in the geographic coordinate system of the corresponding sampling area i , y i ) of each feature data group P included in each of the K different clusters;

[0176] A position calculation unit for calculating the central positions of the K different clustering areas according to the coordinates (x i , y i ); the calculation formula is: The calculation formula is: where n k represents the number of feature data groups P included in each of the K different clusters; i the number of;

[0177] A position determination unit for using the central positions as potential subway stations.

[0178] Optionally, the potential site scoring module 204 includes:

[0179] A potential site data acquisition unit for obtaining the population density ρ of the potential site, the population inflow I during the peak period of the potential site, the population outflow O during the peak period of the potential site, the traffic flow Q of the potential site, the vehicle speed V of the potential site, the number of bus transfer passengers B of the potential site, the number of taxi pick-up and drop-off passengers D of the potential site, the area proportion of different land use types of the potential site, the average slope θ of the potential site, and the average elevation h of the potential site based on the population data, traffic flow data, and geographic information data of the sampling area to which the central position corresponding to each potential subway station belongs; where m represents the serial number of different land use types. k , the population inflow I during the peak period of the potential site k , the population outflow O during the peak period of the potential site k , the traffic flow Q of the potential site to obtain the traffic flow k , the vehicle speed V of the potential site k , the number of bus transfer passengers B of the potential site k , the number of taxi pick-up and drop-off passengers D of the potential site k , the area proportion of different land use types of the potential site the average slope θ of the potential site k , the average elevation h of the potential site k ; among them, m represents the serial number of different land use types.

[0180] A scoring calculation unit for calculating according to the formula S k =α×R k +β×T k +γ×L k+δ×G k , calculate the comprehensive score S of each potential subway station k .

[0181] Among them, R k represents the population score of the sampling area where the central location corresponding to each potential subway station is located, T represents the traffic flow score of the sampling area where the central location corresponding to each potential subway station is located, L k represents the central location corresponding to each potential subway station is located in the land use type score of the sampling area, G k represents the central location corresponding to each potential subway station is located in the geographical terrain score of the sampling area; α is the weight coefficient of R k represents the central location corresponding to each potential subway station is located in the sampling area, β is the weight coefficient of T k is the weight coefficient of, γ is the weight coefficient of L k is the weight coefficient of, δ is the weight coefficient of G k is the weight coefficient of. k

[0182] For R k , represents the central location corresponding to each potential subway station is located in the population density score of the sampling area, represents the central location corresponding to each potential subway station is located in the population flow score of the sampling area, ω 1 is is the weight coefficient of, ω 2 is is the weight coefficient of, ω 1 +ω 2 = 1, ρ max represents the maximum population density of the city, I max is the maximum population inflow of the city, O max is the maximum population outflow of the city;

[0183] For T k , is the central location corresponding to each potential subway station is located in the road traffic flow score of the sampling area, is the central location corresponding to each potential subway station is located in the public transport transfer score of the sampling area, ω 3 is is the weight coefficient of, ω 4 is is the weight coefficient of, ω 3 +ω​4 = 1, Q max is the maximum traffic flow of the city, V max is the maximum vehicle speed of the city, B max is the maximum number of bus transfers in the city, D max is the maximum number of taxi pick-ups and drop-offs in the city;

[0184] For L k , σ m is the corresponding score of different land use types in the sampling area where the central position corresponding to each potential subway station is located ;

[0185] For G k , is the terrain slope score of the sampling area where the central position corresponding to each potential subway station is located , is the elevation score of the sampling area where the central position corresponding to each potential subway station is located , ω 5 is 's weight coefficient, ω 6 is 's weight coefficient, ω 5 + ω 6 = 1, θ max is the maximum slope of the city, h ref is the reference elevation of the city, Δh max is the maximum elevation difference of the city.

[0186] Optionally, the subway station planning device based on multi-source data further includes:

[0187] A potential site screening module, configured to calculate the feasibility score of each potential subway station based on the geographic information data before the potential site scoring module 204 calculates the comprehensive score of each potential subway station based on the traffic flow data, population data, and geographic information data of the sampling area where each potential subway station is located, and remove the potential subway stations with the feasibility score lower than the second preset value.

[0188] Optionally, the potential site screening module is specifically configured to perform the following operations:

[0189] Based on the geographic information data of the sampling area where the central position corresponding to each potential subway station is located obtain the priority of the geological conditions of the potential site W k , the density of underground pipelines of the potential site H k , and the protection level of surrounding buildings of the potential site A k ;

[0190] According to the formula F k = λ 1 ×W k - λ 2 ×H k - λ 3 ×A k , calculate the feasibility score of each potential subway station; where λ 1 is the weight coefficient of W k , λ 2 is the weight coefficient of H k , λ 3 is the weight coefficient of A k .

[0191] An embodiment of the present application also provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps in the foregoing method embodiments are implemented.

[0192] An embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the foregoing method embodiments are implemented.

[0193] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The device embodiments described above are only illustrative. The components described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present disclosure solution. Those of ordinary skill in the art can understand and implement it without creative work.

[0194] The specific embodiments described above further elaborate the purpose, technical solution and beneficial effects of the present invention. It should be understood that the above is only the specific embodiment of the present invention and is not used to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A subway station planning method based on multi-source data, characterized in that: include: Obtain traffic flow data, population data and geographic information data for each sampling area in the city; Based on traffic flow data and population data, a clustering algorithm is used to divide the city into different cluster areas; Calculate the center position of each cluster area and identify each center position as a potential subway station; Calculate the comprehensive score of each potential subway station based on the traffic flow data, population data and geographic information data of the sampling area where each potential subway station is located; The potential subway station sites with comprehensive scores higher than the first preset value are determined as target subway stations.

2. The subway station planning method based on multi-source data according to claim 1 is characterized in that: Based on traffic flow data and population data, a clustering algorithm is used to divide the city into different cluster areas, including the following steps: Step 1: Based on the population data of n sampling areas, process and obtain the population density p corresponding to the n sampling areas i1 , population inflow during peak hours p i2 and the population outflow during peak hours p i3 Based on the traffic flow data of n sampling areas, the traffic flow p corresponding to the n sampling areas is obtained i4 , vehicle speed i5 、Number of bus transfer passengers p i6 and the number of taxi passengers p i7 ; Where i = 1, 2, ..., n, represents the serial number of the sampling area; p i1 、p i2 、p i3 、p i4 、p i5 、p i6 and p i7 are all represented by n-dimensional vectors; Step 2: i1 、p i2 、p i3 、p i4 、p i5 、p i6 and p i7 As feature data, each sampling area corresponds to a feature data set P i =(p i1 ,p i2 ,…,p ij ), where j = (1, 2, ..., 7), represents the feature sequence number of the feature data; Step 3: Based on each feature data group P i , K-means clustering algorithm is used to obtain multiple clustering areas.

3. The subway station planning method based on multi-source data according to claim 2 is characterized in that: Based on each feature data set P i , using the K-means clustering algorithm to obtain multiple clustering regions, including the following steps: Step 1: Determine the number of clusters K; the number of clusters K is used to group each feature data group P i Divide into K clusters, L = (1, 2, ..., n); among the K clusters, the feature data set P contained in the lth cluster i The set is C k , the cluster center of the kth cluster is E k =(μ k1 ,μ k2 ,…,μ kj ), k=(1,2,…,K); Step 2: From n feature data sets P i K are randomly selected as cluster centers; the initial cluster centers are E1, E2, …, E K ; Step 3: Group each feature data into P i =(p i1 ,p i2 ,…,p ij ) is normalized to Z i =(z i1 ,z i2 ,…,z ij ), according to Z i Calculate each feature data set P i To each cluster center E k The distance d(Z i ,E k ), the calculation formula is Step 4: Group each feature data into P i Assign to the distance d(Z i ,E k ) The nearest cluster center E k The cluster C k ; where k = arg min l=1,2,…,K d(Z i ,E l ); Step 5: For each cluster C k , recalculate and update its cluster center E k =(μ k1 ,μ k2 ,…,μ kj ), where if cluster C k There are n k Feature data set P i , for the jth feature, Represents cluster C k Contains characteristic data set P i The mean value on the jth feature; Step 6: Repeat steps 3 to 5 until for all k = (1, 2, ..., K), We get K different clusters, among which, is the updated cluster center, is the cluster center before updating, ∈ is the threshold; Step 7: Group all feature data contained in each cluster into P i The total area composed of the corresponding sampling areas is used as the clustering area of ​​each cluster, and K different clustering areas are obtained.

4. The subway station planning method based on multi-source data according to claim 1 is characterized in that: Calculate the center position of each cluster area and determine each center position as a potential subway station, including the following steps: Get each feature data group P contained in each cluster of K different clusters i The center coordinates of the corresponding sampling area in the geographic coordinate system (x i ,y i ); According to the coordinates (x i ,y i ) Calculate the center positions of K different clustering areas The calculation formula is: Among them, n k Represents the feature data set P contained in each of the K different clusters i the number of The center locations as a potential subway station site.

5. The subway station planning method based on multi-source data according to claim 1 is characterized in that: Based on the traffic flow data, population data and geographic information data of the sampling area where each potential subway station is located, the comprehensive score of each potential subway station is calculated, including the following steps: Based on the central location of each potential subway station The population data, traffic flow data and geographic information data of the sampling area to obtain the population density ρ of the potential site k , Population inflow at potential sites during peak hours I k , the population outflow of potential sites during peak hours O k , traffic flow data of potential sites to obtain vehicle flow Q k , potential station speed V k 、The number of bus transfer passengers at potential stations B k , the number of taxi passengers picking up and dropping off at potential stations D k , Area proportion of different land use types of potential sites Average slope of potential sites θ k , potential site average elevation h k ;in, m represents the serial number of different land use types; According to the formula S k =α×R k +β×T k +γ×L k +δ×G k , calculate the comprehensive score S of each potential subway station k ; Among them, R k Indicates the central location of each potential subway station The population score of the sampling area to which it belongs, T k Indicates the central location of each potential subway station The traffic flow score of the sampling area, L k Indicates the central location of each potential subway station The land use type score of the sampling area, G k Indicates the central location of each potential subway station The geographical terrain score of the sampling area; α is R k The weight coefficient of T k The weight coefficient of L k The weight coefficient of G k The weight coefficient of .

6. The subway station planning method based on multi-source data according to claim 1 is characterized in that: Based on the traffic flow data, population data and geographic information data of the sampling area where each potential subway station is located, before calculating the comprehensive score of each potential subway station, the following steps are also included: Calculate the feasibility score of each potential subway station based on geographic information data; Potential subway stations with feasibility scores lower than a second preset value are removed.

7. The subway station planning method based on multi-source data according to claim 6 is characterized in that: Based on geographic information data, the feasibility score of each potential subway station is calculated, including the following steps: Based on the central location of each potential subway station Geographic information data of the sampling area to obtain the geological conditions priority of potential sites W k , density of underground pipelines at potential sites H k 、Protection level of buildings around potential sites A k ; According to formula F k =λ1×W k -λ2×H k -λ3×A k , calculate the feasibility score of each potential subway station; where λ1 is W k The weight coefficient of H k The weight coefficient of A k The weight coefficient of .

8. A subway station planning device based on multi-source data, characterized in that: include: A data acquisition module is used to obtain traffic flow data, population data and geographic information data of each sampling area in the city; A region division module, used for dividing the city into different cluster regions by using a clustering algorithm based on the traffic flow data and the population data; A potential site determination module, used to calculate the center position of each of the clustering areas and determine each of the center positions as a potential subway station; A potential site scoring module, used to calculate a comprehensive score of each potential subway station based on traffic flow data, population data and geographic information data of the sampling area where each potential subway station is located; The target site confirmation module is used to determine the potential subway station whose comprehensive score is higher than a first preset value as a target subway station.

9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the subway station planning method based on multi-source data as described in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the subway station site planning method based on multi-source data as described in any one of claims 1 to 7 is implemented.