Hub city correlation feature tracking analysis method and system based on network analysis

By collecting multi-source, multi-dimensional data to construct a multi-level network model, calculating urban correlation and centrality indicators, and combining time series analysis, the problem of insufficient multi-dimensionality in urban correlation analysis in existing technologies is solved, enabling accurate prediction and planning support for urban development trends.

CN121599287BActive Publication Date: 2026-04-17TRANSPORT PLANNING & RES INST MINIST OF TRANSPORT
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TRANSPORT PLANNING & RES INST MINIST OF TRANSPORT
Filing Date
2025-11-26
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing technologies typically rely on a single data source for urban correlation analysis, which fails to fully reflect the multi-dimensional interactions between hub cities, resulting in a lack of accuracy and foresight in predicting future urban development trends.

Method used

Collect multi-source, multi-dimensional data, construct a multi-dimensional database, build a multi-level network model, calculate urban connectivity, network density, and centrality indicators through network analysis methods, combine historical and real-time data for time series analysis, track inter-city connectivity characteristics, and predict future development trends.

Benefits of technology

It provides an accurate data foundation, truly reflects the correlation and interaction intensity between cities, enhances the forward-looking prediction capability of urban interaction relationships, and supports urban planning and policy adjustments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121599287B_ABST
    Figure CN121599287B_ABST
Patent Text Reader

Abstract

The application provides a hub city correlation feature tracking analysis method and system based on network analysis, relates to the technical field of data analysis, and the method comprises the following steps: collecting multi-source multi-dimensional data of a hub city, constructing a multi-dimensional database, including traffic flow, economic flow, information flow and population flow; setting edge weights to construct a multi-level network model of the hub city by taking cities as nodes and inter-city flow data as edges; performing network index calculation based on a network analysis method, establishing multi-dimensional interactive relationships among cities; and performing time sequence reverse tracking and time sequence forward trend prediction through time sequence analysis on historical data and real-time data, and analyzing multi-dimensional feature correlation strength and development direction. The application solves the technical problem in the prior art that city correlation analysis is usually based on a single data source, the interactive relationships among hub cities in multiple dimensions cannot be comprehensively reflected, and the prediction of future city development trends lacks accuracy and foresight.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data analysis technology, specifically to a method and system for tracking and analyzing the correlation characteristics of hub cities based on network analysis. Background Technology

[0002] Hub cities not only bear the core functions of population flow, transportation hubs, and economic activities, but also play a guiding and radiating role in information flow, logistics, and talent flow. Therefore, accurately identifying and quantifying the interconnected characteristics between hub cities is of great significance for urban planning, regional coordinated development, and optimal resource allocation. Currently, many urban connection analysis methods, especially in urban network research, often rely on a single data source to analyze relationships between cities. Single-data-source analysis can only reveal the interaction between cities in a certain aspect or field, but cannot comprehensively demonstrate the interconnections between cities across different dimensions. In reality, the connections between modern cities are multi-dimensional and multi-layered. Simple single-data-source analysis easily overlooks or underestimates the interactions and feedback between various dimensions, leading to an inability to comprehensively consider multiple dimensions in analysis and decision-making. This affects the comprehensive assessment of inter-city interactions, resulting in a lack of accuracy and foresight in predicting future urban development trends. Summary of the Invention

[0003] This application provides a method and system for tracking and analyzing the association characteristics of hub cities based on network analysis. It aims to solve the technical problem that existing technologies usually conduct urban association analysis based on a single data source, which cannot fully reflect the multi-dimensional interaction relationships of hub cities, resulting in a lack of accuracy and foresight in predicting future urban development trends.

[0004] The first aspect disclosed in this application provides a method for tracking and analyzing the correlation characteristics of hub cities based on network analysis. The method includes: collecting multi-source, multi-dimensional data of hub cities; integrating the multi-source data to construct a multi-dimensional database, wherein the multi-dimensional database includes traffic flow, economic flow, information flow, and population flow; based on the multi-dimensional database, constructing a multi-level network model of hub cities with cities as nodes and inter-city flow data as edges, and setting corresponding edge weights reflecting the correlation strength between cities; according to the multi-level network model, calculating network indicators such as city correlation, network density, and centrality based on network analysis methods to establish multi-dimensional interactive relationships between cities; and based on the network indicators, performing time-series reverse tracking and time-series positive trend prediction of the correlation characteristics between hub cities through time-series analysis of historical and real-time data, and analyzing the correlation strength and development direction of multi-dimensional characteristics.

[0005] The second aspect of this application discloses a hub city association feature tracking and analysis system based on network analysis. This system is used in the aforementioned hub city association feature tracking and analysis method based on network analysis. The system includes: a multi-source data integration module for collecting multi-source, multi-dimensional data from hub cities, integrating the multi-source data, and constructing a multi-dimensional database, wherein the multi-dimensional database includes traffic flow, economic flow, information flow, and population flow; a network model construction module for constructing a multi-level network model of hub cities based on the multi-dimensional database, using cities as nodes and inter-city flow data as edges, and setting corresponding edge weights reflecting the strength of inter-city associations; a network indicator calculation module for calculating network indicators such as city association degree, network density, and centrality based on the multi-level network model and network analysis methods, establishing multi-dimensional interactive relationships between cities; and a time-series analysis module for performing time-series reverse tracking and time-series positive trend prediction of the association features between hub cities based on the network indicators and through time-series analysis of historical and real-time data, analyzing the multi-dimensional feature association strength and development direction.

[0006] One or more technical solutions provided in this application have at least the following beneficial effects:

[0007] By collecting data from different sources and dimensions, and integrating this multi-source, multi-dimensional data, a comprehensive multi-dimensional database is established. This database provides data support for the interaction of hub cities at multiple levels, providing an accurate and reliable data foundation for subsequent analysis. A multi-level network model of hub cities is constructed using cities as nodes, inter-city flow data as edges, and corresponding edge weights reflecting the strength of inter-city connections. This model accurately reflects the connections and interaction strength between cities. Through analysis of the multi-level network model, key network indicators such as inter-city connectivity, network density, and centrality are calculated. These indicators effectively describe the interaction strength, network density, and the importance and influence of cities within the network, providing quantitative basis for urban management, regional development, and policy design. Through time-series analysis of historical and real-time data, the connection characteristics between hub cities are traced backward, analyzing the causes and impact of historical changes. Simultaneously, positive trends are used to predict future connection characteristics and development directions. This method not only reveals the evolutionary patterns of inter-city relationships but also predicts possible future interaction trends. This analytical approach enhances the forward-looking predictive ability of hub city interactions, supporting more precise urban development planning, regional strategic layout, and policy adjustments.

[0008] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0009] Figure 1 A schematic diagram of the process for tracking and analyzing the association features of hub cities based on network analysis, provided in an embodiment of this application.

[0010] Figure 2 A schematic diagram of the structure of a hub city association feature tracking and analysis system based on network analysis provided in this application embodiment.

[0011] Figure labeling: Multi-source data integration module 10, network model construction module 20, network index calculation module 30, time series analysis module 40. Detailed Implementation

[0012] This application provides a method and system for tracking and analyzing the association features of hub cities based on network analysis. It solves the technical problem that existing technologies usually conduct urban association analysis based on a single data source, which cannot fully reflect the multi-dimensional interactive relationships of hub cities, resulting in a lack of accuracy and foresight in predicting future urban development trends.

[0013] After introducing the basic principles of this application, various non-limiting embodiments of this application will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.

[0014] Example 1, as Figure 1 As shown in the embodiments of this application, a method for tracking and analyzing the association features of hub cities based on network analysis is provided. The method includes:

[0015] Collect multi-source, multi-dimensional data from hub cities, integrate the multi-source data, and construct a multi-dimensional database, which includes traffic flow, economic flow, information flow, and population flow.

[0016] Traffic flow includes aviation, railway, and highway information. Specifically, flight schedules and actual flight data are obtained from platforms such as the Civil Aviation Administration of China, OAG, and VariFlight, including departure and arrival cities, flight frequency, number of seats, load factor, and freight volume; train schedules and ticket sales data (OD flow) are obtained from the railway passenger transport system to calculate the intensity of passenger transport connections between cities; cross-city traffic OD matrix is ​​extracted using highway toll data and GPS floating car data (such as freight truck and taxi trajectories); and container throughput and ship AIS data are obtained from port authorities to analyze cargo flows between port cities.

[0017] Economic flows include enterprise flow information and trade flow information. Specifically, it involves crawling business registration information to identify the network relationships of large enterprises (such as listed companies) that have established headquarters and branches in different cities. The quantity and quality of headquarters and branches constitute economic control. It also involves using import and export data, UnionPay intercity transaction data, and logistics data from e-commerce platforms to depict the flow of goods and funds between cities.

[0018] Information flow includes: obtaining geotagged posts and interaction (forwarding, commenting) data through social media APIs (such as Weibo and Douyin) to analyze the path and intensity of information dissemination between cities; and using search engines (Baidu Index, Google Trends) to obtain the popularity index of mutual searches between cities, reflecting the connection of attention.

[0019] Population flow includes long-term migration information and short-term mobility information. Specifically, it utilizes population census and social security payment data changes; and mobile signaling data (China Mobile, China Unicom, China Telecom) and LBS (Location-Based Service) data (such as Baidu Maps migration big data) to accurately capture population movement OD matrix during holidays, daily commutes, and other periods.

[0020] The above-mentioned multi-dimensional information about hub cities is collected from different data sources. The collected data from different dimensions are integrated, and a unified multi-dimensional database is generated through correlation evaluation, alignment and compensation for subsequent analysis.

[0021] Based on the multi-dimensional database, a multi-level network model of hub cities is constructed, with cities as nodes, inter-city flow data as edges, and corresponding edge weights reflecting the strength of inter-city connections.

[0022] Each city is treated as a node in the network, and the flow data between cities is represented as the network edges. For example, traffic flow, economic flow, information flow, and population flow connect cities, forming the edges. To reflect the strength of the connections between cities, a corresponding weight is assigned to each edge, obtained by quantifying data across different dimensions. For each data type, a single-layer network is established, containing both the city and its corresponding flow data. These different single-layer networks are then hierarchically connected to reflect the interaction relationships between data across different dimensions, forming a multi-layered network model of hub cities.

[0023] Based on the multi-level network model, network indicators such as city correlation, network density, and centrality are calculated using network analysis methods to establish multi-dimensional interactive relationships among cities.

[0024] City connectivity represents the strength of interaction between cities. Based on a multi-level network model, it quantifies the frequency of interaction by measuring the frequency of interaction based on inter-city flow data. Multiple dimensions are weighted and calculated to arrive at the overall city connectivity. Network density refers to the ratio between the actual number of edges in a network and the maximum possible number of edges, reflecting the tightness of connections between cities. Higher network density indicates more frequent interactions and closer connections between cities. Centrality measures a city's importance in the network, calculated based on the number of data connections between cities. Through these network indicators, multi-dimensional inter-city interaction relationships are established, indicating which cities are more closely connected in terms of transportation, economy, information, and population, and which cities occupy central positions in the network and have a greater influence on the overall network.

[0025] Based on the network indicators, through time-series analysis of historical and real-time data, the correlation characteristics between hub cities are tracked in reverse time series and predicted in forward time series, and the correlation strength and development direction of multi-dimensional features are analyzed.

[0026] Based on historical data, a spatiotemporal network indicator sequence is constructed to reflect the evolution of the network in different times and spaces, including city status, interaction network indicators, and city changes. Time-series reverse tracking refers to analyzing changes in historical data to determine the evolution of urban connectivity characteristics and their influence intensity. For example, by analyzing changes in traffic flow, economic flow, and population flow over the past few years, the reasons for changes in inter-city connectivity can be traced, revealing the historical background and patterns of urban interaction. Time-series forward trend prediction, based on historical and real-time data, predicts the changing trends of inter-city connectivity characteristics over a future period. For example, through trend decomposition, historical data is broken down into long-term trends, seasonal cycles, and random fluctuations to analyze the long-term evolution of inter-city interactions. Based on time-series reverse tracking and time-series forward trend prediction, the correlation strength and development direction of different dimensions of characteristics are analyzed, indicating the historical evolution and future development direction of inter-city connectivity characteristics, providing strong support for urban planning and decision-making.

[0027] In summary, by integrating multi-source data such as traffic flow, economic flow, information flow, and population flow to construct a multi-layered urban network, this method can quantify the interactive relationships and key hub status among cities. Through time-series analysis combining historical and real-time data, it can not only trace the evolution of inter-city connections but also predict future development trends. This method can be directly applied to urban planning, transportation layout optimization, industrial collaboration, and population management, providing a scientific basis for policy-making and enabling more precise, forward-looking, and dynamic urban development decisions. Compared to traditional methods, it integrates multi-dimensional data, network structure, and trend prediction, significantly improving the operability and innovativeness of decision-making.

[0028] Furthermore, the collection of multi-source, multi-dimensional data from the hub city, and the integration of this multi-source data, includes:

[0029] Targeting traffic flow, economic flow, information flow, and population flow, the correlation and confidence of multi-source data collection paths are evaluated to generate multi-source data integration labels; the multi-source data are aligned to identify cross-data and compensation data; based on the multi-source data integration labels, adaptive data integration is performed according to the cross-data and compensation data types to obtain the multi-dimensional database.

[0030] Correlation assessment evaluates the correlation between different data sources to determine their effective integration. For example, based on cohesion indicators, it assesses the closeness of relationships between data sources; traffic flow and economic flow are strongly correlated because transportation is typically associated with the flow of goods and funds. Confidence assessment is based on evaluation indicators such as authority, reliability, spatiotemporal resolution, and coverage. Adaptive parameter correlation calculations are performed based on the correlation and confidence assessment results to establish an integrated mapping relationship and generate multi-source data integration labels, serving as the basis for subsequent adaptive data integration.

[0031] Data alignment ensures consistency between different data sources, especially synchronization in time and space. Specifically, different data sources may have different time frequencies and time ranges. For example, some data may be collected hourly or daily, while others may be updated monthly or quarterly. Time alignment of these data ensures synchronization during analysis. Different cities have different geospatial units. Some data may be based on cities, while others may be based on regions or blocks, requiring spatial matching.

[0032] Cross data refers to redundant information from different data sources that record the same phenomenon and may contain similar information in multiple data sources. During the alignment process, it is necessary to identify and label this redundant data to avoid duplicate calculations. Compensation data refers to data that can be supplemented by another data source when one data source is missing. For example, if traffic flow data for some cities is missing, but economic flow data for these cities can indirectly reflect the traffic situation in the region, compensation data helps to fill the gaps in certain data sources and improve the completeness of the integrated data.

[0033] Based on the generated multi-source data integration labels, an adaptive integration method is used to merge data from different data sources. During this process, when encountering cross-data, redundancy needs to be avoided by merging or deduplication. The cross-data can be weighted according to its weight to avoid the duplication of a data source affecting the final result. If some data sources are missing or have errors, they can be filled with compensating data. The use of compensating data needs to be combined with the overall confidence assessment to ensure that the source of the supplementary data is reliable.

[0034] Furthermore, aligning multi-source data and identifying overlapping and compensating data includes:

[0035] Multi-source data alignment is performed based on the spatiotemporal characteristics of the collected data; based on the data alignment relationship, overlapping and complementary data regions are identified, and cross-data and compensation data are determined. Cross-data refers to redundant data that is recorded by different data sources for the same real-world phenomenon; compensation data refers to data that is missing from one data source but can be supplemented by another data source.

[0036] The spatiotemporal characteristics of collected data refer to the data's performance in both time and space dimensions. Temporal characteristics include the data's time stamp, frequency, and period. For example, some traffic flow data is updated hourly, while some economic flow data is updated daily or monthly. For data with different frequencies, interpolation methods are used to convert low-frequency data to high-frequency data, or vice versa. Interpolation methods include linear interpolation and spline interpolation. Spatial characteristics indicate the spatial dimension of the data, typically referring to the data's geographical scope or location unit. For example, traffic flow data is organized by city, while economic flow data is categorized by region or enterprise. It's necessary to ensure that all data sources have consistent spatial units. This can be achieved through spatial mapping techniques, such as using a geographic information system to convert different spatial units into unified coordinates or geographical areas, or through spatial aggregation as needed, such as consolidating data from smaller areas into larger areas.

[0037] Based on data alignment relationships, we can further identify which data is redundant and which data can complement each other. Among them, data overlap areas refer to repeated records of the same phenomenon or event in different data sources. The key to identifying these overlap areas is to analyze whether the performance of the same city or region is consistent in different data sources. Similarity matching or data overlap analysis can be used to identify the intersection between data sources. Data complement areas refer to the information that other data sources can provide to supplement the missing data when some data sources are missing. Data missing data analysis can be used to identify which regions or time periods have missing data, and then fill them with other data sources.

[0038] Cross data refers to data from multiple data sources that record similar information about the same phenomenon. In other words, cross data is data that appears repeatedly in multiple data sources. This type of data is usually redundant and needs to be processed by deduplication or weighting to avoid it having an excessive impact on the results during the analysis process. Compensating data refers to data from other data sources that can fill the gaps caused by missing or incomplete data in one data source. Compensating data is a substitute for missing data and is used to enhance the completeness and accuracy of the data.

[0039] Furthermore, the correlation and confidence of multi-source acquisition paths are evaluated, and multi-source data integration labels are generated, including:

[0040] The correlation of multi-dimensional target data is evaluated based on the cohesion and update synchronization frequency of the multi-source acquisition path to obtain correlation indicators. The authority, reliability, spatiotemporal resolution, and coverage of the multi-dimensional target data are evaluated quantitatively based on the multi-source acquisition path to obtain confidence indicators. Based on the correlation indicators and the confidence indicators, the correlation of integrated adaptive parameters is calculated to establish an integrated mapping relationship and generate multi-source data integration labels, which serve as a guide for selecting data adaptive integration strategies.

[0041] Cohesion refers to the degree to which different data sources describe the same or related phenomena; that is, whether they are closely connected and can jointly describe the same phenomenon. Cohesion is assessed by comparing the matching degree of the content recorded by different data sources with time, spatial coverage, etc. If the cohesion of two data sources is high, their correlation index is also high. Update synchronization frequency refers to the time consistency of data updates from different data sources. Differences in data update frequency affect the comparability between data. By comparing the update times of multiple data sources, the update time difference between data sources can be determined. For data sources with strong synchronization, their correlation index is high.

[0042] The authority of a data source refers to its reliability. Data sources with higher authority typically have higher credibility and accuracy. For example, data from professional statistical bureaus is assigned a value of 1.0, data from well-known commercial data platforms (such as UnionPay) is assigned a value of 0.8, and social media API data is assigned a value of 0.6. The reliability of a data source is verified through historical data. For example, if a data source's records match the final facts 9 out of 10 times, its reliability index is 0.9. Spatiotemporal resolution refers to the temporal and spatial precision of the data source. Data with high precision (such as city-level or daily-level data) scores higher, while data with low precision (such as provincial-level or annual data) scores lower. Coverage refers to the representativeness of the data source within a specific region or population. The wider the coverage of the data source, the higher the confidence level. For example, mobile signaling data (covering over 70% of the population) scores much higher than LBS check-in data (covering a young and active population, approximately 30%). By evaluating multi-dimensional target data based on the above-mentioned authority, reliability, spatiotemporal resolution, and coverage indicators, and then weighting the scores of each evaluation, a confidence index is obtained.

[0043] Based on the weights of correlation and confidence indices, the integration value for each data source is calculated. For example, a data source with strong correlation and high confidence receives a greater weight, while data sources with weak correlation or low confidence receive a smaller weight. By quantitatively evaluating the correlation and confidence indices of multiple data sources, the importance of each data source in the integration process is calculated. The integration mapping relationship determines the contribution of different data sources to the integrated data based on their characteristics (such as correlation and authority). Finally, through the calculation of these evaluation indicators and the establishment of integration mapping relationships, a multi-source data integration label is generated. This label contains the integration strategy for each data source and serves as a reference for subsequent data integration and processing.

[0044] Furthermore, based on the aforementioned multi-dimensional database, using cities as nodes and inter-city flow data as edges, and setting corresponding edge weights reflecting the strength of inter-city connections, a multi-layered network model of hub cities is constructed, including:

[0045] Based on each data type in the multi-dimensional database, a corresponding single-layer network is established, including a transportation network layer, an economic network layer, an information flow network layer, and a population flow network layer. Each single-layer network uses cities as nodes, corresponding types of flow data as edges, and the normalized result of the quantified value of the flow data as the edge weight. According to the data interaction relationship between layers, the transportation network layer, economic network layer, information flow network layer, and population flow network layer are connected hierarchically to construct the multi-layer network model.

[0046] For each data type, a corresponding independent single-layer network is established. Each single-layer network consists of city nodes and flow data between cities. Each single-layer network uses cities as nodes and flow data of the corresponding type as edges. Specifically, the edges of the transportation network layer represent traffic flows between cities, and the edge weights are determined based on the quantitative data of each traffic flow (such as traffic volume, quantity of transported goods, etc.). These quantitative data are normalized to ensure that they are numerically comparable. The edges of the economic network layer represent economic flows between cities, such as business flows and trade flows. The edge weights are the quantitative values ​​of economic flows, which can be trade volume, investment amount, or other indicators reflecting the intensity of economic interaction. The edges of the information flow network layer represent information flows between cities, such as communication data and internet traffic. The edge weights can use quantitative indicators such as data transmission volume and information exchange frequency. The edges of the population flow network layer represent population flows between cities, including long-term migration and short-term migration. The edge weights can be set according to the number of migrations or the frequency of migration.

[0047] In each single-layer network, the centrality of a node can reflect the importance or activity of a city in that dimension of the network. For example, it can be determined by the number of connections between nodes. The relationship between different dimensions can be evaluated by calculating the correlation coefficient of the node centrality index between different single-layer networks. For example, the centrality of the traffic flow network is related to the centrality of the economic flow network because economic activities are highly dependent on traffic flow. This correlation coefficient can reflect the coupling relationship or the strength of mutual influence between different dimensions.

[0048] Besides node centrality, the correlation between edge weights is also an important aspect of the interaction between layers. By calculating the correlation coefficient of edge weights between different network layers, the coupling degree between different dimensions can be quantified. For example, if the edge weights between traffic flow and economic flow are highly correlated, it means that the flow between these two dimensions is likely to be interdependent. In this case, the coupling relationship between layers is strong.

[0049] Based on the correlation coefficient between node centrality and edge weight, the coupling relationship between different data flow dimensions is evaluated. The stronger the coupling, the stronger the interdependence between the multi-dimensional flows between cities. Based on the calculated coupling relationship, different network layers are connected. These connections will affect the overall structure of the network and the interaction between nodes in different ways.

[0050] Furthermore, constructing a multi-layered network model for hub cities also includes:

[0051] Based on the spatiotemporal characteristics of data from a multi-dimensional database, one or more spatiotemporal index fields are added to each edge in the network; wherein, the spatiotemporal index fields are used to record the effective time range corresponding to the edge and the geospatial unit code involved.

[0052] In traditional network models, edge weights are fixed values ​​representing the strength of a relationship between cities, such as traffic flow, economic activity, or information flow. This model assumes a constant network structure and does not consider temporal and spatial variations. To better reflect dynamic changes in the real world, temporal and spatial variations are incorporated. Thus, each edge not only records the current flow intensity but also the trend of that flow over a certain time period and its spatial expansion. Specifically, one or more spatiotemporal index fields are added to each edge. These spatiotemporal index fields allow the network to reflect dynamic relationships that change over time. For example, traffic flow varies at different times (such as morning and evening rush hours), and the spatiotemporal index records these changes.

[0053] The spatiotemporal index field records the valid time range corresponding to each edge. This means that the flow relationship described by the edge is valid for a certain period of time. The time range can be fixed or dynamic, such as being dynamically updated by hour, month, or quarter. The geospatial unit code indicates the spatial range involved in each edge, represented by spatial unit codes. These codes can be generated using Geographic Information System (GIS) tools or divided according to different geographical regions. By recording the spatial range of edges, the network flow characteristics of a region can be analyzed, such as the difference in traffic flow between the city center and suburbs, or the economic activity of different regions.

[0054] Furthermore, based on the aforementioned multi-level network model, network indicators such as city correlation, network density, and centrality are calculated using network analysis methods to establish multi-dimensional interactive relationships among cities, including:

[0055] Based on the multi-level network model, the frequency of interaction is quantified according to the flow data between cities, and a weighted average is calculated according to the interaction frequency of the multi-level network to obtain the city correlation degree; the ratio of the actual number of data connections in the city network to the preset maximum number of connections is calculated to obtain the network density; the centrality ratio of each level of the network is analyzed according to the number of data connections between cities to obtain the centrality index of each dimension; based on the city correlation degree, network density, and centrality index, a multi-dimensional interaction relationship between hub cities is comprehensively established.

[0056] Interaction frequency measures the frequency or intensity of inter-city flow data. By quantifying the frequency of inter-city flow data, it can reveal which cities have strong interactions and which have weak interactions. By analyzing the flow data between each pair of cities—for example, the total traffic flow between city A and city B over a certain period—the interaction frequency of each pair can be calculated. Interaction frequency is represented by the volume of flow data, such as traffic flow, goods trade volume, and information transmission volume. A weighted average of the interaction frequency is then used to calculate the city correlation. A higher correlation value indicates more frequent interactions and a stronger correlation between the two cities. The weights in the weighted average depend on the type of flow, the volume, and the time period.

[0057] Network density measures the ratio of the actual number of connections in a network to the preset maximum number of connections. It describes the density of the network and reflects the connectivity between cities. The actual number of data connections refers to the number of data connections between cities, indicating which cities actually have data flow within a specific time period. The preset maximum number of connections refers to the possible number of connections between all cities in the network. If there are N cities in the network, then the maximum number of connections is... A higher network density value indicates denser interactions between cities and a more compact network; conversely, a lower value indicates less interaction and a sparser network.

[0058] Centrality measures a city’s importance in a network, typically referring to its influence or key role in the overall city network. A high centrality index for a city means that it interacts more frequently with other cities in that dimension, playing a pivotal role. The centrality index of each city in each dimension of the network is calculated based on the different dimensions of the network. The more direct connections a city has, the higher its centrality.

[0059] By combining urban connectivity, network density, and centrality indicators, a comprehensive multi-dimensional interactive relationship between hub cities can be established. Specifically, a weighted average method can be used to integrate various indicators to identify hub cities that play an important role in different dimensions and provide information on the interaction intensity of different cities in multiple dimensions.

[0060] Furthermore, based on the aforementioned network indicators, through time-series analysis of historical and real-time data, the correlation characteristics between hub cities are analyzed using time-series reverse tracking and time-series positive trend prediction. This process reveals the strength and development direction of multi-dimensional feature correlations, including:

[0061] Based on the historical data, a spatiotemporal network indicator sequence is constructed, including city status, interactive network indicators, and city change. Based on the network indicators in the spatiotemporal network indicator sequence, a time-series reverse tracking analysis is performed to analyze the correlation characteristics and the intensity of their influence. A time-series forward trend prediction is performed to predict the trend direction and intensity of the correlation characteristics on the future development dimension of the city.

[0062] Based on historical data, a spatiotemporal network indicator sequence can be constructed. For example, a Markov model can be used. Markov models are suitable for time series data with memoryless properties, where the state at each time step depends only on the state at the previous time step. For example, changes in urban traffic flow or economic activity exhibit the characteristics of a Markov process. Transition probabilities can be inferred from historical data, and a state transition matrix can be constructed. In this way, possible paths for state changes between cities can be simulated. Alternatively, an LSTM model can be trained to learn the historical trends of inter-city flow data, thereby predicting future flow patterns. This is very effective in capturing complex spatiotemporal relationships, especially when there are strong nonlinear interactions between multiple dimensions.

[0063] City status represents the basic attributes of a city at a certain moment, such as economic conditions, population size, and traffic flow. This data forms the basis of spatiotemporal networks and provides important background for subsequent analysis. Interaction network indicators reflect information such as the frequency and intensity of interactions between cities, such as changes in traffic flow, economic cooperation, information flow, and population flow. This data will serve as input for time series analysis, describing the dynamic relationships between cities. City change quantities are the changes in a city at different points in time, such as economic growth rate, population increase or decrease, and fluctuations in traffic flow. These changes reflect the dynamic changes in city status and can reveal the patterns and trends of change.

[0064] Time-series backward tracking analysis analyzes how historical events and characteristics affect current and future urban development. For example, through causal inference analysis, it analyzes how certain historical events affect the intensity of inter-city flow. For instance, a certain economic policy or major event may have caused significant changes in traffic flow and economic interaction between specific cities. Through time lag effect analysis, it clarifies how past events and policies gradually affect the interaction between cities over time intervals.

[0065] Time-series positive trend forecasting is based on historical data to predict the changing trends of inter-city connectivity characteristics in the future. For example, time-series models such as ARIMA and LSTM are used to predict future inter-city connectivity characteristics, such as traffic flow, economic interaction, or population flow trends in the next few months or years. The forecast results give the changing trend of the intensity of inter-city interaction at a certain time or stage in the future, as well as the intensity of their impact on future development.

[0066] Furthermore, based on the network indicators in the spatiotemporal network indicator sequence, a time-series reverse tracking analysis is performed to analyze the correlation characteristics and the strength of their influence, and a time-series forward trend prediction is performed to predict the trend direction and strength of the correlation characteristics for the future development dimension of the city, including:

[0067] Trend decomposition is performed based on the spatiotemporal network indicator sequence to identify long-term trends, seasonal cycles, and random fluctuation components, and to analyze the historical evolution patterns of related features. Causal relationship analysis is performed based on the historical evolution patterns to obtain temporal causal relationships. Based on the temporal causal relationships, the network indicators are used as influencing variables to perform positive causal prediction analysis and reverse causal tracing on the spatiotemporal network indicator sequence to obtain historical related features and the intensity of feature influence, as well as the direction and intensity of future multi-dimensional trends.

[0068] Trend decomposition based on spatiotemporal network indicator sequences aims to identify long-term trends, seasonal cycles, and random fluctuations in the data. Long-term trends represent the overall direction of data over time, which can be upward (growth), downward (decline), or stable. Long-term trends reflect underlying, persistent changes in the system; for example, certain hub cities continuously attract more economic and population flows due to economic development. Seasonal components represent periodic fluctuations in the data, typically recurring at fixed intervals. For example, tourist traffic in some cities fluctuates significantly during certain seasons each year; this seasonal variation may be related to weather, holidays, or special events. Random fluctuations are those that cannot be explained by trend or seasonal components, usually unpredictable fluctuations caused by external factors (such as sudden events, policy changes, etc.) or internal noise. Trend decomposition methods include: classical decomposition, a classic method that decomposes time series into trend, seasonal, and random components, extracting each component separately through smoothing and filtering; and wavelet transform, suitable for decomposing complex time series data, revealing changes at different time scales by decomposing the signal into different frequency components.

[0069] The core of causal relationship analysis is to identify which factors have influenced the flow between cities. For example, cointegration analysis can be used to test the long-term relationship between two or more time series variables. If two time series variables show a common trend in the long run, it indicates that there is a causal relationship between them. The obtained time series causal relationship can intuitively show the causal chain between different factors.

[0070] Positive causal predictive analysis is based on known temporal causal relationships to predict future correlation characteristics and trends. For example, based on past economic policies, infrastructure construction, and other factors, it can predict the changing trends of traffic flow, population flow, and economic flow between cities in the future. Machine learning models, such as LSTM and random forest, can be used to make positive predictions in combination with temporal causal relationships. The prediction results include future traffic flow, population changes, and economic interactions, which can provide forward-looking basis for urban planning and policy decisions.

[0071] Reverse causal tracing involves tracing key moments in historical data to analyze the impact of specific events or policy changes on current urban network interaction patterns. Based on historical data, time-series backtracking methods are used to analyze the impact of various factors. For example, by tracing back to a large-scale infrastructure construction at a certain moment, its short-term and long-term impacts on urban mobility at that time can be analyzed. Through reverse causal tracing, the intensity of the impact of certain events or decisions on urban networks can be quantified, and their propagation effects over time can be determined.

[0072] Example 2, based on the same inventive concept as the hub city association feature tracking and analysis method based on network analysis in the previous examples, such as... Figure 2 As shown in the embodiment of this application, a hub city association feature tracking and analysis system based on network analysis is provided. The system includes:

[0073] The multi-source data integration module 10 is used to collect multi-source, multi-dimensional data from hub cities, integrate the multi-source data, and construct a multi-dimensional database, which includes traffic flow, economic flow, information flow, and population flow. The network model construction module 20 is used to construct a multi-level network model of hub cities based on the multi-dimensional database, with cities as nodes and inter-city flow data as edges, and setting corresponding edge weights that reflect the strength of inter-city connections. The network index calculation module 30 is used to calculate network indicators such as city correlation, network density, and centrality based on the multi-level network model and network analysis methods, and establish multi-dimensional interactive relationships between cities. The time series analysis module 40 is used to perform time series reverse tracking and time series positive trend prediction of the correlation characteristics between hub cities based on the network indicators and through time series analysis of historical and real-time data, and analyze the multi-dimensional feature correlation strength and development direction.

[0074] Furthermore, the multi-source data integration module 10 is used to perform the following operation steps:

[0075] Targeting traffic flow, economic flow, information flow, and population flow, the correlation and confidence of multi-source data collection paths are evaluated to generate multi-source data integration labels; the multi-source data are aligned to identify cross-data and compensation data; based on the multi-source data integration labels, adaptive data integration is performed according to the cross-data and compensation data types to obtain the multi-dimensional database.

[0076] Furthermore, the multi-source data integration module 10 is used to perform the following operation steps:

[0077] Multi-source data alignment is performed based on the spatiotemporal characteristics of the collected data; based on the data alignment relationship, overlapping and complementary data regions are identified, and cross-data and compensation data are determined. Cross-data refers to redundant data that is recorded by different data sources for the same real-world phenomenon; compensation data refers to data that is missing from one data source but can be supplemented by another data source.

[0078] Furthermore, the multi-source data integration module 10 is used to perform the following operation steps:

[0079] The correlation of multi-dimensional target data is evaluated based on the cohesion and update synchronization frequency of the multi-source acquisition path to obtain correlation indicators. The authority, reliability, spatiotemporal resolution, and coverage of the multi-dimensional target data are evaluated quantitatively based on the multi-source acquisition path to obtain confidence indicators. Based on the correlation indicators and the confidence indicators, the correlation of integrated adaptive parameters is calculated to establish an integrated mapping relationship and generate multi-source data integration labels, which serve as a guide for selecting data adaptive integration strategies.

[0080] Furthermore, the network model construction module 20 is used to perform the following operation steps:

[0081] Based on each data type in the multi-dimensional database, a corresponding single-layer network is established, including a transportation network layer, an economic network layer, an information flow network layer, and a population flow network layer. Each single-layer network uses cities as nodes, corresponding types of flow data as edges, and the normalized result of the quantified value of the flow data as the edge weight. According to the data interaction relationship between layers, the transportation network layer, economic network layer, information flow network layer, and population flow network layer are connected hierarchically to construct the multi-layer network model.

[0082] Furthermore, the network model construction module 20 is used to perform the following operation steps:

[0083] Based on the spatiotemporal characteristics of data from a multi-dimensional database, one or more spatiotemporal index fields are added to each edge in the network; wherein, the spatiotemporal index fields are used to record the effective time range corresponding to the edge and the geospatial unit code involved.

[0084] Furthermore, the network metric calculation module 30 is used to perform the following operation steps:

[0085] Based on the multi-level network model, the frequency of interaction is quantified according to the flow data between cities, and a weighted average is calculated according to the interaction frequency of the multi-level network to obtain the city correlation degree; the ratio of the actual number of data connections in the city network to the preset maximum number of connections is calculated to obtain the network density; the centrality ratio of each level of the network is analyzed according to the number of data connections between cities to obtain the centrality index of each dimension; based on the city correlation degree, network density, and centrality index, a multi-dimensional interaction relationship between hub cities is comprehensively established.

[0086] Furthermore, the timing analysis module 40 is used to perform the following operation steps:

[0087] Based on the historical data, a spatiotemporal network indicator sequence is constructed, including city status, interactive network indicators, and city change. Based on the network indicators in the spatiotemporal network indicator sequence, a time-series reverse tracking analysis is performed to analyze the correlation characteristics and the intensity of their influence. A time-series forward trend prediction is performed to predict the trend direction and intensity of the correlation characteristics on the future development dimension of the city.

[0088] Furthermore, the timing analysis module 40 is used to perform the following operation steps:

[0089] Trend decomposition is performed based on the spatiotemporal network indicator sequence to identify long-term trends, seasonal cycles, and random fluctuation components, and to analyze the historical evolution patterns of related features. Causal relationship analysis is performed based on the historical evolution patterns to obtain temporal causal relationships. Based on the temporal causal relationships, the network indicators are used as influencing variables to perform positive causal prediction analysis and reverse causal tracing on the spatiotemporal network indicator sequence to obtain historical related features and the intensity of feature influence, as well as the direction and intensity of future multi-dimensional trends.

[0090] Through the foregoing detailed description of the hub city association feature tracking and analysis method based on network analysis, those skilled in the art can clearly understand the hub city association feature tracking and analysis system based on network analysis in this embodiment. Since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and relevant parts can be referred to the method section.

[0091] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for tracking and analyzing the association characteristics of hub cities based on network analysis, characterized in that, The method includes: Collect multi-source, multi-dimensional data from hub cities, integrate the multi-source data, and construct a multi-dimensional database, which includes traffic flow, economic flow, information flow, and population flow. Based on the multi-dimensional database, a multi-level network model of hub cities is constructed, with cities as nodes, inter-city flow data as edges, and corresponding edge weights reflecting the strength of inter-city connections. Based on the multi-level network model, network indicators such as urban correlation, network density, and centrality are calculated using network analysis methods to establish multi-dimensional interactive relationships among cities. Based on the network indicators, through time-series analysis of historical and real-time data, the correlation characteristics between hub cities are tracked in reverse time and predicted in forward time, and the correlation strength and development direction of multi-dimensional features are analyzed. Based on the multi-dimensional database, a multi-level network model of hub cities is constructed, using cities as nodes and inter-city flow data as edges, and setting corresponding edge weights to reflect the strength of inter-city connections. This model includes: Based on each data type in the multi-dimensional database, a corresponding single-layer network is established, including a transportation network layer, an economic network layer, an information flow network layer, and a population flow network layer. Each single-layer network uses cities as nodes, corresponding types of flow data as edges, and the normalized result of the quantified value of the flow data as the edge weight. Based on the data interaction relationships between the layers, the transportation network layer, economic network layer, information flow network layer, and population flow network layer are connected hierarchically to construct the multi-level network model; Constructing a multi-layered network model for hub cities also includes: Based on the spatiotemporal characteristics of data from a multi-dimensional database, one or more spatiotemporal index fields are added to each edge in the network; The spatiotemporal index field is used to record the effective time range corresponding to the edge and the geospatial unit code involved. Based on the aforementioned multi-level network model, network indicators such as city connectivity, network density, and centrality are calculated using network analysis methods to establish multi-dimensional interactive relationships among cities, including: Based on the multi-level network model, the frequency of interaction is quantified according to the flow data between cities, and the city correlation degree is obtained by weighted average calculation according to the interaction frequency of the multi-level network. Calculate the ratio of the actual number of data connections in the urban network to the preset maximum number of connections to obtain the network density; Based on the number of data connections between cities, a centrality ratio analysis of networks at each level is conducted to obtain centrality indicators for each dimension. Based on the aforementioned city connectivity, network density, and centrality indicators, a multi-dimensional interactive relationship between hub cities is comprehensively established.

2. The method for tracking and analyzing the association features of hub cities based on network analysis according to claim 1, characterized in that, The collection of multi-source, multi-dimensional data from hub cities involves multi-source data integration, including: Using traffic flow, economic flow, information flow, and population flow as targets, the correlation and confidence of multi-source data collection paths are evaluated, and multi-source data integration labels are generated. Align multi-source data, identify overlapping data, and compensate for data. Based on the multi-source data integration tags, adaptive data integration is performed according to the cross data and compensation data types to obtain the multi-dimensional database.

3. The method for tracking and analyzing the association features of hub cities based on network analysis according to claim 2, characterized in that, Aligning multi-source data, identifying overlapping data, and compensating for data, including: Align multi-source data according to the spatiotemporal characteristics of the collected data; Based on data alignment relationships, overlapping and complementary data regions are identified, and cross-data and compensating data are determined. Cross-data refers to redundant data that is recorded by different data sources for the same real-world phenomenon; compensating data refers to data that is missing from one data source but can be supplemented by another data source.

4. The method for tracking and analyzing the association features of hub cities based on network analysis according to claim 2, characterized in that, The correlation and confidence of multi-source acquisition paths are evaluated, and multi-source data integration labels are generated, including: Based on the coherence and update synchronization frequency of multi-dimensional target data collected from multiple sources, a correlation evaluation is conducted to obtain correlation indicators. The confidence level index is obtained by quantitatively evaluating the authority, reliability, spatiotemporal resolution, and coverage of multi-dimensional target data based on the multi-source acquisition path. Based on the correlation index and the confidence index, the correlation of the integrated adaptive parameters is calculated to establish an integrated mapping relationship and generate the multi-source data integration label, which serves as a guide for selecting the data adaptive integration strategy.

5. The method for tracking and analyzing the association features of hub cities based on network analysis according to claim 1, characterized in that, Based on the aforementioned network indicators, through time-series analysis of historical and real-time data, the correlation characteristics between hub cities are analyzed using time-series reverse tracking and time-series positive trend prediction. This process reveals the strength and development direction of multi-dimensional feature correlations, including: Based on the historical data, a spatiotemporal network indicator sequence is constructed, including city status, interaction network indicators, and city change. Based on the network indicators in the spatiotemporal network indicator sequence, perform time-series reverse tracking analysis to analyze the correlation features and the intensity of their influence, and perform time-series forward trend prediction to predict the trend direction and intensity of the correlation features for the future development dimension of the city.

6. The method for tracking and analyzing the association features of hub cities based on network analysis according to claim 5, characterized in that, Based on the network indicators in the spatiotemporal network indicator sequence, perform time-series backward tracking analysis to analyze the correlation characteristics and the strength of their influence, and perform time-series forward trend prediction to predict the trend direction and strength of the correlation characteristics for the future development dimension of the city, including: Based on the spatiotemporal network indicator sequence, trend decomposition is performed to identify long-term trends, seasonal cycles, and random fluctuation components, and to analyze the historical evolution patterns of related features. Based on the historical evolution pattern, causal relationship analysis is performed to obtain temporal causal relationships; Based on the temporal causal relationship, the network indicators are used as influencing variables. Positive causal prediction analysis and reverse causal tracing are performed on the spatiotemporal network indicator sequence to obtain historical correlation characteristics and the intensity of characteristic influence, as well as future multi-dimensional trend direction and intensity.

7. A hub city association feature tracking and analysis system based on network analysis, characterized in that, The system is used to implement the hub city association feature tracking and analysis method based on network analysis as described in any one of claims 1-6, the system comprising: The multi-source data integration module is used to collect multi-source and multi-dimensional data from hub cities, integrate the multi-source data, and build a multi-dimensional database, which includes traffic flow, economic flow, information flow, and population flow. The network model construction module is used to construct a multi-level network model of hub cities based on the multi-dimensional database, with cities as nodes, inter-city flow data as edges, and setting corresponding edge weights that reflect the strength of inter-city connections. The network indicator calculation module is used to calculate network indicators such as city correlation, network density, and centrality based on the multi-level network model and network analysis methods, and to establish multi-dimensional interactive relationships between cities. The time series analysis module is used to perform time series reverse tracking and time series positive trend prediction of the correlation characteristics between hub cities based on the network indicators and through time series analysis of historical and real-time data, and to analyze the correlation strength and development direction of multi-dimensional features.

Citation Information

Patent Citations

  • Urban road traffic flow state estimation method and system based on block chain

    CN117671961A

  • Data processing method and system for smart city and storage medium

    CN119989224A