Method for analyzing correlation between urban spatial three-dimensional structure and population flow
By constructing an urban construction indicator system and analyzing population flow using base station data, combined with the LightGBM and SHAP algorithms, the problem of in-depth analysis of the relationship between urban spatial structure and population flow was solved, enabling precise guidance for urban development planning.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- NANJING UNIV
- Filing Date
- 2025-02-18
- Publication Date
- 2026-04-30
AI Technical Summary
Existing technologies are insufficient to deeply analyze the specific relationship between urban spatial structure and population flow, and cannot effectively utilize population flow trends to make precise adjustments to urban development plans.
By constructing an urban construction indicator system, using mobile phone signaling data obtained from base stations to identify commuting behavior, and combining GIS spatial analysis methods, the network structure characteristics of urban population flow are constructed. Furthermore, regression model analysis using the LightGBM and SHAP algorithms is conducted to explore the correlation between urban three-dimensional spatial structure and population flow.
Accurately grasping the essential characteristics of urban spatial structure and population flow distribution provides valuable reference for urban development planning, ensuring an accurate understanding of the essential characteristics of urban spatial structure and effectively utilizing population flow trends for urban development planning.
Smart Images

Figure CN2025077767_30042026_PF_FP_ABST
Abstract
Description
A Correlation Analysis Method Between Urban Spatial Three-Dimensional Structure and Population Flow Technical Field
[0001] This invention relates to the field of correlation analysis between urban space and population flow, and more specifically, to a method for correlation analysis between the three-dimensional structure of urban space and population flow. Background Technology
[0002] Currently, my country's urbanization is developing rapidly. Population mobility serves as a crucial carrier for the flow of information, goods, capital, and technology between regions. Large-scale population movement between cities significantly impacts the industrial structure, transportation links, infrastructure, employment situation, and urban scale of both inflow and outflow areas, leading to the reorganization of internal urban elements and causing certain changes in the urban system structure and social development pattern. On the one hand, human activities drive the emergence and decline, expansion and contraction, prosperity and decay of cities, dominating urban production, life, technological innovation, and knowledge iteration. On the other hand, human activities occur within the spatial organizational framework of cities, constraining human behavior. Therefore, deeply identifying the patterns of population flow direction, scale, and network structure, and combining these with geographical environment, socio-economic factors, and historical factors to explore the state of human activities and urban system construction in different regions, can enrich the theoretical system and methods of population urban planning and population geography. This can provide reference and ideas for planning and coordinating future population flow patterns and promoting the healthy and stable development of inflow and outflow areas.
[0003] Current research on urban space and population flow mainly focuses on the following aspects: First, by analyzing the direction, scale, and network structure of population flow, and combining factors such as geographical environment and socio-economic background, it explores the characteristics of human activities in different regions and their impact on the urban system; Second, based on the characteristics of human flow activities or from the perspective of human flow, it uses the network system formed by the activities of people in the city to reflect the spatial distribution of activity centers and the connections between centers, thereby revealing the essential characteristics of urban spatial structure; Third, some studies map urban spatial structure and function through residents' travel behavior patterns, such as establishing system dynamics models to measure the impact of urban spatial density, traffic conditions, and economic conditions on residents' travel costs; In addition, some scholars use big data technologies, such as mobile phone location data and trajectory data, to characterize the spatiotemporal characteristics of individual travel and explore their interaction with urban grids and urban functional structures, further deepening the understanding of the characteristics of urban spatial structure.
[0004] Although current research has comprehensively covered multiple dimensions of urban spatial structure and emphasized the impact of population distribution, density, and human activity on urban structure, it still falls short in deeply analyzing the specific relationship between urban spatial structure and population flow. It cannot fully reveal the characteristics of population distribution through urban spatial layout, nor can it effectively utilize population flow trends to make precise adjustments to urban development planning.
[0005] No effective solutions have yet been proposed to address the problems in the relevant technologies. Summary of the Invention
[0006] In response to the problems in related technologies, this invention proposes a correlation analysis method between the three-dimensional structure of urban space and population flow, so as to overcome the above-mentioned technical problems existing in the existing related technologies.
[0007] Therefore, the specific technical solution adopted by the present invention is as follows:
[0008] A correlation analysis method for urban spatial three-dimensional structure and population flow, comprising the following steps:
[0009] S1. Construct an urban construction indicator system based on the urban construction dataset, and obtain the three-dimensional spatial structural characteristics of urban construction by analyzing the structural features of the urban construction indicator system.
[0010] S2. Use base stations to obtain mobile phone signaling data and identify the commuting behavior of urban populations. Combine GIS spatial analysis methods to analyze and construct the network structure characteristics of urban population flow.
[0011] S3. A dataset is constructed using the three-dimensional spatial structure characteristics of urban construction and the network structure characteristics of urban population flow. A regression model is built and trained using the LightGBM algorithm. The SHAP algorithm is used to interpret the regression model and obtain the correlation analysis results between the three-dimensional spatial structure of urban areas and population flow.
[0012] Furthermore, an urban construction indicator system is constructed based on the urban construction dataset. Through structural feature analysis of this indicator system, the three-dimensional spatial structural features of urban construction are obtained, including the following steps:
[0013] S11. Collect urban building parameters to construct an urban construction dataset, and analyze the distribution indicators and suitability indicators of urban buildings through statistical analysis.
[0014] S12. Based on the distribution index of urban buildings, construct the structural characteristics of urban buildings through index calculation;
[0015] S13. Based on the urban construction suitability index, calculate the urban construction suitability characteristics using expert scoring and entropy weight methods;
[0016] S14. Combine the structural characteristics of urban buildings and the suitability characteristics of urban construction to obtain the three-dimensional spatial structural characteristics of urban construction.
[0017] Furthermore, urban building structure characteristics include weighted average building height, building volume ratio, building coverage ratio, building volume, proportion of high-rise buildings, and standard deviation of building height.
[0018] Furthermore, the formula for calculating the weighted average height of buildings is as follows:
[0019] ;
[0020] In the formula, BH represents the weighted average height of the building, and A i H represents the area occupied by the base of building i. i This represents the height information of the i-th building, and n represents the sum of the number of areas belonging to the building.
[0021] The formula for calculating the building's floor area ratio is:
[0022] ;
[0023] In the formula, BPR represents the building volume ratio, and A... i L represents the area occupied by the base of building i. i Let P represent the number of floors in building i, n represent the sum of the number of areas belonging to the building, and P represent the number of floors in building i. a This refers to the area of the regional plot;
[0024] The formula for calculating building coverage is:
[0025] ;
[0026] In the formula, BC represents the building coverage ratio, and A i Let P represent the area occupied by the base of building i, n represent the sum of the number of areas belonging to the building, and P a Indicates the area of the land parcel;
[0027] The formula for calculating building volume is:
[0028] ;
[0029] In the formula, BV represents the building volume, and A i H represents the area occupied by the base of building i. i The height data of building i is represented by n, and n represents the sum of the number of areas belonging to the building.
[0030] The formula for calculating the proportion of high-rise buildings is:
[0031] ;
[0032] In the formula, HBR represents the percentage of high-rise buildings, and N g This represents the number of high-rise buildings, and N represents the total number of buildings in the area.
[0033] The formula for calculating the standard deviation of building height is:
[0034] ;
[0035] In the formula, BHSD represents the standard deviation of building height, and H i This represents the height information of the i-th building, μ represents the average height of the buildings, and N represents the total number of buildings in the area.
[0036] Furthermore, based on urban development suitability indicators, the calculation of urban development suitability characteristics using expert scoring and entropy weight methods includes the following steps:
[0037] S131. Based on the normalization method, the suitability index for each town is normalized to obtain the normalized index;
[0038] S132. Quantify the obtained normalized index to obtain the quantified index;
[0039] S133. Combining expert scoring and entropy weighting methods, a weight analysis is performed on each quantitative indicator. Based on the weight analysis results, a weighted summation method is used to calculate the comprehensive score of each quantitative indicator, and the comprehensive score is used as the suitability characteristic for urban construction.
[0040] Furthermore, by utilizing base stations to acquire mobile phone signaling data and identify the commuting behavior of the urban population, and combining this with GIS spatial analysis methods to analyze and construct the network structure characteristics of urban population flow, the following steps are included:
[0041] S21. Obtain mobile phone signaling data of urban residents using base stations, and preprocess the mobile phone signaling data.
[0042] S22. Based on the preprocessed mobile phone signaling data, identify the commuting behavior of urban residents by analyzing commuting data between residence and workplace during the morning and evening peak hours.
[0043] S23. Based on the commuting behavior of the urban population, calculate the commuting volume of the urban population during a preset time period;
[0044] S24. Construct a population flow network graph of the town, with building locations in the town as nodes, commuting behavior of the town's population as edges, and commuting volume of the town's population in a preset time period as the weight attribute of the edges.
[0045] S25. Use the Gephi layout algorithm to analyze the population flow network graph of towns and cities, and obtain the network structure characteristics of urban population flow.
[0046] Further preprocessing includes invalid and redundant data removal, ping-pong data processing, and drift data processing.
[0047] Furthermore, a dataset was constructed using the three-dimensional spatial structure characteristics of urban construction and the network structure characteristics of urban population flow. A regression model was built and trained using the LightGBM algorithm, and the SHAP algorithm was used to perform interpretive analysis on the regression model. The results of the correlation analysis between urban three-dimensional spatial structure and population flow were obtained, including the following steps:
[0048] S31. Divide the dataset into training and testing sets, and use the LightGBM algorithm to build a regression model and initialize the model parameters.
[0049] S32. Based on the initialized model parameters, the regression model is trained using the training set, and the model parameters are adjusted and optimized through cross-validation.
[0050] S33. Use the test set to evaluate the effect of the regression model and calculate the performance evaluation index data of the model to obtain the evaluation results of the regression model.
[0051] S34. Use the SHAP algorithm to analyze the feature importance and influence direction of the evaluated regression model, and establish a summary diagram of SHAP based on the analysis results.
[0052] S35. Based on the summary diagram of SHAP, analyze the correlation between the three-dimensional spatial structure of towns and population flow, and obtain the correlation analysis results.
[0053] Furthermore, the model's performance evaluation metrics include the coefficient of determination, mean absolute error, mean absolute percentage error, and root mean square error.
[0054] Furthermore, the SHAP algorithm is used to analyze the feature importance and influence direction of the evaluated regression model. Based on the analysis results, a summary plot of SHAP is constructed, including the following steps:
[0055] S341. Based on the principles of cooperative game theory, calculate the Shapley value using the evaluated regression model.
[0056] S342. Based on the influence of Shapley values on characteristic variables and the direction of characteristic effects, we will interpret and visualize the results, and use the SHAP model to draw visualization charts to obtain a summary chart of SHAP.
[0057] The beneficial effects of this invention are as follows:
[0058] 1. This invention ensures an accurate grasp of the essential characteristics of urban spatial structure by analyzing the three-dimensional features of urban spatial structure. At the same time, it grasps the distribution of urban population flow and the characteristics of human activities by analyzing the network characteristics of population flow, further providing a basis for in-depth analysis of the relationship between urban spatial structure and population flow. Then, by analyzing the spatiotemporal correlation between urban spatial characteristics and population flow, it explores the importance of various features of urban spatial structure to population flow, and thus effectively uses population flow trends to provide valuable reference for the direction of urban development planning.
[0059] 2. This invention analyzes the three-dimensional characteristics of urban spatial structure by obtaining the suitability of urban construction and six indicators based on urban building height datasets and various urban element indicators, in order to describe the architectural form, spatial distribution and urban development level of the city, and ensure an accurate grasp of the essential characteristics of urban spatial structure.
[0060] 3. This invention identifies users' commuting behavior based on mobile phone signaling data, and then analyzes and compares the network characteristics of population flow in various towns and cities using GIS spatial analysis methods. This allows for an understanding of the distribution of population flow and the characteristics of human activities in these towns and cities, thereby gaining insight into the economic development trends, employment opportunities, urban planning, transportation facilities, and the density of economic activities and social exchanges between towns and cities. This provides a basis for a deeper analysis of the relationship between urban spatial structure and population flow.
[0061] 4. This invention constructs a model based on the LightGBM algorithm, taking the comprehensive suitability of urban construction and six indicators as inputs and population flow as output. Based on the model research results, the SHAP algorithm is used to analyze the influence of each input feature on population flow, explore the importance of each feature of urban spatial structure on population flow, accurately grasp the driving factors of population flow, further analyze the correlation between urban spatial features and population flow, and effectively use the correlation results to provide valuable reference for the direction of urban development planning. Attached Figure Description
[0062] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0063] Figure 1 is a flowchart of a method for analyzing the correlation between three-dimensional urban spatial structure and population flow according to an embodiment of the present invention;
[0064] Figure 2 is a schematic diagram of the calculation of the weighted average height (BH) of buildings within a grid in a correlation analysis method between urban spatial three-dimensional structure and population flow according to an embodiment of the present invention.
[0065] Figure 3 is a schematic diagram of the calculation of building volume ratio (BPR) within a grid in a correlation analysis method between urban spatial three-dimensional structure and population flow according to an embodiment of the present invention.
[0066] Figure 4 is a population flow network diagram in a method for correlation analysis of urban spatial three-dimensional structure and population flow according to an embodiment of the present invention.
[0067] Figure 5 is a scatter plot of the LightGBM model SHAP in a correlation analysis method between urban spatial three-dimensional structure and population flow according to an embodiment of the present invention.
[0068] Figure 6 is a bar summary diagram of the LightGBM model SHAP in a correlation analysis method between urban spatial three-dimensional structure and population flow according to an embodiment of the present invention. Detailed Implementation
[0069] To further illustrate the various embodiments, the present invention provides accompanying drawings, which are part of the disclosure of the present invention. These drawings are mainly used to illustrate the embodiments and can be used in conjunction with the relevant descriptions in the specification to explain the operating principles of the embodiments. With reference to these drawings, those skilled in the art should be able to understand other possible implementation methods and the advantages of the present invention.
[0070] According to an embodiment of the present invention, a method for correlation analysis between the three-dimensional structure of urban space and population flow is provided.
[0071] The present invention will now be further described with reference to the accompanying drawings and specific embodiments. As shown in Figure 1, the correlation analysis method between the three-dimensional structure of urban space and population flow according to an embodiment of the present invention includes the following steps:
[0072] S1. Construct an urban construction indicator system based on the urban construction dataset, and obtain the three-dimensional spatial structural characteristics of urban construction by analyzing the structural features of the urban construction indicator system.
[0073] Specifically, urban construction datasets include spatial survey data, Earth observation big data, and so on.
[0074] Spatial survey data: This refers to data from the National Land Survey (referred to as the "Third National Land Survey"). Based on the division of 12 primary categories, land can be classified into six major types: grassland, cultivated land, forest land, water area, construction land, and unused land.
[0075] Geospatial Big Data: Building height data, provided by the National Earth System Science Data Center. This dataset includes all buildings in China with a height exceeding 10 meters. Each building in the data contains its geographical coordinates, height information, building outline, and other attributes. The building height data can be used to construct a three-dimensional urban feature index to reflect the characteristics of the urban spatial structure.
[0076] For example, according to the "2022 Fire Protection Code" and the "General Rules for the Design of Civil Buildings", buildings in a certain southern province of China are classified into 4 categories, namely low-rise buildings (h <= 10m), mid-rise buildings (10m < h <= 24m), mid-high-rise buildings (24m < h <= 50m), and high-rise buildings (h > 50m). Following these rules, based on the building height dataset, using the statistical analysis method of ArcGIS, the distribution profile of the number of buildings in different height ranges is studied, and the analysis results are shown in Table 1. It can be seen that the number of mid-rise buildings is the largest, accounting for 78.65%; followed by mid-high-rise buildings, accounting for 18.15%; the proportions of low-rise buildings and high-rise buildings are close, and the numbers are very small, accounting for 1.68% and 1.52% respectively.
[0077] Table 1 Number and Proportion of Buildings of Different Heights in a Certain Southern Province of China
[0078] Building ClassificationHeight (m)NumberProportion(%)Low-rise Buildingh <= 10m7684621.68Mid-rise Building10m < h <= 24m3607845978.65Mid-high-rise Building24m < h <= 50m832364818.15High-rise Buildingh > 50m6967711.52
[0079] Among them, six indicators are selected, namely the weighted average building height (Buliding Height), building plot ratio (Building Plot Ratio), building coverage ratio (Building Coverage), building volume (Building Volume), high-rise building ratio (High-rise Building Ratio), and building height standard deviation (Building Height Standard Deviation), to describe the three-dimensional structure characteristics of the urban space, so as to describe the building form, spatial distribution, and urban development level of the city. At the same time, to ensure the homogenization of the evaluation unit, a 1000-meter grid is used as the basic evaluation unit, that is, the research area of a certain southern province of China is divided into 148,551 1000*1000m grids.
[0080] Specifically, based on the urban construction dataset, an urban construction index system is constructed. Through the structural characteristic analysis of the urban construction index system, the three-dimensional spatial structure characteristics of urban construction are obtained, including the following steps:
[0081] S11. Collect urban building parameters to construct an urban construction dataset, and analyze the distribution indicators and suitability indicators of urban buildings through statistical analysis.
[0082] Furthermore, urban suitability for development is also one of the characteristics of urban structure. Urban suitability assessment is an integration of factors such as urban land resources, water resources, environment, disasters, and location. It involves a comprehensive and detailed analysis of the specific conditions of various land uses and functional zones within the city, and provides insights into the current state of urban development, further clarifying the city's positioning, nature, and development goals.
[0083] S12. Based on the distribution index of urban buildings, construct the structural characteristics of urban buildings through index calculation;
[0084] Specifically, urban building structural characteristics include the weighted average height of buildings, building volume ratio, building coverage ratio, building volume, proportion of high-rise buildings, and standard deviation of building height.
[0085] Among them, the weighted average building height (BH) reflects the average height level of buildings in the city and is used to characterize building height information and the degree of vertical development of the city. When this index is analyzed on a grid basis, a schematic diagram of the calculation within each grid is shown in Figure 2.
[0086] Building Volume (BV): This index represents the three-dimensional space occupied by buildings and reflects the scale and level of building development in a city.
[0087] Building Coverage (BC): This index represents the horizontal space occupied by buildings, reflecting the building density and land use efficiency of a city.
[0088] Floor Area Ratio (BPR): This index reflects the ratio of the floor area of urban buildings to the land area, and can characterize the density of horizontal building distribution and the intensity of vertical development. When this index is analyzed in units of grid, i.e., the plot area is the grid area, the schematic diagram of which is shown in Figure 3.
[0089] High-rise building ratio (HBR): This index reflects the proportion of high-rise buildings in a city and can characterize the city's building structure. The index is calculated for each district / county within the city.
[0090] Building Height Standard Deviation (BHSD): This index reflects the dispersion of building heights in a city, characterizing the distribution and variation of building heights and architectural forms. The index is calculated for each district / county within the city.
[0091] Specifically, the formula for calculating the weighted average height of buildings is as follows:
[0092] ;
[0093] In the formula, BH represents the weighted average height of the building, and A i H represents the area occupied by the base of building i. i This represents the height information of the i-th building, and n represents the sum of the number of areas belonging to the building.
[0094] The formula for calculating the building's floor area ratio is:
[0095] ;
[0096] In the formula, BPR represents the building volume ratio, and A... i L represents the area occupied by the base of building i. i Let P represent the number of floors in building i, n represent the sum of the number of areas belonging to the building, and P represent the number of floors in building i. a This refers to the area of the regional plot;
[0097] The formula for calculating building coverage is:
[0098] ;
[0099] In the formula, BC represents the building coverage ratio, and A i Let P represent the area occupied by the base of building i, n represent the sum of the number of areas belonging to the building, and P a Indicates the area of the land parcel;
[0100] The formula for calculating building volume is:
[0101] ;
[0102] In the formula, BV represents the building volume, and A i H represents the area occupied by the base of building i. i The height data of building i is represented by n, and n represents the sum of the number of areas belonging to the building.
[0103] The formula for calculating the proportion of high-rise buildings is:
[0104] ;
[0105] In the formula, HBR represents the percentage of high-rise buildings, and N g This indicates the number of high-rise buildings (buildings with a height greater than 24m), where N represents the total number of buildings in the area.
[0106] The formula for calculating the standard deviation of building height is:
[0107] ;
[0108] In the formula, BHSD represents the standard deviation of building height, and H i This represents the height information of the i-th building, μ represents the average height of the buildings, and N represents the total number of buildings in the area.
[0109] S13. Based on the urban construction suitability index, calculate the urban construction suitability characteristics using expert scoring and entropy weight methods;
[0110] Specifically, based on urban development suitability indicators, the calculation of urban development suitability characteristics using expert scoring and entropy weight methods includes the following steps:
[0111] S131. Based on the normalization method, the suitability index for each town is normalized to obtain the normalized index;
[0112] S132. Quantify the obtained normalized index to obtain the quantified index;
[0113] S133. Combining expert scoring and entropy weighting methods, a weight analysis is performed on each quantitative indicator. Based on the weight analysis results, a weighted summation method is used to calculate the comprehensive score of each quantitative indicator, and the comprehensive score is used as the suitability characteristic for urban construction.
[0114] S14. Combine the structural characteristics of urban buildings and the suitability characteristics of urban construction to obtain the three-dimensional spatial structural characteristics of urban construction.
[0115] S2. Use base stations to obtain mobile phone signaling data and identify the commuting behavior of urban populations. Combine GIS spatial analysis methods to analyze and construct the network structure characteristics of urban population flow.
[0116] It should be noted that mobile signaling data is data automatically captured and recorded by the system when a user triggers a relevant communication event. As long as the mobile phone is powered on, it continuously interacts with communication base stations. When the mobile phone user moves to different locations at different times, the communication base station they are connected to will also change accordingly. Because the current base station location information is recorded with each interaction, mobile signaling data is generated accordingly. Each piece of signaling data contains several key fields, covering core information such as the base station identification code, the location area code of the signaling trigger, and the signaling trigger time.
[0117] Specifically, the process of acquiring mobile phone signaling data from base stations and identifying the commuting behavior of urban residents, combined with GIS spatial analysis methods to analyze and construct the network structure characteristics of urban population flow, includes the following steps:
[0118] S21. Obtain mobile phone signaling data of urban residents using base stations, and preprocess the mobile phone signaling data.
[0119] Specifically, preprocessing includes removing invalid and redundant data, ping-pong data processing, and drift data processing.
[0120] Specifically, invalid and redundant data are removed. This data loss occurs because of network problems, delays, and data transmission interruptions that may occur when mobile devices communicate with base stations. This leads to the loss of base station information that records signaling data, including base station codes, Location Area Codes (LACs), and base station latitude and longitude. Redundant data refers to multiple duplicate mobile signal data triggered by the same base station, even though the user has not actually changed their location. This invalid and redundant data affects the accuracy and efficiency of subsequent data analysis. Therefore, it needs to be filtered and cleaned during the data preprocessing stage to retain valid data samples. For missing data, appropriate data filtering is used before deletion. For redundant data generated continuously with complete base station information, a merging and addition method is used. This method adds up the duration of communication between the mobile terminal and the base station and deletes the merged data.
[0121] Ping-pong data processing occurs when mobile devices encounter weak or unstable signals during communication, leading to a decline in connection quality with the current base station. To maintain communication stability, the device attempts to switch to other base stations. Mobile devices may also cross the coverage boundaries of different base stations during movement. When a device moves to a new base station's coverage area, it attempts to establish a connection with the new base station, all of which result in frequent base station handovers. Ping-pong data is generated when a mobile device repeatedly locates itself between different base stations. Common cell coding sequences for ping-pong data handover phenomena are ABCA and ABA. Based on existing research methods, this study first determines whether a cell coding sequence belongs to the ping-pong data handover type, and then merges sequences that meet the criteria. During processing, the start time of the first signaling message and the end time of the last signaling message in the ping-pong data are updated to the corresponding fields of the new signaling message. The total duration of the repeated data is accumulated and recorded as the actual dwell time of the new signaling data at the base station.
[0122] Drift data processing refers to the phenomenon where the location of a mobile device remains constant or changes slowly over a period of time, but data shifts or drifts due to factors such as error accumulation or environmental changes. For example, during the long-term operation of a mobile device, sensors may change due to external environmental factors such as temperature fluctuations, humidity changes, or external pressure, leading to a gradual accumulation of errors in the data output, thus causing location data drift. Simultaneously, positioning system errors, algorithm errors, and signal interference can all introduce drift during data processing, affecting the accuracy of location data. Based on existing research, this paper first screens whether the spatial distance between mobile base stations recorded in signaling data exceeds a certain threshold. To identify drift data, the primary task is to assess whether the handover speed exceeds the normal range. Then, adjacent drift records are integrated, and the start time of the merged data is updated to the start time of the previous signaling and the end time to the end time of the next signaling. The duration of the integrated drift records is accumulated and recorded as the base station dwell time of the new signaling data.
[0123] S22. Based on the preprocessed mobile phone signaling data, identify the commuting behavior of urban residents by analyzing commuting data between residence and workplace during the morning and evening peak hours.
[0124] S23. Based on the commuting behavior of the urban population, calculate the commuting volume of the urban population during a preset time period;
[0125] For example, analyzing personnel movement first requires determining the user's residence and workplace. This paper uses each calendar month as the statistical unit. Each user's monthly trajectory contains numerous stops, and it's necessary to identify their residence and workplace from these points. Generally, working hours are between 7 AM and 7 PM. Since nighttime activities after get off work vary from person to person, some users may not return directly to their residence; therefore, rest time is defined as between 11 PM and 7 AM the following day. This invention defines the location where the user spends the most time between 11 PM and 7 AM the following day (and stays for more than 15 days) as their residence; and the location where the user spends the most time between 7 AM and 7 PM (and stays for more than 10 days) as their workplace.
[0126] First, relevant concepts are defined. This invention's population mobility research primarily focuses on commuting behavior between districts and counties. Commuting behavior is defined as follows: 6:00 AM to 9:00 AM and 5:00 PM to 9:00 PM on weekdays are considered the commute periods for employment. A commuting behavior is defined as a movement between a signaling user's workplace and residence or residence and workplace during these periods. The number of commuters is calculated by deduplicating signaling users who engage in commuting behavior based on their ID card numbers. The monthly average number of commuters is obtained by summing the daily commuter numbers. Commuting volume is obtained by summing single commuting behaviors, distinguishing between the origin and destination; that is, a journey from location A to location B and a journey from location B to location A are counted as two commuting behaviors. The monthly average commuting volume is obtained by summing the daily commuting volume.
[0127] The specific calculation logic is as follows: First, signaling data is read from the base station, and individuals whose workplace and residence do not overlap are identified and filtered based on their work-residence location. The signaling data of these users is analyzed, and to determine commuting behavior, the time range is filtered to the time of their round-trip commute to work. Simultaneously, based on the latitude and longitude information of the base station in their signaling data, their district / county is located, and administrative region information is added. Based on the daily base station recording time sorting, the earliest time of arrival at the residence and the earliest time of arrival at the workplace during the morning and evening peak hours are calculated. The data from the residence and workplace during the morning and evening peak hours are integrated; if the earliest time of arrival at the residence during the morning peak is less than the earliest time of arrival at the workplace, or if the earliest time of arrival at the workplace during the evening peak is less than the earliest time of arrival at the residence, it is counted as one commuting activity. Finally, according to this standard, the daily commuting volume is calculated by grouping commuting workplaces and residences, and the monthly commuting volume is calculated cumulatively. The number of commutes per person can be calculated by grouping users; and the number of daily commuters is calculated by deduplicating user ID numbers, and the monthly commuter count is calculated cumulatively. At the same time, based on the user's ID card information, the user's age information can be obtained, the user's age profile can be read, and the number of people in each age group can be statistically obtained by grouping by age.
[0128] S24. Construct a population flow network graph of the town, with building locations in the town as nodes, commuting behavior of the town's population as edges, and commuting volume of the town's population in a preset time period as the weight attribute of the edges.
[0129] S25. Use the Gephi layout algorithm to analyze the population flow network graph of towns and cities, and obtain the network structure characteristics of urban population flow.
[0130] Specifically, this invention uses complex network analysis combined with GIS spatial analysis methods to study the population flow relationship and node size between towns, analyzes and compares the network characteristics of population flow in each town, and explores the ability of different regions to attract and radiate population flow by analyzing the characteristics of population travel from different regions from the perspective of connectivity.
[0131] A population flow network diagram was drawn using Gephi software (as shown in Figure 4). Gephi is an open-source network visualization tool that can be used to analyze and visualize complex network data, helping to deeply study network structure, relationships, and patterns, and intuitively displaying network data graphically. First, a node table and edge table were created to input into Gephi. The node table contained information on districts and counties in a southern province of my country, supplemented with the latitude and longitude information of each town. This allows for geographic coordinate network visualization when using Gephi's layout algorithm to optimize the network structure. Nodes were also categorized by city level for easier visualization later. The edge table included inflow and outflow information between regions, and the commuting volume of the population in a county of a southern province of my country was used as the edge weight attribute to represent the directionality and strength of the network. Gephi can perform various network analysis tasks, such as calculating node degree, betweenness, clustering coefficient, etc., as well as detecting community structure and identifying important nodes in the network. Here, this invention first uses edge weights to define the edge style; the greater the flow, the wider the edge. To improve visualization, node sizes and edge colors are not heavily customized to avoid clutter. Instead, a provincial-level administrative region in southern my country is used for classification and visualization, with different colors representing different cities. The map shows frequent population movement and strong inter-city connections in the northeastern and central regions, indicating robust economic development, abundant employment opportunities, and well-developed urban planning and transportation infrastructure. The dense economic activities and social exchanges between cities attract significant population inflows. Conversely, weaker connections and the lack of strong population attraction and outreach in other regions suggest relatively lower economic development levels and insufficient influence of central urban areas to effectively mobilize economic exchange and attract talent.
[0132] S3. A dataset is constructed using the three-dimensional spatial structure characteristics of urban construction and the network structure characteristics of urban population flow. A regression model is built and trained using the LightGBM algorithm. The SHAP algorithm is used to interpret the regression model and obtain the correlation analysis results between the three-dimensional spatial structure of urban areas and population flow.
[0133] It should be noted that, in order to capture the complex nonlinear interactions between variables, this invention uses the LightGBM algorithm and the Shapley value interpretation model to further analyze the relationship between urban spatial characteristics and spatiotemporal behavior of population flow, to explain and visualize the effect of urban spatial factors on population flow, and to explore the correlation mechanism.
[0134] LightGBM (GBDT, Gradient Boosting Decision Tree) is an ensemble learning algorithm that uses an ensemble of multiple decision tree models to perform prediction tasks. GBDT is an ensemble learning method that progressively reduces training error by building decision tree models. In each iteration, GBDT learns a new decision tree model that attempts to capture the residuals of previously mispredicted samples in the training data, thus gradually reducing the residuals and ultimately fitting the training data. The core principle of LightGBM is to optimize the speed and efficiency of the gradient boosting algorithm through a leaf-wise growth strategy and a histogram algorithm. The leaf-wise growth strategy always splits the leaf node that provides the highest benefit to the current node, allowing for the construction of more leaf nodes at the same depth, thus improving model accuracy. The histogram algorithm first discretizes continuous feature values into different bins, then uses a histogram at each split to find the optimal split point. This algorithm reduces computation and memory usage, improving the training speed of the model.
[0135] As machine learning models become increasingly complex and their decision-making processes opaque, their interpretability decreases. The proposed Shapley value (SHAP) serves as an interpretability metric for feature values within a model, representing an additive feature attribution method with both local and global explanatory power. This concept originated in game theory to measure the value of each party in a game. In machine learning models, each feature variable of a sample can be viewed as a participant in the game, and the Shapley value of a feature reflects its contribution to the model's outcome. The calculation principle of the Shapley value is based on concepts from cooperative game theory; it measures the degree to which each feature contributes to the model's prediction.
[0136] Specifically, a dataset is constructed using the three-dimensional spatial structure characteristics of urban construction and the network structure characteristics of urban population flow. A regression model is built and trained using the LightGBM algorithm, and the SHAP algorithm is used to interpret the regression model. The results of the correlation analysis between urban three-dimensional spatial structure and population flow are obtained, including the following steps:
[0137] S31. Divide the dataset into training and testing sets, and use the LightGBM algorithm to build a regression model and initialize the model parameters.
[0138] S32. Based on the initialized model parameters, the regression model is trained using the training set, and the model parameters are adjusted and optimized through cross-validation.
[0139] S33. Use the test set to evaluate the effect of the regression model and calculate the performance evaluation index data of the model to obtain the evaluation results of the regression model.
[0140] It should be noted that, based on the LightGBM algorithm, this invention uses various variables of urban spatial structure (three-dimensional spatial structural features of urban construction) as input and population flow (network structural features of urban population flow) as output to construct a dataset. 85% of the dataset is randomly selected as the training set, and 15% as the test set to build the model. Model parameters are set, including the number of training base learners, the maximum leaf node of each base learner, the learning rate, the tree depth, and the number of iterations. Through cross-validation and continuous adjustment and optimization of parameters, the model's accuracy is further improved. Finally, the parameters are set to 70, 50, 0.01, 6, and 10000 respectively. After setting the parameters, the model is trained, and its performance is evaluated using the test set. This invention selects the coefficient of determination (R²). 2 Mean absolute error (MAE), mean absolute percentage error (MAPE), and root mean square error (RMSE) are used as indicators to verify the accuracy of the model.
[0141] In R, data is read, parameters are set, a regression model using the LightGBM algorithm is built and trained, a fitted scatter plot is obtained, and R is calculated. 2 The coefficient of determination (COD) is 0.97, RMSE is 292.04, MAE is 85.55, and MAPE is 0.13. The COD is close to 1, indicating that the model can well explain the changes in the target variable. Meanwhile, the other three error values are relatively small, indicating that the model's prediction error is small and the fit is good. The fitting results demonstrate that the influencing factor system selected in this invention is reasonable and has strong explanatory power for the dependent variable.
[0142] S34. Use the SHAP algorithm to analyze the feature importance and influence direction of the evaluated regression model, and establish a summary diagram of SHAP based on the analysis results.
[0143] Specifically, the performance evaluation metrics for the model include the coefficient of determination, mean absolute error, mean absolute percentage error, and root mean square error.
[0144] Specifically, the SHAP algorithm is used to analyze the feature importance and influence direction of the evaluated regression model, and a summary plot of SHAP is built based on the analysis results, including the following steps:
[0145] S341. Based on the principles of cooperative game theory, calculate the Shapley value using the evaluated regression model.
[0146] S342. Based on the influence of Shapley values on characteristic variables and the direction of characteristic effects, we will interpret and visualize the results, and use the SHAP model to draw visualization charts to obtain a summary chart of SHAP.
[0147] S35. Based on the summary diagram of SHAP, analyze the correlation between the three-dimensional spatial structure of towns and population flow, and obtain the correlation analysis results.
[0148] It's important to note that SHAP summary plots come in two forms: one is a scatter plot displaying the individual features of all samples, and the other is a bar chart based on the absolute mean of each feature's Shapley value, used to reflect the feature's importance. Feature importance indicates the degree of reliance on that feature in model predictions. Generating SHAP summary plots not only determines the importance ranking of all features but also analyzes the positive or negative impact of each feature's Shapley value on the prediction results, thus providing a better understanding of the contribution of different features to model predictions.
[0149] Figure 5 shows a SHAP scatter plot (where each dot represents a sample point, and the color of the dot represents the corresponding feature value; lighter colors indicate higher feature values, and darker colors indicate lower feature values). The vertical axis represents the influencing factors, i.e., the independent variable features, arranged from top to bottom according to their importance in the model. The top three are Building Coverage (BC), High-Rise Building Ratio (HBR), and Urban Suitability (ConsS). The horizontal axis represents the Shapley value of each feature, i.e., the weight of each variable's influence on the dependent variable. A value greater than 0 indicates a positive promoting effect on population mobility, while a value less than 0 indicates a negative inhibiting effect. The larger the absolute value, the deeper the influence. The scatter plot for each feature in each row of the figure represents the influence of that feature. Each dot is a sample, and the color of the dot represents the level of that feature value for each sample; lighter colors indicate higher feature values, and darker colors indicate lower feature values.
[0150] Figure 5 shows that the top three features have similar impacts on population mobility. SHAP values less than 0 are mainly found in darker colored samples, while values greater than 0 are predominantly lighter colored. This suggests that higher SHAP values promote population mobility, and vice versa. However, it's clear that the building coverage (BC) sample distribution is more dispersed, while the high-rise building ratio (HBR) and urban suitability (ConsS) samples are more densely packed, concentrated around the SHAP value of zero. This indicates that the impact of higher SHAP values on population mobility is not as pronounced as that of building coverage (BC). The impact of the remaining four features is more complex. For the weighted average building height (BH), high-valued samples are distributed around the zero SHAP value, with a roughly uniform distribution. Meanwhile, some lower-valued samples are located in the negative SHAP region. This indicates that a lower weighted average building height has a certain inhibitory effect on population flow, but the impact of a higher weighted average building height on population flow is complex and lacks a clear guiding pattern. For building volume (BH), most low-valued samples are located in the positive SHAP region, near the zero value, while low-valued samples are scattered in the negative SHAP region, indicating a complex mechanism by which this feature affects population flow. For building height standard deviation (BHSD), a large number of light-colored (high-valued) samples cluster around the zero SHAP value, indicating that this feature has little effect on the dependent variable. Similarly, building volume ratio (BPR) shows all samples clustering around the zero SHAP value, indicating that volume ratio has almost no impact on population flow in this model.
[0151] Figure 6 shows another type of SHAP bar summary chart. The vertical axis represents the various input features in the model, and the horizontal axis represents the mean of the absolute values of the SHAP of each feature, quantifying the magnitude of the feature's influence on the overall prediction of the model. Similar to the scatter plot, the features are arranged from top to bottom according to their importance in the model. However, unlike the scatter plot, the bar chart reflects the relative importance of features rather than just their ranking. As shown in Figure 6, although the percentage of high-rise buildings (HBR) and urban suitability (ConsS) rank second and third, their importance is very close, and both are only about half that of the first-ranked building coverage (BC). In addition, it is observed that the building volume ratio (BPR) not only ranks last, but its importance is almost zero.
[0152] In summary, the analysis results based on LightGBM and SHAP show that among the features, building coverage (BC), high-rise building ratio (HBR), and urban suitability (ConsS) are the most important. The higher their feature values, the more they promote population mobility. Among these, building coverage is twice as important as the second most important feature, having the greatest impact on pedestrian flow. Therefore, spatial structure is also a driving factor for population mobility characteristics.
[0153] In summary, by utilizing the above-mentioned technical solutions of this invention, the present invention ensures an accurate grasp of the essential characteristics of urban spatial structure through the analysis of the three-dimensional features of urban spatial structure. Simultaneously, through the analysis of the network characteristics of population flow, it grasps the distribution of urban population flow and the characteristics of human activities, further providing a basis for in-depth analysis of the relationship between urban spatial structure and population flow. Then, through the analysis of the spatiotemporal correlation between urban spatial characteristics and population flow, it explores the importance of various features of urban spatial structure to population flow, thereby effectively utilizing population flow trends to provide valuable reference for urban development planning. This invention analyzes the three-dimensional characteristics of urban spatial structure by obtaining urban construction suitability and six indicators based on urban building height datasets and various urban element indicators, thereby describing the architectural form, spatial distribution, and urban development level of the city, ensuring an accurate grasp of the essential characteristics of urban spatial structure. This invention identifies users' commuting behavior based on mobile phone signaling data and then analyzes and compares the network characteristics of population flow in various towns using GIS spatial analysis methods. This allows for an understanding of the distribution of population flow and the characteristics of human activities in these towns, thereby gaining insight into the economic development trends, employment opportunities, urban planning, transportation infrastructure, and the density of economic activities and social exchanges between towns. This provides a basis for further analysis of the relationship between urban spatial structure and population flow. Based on the LightGBM algorithm, this invention constructs a model using the comprehensive suitability of urban construction and six indicators as inputs, with population flow as the output. The SHAP algorithm is then used to analyze the impact of each input feature on population flow, exploring the importance of various features of urban spatial structure on population flow. This accurately grasps the driving factors of population flow, further analyzes the correlation between urban spatial characteristics and population flow, and effectively utilizes the correlation results to provide valuable references for urban development planning.
[0154] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for correlation analysis between urban spatial three-dimensional structure and population flow, characterized in that, This correlation analysis method includes the following steps: S1. Construct an urban construction indicator system based on the urban construction dataset, and obtain the three-dimensional spatial structural characteristics of urban construction by analyzing the structural features of the urban construction indicator system. S2. Use base stations to obtain mobile phone signaling data and identify the commuting behavior of urban populations. Combine GIS spatial analysis methods to analyze and construct the network structure characteristics of urban population flow. S3. A dataset is constructed using the three-dimensional spatial structure characteristics of urban construction and the network structure characteristics of urban population flow. A regression model is built and trained using the LightGBM algorithm. The SHAP algorithm is used to interpret the regression model and obtain the correlation analysis results between the three-dimensional spatial structure of urban areas and population flow.
2. The method for correlation analysis of urban spatial three-dimensional structure and population flow according to claim 1, characterized in that, The process of constructing an urban construction indicator system based on an urban construction dataset, and obtaining the three-dimensional spatial structural features of urban construction through structural feature analysis of the urban construction indicator system, includes the following steps: S11. Collect urban building parameters to construct an urban construction dataset, and analyze the distribution indicators and suitability indicators of urban buildings through statistical analysis. S12. Based on the distribution index of urban buildings, construct the structural characteristics of urban buildings through index calculation; S13. Based on the urban construction suitability index, calculate the urban construction suitability characteristics using expert scoring and entropy weight methods; S14. Combine the structural characteristics of urban buildings and the suitability characteristics of urban construction to obtain the three-dimensional spatial structural characteristics of urban construction.
3. The method for correlation analysis of urban spatial three-dimensional structure and population flow according to claim 2, characterized in that, The urban building structure characteristics include the weighted average building height, building volume ratio, building coverage ratio, building volume, proportion of high-rise buildings, and standard deviation of building height.
4. The method for correlation analysis of urban spatial three-dimensional structure and population flow according to claim 2, characterized in that, The formula for calculating the weighted average height of the buildings is as follows: ; In the formula, BH represents the weighted average height of the building, and A i H represents the area occupied by the base of building i. i This represents the height information of the i-th building, and n represents the sum of the number of areas belonging to the building. The formula for calculating the building's floor area ratio is: ; In the formula, BPR represents the building volume ratio, and A... i L represents the area occupied by the base of building i. i Let P represent the number of floors in building i, n represent the sum of the number of areas belonging to the building, and P represent the number of floors in building i. a This refers to the area of the regional plot; The formula for calculating the building coverage ratio is: ; In the formula, BC represents the building coverage ratio, and A i Let P represent the area occupied by the base of building i, n represent the sum of the number of areas belonging to the building, and P a Indicates the area of the land parcel; The formula for calculating the building volume is: ; In the formula, BV represents the building volume, Ai represents the area occupied by the base of building i, and H... i The height data of building i is represented by n, and n represents the sum of the number of areas belonging to the building. The formula for calculating the proportion of high-rise buildings is as follows: ; In the formula, HBR represents the percentage of high-rise buildings, and N g This represents the number of high-rise buildings, and N represents the total number of buildings in the area. The formula for calculating the standard deviation of building height is: ; In the formula, BHSD represents the standard deviation of building height, and H i This represents the height information of the i-th building, μ represents the average height of the buildings, and N represents the total number of buildings in the area.
5. The method for correlation analysis of urban spatial three-dimensional structure and population flow according to claim 2, characterized in that, The calculation of urban construction suitability characteristics based on urban construction suitability indicators, using expert scoring and entropy weight methods, includes the following steps: S131. Based on the normalization method, the suitability index for each town is normalized to obtain the normalized index; S132. Quantify the obtained normalized index to obtain the quantified index; S133. Combining expert scoring and entropy weighting methods, a weight analysis is performed on each quantitative indicator. Based on the weight analysis results, a weighted summation method is used to calculate the comprehensive score of each quantitative indicator, and the comprehensive score is used as the suitability characteristic for urban construction.
6. The method for correlation analysis of urban spatial three-dimensional structure and population flow according to claim 2, characterized in that, The process of acquiring mobile phone signaling data from base stations and identifying the commuting behavior of urban residents, combined with GIS spatial analysis methods to analyze and construct the network structure characteristics of urban population flow, includes the following steps: S21. Obtain mobile phone signaling data of urban residents using base stations, and preprocess the mobile phone signaling data. S22. Based on the preprocessed mobile phone signaling data, identify the commuting behavior of urban residents by analyzing commuting data between residence and workplace during the morning and evening peak hours. S23. Based on the commuting behavior of the urban population, calculate the commuting volume of the urban population during a preset time period; S24. Construct a population flow network graph of the town, with building locations in the town as nodes, commuting behavior of the town's population as edges, and commuting volume of the town's population in a preset time period as the weight attribute of the edges. S25. Use the Gephi layout algorithm to analyze the population flow network graph of towns and cities, and obtain the network structure characteristics of urban population flow.
7. The method for correlation analysis of urban spatial three-dimensional structure and population flow according to claim 6, characterized in that, The preprocessing includes invalid and redundant data removal, ping-pong data processing, and drift data processing.
8. The method for correlation analysis of urban spatial three-dimensional structure and population flow according to claim 1, characterized in that, The process of constructing a dataset using the three-dimensional spatial structure characteristics of urban construction and the network structure characteristics of urban population flow, building and training a regression model using the LightGBM algorithm, and interpreting the regression model using the SHAP algorithm to obtain the correlation analysis results between urban three-dimensional spatial structure and population flow includes the following steps: S31. Divide the dataset into training and testing sets, and use the LightGBM algorithm to build a regression model and initialize the model parameters. S32. Based on the initialized model parameters, the regression model is trained using the training set, and the model parameters are adjusted and optimized through cross-validation. S33. Use the test set to evaluate the effect of the regression model and calculate the performance evaluation index data of the model to obtain the evaluation results of the regression model. S34. Use the SHAP algorithm to analyze the feature importance and influence direction of the evaluated regression model, and establish a summary diagram of SHAP based on the analysis results. S35. Based on the summary diagram of SHAP, analyze the correlation between the three-dimensional spatial structure of towns and population flow, and obtain the correlation analysis results.
9. The method for correlation analysis of urban spatial three-dimensional structure and population flow according to claim 8, characterized in that, The performance evaluation metrics of the model include the coefficient of determination, mean absolute error, mean absolute percentage error, and root mean square error.
10. The method for correlation analysis of urban spatial three-dimensional structure and population flow according to claim 9, characterized in that, The process of using the SHAP algorithm to analyze the feature importance and influence direction of the evaluated regression model, and then establishing a summary plot of SHAP based on the analysis results, includes the following steps: S341. Based on the principles of cooperative game theory, calculate the Shapley value using the evaluated regression model. S342. Based on the influence of Shapley values on characteristic variables and the direction of characteristic effects, we will interpret and visualize the results, and use the SHAP model to draw visualization charts to obtain a summary chart of SHAP.
Citation Information
Patent Citations
Pedestrian flow simulation display method and system for digital twin cities
CN118261763A
Method for analyzing correlation between urban space three-dimensional structure and population flow
CN119443378A
System, apparatus and method for mapping
US20060136126A1
Method for creating a map relating to location-related data on the probability of future movement of a person
US20120232795A1