Multi-time-scale historical population distribution simulation method and system
Patent Information
- Application Number
- CN202510925018.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-11-14
AI Technical Summary
Existing technologies struggle to acquire population distribution data across multiple time scales. Limited data acquisition and difficulty in integrating data from different technological sources result in insufficient data accuracy and reliability, thus limiting their application in different regions and scenarios.
By acquiring historical POI data for the target area, a population distribution simulation model is trained using machine learning algorithms. Based on sample data at different time scales, a multi-time-scale population distribution simulation method and system are constructed, including a POI data acquisition module and a population distribution simulation module. The XGBoost machine learning algorithm is used for training and prediction.
It enables detailed simulation of population distribution in target areas within a preset historical period, providing more accurate population distribution data across multiple time scales and supporting scientific decision-making in areas such as urban planning and resource allocation.
Smart Images

Figure CN120950753A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method and system for simulating historical population distribution across multiple time scales. Background Technology
[0002] Accurate understanding of population distribution is of vital importance for many areas such as urban planning, resource allocation, and business layout.
[0003] To obtain population distribution data, researchers and related institutions have developed and applied several technological methods. Common methods include analysis techniques based on mobile phone signaling data. For example, by counting the number of mobile phone users in a certain area at different times, one can roughly understand the changes in population density in that area. In addition, there is pedestrian flow monitoring technology based on smart cameras. Smart cameras are installed in public places such as main streets, shopping malls, and stations in cities. Using image recognition and artificial intelligence algorithms, the images captured by the cameras are analyzed to count the number of people passing through the area at different times, thereby obtaining relevant information on population distribution.
[0004] While existing technologies have played a role in acquiring population distribution data, they all have significant drawbacks. First, data acquisition is limited and availability is low; this is especially true for population distribution data from specific historical periods, which is particularly difficult to obtain. Second, due to the different data sources and collection methods used by various technologies, the acquired population distribution data varies in format and accuracy, posing significant challenges to data integration and analysis and limiting the application of the data in different regions and scenarios. Therefore, there is an urgent need for a multi-timescale historical population distribution simulation method and system to address these issues. Summary of the Invention
[0005] To address the problems existing in the prior art, this invention provides a method and system for simulating historical population distribution across multiple time scales.
[0006] This invention provides a method for simulating historical population distribution across multiple time scales, comprising: Obtain historical POI data for the target area within a preset historical period; The historical data of the POI is input into population distribution simulation models at different time scales to obtain population distribution simulation data of the target area at the corresponding time scale of the preset historical period, output by the population distribution simulation model. The real-time population distribution simulation models at different time scales are trained using machine learning algorithms based on sample data.
[0007] According to the present invention, a method for simulating historical population distribution at multiple time scales includes inputting the historical data of the Point of Interest (POI) into population distribution simulation models at different time scales to obtain population distribution simulation data of the target area at the corresponding time scale of the preset historical period, output by the population distribution simulation models. Determine the target time scale type corresponding to the historical data of the POI; According to the target time scale type, a target population distribution simulation model is determined from multiple population distribution simulation models of different time scales. Each population distribution simulation model is trained using the machine learning algorithm based on population distribution sample data of different time scale types and POI sample data corresponding to the population distribution sample data. The historical data of the POI is input into the target population distribution simulation model to obtain the population distribution simulation data of the target area under the target time scale type within the preset historical period, output by the target population distribution simulation model.
[0008] According to the present invention, a method for simulating historical population distribution across multiple time scales is provided, wherein the population distribution simulation model is trained through the following steps: Based on preset spatial boundary information, obtain basic population data of the sample area within a preset period; Based on the spatiotemporal resolution corresponding to the base population data, the base population data is divided into population distribution sample data of different time scale types, wherein the time scale types include at least daytime time scale, nighttime time scale, weekday time scale and weekend time scale under non-holiday conditions; The population distribution sample data is obtained by acquiring POI sample data within the spatial boundary of the preset period, and the type of the acquired POI sample data and the number of each type of POI are used as the label data of the population distribution sample data to obtain the sample data. Based on the sample data of different time scale types, a training dataset is constructed for the population distribution simulation model corresponding to each time scale type; Based on the training datasets of the population distribution simulation models corresponding to each of the aforementioned time scale types, the XGBoost machine learning algorithm is used for training, and the population distribution simulation models for different time scale types are obtained based on the optimal parameter tuning technique.
[0009] According to the present invention, a method for simulating historical population distribution at multiple time scales, wherein obtaining base population data of a sample area within a preset period based on preset spatial boundary information includes: The sample region range is determined based on the preset spatial boundary information; The sample area is divided into multiple sub-regions, and the population distribution corresponding to each time scale is obtained based on the total number of people in each sub-region at different time scales within the preset period. Based on the population distribution corresponding to all the sub-regions, the base population data within the sample region during the preset period is obtained. According to a multi-timescale historical population distribution simulation method provided by the present invention, obtaining the population distribution corresponding to each time scale based on the total number of people in each sub-region at different time scales within the preset period includes: Based on a preset buffer range, a population distribution statistics buffer is set in each of the sub-regions; Based on the total number of people at different time scales within the preset period, according to the population distribution statistical buffer corresponding to each grid area and each sub-area, the population data corresponding to each time scale is obtained.
[0010] This invention also provides a multi-timescale historical population distribution simulation system, comprising: The POI data acquisition module is used to acquire historical POI data for the target area within a preset historical period. The population distribution simulation module is used to input the historical data of the POI into population distribution simulation models at different time scales to obtain population distribution simulation data of the target area range at the corresponding time scale of the preset historical period, output by the population distribution simulation model. The population distribution simulation model is trained based on sample data using machine learning algorithms.
[0011] According to the present invention, a historical population distribution simulation system with multiple time scales is provided, wherein the population distribution simulation module includes: The time scale determination unit is used to determine the target time scale type corresponding to the POI historical data; The target model determination unit is used to determine a target population distribution simulation model from multiple population distribution simulation models of different time scales according to the target time scale type, wherein each population distribution simulation model is trained using the machine learning algorithm based on population distribution sample data of different time scale types and POI sample data corresponding to the population distribution sample data; The processing unit is used to input the historical data of the POI into the target population distribution simulation model to obtain the population distribution simulation data of the target area range under the target time scale type within the preset historical period, output by the target population distribution simulation model.
[0012] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the multi-timescale historical population distribution simulation method as described above.
[0013] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the multi-timescale historical population distribution simulation method as described above.
[0014] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the multi-timescale historical population distribution simulation method as described above.
[0015] The present invention provides a multi-timescale historical population distribution simulation method and system, which acquires historical POI data of a target area that reflects socio-economic activities during a preset historical period, and then inputs the historical POI data into a population distribution simulation model for different time periods trained based on sample data. The model captures the relationship between population distribution and the socio-economic activities represented by POIs, and outputs population distribution simulation data of the target area at different time periods during the preset historical period, thereby achieving a more refined population distribution simulation within the region. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0017] Figure 1 A flowchart illustrating the multi-timescale historical population distribution simulation method provided by this invention; Figure 2 A schematic diagram of population distribution data during daytime and nighttime periods provided by this invention; Figure 3 A schematic diagram of population distribution simulation data for different periods of the year provided by this invention; Figure 4 A schematic diagram illustrating the explanatory results of the population distribution simulation model provided by this invention; Figure 5 A schematic diagram showing the results of the SHAPImportance value and Direction attribute in the population distribution simulation model of various features provided by the present invention at different time periods; Figure 6A schematic diagram of the structure of the multi-timescale historical population distribution simulation system provided by the present invention; Figure 7 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0019] With rapid urbanization and increasingly frequent population movement, population distribution at different times exhibits dynamic changes. For example, during the daytime on weekdays, population density in urban central business districts increases significantly, while at night, the population in residential areas becomes relatively concentrated. Therefore, obtaining population distribution data across multiple time scales can provide relevant departments and enterprises with more comprehensive and accurate information, enabling them to make more scientific and rational decisions. However, obtaining multi-timescale population distribution data currently faces numerous difficulties, which to some extent restricts the development of related fields and the scientific nature of decision-making.
[0020] To obtain population distribution data across multiple time scales, common methods include analysis techniques based on mobile signaling data. Mobile signaling data is generated by mobile phone users during communication, containing information such as user location and call duration. By collecting and analyzing large amounts of mobile signaling data, population distribution at different times can be inferred. For example, by counting the number of mobile phone users in a certain area at different times, one can roughly understand the changes in population density in that area.
[0021] In addition, there is pedestrian monitoring technology based on smart cameras. Smart cameras are installed in public places such as main streets, shopping malls, and stations in cities. Using image recognition and artificial intelligence algorithms, the images captured by the cameras are analyzed to count the number of people passing through the area at different times, thereby obtaining relevant information on population distribution. Furthermore, some cities also collect population distribution data through questionnaires and community registration. Although the data obtained through these methods is relatively coarse, it can still reflect the population distribution to some extent.
[0022] While existing technologies have played a role in acquiring population distribution data across multiple time scales, they all have significant limitations. First, data acquisition is geographically restricted. Acquisition of mobile signaling data depends on the network coverage of telecommunications operators; in remote areas or places with poor signal, data may be missing or inaccurate. The installation of smart cameras is also limited by factors such as space and funding, making it impossible to cover all areas of the city, resulting in the inability to obtain population distribution data for some areas.
[0023] Secondly, the accuracy and reliability of the data need improvement. Mobile phone signaling data may be affected by factors such as users turning off their phones or enabling airplane mode, leading to incomplete data. Smart cameras may make identification errors in densely populated areas, affecting population statistics. Questionnaires and community registration methods may be subject to sampling bias and cannot fully and accurately reflect the true population distribution.
[0024] Finally, discrepancies exist between data acquired using different technologies, making integration and analysis difficult. Because various technologies rely on different data sources and collection methods, the resulting population distribution data vary in format and accuracy, posing significant challenges to data integration and analysis and limiting its application in different regions and scenarios.
[0025] Figure 1 A flowchart illustrating the multi-timescale historical population distribution simulation method provided by this invention is shown below. Figure 1 As shown, the present invention provides a method for simulating historical population distribution, comprising: Step 101: Obtain historical POI data for the target area within a preset historical period.
[0026] In this invention, an API interface for Point of Interest (POI) data can be provided by a map service provider. For example, to obtain all restaurant-related POI data for a city in 2020, the corresponding parameters can be set, and the API can be called to obtain JSON format data, including fields such as name, coordinates (latitude and longitude), address, and category.
[0027] Step 102: Input the historical data of the POI into population distribution simulation models at different time scales to obtain population distribution simulation data of the target area at the time scale corresponding to the preset historical period, output by the population distribution simulation model. The population distribution simulation models at different time scales are trained using machine learning algorithms based on sample data.
[0028] In this invention, a pre-trained machine learning algorithm (i.e., a population distribution simulation model) is used. Simultaneously, the target area is divided into multiple sub-regions (e.g., in a grid format). The number of each type of POI within each sub-region is counted and used as input variables to the population distribution simulation model, thereby outputting the population distribution simulation results within each grid. Specifically, in this invention, there are many types of POIs, and they need to be classified for ease of analysis and research. For example, POIs can be classified into commercial (shopping malls, supermarkets, specialty stores, etc.), educational (schools, training institutions, etc.), medical (hospitals, clinics, etc.), leisure and entertainment (parks, cinemas, amusement parks, etc.), and public service (government agencies, libraries, post offices, etc.). Classification allows for a clearer understanding of the distribution of different types of POIs within each sub-region. After determining the POI types within each grid sub-region, it is necessary to count the specific number of each type of POI within that sub-region. For example, in a certain grid sub-region, there may be 5 shopping malls, 3 schools, 2 hospitals, etc. These numbers can intuitively reflect the density and distribution characteristics of various types of POIs within that sub-region.
[0029] In this invention, the sample data includes known population distribution information at different time scales and corresponding POI data, etc. For example, the number of POIs and population data of various types in multiple administrative regions of a certain area are collected, and the samples are randomly divided into training set and test set. The training set is used to train the machine learning algorithm, and the test set is used to test the prediction accuracy of the model.
[0030] In this invention, before inputting historical POI data into the population distribution simulation model, data preprocessing is required, including handling missing values and coordinate transformation. Further, the preprocessed historical POI data is input into the population distribution simulation model. The model analyzes and processes the input data based on learned relationships, resulting in population distribution simulation data that can include information such as population density and population size at different times and locations. For example, by analyzing the simulation data, one can understand the differences in population distribution during the day and night in a certain area, as well as the population aggregation around different types of POIs.
[0031] In this invention, these simulation data can provide a reference for urban planning, resource allocation, and other purposes. For example, urban planning departments can use population distribution simulation data to rationally plan the layout of commercial areas, residential areas, and public facilities; traffic management departments can optimize the setting of traffic lights and the planning of bus routes based on population flow.
[0032] The present invention provides a multi-timescale historical population distribution simulation method that obtains historical POI data of a target area that reflects socio-economic activities during a preset historical period, and then inputs the historical POI data into a population distribution simulation model trained based on sample data. The model captures the relationship between population distribution and the socio-economic activities represented by POIs, and outputs population distribution simulation data of the target area at different time periods during the preset period, thereby achieving a more refined population distribution simulation within the region.
[0033] Based on the above embodiments, the step of inputting the historical POI data into population distribution simulation models at different time scales to obtain population distribution simulation data of the target area at the corresponding time scale of the preset period, output by the population distribution simulation model, includes: Determine the target time scale type corresponding to the historical data of the POI; According to the target time scale type, a target population distribution simulation model is determined from multiple population distribution simulation models of different time scales, wherein each population distribution simulation model is trained using the machine learning algorithm based on population distribution sample data of different time scale types and POI sample data corresponding to the population distribution sample data; The historical data of the POI is input into the target population distribution simulation model to obtain the population distribution simulation data of the target area under the target time scale type within the preset historical period, output by the target population distribution simulation model.
[0034] In this invention, population distribution exhibits different characteristics at different time scales. For example, on a daytime time scale, people typically congregate in places of work, study, and shopping, resulting in dense populations around commercial areas, office buildings, and schools; while on a nighttime time scale, people mostly return to residential areas, leading to a significant increase in population density in residential communities. On a weekday time scale, population movement patterns are relatively fixed, primarily commuting; on a weekend time scale, people engage in more leisure and entertainment activities, resulting in population concentrations in places such as parks, shopping malls, and cinemas. Therefore, this invention requires selecting an appropriate population distribution simulation model based on the target time scale type corresponding to the historical POI data.
[0035] In this invention, each population distribution simulation model is trained using machine learning algorithms based on population distribution sample data and corresponding POI sample data at different time scales. The machine learning algorithms establish corresponding mapping relationships by learning the inherent connections between POI data and population distribution data at different time scales. For example, for a daytime time scale model, the training data may contain the correspondence between the number of POIs (such as shopping malls and restaurants) in commercial areas during weekdays and the population density of that area. After learning this relationship, the machine learning algorithm can predict the daytime population distribution of that area based on new POI data.
[0036] During model training, the first step is to analyze the temporal characteristics reflected in the historical POI data. For example, if the historical POI data contains a large amount of information about locations with high daytime activity, such as office buildings and factories, and the data collection time is concentrated during weekday working hours, then the target time scale type can be initially determined to be a daytime time scale (weekdays). Then, the target time scale type is determined based on the specific application scenario and business needs. For instance, if a city traffic management department wants to understand population flow during weekday morning and evening rush hours to optimize traffic light settings, then the target time scale type should be determined to be a daytime time scale (weekday morning and evening rush hours).
[0037] Furthermore, a model library is established containing population distribution simulation models of multiple different time scales. Each model has been trained on sample data at a specific time scale and has the ability to predict population distribution for that time scale. In this invention, based on the determined target time scale type, a target population distribution simulation model corresponding to that time scale is selected from the model library. For example, if the target time scale type is the weekend time scale, a model trained on weekend time scale sample data is selected from the model library.
[0038] Similarly, before inputting historical POI data into the target population distribution simulation model, the data needs to be preprocessed, including data cleaning to remove duplicate, erroneous, or missing key information from POI data. Data standardization unifies the processing of POI data of different types and magnitudes, enabling the model to better identify them. Then, the preprocessed historical POI data is input into the target population distribution simulation model. Based on its internally learned mapping relationships, the model outputs simulated population distribution data for the target area within a preset historical time period at the target time scale. For example, after inputting POI data from the vicinity of a commercial area on weekends, the target population distribution simulation model outputs simulated data such as population density and pedestrian traffic in that commercial area at different times of the weekend.
[0039] In one embodiment, suppose a city planning department wants to understand the population distribution of a commercial district in the city at different time scales in order to conduct reasonable business planning and infrastructure construction. This requires analyzing the population distribution of the commercial district at three time scales: daytime (weekdays), nighttime (weekdays), and weekends. Therefore, the target time scale types are determined sequentially as daytime (weekdays), nighttime (weekdays), and weekend. In addition, this invention can construct more time scale types, such as a specific time period (e.g., 10:00 AM to 2:00 PM), based on the actual population distribution simulation accuracy, to conform to population distribution simulations in more scenarios.
[0040] Then, population distribution simulation models for different time scales were established. For the daytime time scale (weekdays), the model was trained based on POI data (such as the number of office buildings, shopping malls, and restaurants) and corresponding population distribution data for multiple business districts in the city during weekdays. For the nighttime time scale (weekdays), the model was trained based on POI data (such as the number of bars and convenience stores) and population distribution data for weekday nights. For the weekend time scale, the model was trained based on POI data and population distribution data for each time period on weekends. Based on the determined target time scale type, the corresponding model was selected from the model library as the target population distribution simulation model.
[0041] Furthermore, historical POI data for the business district within a preset historical period (such as the past month) is collected, including information such as the number and location of various venues at different times. After preprocessing the data, it is input into the corresponding target population distribution simulation models. For example, weekday daytime POI data is input into the daytime timescale (weekday) model to obtain simulated population distribution data for the business district during weekdays, such as customer traffic and population density in different areas; weekday nighttime POI data is input into the nighttime timescale (weekday) model to obtain simulated population distribution data at night; and weekend time-scale POI data is input into the weekend timescale model to obtain simulated population distribution data for the weekend.
[0042] Based on the above embodiments, the population distribution simulation model is trained through the following steps: Based on preset spatial boundary information, obtain basic population data of the sample area within a preset period; Based on the spatiotemporal resolution corresponding to the base population data, the base population data is divided into population distribution sample data of different time scale types, wherein the time scale types include at least daytime time scale, nighttime time scale, weekday time scale and weekend time scale under non-holiday conditions; The population distribution sample data is obtained by acquiring POI sample data within the spatial boundary of the preset period, and the type and number of each type of POI sample data are used as the label data of the population distribution sample data to obtain the sample data. Based on the sample data of different time scale types, a training dataset is constructed for the population distribution simulation model corresponding to each time scale type; Based on the training datasets of the population distribution simulation models corresponding to each of the aforementioned time scale types, the XGBoost machine learning algorithm is used for training, and the population distribution simulation models for different time scale types are obtained based on the optimal parameter tuning technique.
[0043] In this invention, the preset spatial boundary information is the administrative boundary of a certain region, and subsequent population data collection and analysis will revolve around this region. The preset period is four seasons, with data collected for 7 days in each season, totaling 672 hours. By collecting data over a longer time span (covering all four seasons), the population distribution characteristics of different seasons and time periods can be more comprehensively reflected, avoiding biases caused by seasonal factors.
[0044] Furthermore, using the divided sub-regions as a grid for illustration, population location data from a map service provider in 2024 was used. This data effectively represents the hourly effective population dwell time within a 200m × 200m grid. This data meticulously records the population's dwell time at specific times and spatial locations, providing a foundation for subsequent analysis. During the data collection process, population distribution information was collected from 203,927 record points within the city's administrative boundaries. This information constitutes the base population data. Simultaneously, the impact of holidays and other special circumstances on population distribution was excluded to ensure the accuracy and representativeness of the data.
[0045] In this invention, the base population data has a specific spatiotemporal resolution, namely a 200m × 200m grid area and hourly time intervals. This resolution allows the data to accurately reflect the spatial and temporal distribution changes of the population. Then, by calculating the total number of active people at each point over 24 hours, the fluctuation trend of the city's population is captured to determine the day-night division standard that conforms to the local context. It should be noted that in this invention, the specific spatiotemporal resolution can be determined in the early stages of the population distribution simulation process according to actual needs, and then the pre-set spatiotemporal resolution can be directly used for division when performing time scale type division.
[0046] Figure 2 The schematic diagram of population distribution data during the day and night periods provided by this invention can be used as a reference. Figure 2As shown, statistical analysis revealed that the number of visitors to the city dynamically increased from 5:00 AM to 5:00 PM and continuously decreased from 6:00 PM to 4:00 AM the following day. Therefore, this invention defines 5:00 AM to 5:00 PM as the daytime time scale and 6:00 PM to 4:00 AM the following day as the nighttime time scale. Based on this classification, data for the corresponding time periods were extracted from the baseline population data to form population distribution sample data for the daytime and nighttime time scales.
[0047] Furthermore, based on the specific collection time, weekdays and weekends are clearly distinguished. Of the 7 days collected in each season, 5 are weekdays and 2 are weekends. By extracting the corresponding weekday and weekend data from the base population data, population distribution sample data at weekday and weekend time scales are formed. Specific collection dates are shown in Table 1. Table 1
[0048] Then, based on the time periods of Workday / Weekend and Daytime / Nighttime, the nearest neighbor aggregation tool in ArcGIS Pro is used to count the total number of people within each grid as the population distribution data for each time period. The statistical unit for the Workday / Weekend period is "total number of people per day," and the statistical unit for the Daytime / Nighttime period is "total number of people per day / night." Preferably, a 1km buffer zone is set around each grid, and then the total number of people within each grid is counted as the population distribution data for each time period.
[0049] In this invention, POI data includes various types of location information related to people's lives and activities, such as shopping malls, schools, hospitals, and parks. To accurately reflect the potential correlation between socioeconomic activities and population distribution in different regions, the specific categories and codes of the POI data provided by this invention are shown in Table 2: Table 2 Specific Categories and Codes of POI Data
[0050] Furthermore, POI sample data is used as label data for population distribution sample data to explain the reasons for population distribution through POI data. For example, areas with a high density of POIs, such as shopping malls and restaurants, often have a relatively concentrated population. Specifically, in this invention, for each grid sub-region, POIs falling within that grid are classified and statistically analyzed according to a predefined POI type classification standard. Common POI type classifications may include commercial (e.g., shopping malls, supermarkets, restaurants), educational (e.g., schools, training institutions), medical (e.g., hospitals, clinics), transportation (e.g., subway stations, bus stops), and public service (e.g., government agencies, libraries), etc. This invention generates a statistical result on the POI type distribution of that grid by counting the number of each type of POI within each grid sub-region. For example, a certain grid may contain 5 commercial POIs, 2 educational POIs, and 1 medical POI.
[0051] Furthermore, by using POI data as labels, a correlation between population distribution and POIs can be established, providing a basis for subsequent model training. For example, when training the model, population distribution sample data can be used as input features, and POI sample data as output labels, allowing the machine learning algorithm to learn the mapping relationship between the two.
[0052] In this invention, sample data of different time scales are organized into training datasets. Each training dataset contains population distribution sample data and corresponding POI sample data (label data) for the corresponding time scale. For example, the daytime time scale training dataset contains population distribution data and corresponding POI data during the daytime; the weekday time scale training dataset contains population distribution data and corresponding POI data for weekdays. By constructing multiple training datasets, population distribution characteristics at different time scales can be independently modeled, improving the accuracy and adaptability of the model.
[0053] XGBoost is a high-efficiency machine learning algorithm with advantages such as handling large-scale data, preventing overfitting, and supporting parallel computing. In population distribution simulation, the XGBoost machine learning algorithm can learn the complex nonlinear relationship between population distribution and POIs, improving the model's prediction accuracy. This invention uses training datasets corresponding to various time scales to train the XGBoost machine learning algorithm. During training, the XGBoost machine learning algorithm continuously adjusts its internal parameters based on the input population distribution sample data and corresponding POI label data to minimize prediction error. For example, for the daytime time scale training dataset, the XGBoost machine learning algorithm learns the relationship between daytime population distribution and POIs; for the weekday time scale training dataset, the XGBoost machine learning algorithm learns the relationship between weekday population distribution and POIs.
[0054] After training, population distribution simulation models of different time scales are obtained. These models can simulate population distribution at corresponding time scales based on new POI data. For example, by inputting POI data for a region on weekdays, the daytime time scale model can simulate the population distribution of that region during the weekdays; by inputting POI data for weekends, the weekend time scale model can simulate the population distribution of that region on weekends.
[0055] Machine learning algorithms are advantageous for handling nonlinear relationships. This invention uses POI data as a large-scale network dataset capable of characterizing urban socioeconomic features, thereby representing urban vitality and human activity. This invention integrates the number of various POIs and their population within a 1000-meter buffer zone in 2024, and utilizes machine learning algorithms to construct a population distribution simulation model based on POI distribution characteristics.
[0056] This invention compares several machine learning algorithms, including Random Forest (RF), Extreme Gradient Boosting (XGBoost), and Lightweight Gradient Boosting Machine (LightGBM), and selects the algorithm with the best performance for subsequent historical population prediction. During the training phase, hyperparameters of each machine learning algorithm are tuned using random search. The optimal parameters obtained from the hyperparameter tuning are shown in Table 3. Table 3 Optimal parameters for three machine learning algorithms
[0057] In one embodiment, for relevant sample data of a certain city, the optimal model (XGBoost) and its optimal parameters are: colsample_bytree: 0.8, learning_rate: 0.05, max_depth: 12, min_child_weight: 3, n_estimators: 975, reg_alpha: 0.5, reg_lambda: 1, subsample: 0.5.
[0058] This invention introduces RMSE, MAE, and R 2 and Adj-R 2 A comprehensive evaluation of model performance was conducted. RMSE is more sensitive to potentially large errors, while MAE can robustly assess the overall accuracy of the model. (Adj-R) 2 Then it can be in R 2 Based on the initial model evaluation, adjustments were made to mitigate the impact of the number of variables and prevent overfitting. Specific performance comparison results are shown in Table 4. Table 4 Performance Comparison of Various Machine Learning Algorithms at Different Time Periods
[0059] Preliminary experiments showed that the population distribution simulation model trained using the XGBoost machine learning algorithm maintained the best performance across four time periods.
[0060] After the model training is completed, the historical POI statistics of each 200m×200m grid (i.e., historical POI data) are used as input and substituted into the population distribution simulation model trained based on data from a certain year (e.g., 2024) to obtain the historical population prediction results of each grid at different time periods (population distribution simulation data). Figure 3 The diagram illustrates population distribution simulation data for different time periods over the years provided by this invention. It shows population distribution simulation data output from corresponding population distribution simulation models, such as population distribution simulation data at daytime, nighttime, weekday, and weekend time scales. (For reference...) Figure 3 As shown.
[0061] Furthermore, to explain the contribution of different types of POI data to population distribution simulation, this invention combines machine learning techniques with the SHAP (Shapley Additive Explanations) method, introducing Shapley values to reveal the relative importance of variables, nonlinear relationships, thresholds, and variable interactions, thereby enhancing the interpretability of the population distribution simulation model and the reliability of the output results.
[0062] Specifically, this invention integrates machine learning and interpretability analysis techniques, and combines multiple parameters such as SHAP_Importance, XGB_Gain / Weight, and Direction to quantify the impact mechanism of urban socio-economic factors on population distribution.
[0063] Here, SHAP_Importance is the global feature importance based on SHAP values. In this invention, the population distribution simulation model trained based on the XGBoost machine learning algorithm can obtain the Importance of each feature by calculating and normalizing the global absolute mean of the SHAP value of each feature (i.e., the marginal contribution of the feature to the predicted value). This Importance reflects the overall strength of the feature's influence on population distribution.
[0064] XGB_Gain is a built-in gain importance in XGBoost. It calculates the average gain (Gini purity improvement) of each feature used for splitting across all trees, and then normalizes it to the range of 0-1 to obtain the Gain value for each feature. This value is used to characterize the contribution of each feature to the optimization of the model's prediction accuracy.
[0065] XGB_Weight is also a built-in weight importance in XGBoost. By counting and normalizing the frequency of each feature being used as a split node in the decision tree, the XGB_Weight value of each feature can be obtained, which is used to characterize the frequency of each feature in the model rules.
[0066] The direction is determined by calculating the sign (positive / negative) of the Pearson correlation coefficient between the SHAP value and the original feature value, revealing the spatial correlation direction between each feature and the population distribution at a specific time period.
[0067] In this invention, the four indicators mentioned above are used to construct a multi-dimensional evaluation system based on the impact intensity (SHAP_Importance), model dependence (XGB_Gain / Weight), and direction of action. For example, the high SHAP_Importance (1.00) and positive direction of commercial and residential services (CH) indicate that it significantly promotes population agglomeration, while the high XGB_Gain (1.00) verifies that this feature is a key variable for optimizing prediction accuracy, providing a quantitative basis for analyzing the coupling relationship between urban function and population. The roles of various socioeconomic characteristics in the population distribution simulation model at different time periods can be seen in Tables 5.1 and 5.2. Table 5.1 The Role of Various Socioeconomic Characteristics in Population Forecasting Models at Different Time Periods
[0068] Table 5.2 The Role of Various Socioeconomic Characteristics in Population Forecasting Models at Different Time Periods
[0069] Figure 4 This is a schematic diagram illustrating the explanatory results of the population distribution simulation model provided by the present invention. Figure 5 This diagram illustrates the SHAP Importance values and Direction attributes of various features provided by this invention in population distribution simulation models at different time periods. It should be noted that the Pearson correlation is based on the 2024 truth test (see reference). Figure 4 As shown, the Direction index output by the population distribution simulation model shows positive and negative correlations at different time periods (see reference). Figure 5 As shown in the diagram, this is mainly because the SHAP value reflects the marginal contribution of a feature to the model's prediction, while Pearson correlation is a measure of linear relationships. The model may capture non-linear relationships, causing the direction of SHAP to differ from that of linear correlation. For example, a feature may have a non-linear relationship with the target variable (such as a quadratic relationship), in which case the sign of the linear correlation may differ from the direction of the marginal contribution of SHAP. Furthermore, the SHAP value considers the interaction effects between features, while the individual Pearson correlation only considers the linear relationship between a single feature and the target variable. For example, a feature may have a positive impact in the model through interactions with other features, but when viewed alone, its correlation with the target variable may be negative or the opposite.
[0070] Based on the above embodiments, the acquisition of basic population data for a sample area within a preset period includes: The sample region range is determined based on the preset spatial boundary information; The sample area is divided into multiple sub-regions, and the population distribution corresponding to each time scale is obtained based on the total number of people in each sub-region at different time scales within the preset period. Based on the population distribution corresponding to all the sub-regions, the base population data within the sample area during the preset period is obtained.
[0071] In this invention, the sample area is determined based on preset spatial boundary information, which ensures the relevance and accuracy of the research. This allows the subsequent analysis results to truly reflect the population distribution of a city, providing a reliable basis for practical applications such as urban planning and resource allocation.
[0072] Then, using Geographic Information System (GIS) software, such as ArcGIS, the sample area is divided into sub-regions according to preset sizes. Taking a grid as an example, this invention uses the spatial analysis function of the software to cut the map of a city's administrative region into multiple grid sub-regions according to a specified grid size. Each grid sub-region has a unique identifier and coordinate information. For example, a 200m × 200m grid size is used. However, depending on actual needs, the city's administrative region can be divided into grid sub-regions of other sizes, such as 100m × 100m, 500m × 500m, etc.
[0073] In this invention, the preset period is four seasons, with 7 days of data collection in each season, totaling 672 hours. During this period, population distribution information was collected from 203,927 record points. By calculating the average number of active people at each point over 24 hours, the fluctuation trend of the population in a city was captured, and the local day-night division standard was determined. Then, ArcGIS Pro's proximity aggregation tool was used to count the total number of people in each grid sub-region at different times within the preset period. For the Workday / Weekend period, the statistical unit is "total number of people per day," that is, calculating the total number of people in each grid sub-region on weekdays or weekends; for the Daytime / Nighttime period, the statistical unit is "total number of people per day / night," that is, calculating the total number of people in each grid sub-region during daytime or nighttime. In this way, the population distribution corresponding to each time scale is obtained, and this data can reflect the distribution of the population in each grid sub-region at different time scales.
[0074] Furthermore, population distribution data for all grid sub-regions at different time scales are integrated. For example, population distribution data for each grid sub-region during different time periods, such as weekday daytime, weekday nighttime, weekend daytime, and weekend nighttime, are aggregated to form a complete dataset. This integrated dataset constitutes the baseline population data for the sample area within a predetermined period. The baseline population data contains population distribution information for each grid sub-region within a city's administrative area at different time scales, providing a comprehensive and detailed description of the city's population distribution. This data can serve as the basis for subsequent analyses, such as studying the relationship between population distribution and green space distribution, and assessing the actual exposure level of the population within the region. Simultaneously, the baseline population data can also provide important reference data for urban planning, traffic management, public safety, and other fields.
[0075] Based on the above embodiments, the step of obtaining the population distribution corresponding to each time scale based on the total number of people in each sub-region within the preset period includes: Based on a preset buffer range, a population distribution statistics buffer is set in each of the sub-regions; Based on the total number of people at different time scales within the preset period for each of the sub-regions and the corresponding population distribution statistical buffers for each sub-region, the population distribution at each time scale is obtained.
[0076] In this invention, determining the predetermined buffer zone requires comprehensive consideration of multiple factors. For example, the specific purpose of the study varies. If the study focuses on the impact of population on surrounding commercial facilities, the buffer zone may be set relatively small to concentrate on the directly affected areas. Conversely, if the study focuses on the impact of population on the urban ecological environment, the buffer zone may be set larger to cover a wider ecological area. Furthermore, factors such as the topography and building distribution of the study area must be considered. For instance, in mountainous or densely built-up areas, the activity range of population may be limited, and the buffer zone may need to be adjusted accordingly.
[0077] In practice, buffer zones can be set using Geographic Information System (GIS) software. For example, using ArcGIS software, buffer zones can be drawn around each grid sub-region based on preset distances (such as 500 meters, 1 kilometer, etc.). These buffer zones can be circular, rectangular, or other irregular shapes, depending on the research needs and actual conditions. By setting population distribution statistical buffer zones, the statistical scope can be expanded, capturing population movement around the grid sub-regions and including population activities within a certain range around the grid sub-regions in the statistics, thus more accurately reflecting the distribution of the population.
[0078] Within a predetermined period, population distribution data for each grid sub-region and its corresponding population distribution buffer zone is collected at different time scales. This data can come from various sources, such as mobile phone signaling data, smart camera data, and questionnaire survey data. For example, mobile phone signaling data can be used to obtain users' location information at different time periods, thereby calculating the total number of active people in each grid sub-region and buffer zone.
[0079] For each grid sub-region and its corresponding population distribution statistical buffer, the total number of active people is counted within each time period. For example, during weekday daytime hours, the total number of active people within a certain grid sub-region and its buffer is counted to obtain the total number of people for that time period. When calculating the total number of people, the collected data needs to be preprocessed, such as removing outliers and imputing missing values, to ensure the integrity and accuracy of the data.
[0080] Finally, the total number of people in each grid sub-region and its corresponding population distribution statistical buffer zone at different time scales is integrated to form the population distribution for each time scale. This data can be presented in the form of tables, charts, etc., to intuitively reflect the population distribution of each region in different time periods. For example, a heat map can be created, using different colors to represent the total number of people in different regions at different time scales, thus clearly showing the spatiotemporal changes in population distribution.
[0081] In one embodiment, the multi-timescale historical population distribution simulation method based on interpretable machine learning and socioeconomic characteristics provided by the present invention will be described in general, and the specific process is as follows: Data acquisition steps: When conducting population distribution simulation studies, the first step is to determine a clear spatial boundary. This boundary can be an administrative boundary such as a district / county, city, province, or even country. Taking a city as an example, if the research objective is to study the population distribution patterns of a specific city, then the city's administrative boundaries become the spatial boundaries of the study. Clearly defining the spatial boundaries helps to focus the research area, making subsequent data collection and analysis more targeted and feasible, and avoiding inaccurate data or unrepresentative analytical results due to an overly large or small research scope.
[0082] Then, high spatiotemporal resolution population data (baseline population data) for specific periods are collected. Specifically, high spatiotemporal resolution means that the data is accurate to small time intervals (such as hourly intervals) and spatial units (such as grids of several hundred meters square). The collected baseline population data will provide basic information for subsequent population distribution simulations.
[0083] Simultaneously, it is necessary to collect POI data within the same spatial boundary for the corresponding year. POI data covers various locations related to people's lives and activities, such as shopping malls, schools, hospitals, parks, and factories. This data can be obtained from the open API interface of map service providers or from third-party data providers. POI data reflects the socio-economic activities of a region and is closely related to population distribution. For example, areas with a high density of POIs, such as shopping malls and restaurants, often have a relatively concentrated population. Obtaining base POI data helps analyze the relationship between population distribution and various socio-economic activities, providing important independent variables for population distribution simulation.
[0084] Data processing steps: First, the size of the statistical pixels is determined by referring to the spatiotemporal resolution of the baseline population data. A statistical pixel is a small spatial unit into which a region is divided, and its size should be determined based on the data accuracy and simulation requirements. For example, if the spatial resolution of the baseline population data is 200 meters × 200 meters, then the size of the statistical pixels can also be set to 200 meters × 200 meters. Simultaneously, different time scales are defined based on people's daily activity patterns and significant changes in population distribution, such as daytime (T1) and nighttime (T2). The specific time division can be determined based on the activity characteristics of the local population. For example, by analyzing the average number of activity times at each point in the baseline population data over 24 hours to capture population fluctuation trends, it is found that the local population dynamically increases from 5:00 AM to 5:00 PM and continuously decreases from 6:00 PM to 4:00 AM the next day. Therefore, 5:00 AM to 5:00 PM can be defined as daytime (T1), and 6:00 PM to 4:00 AM the next day as nighttime (T2).
[0085] After determining the statistical pixels and time scales, the number of people in each base population data pixel is counted at different time scales. For the daytime period (T1), the average number of people in each pixel during the day (Y1) is counted; for the nighttime period (T2), the average number of people in each pixel during the nighttime (Y2), and so on. During the statistical process, data preprocessing is required, such as removing outliers and imputing missing values. For example, if data for a certain pixel at a certain time scale is missing due to equipment failure or other reasons, interpolation or other appropriate methods can be used to impute the missing data to ensure its completeness and accuracy.
[0086] In addition to counting the number of people, it is also necessary to count the number of various POIs within each base population data cell. This invention categorizes POIs according to their type, such as business, education, and healthcare, and then counts the number of each type of POI within each cell at the corresponding time scale. These POI counts will be used as independent variables. Xa , Xb , Xc Together with the number of people in the corresponding pixel (dependent variables Y1 and Y2), they are used for subsequent model construction.
[0087] Model building steps: First, a population distribution simulation model M1 for the daytime period is constructed. This is based on the baseline population data (dependent variable Y1) and baseline POI data (independent variable for the daytime period) for each pixel during the daytime (T1) period. Xa , Xb , XcBased on the statistical results of (etc.), a daytime population distribution simulation model M1 was constructed. When constructing the population distribution simulation model M1, the dataset was divided into a training set and a test set. The model was trained using the training set and its performance was evaluated using the test set. By continuously adjusting the model's parameters, the prediction accuracy and generalization ability of the model were improved.
[0088] Similarly, based on the nighttime (T2) period, the baseline population data (dependent variable Y2) and baseline POI data (independent variable for the nighttime period) of each pixel are used. Xa , Xb , Xc Based on the statistical results of (etc.), a nighttime population distribution simulation model M2 is constructed. The nighttime population distribution differs from the daytime distribution; for example, the nighttime population decreases in commercial areas while increasing in residential areas. Therefore, a separate nighttime model is needed to more accurately simulate the nighttime population distribution.
[0089] When counting Points of Interest (POIs), a buffer distance can be appropriately increased. This is because the influence of POI data may extend beyond its location, affecting a surrounding area. For example, a large shopping mall might attract residents from a radius of several hundred meters or even kilometers. By increasing the buffer distance and including POIs within this surrounding area in the statistics, the impact of POIs on population distribution can be more comprehensively reflected, thereby improving model performance.
[0090] Steps for simulating historical population distribution: After obtaining the multi-period population distribution simulation model through the above steps, the historical years to be simulated (t1, t2, t3, etc.) are determined according to the needs. These historical years can be the past few years or even decades. By simulating the population distribution of historical years, we can understand the evolution trend of population distribution and provide historical reference for urban planning, resource allocation, etc.
[0091] First, obtain the POI dataset for the corresponding historical years within the spatial boundary. Then, count the number of each type of POI data within each base population data cell for each historical year (t1, t2, t3, etc.) (corresponding to different historical years, namely Xa', Xb', Xc', etc.; Xa'', Xb'', Xc'', etc.; Xa''', Xb''', Xc''', etc.). The statistical method is similar to the method used to count the number of POIs in the current year in the above steps, requiring preprocessing and classification of the POI data from historical years.
[0092] Then, the number of POIs of various types, statistically analyzed from each pixel in historical years, is used as the independent variable and substituted into the daytime population distribution simulation model M1. The model will output simulated population distribution data for daytime in each historical year. In this way, the population distribution during the daytime in different historical years can be simulated, and the changing trends of daytime population distribution can be analyzed.
[0093] Similarly, by using the number of POIs of various types from each pixel in historical years as independent variables and substituting them into the population distribution simulation model M2 for nighttime periods, the model will output the population distribution for nighttime periods in historical years. By comparing the population distribution for nighttime periods in different historical years, we can understand the evolution of nighttime population distribution and provide a basis for planning in areas such as urban nighttime economy and public safety.
[0094] This invention, based on population distribution data within a specific time context and corresponding year POI data (used to characterize various socioeconomic activity elements), constructs a nonlinear model between population distribution and socioeconomic characteristics by introducing machine learning algorithms. Finally, by inputting historical POI data into the trained model, historical population distribution simulation data can be obtained. This addresses the current difficulty in obtaining multi-timescale population distribution data, and the limitations of acquiring historical multi-timescale population distribution data due to various factors. This invention utilizes readily available POI data to characterize various socioeconomic activities in cities, captures the relationship between population distribution and various socioeconomic activities, and further inversely predicts high-precision historical multi-period population distribution data based on limited multi-timescale population distribution data and available historical POI data. Furthermore, this invention is not geographically limited and can be applied to different regions.
[0095] The following describes the multi-timescale historical population distribution simulation system provided by the present invention. The multi-timescale historical population distribution simulation system described below can be referred to in correspondence with the multi-timescale historical population distribution simulation method described above.
[0096] Figure 6 A schematic diagram of the structure of the multi-timescale historical population distribution simulation system provided by the present invention is shown below. Figure 6As shown, the present invention provides a historical population distribution simulation system with multiple time scales, including a POI data acquisition module 601 and a population distribution simulation module 602. The POI data acquisition module 601 is used to acquire historical POI data of a target area within a preset historical period. The population distribution simulation module 602 is used to input the historical POI data into population distribution simulation models at different time scales to obtain population distribution simulation data of the target area at the corresponding time scale of the preset historical period, output by the population distribution simulation models. The population distribution simulation models at different time scales are trained using machine learning algorithms based on sample data.
[0097] The multi-timescale historical population distribution simulation system provided by this invention acquires historical POI data reflecting socio-economic activities in a target area within a preset historical period. This historical POI data is then input into a population distribution simulation model trained based on sample data. The model captures the relationship between population distribution and the socio-economic activities represented by POIs, and outputs population distribution simulation data for different time periods within the preset historical period, thus achieving a more refined population distribution simulation within the region.
[0098] Based on the above embodiments, the population distribution simulation module includes a time scale determination unit, a target model determination unit, and a processing unit. The time scale determination unit determines the target time scale type corresponding to the historical POI data. The target model determination unit determines a target population distribution simulation model from multiple population distribution simulation models at different time scales based on the target time scale type. Each population distribution simulation model is trained using the machine learning algorithm based on population distribution sample data at different time scales and corresponding POI sample data. The processing unit inputs the historical POI data into the target population distribution simulation model to obtain the population distribution simulation data output by the target population distribution simulation model, corresponding to the target area range within the preset historical period at the target time scale type.
[0099] The system provided in this embodiment of the invention is used to execute the above-described method embodiments. For specific processes and details, please refer to the above embodiments, which will not be repeated here.
[0100] Figure 7 This is a schematic diagram of the structure of the electronic device provided by the present invention, such as... Figure 7As shown, the electronic device may include: a processor 701, a communications interface 702, a memory 703, and a communication bus 704, wherein the processor 701, the communications interface 702, and the memory 703 communicate with each other via the communication bus 704. The processor 701 can call logical instructions in the memory 703 to execute a multi-timescale historical population distribution simulation method. This method includes: acquiring historical POI data of a target area within a preset historical period; inputting the historical POI data into population distribution simulation models at different time scales to obtain population distribution simulation data of the target area at the corresponding time scale of the preset historical period, output by the population distribution simulation models. The population distribution simulation models at different time scales are trained using machine learning algorithms based on sample data.
[0101] Furthermore, the logical instructions in the aforementioned memory 703 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0102] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program stored on a non-transitory computer-readable storage medium, the computer program comprising program instructions, wherein when the program instructions are executed by a computer, the computer is able to execute the multi-time-scale historical population distribution simulation method provided by the above methods, the method comprising: acquiring historical POI data of a target area in a preset historical period; inputting the historical POI data into population distribution simulation models at different time scales to obtain population distribution simulation data of the target area at the corresponding time scale of the preset historical period output by the population distribution simulation models, wherein the population distribution simulation models at different time scales are trained based on sample data using machine learning algorithm models.
[0103] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the multi-timescale historical population distribution simulation method provided in the above embodiments. The method includes: acquiring historical POI data of a target area in a preset historical period; inputting the historical POI data into population distribution simulation models at different time scales to obtain population distribution simulation data of the target area at the corresponding time scale of the preset historical period, output by the population distribution simulation models, wherein the population distribution simulation models at different time scales are trained using machine learning algorithm models based on sample data.
[0104] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0105] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0106] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for simulating historical population distribution across multiple time scales, characterized in that, include: Obtain historical POI data for the target area within a preset historical period; The historical data of the POI is input into population distribution simulation models at different time scales to obtain population distribution simulation data of the target area at the corresponding time scale of the preset historical period, output by the population distribution simulation model. The population distribution simulation models at different time scales are trained using machine learning algorithms based on sample data.
2. The historical population distribution simulation method with multiple time scales according to claim 1, characterized in that, The step of inputting the historical POI data into population distribution simulation models at different time scales to obtain population distribution simulation data of the target area at the corresponding time scale of the preset historical period, output by the population distribution simulation model, includes: Determine the target time scale type corresponding to the historical data of the POI; According to the target time scale type, a target population distribution simulation model is determined from multiple population distribution simulation models of different time scales, wherein each population distribution simulation model is trained using the machine learning algorithm based on population distribution sample data of different time scale types and POI sample data corresponding to the population distribution sample data; The historical data of the POI is input into the target population distribution simulation model to obtain the population distribution simulation data of the target area under the target time scale type within the preset historical period, output by the target population distribution simulation model.
3. The multi-timescale historical population distribution simulation method according to claim 1 or 2, characterized in that, The population distribution simulation model is trained through the following steps: Based on preset spatial boundary information, obtain basic population data of the sample area within a preset period; Based on the spatiotemporal resolution corresponding to the base population data, the base population data is divided into population distribution sample data of different time scale types, wherein the time scale types include at least daytime time scale, nighttime time scale, weekday time scale, and weekend time scale under non-holiday conditions; The population distribution sample data is obtained by acquiring POI sample data within the spatial boundary of the preset period, and the type and number of POIs of each type of the acquired POI sample data are used as the label data of the population distribution sample data to obtain the sample data. Based on the sample data of different time scale types, a training dataset is constructed for the population distribution simulation model corresponding to each time scale type; Based on the training datasets of the population distribution simulation models corresponding to each of the aforementioned time scale types, the XGBoost machine learning algorithm is used for training, and the population distribution simulation models for different time scale types are obtained based on the optimal parameter tuning technique.
4. The multi-timescale historical population distribution simulation method according to claim 3, characterized in that, The acquisition of basic population data for a sample area within a preset period based on preset spatial boundary information includes: Based on the preset spatial boundary information, the sample area range is determined; the sample area range is divided into sub-regions, and based on the total number of people in each sub-region at different time scales within the preset period, the population distribution corresponding to each time scale is obtained; Based on the population distribution corresponding to all the sub-regions, the base population data within the sample area during the preset period is obtained.
5. The historical population distribution simulation method with multiple time scales according to claim 4, characterized in that, The step of obtaining the population distribution corresponding to each time scale based on the total number of people in each of the sub-regions within the preset period includes: Based on a preset buffer range, a population distribution statistics buffer is set in each of the sub-regions; Based on the total number of people at different time scales within the preset period for each of the sub-regions and the corresponding population distribution statistical buffers for each sub-region, the population distribution at each time scale is obtained.
6. A multi-timescale historical population distribution simulation system, characterized in that, include: The POI data acquisition module is used to acquire historical POI data for the target area within a preset historical period. The population distribution simulation module is used to input the historical data of the POI into population distribution simulation models at different time scales to obtain population distribution simulation data of the target area at the time scale corresponding to the preset historical period, output by the population distribution simulation model. The population distribution simulation models at different time scales are trained using machine learning algorithms based on sample data.
7. The multi-timescale historical population distribution simulation system according to claim 6, characterized in that, The population distribution simulation module includes: The time scale determination unit is used to determine the target time scale type corresponding to the POI historical data; The target model determination unit is used to determine a target population distribution simulation model from multiple population distribution simulation models of different time scales according to the target time scale type, wherein each population distribution simulation model is trained using the machine learning algorithm based on population distribution sample data of different time scale types and POI sample data corresponding to the population distribution sample data; The processing unit is used to input the historical data of the POI into the target population distribution simulation model to obtain the population distribution simulation data of the target area range under the target time scale type within the preset historical period, output by the target population distribution simulation model.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the multi-timescale historical population distribution simulation method as described in any one of claims 1 to 5.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the multi-timescale historical population distribution simulation method as described in any one of claims 1 to 5.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the multi-timescale historical population distribution simulation method as described in any one of claims 1 to 5.