A Crowd Flow Prediction Method Based on Geographic Similarity and Federated Learning

Through the population flow prediction method based on geographical similarity and federated learning, the geospatial data of the target city and federated learning of similar source cities are solved, and the problems of personal information leakage and data silos are achieved.

CN118735064BActive Publication Date: 2025-07-25SHENZHEN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410859563.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-28
Publication Date
2025-07-25
Estimated Expiration
2044-06-28

AI Technical Summary

Technical Problem

Existing methods for predicting population flow distribution data are prone to leakage of personal information and cannot obtain complete population flow data.

Method used

The population flow prediction method based on geographical similarity and federated learning is adopted to obtain the geospatial data of the target city for preprocessing, use the trained neural network model to make predictions, and personalized model training is carried out on the terminals of similar source cities to avoid sharing personal travel data.

Benefits of technology

It realizes accurate prediction of population flow distribution data without revealing personal privacy, solves the data island problem, and improves prediction accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118735064B_ABST
    Figure CN118735064B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of crowd flow prediction, and specifically relates to a crowd flow prediction method based on geographical similarity and federated learning. First, the present invention performs federated learning under a deep learning model on multiple source city terminals by using geospatial data and crowd flow data. The source city terminals upload the model parameters to the central server to update the global neural network model. Then, a geographical similarity measurement is performed on the geospatial data of all source cities and target cities to determine the set of source cities with the highest similarity. Finally, personalized federated learning is performed on these sets of source cities with high geographical similarity, and the final model is applied to the target cities lacking crowd flow data to complete their crowd flow prediction. The present invention not only solves the prediction difficulty caused by the lack of data in the target cities and prevents the leakage of personal privacy data, but also improves the accuracy of crowd flow prediction in personalized federated learning through geographical similarity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of crowd flow prediction, and specifically relates to a crowd flow prediction method based on geographical similarity and federated learning. Background Art

[0002] With the popularization of mobile devices and the rapid development of the Internet, human mobility has become more complex and diverse. Accurately predicting the location information and flow information of crowd movement is helpful for urban planning, traffic management, public health and other fields. Due to incomplete data collection, limited device coverage or defects in the data storage system, many cities cannot obtain complete crowd movement data. Existing crowd flow prediction methods mainly rely on statistical models, using historical travel data to predict future travel patterns, that is, predicting future crowd flow data. Since a large amount of personal travel data needs to be collected and shared, it leads to the leakage of personal privacy information.

[0003] In summary, the method for predicting crowd flow distribution data in the prior art is prone to cause the leakage of personal information.

[0004] Therefore, the prior art still needs to be improved and enhanced. Summary of the Invention

[0005] To solve the above technical problems, the present invention provides a crowd flow prediction method based on geographical similarity and federated learning, which solves the problem that the method for predicting crowd flow distribution data in the prior art is prone to cause the leakage of personal information.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] In a first aspect, the present invention provides a crowd flow prediction method based on geographical similarity and federated learning, which includes:

[0008] Obtain the target original geographical space data of the target city, and preprocess the target original geographical space data to obtain standardized geographical space data;

[0009] Apply the trained neural network model to the standardized geographical space data to obtain the crowd flow distribution prediction data of the target city, and the crowd flow distribution prediction data is used to represent the flow data of the crowd on each grid space in the target city.

[0010] In one implementation, the preprocessing the target original geographical space data to obtain standardized geographical space data includes:

[0011] Divide the target original geographical space data according to time to obtain local data information within each time step;

[0012] The local data information within each time step is partitioned into grids according to the geographical space, obtaining the local geographical space data information within each time step, and the local geographical space data information corresponds to the grid space;

[0013] The local geographical space data information within each time step is normalized to obtain the standardized local space data;

[0014] Based on the standardized local space data within all the time steps, the standardized geographical space data is obtained.

[0015] In one implementation, applying the trained neural network model to the standardized geographical space data to obtain the predicted data of the population flow distribution in the target city includes:

[0016] Applying the trained neural network model to the standardized geographical space data to obtain the predicted data of the population flow spatio-temporal distribution on each grid space within each time step output by the neural network model, and using the predicted data of the population flow spatio-temporal distribution as the predicted data of the population flow distribution in the target city.

[0017] In one implementation, the training method of the trained neural network model includes:

[0018] Obtain the local model parameters of each pre-trained neural network model, where each pre-trained neural network model is obtained by each source city terminal training a neural network based on a training dataset, and the training dataset includes the geographical space data of the source city and the actual population flow distribution data corresponding to the geographical space data;

[0019] Based on each of the local model parameters, obtain the global final parameters of the model;

[0020] Determine the similarity between each source city and the target city, and based on each similarity, screen out the similar source city terminals from each source city terminal;

[0021] Send the global final parameters of the model to the similar source city terminals, and obtain the personalized model parameters of the personalized neural network model sent by the similar source city terminals. The personalized neural network model is a model obtained by the similar source city terminals continuing to train the neural network model based on the global final parameters of the model;

[0022] Based on the personalized model parameters, obtain the trained neural network model.

[0023] In one implementation, determining the similarity between each of the source cities and the target city, and screening out similar source city terminals from each of the source city terminals according to each of the similarities includes:

[0024] Successively dividing the geographical space data of each of the source cities by time and into grids according to the geographical space to obtain each standardized source geographical space data;

[0025] Calculating the cosine similarities between each of the standardized source geographical space data and the standardized geographical space data, and using each of the cosine similarities as the similarity between each of the source cities and the target city;

[0026] According to each of the similarities, screening out several similar source cities from each of the source cities, so as to screen out similar source city terminals from each of the source city terminals. The similarity of each of the similar source cities is greater than that of the remaining source cities, and the remaining source cities are the cities other than the similar source cities among each of the source cities.

[0027] In one implementation, obtaining the global final model parameters according to each of the local model parameters includes:

[0028] Calculating the mean of each of the local model parameters to obtain the global final model parameters.

[0029] In one implementation, when the similar source city terminal trains the neural network, the personalized model parameters are updated. The update method is to update according to the loss function and the global feature vector of the similar source city. The loss function is the loss function composed of the actual population flow distribution data and the predicted population flow distribution data of the similar cities. The predicted population flow distribution data is the prediction data output by the neural network based on the geographical space data of the similar cities, and the global feature vector is the vector form of the geographical space data of the similar cities.

[0030] In a second aspect, an embodiment of the present invention further provides a crowd flow prediction device based on geographical similarity and federated learning. The device includes the following components:

[0031] A data preprocessing module, configured to obtain the target original geographical space data of the target city, and preprocess the target original geographical space data to obtain standardized geographical space data;

[0032] A prediction module, configured to apply the trained neural network model to the standardized geographical space data to obtain the crowd flow distribution prediction data of the target city, where the crowd flow distribution prediction data is used to characterize the flow data of people on each grid space in the target city.

[0033] In a third aspect, an embodiment of the present invention further provides a terminal device, where the terminal device includes a memory, a processor, and a population flow prediction program based on geographical similarity and federated learning stored in the memory and executable on the processor. When the processor executes the population flow prediction program based on geographical similarity and federated learning, the steps of the above-mentioned population flow prediction method based on geographical similarity and federated learning are implemented.

[0034] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a population flow prediction program based on geographical similarity and federated learning is stored. When the population flow prediction program based on geographical similarity and federated learning is executed by a processor, the steps of the above-mentioned population flow prediction method based on geographical similarity and federated learning are implemented.

[0035] Beneficial effects: The present invention first obtains the target original geographical space data of the target city, then preprocesses the target original geographical space data to obtain the standardized geographical space data. Finally, the trained neural network model is applied to the standardized geographical space data to obtain the predicted data of the population flow distribution in the target city. From the above analysis, it can be seen that since the geographical space data of the target city is closely related to the population flow, the population flow distribution data can be predicted using the geographical space data, and the geographical space data does not involve personal information of the population. Therefore, the present invention can not only predict the population flow distribution data but also does not disclose personal privacy data. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 is the overall flowchart of the present invention;

[0037] Figure 2 is the flowchart of the population distribution prediction in the target city in the embodiment of the present invention;

[0038] Figure 3 is the structural diagram of the population flow prediction device based on geographical similarity and federated learning provided by the present invention;

[0039] Figure 4 is the internal structure principle block diagram of the terminal device provided in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] The following combines the embodiments and the accompanying drawings of the specification to clearly and completely describe the technical solutions in the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.

[0041] It has been found through research that with the popularization of mobile devices and the rapid development of the Internet, human mobility has become more complex and diverse. Accurately predicting the location information and flow information of crowd mobility is helpful for fields such as urban planning, traffic management, and public health. Due to incomplete data collection, limited device coverage, or defects in data storage systems, many cities are unable to obtain complete crowd mobility data. Existing crowd mobility prediction methods mainly rely on statistical models and use historical travel data to predict future travel patterns, that is, to predict future crowd mobility data. Since a large amount of personal travel data needs to be collected and shared, it leads to the leakage of personal privacy information.

[0042] To solve the above technical problems, the present invention provides a crowd mobility prediction method based on geographical similarity and federated learning, which solves the problem that the existing crowd mobility distribution data prediction method is prone to personal information leakage. Specifically in implementation, first obtain the target original geographical space data of the target city, and preprocess the target original geographical space data to obtain standardized geographical space data; finally, apply the trained neural network model to the standardized geographical space data to obtain the crowd mobility distribution prediction data of the target city, and the crowd mobility distribution prediction data is used to represent the flow data of the crowd on each grid space within the target city.

[0043] For example, the target city includes Area A, Area B, and Area C (these three areas are three grid spaces). Obtain the geospatial data A1 of Area A for the last month and the geospatial data A2 of Area A for the month before last. Obtain the geospatial data B1 of Area B for the last month and the geospatial data B2 of Area B for the month before last. Obtain the geospatial data C1 of Area C for the last month and the geospatial data C2 of Area C for the month before last. Here, the last month and the month before last are time steps. The geospatial data A1, geospatial data A2, geospatial data B1, geospatial data B2, geospatial data C1, and geospatial data C2 together are the target original geospatial data. Performing standardization processing on the target original geospatial data yields the standardized geospatial data. That is, the standardized geospatial data includes standardized geospatial data a1, standardized geospatial data a2, standardized geospatial data b1, standardized geospatial data b2, standardized geospatial data c1, and standardized geospatial data c2. Among them, a1, a2, b1, b2, c1, and c2 respectively correspond to A1, A2, B1, B2, C1, and C2. Input a1, a2, b1, b2, c1, and c2 into the trained neural network model. The trained neural network model outputs the predicted data of the population flow distribution in the target city. The predicted data of the population flow distribution includes the population quantity of Area A for the next month and the population quantity of Area A for the month after next, the population quantity of Area B for the next month and the population quantity of Area B for the month after next, and the population quantity of Area C for the next month and the population quantity of Area C for the month after next. The population quantity of each area each month is caused by the population flowing out of this area to other areas or the population flowing into this area from other areas. That is, the predicted data of the population flow distribution of the present invention is used to ensure the population quantity of each area in the target city in the next few months.

[0044] The population flow prediction method based on geographical similarity and federated learning in this embodiment can be applied to a terminal device. The terminal device can be a terminal product with data processing functions, such as a computer, etc. In this embodiment, as Figure 1 shown, the population flow prediction method based on geographical similarity and federated learning specifically includes the following steps:

[0045] S100, train a neural network model.

[0046] S200, obtain the target original geospatial data of the target city, and perform preprocessing on the target original geospatial data to obtain standardized geospatial data;

[0047] S300, apply the trained neural network model to the standardized geospatial data to obtain the predicted data of the population flow distribution in the target city. The predicted data of the population flow distribution is used to characterize the flow data of the population on each grid space in the target city.

[0048] In the first embodiment, in step S100, the neural network model is trained by first training the neural network model at the source city terminals in each source city. As Figure 2 shown, each source city terminal uploads the model parameters obtained from its training to the central server. The central server performs aggregation processing on the respective model parameters, and then the central server distributes the aggregated model parameters to each source city terminal. The source city terminal continues to train the neural network model until the model converges or reaches a preset number of iterations, at which point the federated model joint training is completed. Subsequently, based on the similarity between the geospatial data of each source city and the geospatial data of the target city, similar source cities are selected from each source city, and the neural network model is continuously trained on the terminals of the similar source cities.

[0049] In this embodiment, step S100 includes the following specific steps:

[0050] S101, obtaining the local model parameters of each pre-trained neural network model, where each of the pre-trained neural network models is obtained by a source city terminal training a neural network based on a training dataset, and the training dataset includes the geospatial data of the source city and the actual population flow distribution data corresponding to the geospatial data.

[0051] The geospatial data in the training dataset includes multiple types of geospatial data such as population density, GDP (economic production capacity data), transportation road network, transportation hubs, elevation, POI (infrastructure data), etc. For each type of geospatial data of each source city, standardization processing is performed to obtain

[0052]

[0053] where is the standardized source geospatial data of the i-th source city, is the local data information within the j-th time step of the i-th source city, and N is the total number of time steps. Each includes population density, GDP (economic production capacity data), transportation road network, transportation hubs, elevation, POI (infrastructure data), etc.

[0054]

[0055] where F(x, y, j) is the local geospatial data information within the grid space (the coordinates of the grid space are (x, y)) at the j-th time step, W is the maximum abscissa, and H is the maximum ordinate.

[0056] The total number of types of geospatial data includes M classes. Each class of data is processed by formulas (1) and (2). The standardized source geospatial data corresponding to the M classes of data in the i-th source city are Using the matrix to represent the multi-class standardized geospatial data of the i-th source city, then:

[0057]

[0058] The actual population flow distribution data of the i-th source city is represented by

[0059]

[0060]

[0061] f(x, y, j) is the population quantity in the grid space (the coordinates of the grid space are (x, y)) at the j-th time step.

[0062] On the central server, a convolutional neural network (CNN) is selected to initialize a global model as the training starting point. The initial parameters of this global model are W0. Subsequently, the initialized global model will be downloaded to each source city terminal, and each source city terminal will perform local model training on this basis. The specific training process is as follows:

[0063] The terminal of the i-th source city divides D i g M into a training set and a test set. The training set is input into the neural network model. Multiple convolutional kernels of the neural network model perform feature extraction on D i g M to generate multiple feature maps and capture local patterns in the data. The feature maps are processed through a series of non-linear transformations and activation functions to increase the non-linear expression ability of the model. The data in the test set is used to evaluate the performance of the model. The population flow data is used as the model truth value, and the prediction result is compared with the truth value to perform model parameter tuning and loss calculation.

[0064] During the process of training the model, each source city terminal updates the model parameters, and the update rule is as shown in formula (6).

[0065]

[0066] The model parameters obtained at the (r + 1)-th update in the t-th round for the terminal in the source city i. (One round includes several trainings of the model by the terminal in the source city. After each round of training is completed by the terminal in the source city, the model parameters obtained from the training are uploaded to the central server, and the central server aggregates the model parameters of each terminal in the source city. After aggregation, they are sent back to each terminal in the source city for the next round of model training.) The model parameters obtained at the r-th update in the t-th round for the terminal in the i-th source city Use batch gradient descent to Calculate the gradient; K represents the total number of terminals in the source city Is the loss function And F(x, y, n) represent the predicted value (i.e., the predicted population flow distribution data) and the actual value (i.e., the actual population flow distribution data) of the population flow in the source city respectively. η is the learning rate, which is used to control the size of the update step

[0067] S102: Obtain the global final parameters of the model based on each of the local model parameters

[0068] After each round of training is completed, all terminals in the source city upload the model parameters obtained from the training to the central server for aggregation to obtain

[0069]

[0070] Represents the global model parameters after the (t + 1)-th update Is the model parameter (i.e., the local model parameter) after the last local update (i.e., the R-th update in the t-th round) in the t-th round for the source city

[0071] The central server then sends To each terminal in the source city, and repeats the adjustment step until the model converges or reaches the preset number of iteration rounds. That is, when it converges or reaches the preset number of iteration rounds, the global final parameters of the model are obtained

[0072] S103: Divide the geospatial data of each of the source cities by time and then by geospatial grid to obtain each standardized source geospatial data

[0073] That is, use formula (1) to obtain the standardized source geospatial data of each source city

[0074] S104: Calculate the cosine similarity d between each of the standardized source geospatial data and the standardized geospatial data iT ​, and use each of the cosine similarities as the similarity between each of the source cities and the target city.

[0075] Replace the in formula (1) with the geospatial data of the target city to obtain

[0076]

[0077] d iT is the cosine similarity, i.e., the similarity, between the i-th source city and the target city T. For d iT Based on the similarities calculated from one type of geospatial data, applying formula (8) to all of the above M types of geospatial data gives M d's iT , for the M d's iT calculate the mean value to obtain corresponding to all source cities to form the matrix S T :

[0078]

[0079] S105. According to each of the similarities, select several similar source cities from each of the source cities to screen out similar source city terminals from each of the source city terminals. The similarity of each of the similar source cities is greater than the similarity of the remaining source cities, and the remaining source cities are the cities other than the similar source cities among all the source cities.

[0080] Arrange all T in S in descending order to obtain the ranking of all source city terminals, and select the first c source city terminals. The S' corresponding to the first c source city terminals T :

[0081] is the mean value of the cosine similarities corresponding to the first source city terminal, is the mean value of the cosine similarities corresponding to the second source city terminal. S' T is weighted to obtain the global feature vector

[0082]

[0083] In the formula, d i′T is the i'-th mean value of the cosine similarities in S' T .

[0084] S106. Send the global final parameters of the model to the similar source city terminals, and obtain the personalized model parameters of the personalized neural network model sent by the similar source city terminals. The personalized neural network model is the model obtained by the similar source city terminals continuing to train the neural network model based on the global final parameters of the model.

[0085] The central processing unit sends and the global final parameters of the model to the first c source city terminals respectively, and sends and for each of the first c source city terminals as the training data set. Each of the first c source city terminals continues to train the neural network model on the basis that the neural network model has the global final parameters of the model, so that each of the first c source city terminals trains a set of personalized model parameters.

[0086] When the first c source city terminals train the model, the update rule of the model parameters is as follows:

[0087]

[0088] is the model parameter when the (r + 1)-th update of the i-th source city terminal among the first c source city terminals at the t-th round, represents the model parameter when the r-th update of the i-th source city terminal among the first c source city terminals at the t-th round, is to take the gradient, is the regularization term, used to limit the size of the model parameters, and λ is the regularization parameter, used to limit the size of the model parameters to prevent overfitting.

[0089] Send the globally model parameters after geographical similarity measurement and the model parameters of the first c source cities to the central server for aggregation of the globally model parameters. By using different geographical similarity weights S T assign different weights to the updates from different source cities, and then aggregate the updates after weight calculation with the globally model parameters. Update the weight parameters of the globally model to as shown in formula (12). Send the updated weight parameters to each client, and repeat the adjustment steps until the model converges or reaches the preset number of iterations. When the model converges or reaches the preset number of iterations, the personalized model parameters of each source city terminal are obtained.

[0090]

[0091] Among them, Denote the global model parameters after the (t + 1)-th round of personalized model training update; is the model parameter after the last (i.e., the R-th) local update of the source city i in the t-th round.

[0092] S107. Obtain a trained neural network model according to the personalized model parameters.

[0093] The training of the model is completed by weighting all the personalized model parameters, and thus a trained neural network model is obtained.

[0094] Example 2. Based on Example 1, each time the model is trained, the model is also evaluated. In this example, the trained model is evaluated using three evaluation metrics: mean absolute error (MAE), root mean square error (RMSE), and correlation coefficient (CC).

[0095] The mean absolute error (MAE) is the average of all the absolute errors between the predicted value and its corresponding actual value:

[0096]

[0097] In the formula, is the predicted value of the population flow in the source city, and F(x, y, n) is the actual value of the population flow in the source city.

[0098] The root mean square error (RMSE) measures the deviation between the predicted value and its respective actual value:

[0099]

[0100] The correlation coefficient (CC) is used to verify the correlation between variables:

[0101]

[0102] Example 3. Based on Example 1, in this example, step S200 includes the following specific steps S201 to S204:

[0103] S201. Divide the target original geospatial data by time to obtain local data information within each time step

[0104] is the local data information within the k-th time step, where k ∈ N and N is the total number of time steps.

[0105] S202. Grid-divide the local data information within each time step by geospatial space to obtain local geospatial data information within each time step The local geospatial data information corresponds to the grid space.

[0106] For the local geospatial data information with grid space coordinates (x, y) in the k-th time step, x ∈ [1, W], y ∈ [1, H], where W is the maximum abscissa and H is the maximum ordinate.

[0107] S203. Normalize the local geospatial data information in each time step to obtain the standardized local space data.

[0108] S204. Obtain the standardized geospatial data based on the standardized local space data in all time steps.

[0109]

[0110]

[0111] In this embodiment, the specific process of step S300 is as follows: Apply the trained neural network model to the standardized geospatial data to obtain the predicted data of the spatio-temporal distribution of population flow on each grid space in each time step output by the neural network model, and use the predicted data of the spatio-temporal distribution of population flow as the predicted data of the population flow distribution of the target city.

[0112] That is, input the standardized geospatial data into the neural network model trained in Embodiment 1

[0113] The neural network model outputs the predicted data of the spatio-temporal distribution of population flow.

[0114]

[0115]

[0116] For the predicted data of the spatio-temporal distribution of population flow on the grid space with position (x, y) in the k-th time step.

[0117] In summary, the present invention adjusts the model by considering geographical factors to adapt to the characteristics of different regions, makes full use of the knowledge of the source city, adapts to the specific characteristics of the target city, and specifically improves the model prediction ability. By performing federated learning in multiple source cities and uploading the model parameters of each city to the central server for aggregation, the problem of privacy protection is solved, the sharing of a large amount of individual travel data is avoided, and the problem of data silos is effectively solved. The present invention can effectively predict population mobility even in an environment lacking complete data by using an advanced model training method. This method does not rely on data filling or data interpolation, but effectively predicts by directly analyzing and using other existing geographically relevant data.

[0118] This embodiment also provides a population mobility prediction device based on geographical similarity and federated learning, as Figure 3 shown. The device includes:

[0119] A training module 01 for training a neural network model.

[0120] A data preprocessing module 02 for obtaining the target original geospatial data of the target city and preprocessing the target original geospatial data to obtain standardized geospatial data;

[0121] A prediction module 03 for applying the trained neural network model to the standardized geospatial data to obtain the population mobility distribution prediction data of the target city, where the population mobility distribution prediction data is used to characterize the mobility data of the population on each grid space in the target city.

[0122] Based on the above embodiment, the present invention also provides a terminal device, and its principle block diagram can be as Figure 4 shown. The terminal device includes a processor, a memory, a network interface, and a display screen connected by a system bus. Among them, the processor of the terminal device is used to provide computing and control capabilities. The memory of the terminal device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the terminal device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a population mobility prediction method based on geographical similarity and federated learning. The display screen of the terminal device can be a liquid crystal display screen or an electronic ink display screen.

[0123] Those skilled in the art can understand, Figure 4The principle block diagram shown only shows the block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the terminal device to which the solution of the present invention is applied. The specific terminal device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0124] In one embodiment, a terminal device is provided. The terminal device includes a memory, a processor, and a crowd flow prediction program based on geographical similarity and federated learning stored in the memory and executable on the processor. When the processor executes the crowd flow prediction program based on geographical similarity and federated learning, the following operation instructions are implemented:

[0125] Obtain the target original geospatial data of the target city, and preprocess the target original geospatial data to obtain standardized geospatial data;

[0126] Apply the trained neural network model to the standardized geospatial data to obtain the crowd flow distribution prediction data of the target city, and the crowd flow distribution prediction data is used to characterize the flow data of the crowd on each grid space in the target city.

[0127] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to the memory, storage, database or other media used in the embodiments provided by the present invention can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0128] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for predicting population mobility based on geographical similarity and federated learning, characterized in that, Including: Obtain the target original geospatial data of the target city, and preprocess the target original geospatial data to obtain standardized geospatial data. The geospatial data includes population density, transportation road network, and economic production capacity data; Apply the trained neural network model to the standardized geospatial data to obtain the predicted data of the population flow distribution in the target city. The predicted data of the population flow distribution is used to represent the flow data of the population on each grid space in the target city, and the predicted data of the population flow distribution includes the population quantity on each grid space; The preprocessing of the target original geospatial data to obtain standardized geospatial data includes: Divide the target original geospatial data according to time to obtain the local data information within each time step; Perform grid division on the local data information within each time step according to the geospatial area to obtain the local geospatial data information within each time step. The local geospatial data information corresponds to the grid space; Perform normalization processing on the local geospatial data information within each time step to obtain standardized local space data; Based on the standardized local space data within all the time steps, obtain the standardized geospatial data; The training method of the trained neural network model includes: Obtain the local model parameters of each pre-trained neural network model. Each pre-trained neural network model is obtained by a source city terminal training a neural network based on a training dataset. The training dataset includes the geospatial data of the source city and the actual population flow distribution data corresponding to the geospatial data; Based on the local model parameters of each, obtain the global final parameters of the model; Determine the similarity between each source city and the target city, and based on each similarity, screen out the similar source city terminals from each source city terminal; Send the global final parameters of the model to the similar source city terminals, and obtain the personalized model parameters of the personalized neural network model sent by the similar source city terminals. The personalized neural network model is a model obtained by the similar source city terminals continuing to train the neural network model based on the global final parameters of the model; Based on the personalized model parameters, obtain the trained neural network model.

2. The method for predicting population flow based on geographical similarity and federated learning according to claim 1, wherein The application of the trained neural network model to the standardized geospatial data to obtain the predicted data of the population flow distribution in the target city includes: Apply the trained neural network model to the standardized geospatial data to obtain the predicted data of the population flow spatio-temporal distribution on each grid space within each time step output by the neural network model, and use the predicted data of the population flow spatio-temporal distribution as the predicted data of the population flow distribution in the target city.

3. The method for predicting population flow based on geographical similarity and federated learning according to claim 1, wherein, The determination of the similarity between each source city and the target city, and based on each similarity, the screening out of the similar source city terminals from each source city terminal includes: The geographical spatial data of each of the source cities are sequentially divided according to time and gridded according to geographical space to obtain each standardized source geographical spatial data; Calculate the cosine similarities between each of the standardized source geographical spatial data and the standardized geographical spatial data, and use each of the cosine similarities as the similarity between each of the source cities and the target city; Based on each of the similarities, several similar source cities are selected from each of the source cities, so as to select similar source city terminals from each of the source city terminals, and the similarity of each of the similar source cities is greater than that of the remaining source cities.

4. The method for predicting population mobility based on geographical similarity and federated learning according to claim 1, characterized in that The obtaining of the model global final parameters based on each of the local model parameters includes: Calculate the mean of each of the local model parameters to obtain the model global final parameters.

5. The method for predicting population flow based on geographical similarity and federated learning according to claim 1, characterized in that, When the similar source city terminal trains the neural network, the personalized model parameters are updated, and the update method is to update according to the loss function and the global feature vector of the similar source city. The loss function is the loss function composed of the actual population flow distribution data and the predicted population flow distribution data of the similar source city. The predicted population flow distribution data is the prediction data output by the neural network based on the geographical spatial data of the similar source city, and the global feature vector is the vector form of the geographical spatial data of the similar source city.

6. A crowd flow prediction device based on geographical similarity and federated learning, characterized in that, The device includes the following components: A data preprocessing module, configured to obtain the target original geographical spatial data of the target city, and preprocess the target original geographical spatial data to obtain standardized geographical spatial data. The geographical spatial data includes population density, traffic road network, and economic production capacity data; A prediction module, configured to apply the trained neural network model to the standardized geographical spatial data to obtain the predicted population flow distribution data of the target city. The predicted population flow distribution data is used to represent the flow data of people in each grid space within the target city, and the predicted population flow distribution data includes the population quantity in each grid space; The preprocessing of the target original geographical spatial data to obtain standardized geographical spatial data includes: Divide the target original geographical spatial data according to time to obtain local data information within each time step; Grid the local data information within each time step according to geographical space to obtain local geographical spatial data information within each time step. The local geographical spatial data information corresponds to the grid space; Normalize the local geographical spatial data information within each time step to obtain standardized local spatial data; Based on the standardized local spatial data within all time steps, obtain standardized geographical spatial data; The training method of the trained neural network model includes: Obtain the local model parameters of each pre-trained neural network model, where each of the pre-trained neural network models is obtained by each source city terminal training a neural network based on a training dataset, and the training dataset includes the geospatial data of the source city and the actual population flow distribution data corresponding to the geospatial data; Obtain the global final model parameters based on each of the local model parameters; Determine the similarity between each of the source cities and the target city, and based on each of the similarities, screen out the similar source city terminals from each of the source city terminals; Send the global final model parameters to the similar source city terminals, and obtain the personalized model parameters of the personalized neural network model sent by the similar source city terminals, where the personalized neural network model is a model obtained by the similar source city terminals continuing to train the neural network model based on the global final model parameters; Obtain the trained neural network model based on the personalized model parameters.

7. A terminal device, characterized in that, The terminal device includes a memory, a processor, and a population flow prediction program based on geographical similarity and federated learning stored in the memory and executable on the processor. When the processor executes the population flow prediction program based on geographical similarity and federated learning, the steps of the population flow prediction method based on geographical similarity and federated learning according to any one of claims 1-5 are implemented.

8. A computer-readable storage medium, characterized in that, A population flow prediction program based on geographical similarity and federated learning is stored on the computer-readable storage medium. When the population flow prediction program based on geographical similarity and federated learning is executed by a processor, the steps of the population flow prediction method based on geographical similarity and federated learning according to any one of claims 1-5 are implemented.

Citation Information

Patent Citations

  • Urban people flow prediction method based on space-time dynamic neural network

    CN112257934A

  • Cross-city federal migration model training method, device, system and equipment

    CN115935189A

  • Cross-city crowd flow trend prediction method based on city function area matching

    CN117649028A